How We Remove Watermarks from Video Fast: Motion-Adaptive Frame Skipping + Temporal EMA
(and the bug that froze our masks — a regression story about wiring performance switches to product options) Running an inpainting model on every frame of a video is the obvious way to do video watermark removal — and unaffordable in a browser. This is the write-up of how we got 58–67% of frames to

(and the bug that froze our masks — a regression story about wiring performance switches to product options) Running an inpainting model on every frame of a video is the obvious way to do video watermark removal — and unaffordable in a browser. This is the write-up of how we got 58–67% of frames to skip the model entirely: a motion-adaptive, anchor-based frame-skipping gate, a temporal EMA for flicker-free output, plus the failures behind the design — a chain-freeze bug (PSNR 38.7 → 16.7), an anchor-age regression (38.30 → 34.69), and a UX "simplification" that silently disabled the whole optimization and cost us 3x. All numbers are from our own benchmarks; all code is from the shipping pipeline. ClearPix removes watermarks entirely client-side: ONNX-runtime inpainting on WebGPU or WASM, fed by WebCodecs decode, encoded back with MediaBunny. No server, no upload. The profile was blunt. Inside one inpaint call, planning cost about 4.4ms while the model run() segment took 97% or more. Per frame, inpainting took 2100–3500ms on WASM; the WebGPU execution provider brought that to roughly 400ms — still the dominant line item. Pipeline overhead was already cut elsewhere (ffmpeg.wasm → WebCodecs + MediaBunny: 95s → 9s per job). The biggest lever left was not a faster model — it was not running the model at all. The observation that makes skipping viable: most watermarks sit on content that barely changes — a corner logo over a talking head, a username stamp over a screen recording. Under a static mask, the pixels inside the mask bounding box are nearly identical frame to frame. If nothing in the region moved since the last inference, that output is already the correct answer — composite it onto the current frame instead of inferring again. The per-frame gate, trimmed from RemoveVideoWatermark.tsx: const small = makeRegionGray(img, maskBbox); // bbox region → ~64px-wide grayscale // Compare against the anchor (last real inference frame); drift measured // exactly against the anchor too const mae = st.anchorSmall ? regionMAE(small, st.anchorSmall) : Infinity; const deltaAnchor = st.anchorRaw ? estimateRegionDelta(img, st.anchorRaw, maskData, maskBbox) : [0, 0, 0]; if (mae < st.skipThreshold && st.anchorOut) { // Skip: composite the anchor's feathered mask region onto the current frame out = compositePrevRegion(img, st.anchorOut, maskData, maskBbox, deltaAnchor); st.skipped++; } else { out = await runInpaint(); st.anchorRaw = img; // a real inference becomes the new anchor st.anchorOut = out; st.anchorSmall = small; } The cheap part matters: the diff runs on the mask bbox downsampled to ~64px grayscale — under 5ms per frame, against hundreds for inference. On a skip, compositePrevRegion pastes the anchor's mask region back with the inpainter's own feathering semantics (100% inside the mask, transition band on clean pixels outside), so a reused frame looks like an inferred one. The current build skips 58% of frames on our watermark e2e clip (threshold 5.00) and 67% on the subtitle e2e (threshold 4.21); the earlier gate measured 94 of 144 (65%). The threshold is not a magic constant. The first five frames always run full inference; we collect their MAEs and set skipThreshold = median × 3, clamped to [2, 5]/255 — dark and noisy footage get different baselines for free. (The ceiling of 5 is itself a measured fix — keep reading.) The first version (v1) compared each frame to its immediate neighbor and reused the previous output. It passed screen-recording tests, then failed catastrophically on a slow-gradient clip. On gradient footage each frame drifts ~1/255, so per-frame MAE stays below threshold and the gate skips forever — the mask region freezes at the last inference while the rest of the frame changes around it: a visible, motionless rectangle. E2e masked-region PSNR collapsed from 38.7 to 16.7. Patch one was a cumulative drift budget: accumulate skipped-frame MAE since the last inference and force re-inference when the total crosses the threshold. That recovered PSNR only to 22.9 — reused pixels still lagged the scene's exposure on gradients. Patch two added delta compensation (estimateRegionDelta): the per-channel mean difference between current and previous raw frames over the non-mask pixels in the bbox, added to historical pixels before blending. Budget + delta landed at 38.44 (vs a 38.74 no-skip reference), inside tolerance, with a clean three-clip visual check in real Chrome. That version shipped. But the budget was always a proxy: the error we care about is "how far is this frame from the frame whose output we're reusing," and chained per-step MAEs only approximate that. So we refactored (v2): compare each frame directly against the anchor — the most recent frame that really ran inference — and reuse the anchor's output on skips. Skip error is now exactly what the threshold bounds; the drift budget was deleted outright. Delta compensation became exact too — measured against the anchor itself. The first anchored run gave us a proper scare: subtitles gained +3–4dB on all four e2e segments, but the gradient watermark clip dropped from 38.30 to 34.69dB masked-region. Root cause: that clip's drift varies with y (a sinusoidal exposure gradient), so one mean delta can't fully compensate it, and the residual grows with anchor age. Two measured fixes: Tighter threshold ceiling: 8 → 5. Anchors refresh more often, bounding how stale a reuse can get. Split the deltas. Compositing reuses the anchor → to-anchor delta; the EMA blends against the previous frame's output → to-neighbor delta. One mean delta serving both timebases was the bug. Final numbers: e2e-video masked region 39.84dB (+1.5dB over the chained baseline), 58% skipped (threshold 5.00); subtitles 37.3 / 32.3 / 28.7 / 30.1dB across four segments, 67% skipped (threshold 4.21). Skipping frames solves cost, but video inpainting temporal consistency needs one more layer. Even frames that do run inference can jitter: the model hallucinates slightly different textures each time, and at 30fps that reads as flicker inside the repaired region. Our answer: a per-pixel exponential moving average, applied only inside the mask dilated by 4px, with a motion-adaptive blend factor (applyTemporalEMA, trimmed): // Per-pixel motion from raw frames (0..1) const m = (|dr| + |dg| + |db|) / (3 * 255); const t = Math.min(1, m / 0.08); // motion 0→8% const s = t * t * (3 - 2 * t); // smoothstep const a = 0.4 + 0.6 * s; // alpha 0.4→1.0 // History term: to-neighbor deltaPrev (composite uses to-anchor deltaAnchor) out = curr * a + (prev + delta) * (1 - a); Low motion → alpha 0.4 (lean on history, flicker dies); high motion → alpha 1.0 (pure current frame, no ghosting), with a smoothstep between. Three boundaries are deliberate: It runs only inside the dilated mask. The rest of the frame is already bit-identical to the source; touching it would only blur real motion. It runs on both inferred and skipped frames, so reused composites get the same treatment as fresh inferences. Its history term is compensated against the previous frame, not the anchor — the split described above. Everything so far assumes a fixed mask — wrong for a watermark that moves. Our "Detect sampling" setting has three rates: First (detect once, reuse for the whole video), Med (every 5 frames), High (every frame). Moving watermarks need Med or High. A rebuilt mask means all temporal state belongs to the old mask position — reusing it would smear content across the frame — so a refresh resets everything: the anchor group (anchorRaw/anchorOut/anchorSmall), the prev group (prevRaw/prevOut), and the dilation cache. No special-casing needed: anchorSmall = null makes the next MAE Infinity, forcing full inference on the re-detect frame (which becomes the new anchor), and prevOut = null makes the EMA skip itself. Between re-detects the watermark usually holds still, so skipping stays fully active on Med/High — the anchor gate handles the motion. (If detection fails or finds nothing, we keep the previous mask and leave temporal state untouched.) The same machinery powers our subtitle-removal tool, which defaults to Med because burned-in subtitles change every few seconds. Now the embarrassing one. In a September UX pass, the "First" sampling option looked like confusing jargon, so we removed it and forced per-frame re-detection for everyone. What we didn't trace: skipping and the EMA are keyed on the mask's bounding box, and with per-frame re-detection the mask is rebuilt every frame, so maskBbox was never stable — the entire optimization path silently disabled itself, and video removal got ~3x slower. Users reported "streaming feels slower" — wrong; comparing the complaint timeline with the commit timeline pointed at the sampling default. We rolled back the next day. The never-again rule: before deleting a "confusing" option, check whether a performance optimization is keyed to it. Performance switches must be explicitly bound to product options — the dependency belongs in a comment, a test, or the option's name, not in someone's memory. Our detectEvery state now carries exactly that comment. Change Measurement Result Skip gate cost mask-bbox frame diff <5ms per frame v1 chain-freeze bug (single-frame MAE) e2e masked-region PSNR 38.7 → 16.7 v1 drift budget only e2e masked-region PSNR 22.9 v1 drift budget + delta (first shipped version) e2e masked-region PSNR / skip rate 38.74 → 38.44; 94/144 skipped (65%), pipeline 20s vs 170–370s WASM v2 anchored, first run gradient clip masked-region PSNR 38.30 → 34.69 (subtitles +3–4dB on all four segments) v2 anchored, final (ceiling 5, split deltas) e2e-video masked-region PSNR / skip 39.84dB (+1.5dB vs chained baseline), 58% skipped (thr 5.00) v2 anchored, final e2e-subtitles, four segments 37.3 / 32.3 / 28.7 / 30.1dB, 67% skipped (thr 4.21) Removing the "First" option (UX pass) video removal speed ~3x regression, rolled back next day Everything here ships in our free video watermark remover — in-browser, no upload, the model runs on your GPU. If your real problem is burned-in text, the same pipeline drives our subtitle removal tool. Feedback on weird footage is how most of these fixes happened. Part 4 of the ClearPix engineering series — how we build free, private, in-browser media tools at clearpix.org.
Key Takeaways
- •(and the bug that froze our masks — a regression story about wiring performance switches to product options) Running an inpainting model on every frame of a video is the obvious way to do video watermark removal — and unaffordable in a browser
- •This story was reported by Dev.to, covering developments in the dev space.
- •AI advancements continue to reshape industries — read the full article on Dev.to for complete coverage.
📖 Continue reading the full article:
Read Full Article on Dev.to →


