ImgIng · Browser Image EngineOpen Video Beta
ImgIng / Replace video backgrounds with temporal AI in Chrome
TEMPORAL AI · BETA

Replace video backgrounds with temporal AI in Chrome — Beta

Your video is not uploaded. This Beta primarily supports desktop Chrome and Edge. Models download on demand from ModelScope CDN with verified site fallback; decoding, matting, compositing, encoding and audio muxing stay on-device.

Beta releaseNo video uploadThree temporal model tiersAspect ratio and audio retained

Capability quick facts

These facts describe the current product, not unshipped roadmap work.

Status
Beta; desktop Chrome and Edge are the primary targets
Models
RVM 7.2 / 51.3 MB; MatAnyone2 plus first-frame calibration about 345 MB (non-commercial)
Processing
Local aspect-preserving decode, calibrated temporal matting, composite and audio mux
Outputs
MP4 by default for opaque backgrounds; WebM alpha for transparency

Reviewed by the DataDance product and engineering team · Published · Updated

Finish in three steps

Confirm important settings before export; the processing location is always disclosed.

01

Import a video and choose a background

Import MP4, WebM or MOV, then choose transparent, solid, local image or blurred-source background without stretching or cropping the frame.

02

Choose a model and preview

Pick 7.2 MB RVM compatible, 51.3 MB RVM quality, or the MatAnyone2 non-commercial test tier calibrated on its first frame by RVM ResNet50.

03

Process locally and export

WebCodecs encodes the frame-by-frame result and remuxes source audio. Opaque backgrounds default to MP4; transparency uses WebM.

Why not run a still-image cutout model on every frame?

Independent frames can make hair and translucent edges flicker. ImgIng’s RVM path carries recurrent temporal state and resets it on detected scene cuts. The experimental tier first uses RVM ResNet50 to calibrate first-frame hair and the complete silhouette, then passes that seed to MatAnyone2 memory propagation to reduce initial omissions, jitter and previous-shot residue.

Which of the three models should I choose?

RVM MobileNetV3 FP16 is about 7.2 MB and suits quick previews, typical devices and longer video. RVM ResNet50 FP16 is about 51.3 MB and trades speed for steadier detail. MatAnyone2 automatically selects a landscape, square or portrait profile of about 294 MB and reuses the 51.3 MB RVM ResNet50 model for first-frame calibration, about 345 MB in total; it is for non-commercial testing only. Every model uses ModelScope CDN first, falls back to the same pinned version on ImgIng, and enters browser cache only after verification.

Will unusual video ratios be stretched?

No. The editor preserves the original frame ratio and export dimensions. It does not force the source into a square or crop the visible frame for inference. MatAnyone2 only selects the internal landscape, square or portrait input whose usable analysis area is largest. First-frame calibration proportionally caps ultra-high-resolution input at a 1920-pixel long edge to control memory, without changing final output dimensions. Background images independently support cover or contain.

What if the subject is not present at the start?

The experimental tier no longer aborts a full export merely because the opening frame has no subject. Those frames output the selected background while the pipeline keeps looking; temporal memory starts when a person is detected. A scene cut clears the previous memory and runs calibration again.

Can a normal Chrome browser run it?

Current desktop Chrome and Edge are the primary targets. The page needs HTTPS or localhost, WebCodecs and WebAssembly. RVM uses the stable WASM inference path; MatAnyone2 prefers WebGPU and can fall back to WASM. Older devices, phones and long 4K video may be impractical, so the workspace detects capabilities and reports download and processing ETA.

Beta limits: preview the current frame before processing the full video. Complex occlusion, fast motion, translucent objects and scene cuts can still produce edge errors. “Any aspect ratio” means no stretching or cropping; it does not promise identical matting accuracy for every extreme composition. Transparent video exports as WebM VP9 alpha; MP4 is for solid, image or blurred opaque backgrounds.

Updated 2026-08-20 · Live capability detection inside the tool is authoritative

Frequently asked questions

These visible answers match the current product behaviour and structured data.

Does the video upload to a server?

No. The complete video editing pipeline stays in the browser; the network is used only to download the runtime and selected model on first use.

Why is the MatAnyone2 download much larger?

It loads an approximately 294 MB temporal profile for the current aspect class and reuses the roughly 51.3 MB RVM ResNet50 model for first-frame calibration, about 345 MB total. It takes longer and is for non-commercial testing only.

Is the original audio retained?

Yes by default. Compatible audio is copied; otherwise it is encoded to Opus for WebM or AAC for MP4 and remuxed.

Explore ImgIng

Each task page documents real settings, limits and format advice—not keyword-swapped duplicates.