π Accepted to SIGGRAPH Asia 2026
Yuancheng Xu1,2, Mingming He2, Pablo Salamanca1,2, Li Ma2, Yash Kant1,2, Emmett Steven1, Paul Debevec1,2, Ning Yu1,2
1Netflix, 2Eyeline Labs
ID-V2V is a research exploration for identity-preserving video restylization, and is for demonstration and inspiration purposes only. Given a source video plus a stylized keyframe frame (and optionally additional keyframes), it generates a new video where the scene, lighting, and style follow the keyframe(s) while strictly preserving the source characters' identity and performance such as subtle expressions, eye gaze, and body movement β enabling a shoot first, restyle later workflow for visual storytelling.
restylization_demo.mp4
This repository contains:
- Environment setup β install the environment and download checkpoints
- Model inference β the general preprocess β generate flow, with input/output structure
- Use cases in detail β recipes for restylization, imperfect keyframe, longer video, multiple keyframes, relighting
- Two model variants β
idv2vand an optional normal-depth-augmented variant - Repository layout
A single unified uv environment covers preprocessing and inference.
git clone https://github.com/Eyeline-Labs/ID-V2V && cd ID-V2V
uv sync && source .venv/bin/activatePython 3.10, torch 2.6+cu118, transformers 5.9, xfuser, flash-attn 2.7. Tested on 8Γ A100-80GB. ffmpeg is auto-installed on first run.
SAM3 is gated on Hugging Face, so log in once first, then download the checkpoints into checkpoints/:
hf auth login # paste a token with read access to facebook/sam3
bash scripts/download_checkpoints.sh # idv2v.pth + SAM3 + Wan2.1's T5 + VAE + tokenizer + CLIP (~96 GB)This fetches the idv2v.pth checkpoint (from Eyeline-Labs/ID-V2V β checkpoints/idv2v.pth) plus SAM3 and Wan2.1's T5 + VAE + tokenizer + CLIP β everything the default model needs. (Add --skip-idv2v if you already have the checkpoint, and set MODEL_CHECKPOINT to point at it.) See checkpoints/README.md for per-file sources.
The ID-V2V model takes a source video, a stylized first frame (the look you want), an optional set of keyframes, and a text prompt, and generates a stylized video that follows the keyframe(s) while preserving the source's identity and performance. Running it is two steps: a preprocessing pass that derives the control signal from the source, followed by the generation step. Generation runs at 720p.
Every run reads its inputs from a single sample directory. A full example (with multiple keyframes and a longer source) looks like:
my_sample/
βββ source.mp4 (REQUIRED) source video (any length)
βββ stylized_first_frame.png (REQUIRED) the stylized first frame β the look you want (frame 0)
βββ prompt.txt (REQUIRED) text prompt
βββ keyframes/ (OPTIONAL) extra frame-level anchors at chosen indices
βββ 40.png filename N = inject at 0-based frame N (N β₯ 1)
βββ 80.png
stylized_first_frame.png pins the first frame; each keyframes/<N>.png pins frame N. You create these stylized frames with any image-editing tool (e.g. NanoBanana). The model generates clip by clip if the source video is longer than 81 frames.
The default idv2v model conditions on a single control signal β foreground-on-gray pixels (the person segmented by SAM3, background grayed out). Derive it with:
# SAMPLE_DIR is the input directory (here the bundled restylization example; use your own laid out like my_sample/ above).
# SAM_PROMPT is the SAM3 text prompt for what to segment (default "person"; e.g. "head", "dog"):
SAMPLE_DIR=test_samples/restylization/two_sitting_woman SAM_PROMPT=person bash scripts/preprocess.shThis writes the condition video to <SAMPLE_DIR>/preprocessing/orig_pixel.mp4. Relighting is the one exception β it needs no preprocessing; see Use cases in detail.
Run it on the same SAMPLE_DIR (uses checkpoints/idv2v.pth from Environment setup):
SAMPLE_DIR=test_samples/restylization/two_sitting_woman bash scripts/infer.sh- Multi-GPU auto-enables when
GPUis a comma-separated list with >1 id (e.g."0,1,2,3,4,5,6,7") β it launchestorchrunwith USP sequence parallel. For a single GPU setGPU="0"(plainpythonwith CPU-offload; slower but fits one card). - Keyframes are auto-discovered from
<SAMPLE_DIR>/keyframes/<N>.png. MAX_NUM_FRAMEScaps the output length. The video is generated in overlapping 81-frame clips (NUM_FRAMES_PER_CLIP=81), so a longer source simply produces more clips β no separate long-video mode.- The checkpoint is memory-mapped on load, so all GPUs share one copy.
STAGE_CHECKPOINT_TO_SHM=true(default) first copies the.pthinto/dev/shm(RAM) for a big speedup on slow/network storage; it needs free RAM β₯ the checkpoint size β set itfalseif you hit CPU out-of-memory.
The run prints a benign
diffsynthwarning (using --use_multi_control_vace but only 1 control conditions is provided). This is expected and safe to ignore.
my_sample/outputs/<run_name>/
β
β βββ always saved βββ
βββ run_config.json resolved config (incl. prompt)
βββ first_frame.png cropped stylized first frame actually used
βββ original_video.mp4 source re-encoded at output size, trimmed to output length
βββ generated_video.mp4 β FINAL stitched output
βββ flip_test.mp4 plays generated, freezes periodically to A/B vs. source
β
β βββ only when SAVE_VERBOSE=true (default) βββ
βββ condition_0.mp4 the control signal as fed to the model
βββ generated_clip0.mp4 generated_clip1.mp4 ... per-clip generations
βββ generated_video_withOrig.mp4 side-by-side (generated | source) viz
<run_name> defaults to r{W}x{H}_f{F}_kf{N}_idv2v (e.g. r1280x720_f81_kf0_idv2v; F = frames per clip = NUM_FRAMES_PER_CLIP, not the total length, so a multi-clip 240-frame run still shows f81; N = number of keyframes). Set SAVE_VERBOSE=false for the minimal output set. flip_test.mp4 is the headline viz: it plays the generated video and pauses at up to 5 uniformly-spaced frames, alternating Generated (green pill) β Source (red pill) for a quick A/B.
ID-V2V is driven by keyframes β stylized frames that you supply. The five recipes below cover the common ways to use the model. Each has a ready-to-run script in scripts/examples/ and a matching sample under test_samples/ that doubles as the input template: open a recipe's sample folder to see how to lay out your own inputs, then point SAMPLE_DIR at your own directory organized the same way and run that one script. Every sample also ships a generated.mp4 showing the expected result. Finish Environment setup first (that step also fetches checkpoints/idv2v.pth).
What it is. Relight the human subject and regenerate the rest of the scene β background, objects, and overall style.
Why you'd use it. The most general use β reshape the world around a performance without re-filming. You supply a stylized first frame that changes the scene and background while keeping the character's pose and subtle expression, and ID-V2V carries that new look across the whole video while preserving the source's identity, expression, gaze, and lip-sync.
Run it:
# SAMPLE_DIR points at the input directory (swap it for your own, laid out like the sample):
SAMPLE_DIR=test_samples/restylization/two_sitting_woman bash scripts/examples/restylization.shThis runs SAM3 preprocessing (foreground-on-gray condition) using "person" as the SAM3 prompt by default, then generates the video. Samples: test_samples/restylization/.
What it is. The same restylization flow, but with a first frame that does not perfectly match the source.
Why you'd use it (and why it matters). Image-editing models like NanoBanana often can't hold the exact pose or expression when they restyle a frame, so the stylized first frame ends up non-aligned with the source video's first frame. That's fine β you don't need a pixel-perfect keyframe. Frame 0 follows your imperfect keyframe, but from the second frame onward ID-V2V re-aligns to the source video's identity and performance, correcting the mismatch automatically.
Run it β identical command and script to restylization; only the sample differs:
SAMPLE_DIR=test_samples/non_aligned_keyframe/woman_phone bash scripts/examples/restylization.shSamples: test_samples/non_aligned_keyframe/.
What it is. Generation for a source video longer than a single 81-frame clip. ID-V2V generates the video clip by clip, conditioning each new clip on the end of the previous clip (the frame where they overlap), so the clips join seamlessly into one continuous, drift-controlled video.
Run it:
SAMPLE_DIR=test_samples/longer_video/woman_dancing bash scripts/examples/longer_video.shThe example defaults to MAX_NUM_FRAMES=240 (β 3 clips); set it to whatever length you want. Preprocessing and generation are the same as restylization. Samples: test_samples/longer_video/. For the best performance, we recommend using more keyframes when the source video is longer.
What it is. Pinning more than just the first frame.
Why you'd use it. Sometimes one keyframe isn't enough control β e.g. you want to fix both the first and last frame (or a mid-point) so the video lands exactly where you intend. Drop edited frames into keyframes/<N>.png and the model pins each one. The keyframe can be any frame.
Run it:
SAMPLE_DIR=test_samples/first_last_frame/two_women_spotlight bash scripts/examples/first_last_frame.shThe sample pins both ends β stylized_first_frame.png (first frame) + keyframes/80.png (last frame of an 81-frame clip), which infer.sh auto-discovers. Samples: test_samples/first_last_frame/.
What it is. Change only the lighting, keeping the full scene β both the foreground human subjects and the background β exactly as shot.
Why you'd use it (and how it differs from restylization). Relighting is a more restricted case of restylization: whereas restylization can edit the background and the overall scene and substantially change the video's appearance, relighting alters nothing but the illumination. To do this, use an image-editing model to relight the first frame while keeping everything, including the background, then feed it to ID-V2V directly β with no SAM3 mask, no grayed-out background, and no preprocessing at all. The raw source video is used as the control, so the whole scene is preserved and only the lighting changes.
Run it:
SAMPLE_DIR=test_samples/relighting/two_sitting_woman bash scripts/examples/relighting.shThis skips preprocessing entirely and uses scripts/infer_relighting.sh (the raw source video is the condition). Samples: test_samples/relighting/.
Everything above uses the default idv2v model, which is driven by a single per-frame VACE control signal. There is also an alternate variant that adds two more conditions:
| Variant | VACE conditions | Preprocess | Inference | Checkpoint |
|---|---|---|---|---|
idv2v β default, recommended |
1: foreground-on-gray pixels (SAM3-segmented person) | scripts/preprocess.sh |
scripts/infer.sh |
idv2v.pth |
idv2v_with_normal_depth β alternate |
3: foreground-on-gray pixels + surface-normals (DAViD) + depth (DepthAnything-V2) | scripts/idv2v_with_normal_depth/preprocess_with_depth.sh |
scripts/idv2v_with_normal_depth/infer_with_depth.sh |
idv2v_with_normal_depth.pth |
The default idv2v keeps (relights) whatever is inside the SAM3 mask (the segmented person) and freely regenerates the rest of the frame from the text prompt. idv2v_with_normal_depth additionally constrains those other regions with the source video's depth β reach for it when you want that depth control, but only if the source video's depth already matches the scene you intend to generate. Its checkpoint plus two extra preprocessing models (DAViD, DepthAnything-V2) are fetched by adding --with-depth to the download:
bash scripts/download_checkpoints.sh --with-depth # + idv2v_with_normal_depth.pth + DAViD + DepthV2 (~75 GB)Its preprocessing and inference mirror the default flow via scripts/idv2v_with_normal_depth/preprocess_with_depth.sh and scripts/idv2v_with_normal_depth/infer_with_depth.sh; it saves condition_0/1/2.mp4 (pixels + normals + depth) and uses the ..._idv2v_with_normal_depth run-name suffix. The two checkpoints are different, non-interchangeable weights with the same architecture, so loading the wrong .pth into a script does not error β it silently produces poor output. Match the checkpoint to the script. See checkpoints/README.md.
ID-V2V_public/
βββ scripts/
β βββ download_checkpoints.sh fetch public weights (default: SAM3 + Wan; --with-depth adds DAViD + DepthV2)
β βββ preprocess.sh idv2v: SAM3 + origPixel (foreground-on-gray condition)
β βββ infer.sh idv2v: single-condition video generation
β βββ infer_relighting.sh idv2v relighting: raw source as the condition (no preprocessing)
β βββ examples/ one runnable script per recipe (restylization, relighting, longer_video, first_last_frame)
β βββ idv2v_with_normal_depth/ ALTERNATE idv2v_with_normal_depth variant (secondary):
β βββ preprocess_with_depth.sh SAM3 + origPixel + DAViD + DepthV2
β βββ infer_with_depth.sh three-condition video generation
βββ src/idv2v/
β βββ preprocess/ SAM3, origPixel, DAViD, DepthV2
β βββ inference/ pipeline.py (shared 1β3 condition engine) + flip_test.py
βββ diffsynth_studio/ Wan pipeline
βββ test_samples/ ready-to-run examples grouped by use case (source.mp4 + stylized_first_frame.png + prompt.txt + generated.mp4)
βββ checkpoints/ gitignored β populated by download_checkpoints.sh
Builds on DiffSynth-Studio (Wan series), SAM3 (segmentation), DAViD (surface normals, idv2v_with_normal_depth only), DepthAnything v2 (depth, idv2v_with_normal_depth only), and Stable-Video-Infinity (long-video generation).
