Skip to content

[Task Proposal] Editing-magic: produce cut-based magic effects from raw takes (vanish / outfit swap / reverse restore) #89

Description

@lijrjyan

[Task Proposal — revised] Editing-magic: produce cut-based magic effects from raw takes (disappear / outfit swap / reverse restore)

Revision note. This proposal originally described a covert-splice forensic timeline task (magic performance, machine-truth GT, deterministic F1). We built it end to end — generator, stdlib judge, six measured anti-shortcut gates, and three adversarial hardening rounds against real agent rollouts. We are revising rather than continuing because the calibration data killed the shape, not the implementation: localize-an-anomaly tasks fall to scan-then-verify (Codex medium reached 0.86, xhigh 0.95 on our hardest instance, raw trajectories kept), and a task human editors cannot do either sits outside this benchmark's "expert-strong, agent-weak" premise. Full attack/defense data is available on request and will ship with the PR. The revised proposal keeps the magic domain and the deterministic-verification machinery, and moves the task to what the benchmark actually measures: doing post-production work.

One-line

The agent receives raw locked-camera takes and a production brief, and must edit the magic effect into existence — the cut-based illusions that editing tutorials teach (jump-cut vanish, beat-matched outfit swap, reverse-motion restore) — delivering final.mp4 + a source manifest, scored by deterministic code only.

Editing magic is post-production in its purest form: the trick does not exist in the footage; the cut is the method. Tutorials for these effects are a standard editing exercise (60+ tutorials across 40+ channels verified via search + yt-dlp), so each task mirrors a real, teachable editing workflow with a natural human baseline.

Task 1 — magic-vanish-cut

Brief: from the provided takes (subject present; clean plate) cut a 3–5 s clip in which the subject fully vanishes on the clap. Background and ambience stay continuous; no black frames, dissolves, or camera jumps.

Task 2 — magic-outfit-swap

Brief: from two takes of the same repeated jump in different outfits, cut a 4–6 s clip in which the outfit changes at the jump apex while pose, background, and ambience stay continuous.

Task 3 — magic-reverse-restore

Brief: from a single take of papers scattering, cut a 5–7 s clip in which the papers fly back into the hand (normal intro, reversed action segment, normal outro), with the specified audio treatment.

The three cover orthogonal editing skills: single-cut deletion with continuity, two-take alignment, and time-direction manipulation.

Submission contract & deterministic scoring

Submission = outputs/final.mp4 + manifest.json (which source segments were used, and the claimed cut points).

Scoring pipeline (pure Python + ffmpeg/numpy; no LLM/VLM judge anywhere):

  1. Format/duration checks.
  2. Source honesty — output frames fingerprint-matched to the declared take segments (the assembly family's approach).
  3. Manifest honesty — claimed cut points must match the rendered file (the repair family's honesty-check pattern, H.1.3): mismatch → 0.
  4. Effect criteria (per task, all measurable): cut lands within tolerance of the clap/apex beat (beat located from the audio track / pose trajectory); background continuity across the cut via masked SSIM on the static region; subject-ROI state change in the required direction (occupancy drop for vanish; torso appearance change with pose-keypoint continuity for swap; strictly decreasing source-time mapping for the reversed segment); audio continuity via cross-correlation, or the brief's specified treatment.
  5. Weighted 0–1 reward.

Gate suite (measured before PR, all programmatic): oracle reference edit = 1.0 and ≥ 0.98 after H.264 re-encode; passthrough / empty / JSON-only / manifest-mismatch = 0; wrong take, wrong order, cut offset ±5/±10 frames, global color-grade masquerading as outfit swap, forward-clip masquerading as reverse — each must score 0 or degrade monotonically. Judge robustness across ffmpeg/NLE exports will be reported.

Media

Raw takes are contributor-shot to a published shooting spec (locked camera, locked exposure/focus, slate beat separated from the effect beat), released CC0/CC-BY, hosted with pinned sha256. This makes provenance trivial, gives every task a clean plate and controlled masks (which the scorer depends on), and means the benchmark can re-shoot or extend instances freely. No third-party footage is used in the shipped tasks.

Why this fits

  • Production realism: the brief is a realistic editing request — these effects are taught as editing exercises; conform-level continuity is exactly review-and-finishing work.
  • Finished media artifact: the deliverable is a rendered final.mp4, verified against the agent's own claims — not a JSON answer about a video.
  • Deterministic verification with oracle-1.0 / baseline-0 by construction, in the spirit of the repair/assembly verifiers.
  • Human baseline exists and is strong: an editing student can do each task well; the interesting question is whether agents can.
  • Long-horizon, tool-heavy: locate beats, align takes, cut, composite audio, render, self-check — a genuine multi-step ffmpeg/NLE workflow.

Status

  • Feasibility scan: tutorial ecosystem, effect-type × verifiability matrix (levitation/morph-class effects excluded as not decidable without perceptual judges)
  • Judge suite built and gate matrix measured on seeded synthetic takes: oracles 1.000/1.000/1.000 (and ≥0.98 after CRF-28 H.264 re-encode — measured 1.000); every hard negative (passthrough, empty, JSON-only, manifest-mismatch, wrong take, wrong order, forward-as-reverse, color-grade-as-swap, audio mismatch) = 0; cut offsets degrade monotonically (±5 fr → 0.95/0.75/0.95, ±10 fr → 0.825/0.625/0.825); bit-reproducible material generation; threshold stability spot-checked on real CC footage (masked SSIM 0.93, audio xcorr 0.95)
  • Shooting spec finalized; takes shot and hosted
  • Oracle reference edits; gate measurements
  • Agent calibration with raw trajectories
  • PR

Questions for maintainers

  1. Placement: these are production-shaped tasks (rendered deliverable + deterministic verifier, closest kin to repair/assembly) rather than understanding-QA. Should they target agentic_vbench_understanding anyway, a future versioned post-production suite, or a new community area? We will build to whichever contract you point at.
  2. Human baseline: v1.0 reports expert medians per task; the understanding README does not mention one. For community tasks of this shape, do you want a human-editor baseline included in calibration?
  3. Is contributor-shot CC0 media with a published shooting spec the provenance you prefer for tasks whose scorer depends on controlled capture conditions?

(The original splice-forensics thread below is preserved; its generator, judge, and measured gates are reusable and available.)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions