A curated list of AI agents and agentic frameworks for video editing, video production, and video understanding-for-production.
Video creation is rapidly shifting from single-shot model inference toward agentic systems that plan, decompose tasks, use tools, orchestrate specialized roles (director, screenwriter, editor, cinematographer), and self-correct. This list tracks the projects driving that shift.
A project qualifies if it is an agent or agentic framework whose primary purpose is video editing, video production, or video understanding-in-service-of-production. It must satisfy at least one of:
- Multi-step planning / tool-use for video tasks (storyboarding, shot selection, editing decisions, post-production)
- Multi-agent orchestration for video pipelines (director / writer / editor / producer roles)
- Agents that interact with video editing software, video APIs, or composition primitives
- LLM/VLM-driven controllers wrapping video generation, editing, or understanding models
Out of scope (and curated elsewhere):
- Pure video-generation foundation models with no agent layer (Sora, Veo, Kling, Runway core models) — see awesome-any2any-models for unified multimodal models
- Benchmarks and evaluation suites for multimodal agents — see awesome-multimodal-agent-benchmarks
- Generic NLE software, video diffusion architectures, or tracking/segmentation libraries
- All-in-One Agentic Frameworks
- Multi-Agent Pipelines (Director / Writer / Editor)
- Video Editing Agents
- Video Generation / Production Agents
- Video Understanding Agents (for Production)
- NLE & Software-Control Integrations (MCP & Tools)
- Related Surveys & Papers
- Contributing
- License
End-to-end systems that span understanding, editing, and generation under a single agentic controller.
- HKUDS/VideoAgent — All-in-one agentic framework for video understanding, editing, and remaking with intent-to-agent workflow routing.
code - video-db/Director — Open-source framework for building video agents that reason across search, edit, compile, and generate over a VideoDB backend.
code - HKUDS/ViMax — 12 specialized agents (director, screenwriter, producer, etc.) for end-to-end multi-shot video generation with RAG long-script design.
code - diffusionstudio/agent — Agentic video editing framework built on a browser-based WebCodecs compositing engine.
code
Systems that explicitly simulate film-crew roles via multi-agent collaboration.
- showlab/MovieAgent — Automated movie generation via multi-agent CoT planning across director, screenwriter, storyboard, and location agents.
papercode - HITsz-TMG/FilmAgent — LLM multi-agent collaboration for end-to-end film automation in virtual 3D spaces (SIGGRAPH Asia 2024).
papercode - HITsz-TMG/Anim-Director — Large multimodal model agent that autonomously directs controllable animation video generation (SIGGRAPH Asia 2024).
papercode - Anim-Director/AniMaker — Multi-agent animated storytelling with MCTS-driven candidate clip generation (SIGGRAPH Asia 2025).
papercode - Vchitect/Vlogger — LLM-as-director decomposes vlog generation into Script, Actor, ShowMaker, and Voicer roles (CVPR 2024).
papercode - HL-hanlin/VideoDirectorGPT — LLM-guided planning for consistent multi-scene video generation with a layout-grounded video module (COLM 2024).
papercode - X-PLUG/MM_StoryAgent — Open-source multi-agent paradigm for immersive narrated storybook video across text, image, and audio.
papercode - DreamFactory — Multi-agent framework with director / art-director / screenwriter / artist roles for multi-scene long video generation.
paper - AesopAgent — Agent-driven evolutionary RAG system from DAMO that turns story proposals into scripted, scored, voiced videos.
paper - multimodal-art-projection/AutoMV — Multi-agent music video generation from raw audio plus lyrics with screenwriter, director, and verifier agents.
papercode - Co-Director — Hierarchical multi-agent framework for generative video storytelling along creative-strategy / narrative-mode / aesthetic axes.
paper
Agents focused on editing existing footage — cut planning, trimming, montage, color, captions.
- browser-use/video-use — Edit videos with Claude Code: word-boundary cuts, color grading, subtitles, self-evaluating output. Designed as a Claude Code skill.
code - GVCLab/CutClaw — Autonomous multi-agent framework for hours-long montage editing with music synchronization (Playwriter / Editor / Reviewer agents).
papercode - EditDuet — Editor + Critic multi-agent system for non-linear video editing from natural language (SIGGRAPH 2025, Adobe Research).
paper - LAVE — LLM-powered plan-and-execute agent for video editing with language-augmented UI (IUI 2024).
paper - GLANCE — Global-local coordination multi-agent framework for music-grounded non-linear mashup editing.
paper - Prompt-Driven Agentic Video Editing — Modular pipeline using hierarchical semantic indexing for long-form, story-driven editing.
paper
Agents that wrap or compose generative video models for controllable, multi-shot, or long-form output.
- lichao-sun/Mora — Generalist video generation via a multi-agent framework that composes specialized visual agents to cover T2V / I2V / editing tasks.
papercode - DuNGEOnmassster/VideoGen-of-Thought — Step-by-step multi-shot video synthesis from one sentence via dynamic storyline modeling (NeurIPS 2025 Workshop).
papercode - GenMAC — Iterative DESIGN / GENERATION / REDESIGN multi-agent loop for compositional text-to-video (AAAI 2025).
paper - Video-as-Agent/VideoAgent — Self-improving video generation that refines plans with VLM and execution feedback (for embodied planning).
papercode - Vibe AIGC — Paradigm paper on agentic orchestration as a bridge between high-level creator intent and stochastic generative models.
paper
Understanding agents whose outputs feed editing, search, indexing, or selection — included when used agentically (planning, tool-use, memory), not as static VLM inference.
- wxh1996/VideoAgent — LLM as central agent that iteratively calls VLM tools to answer queries about long-form video (ECCV 2024, Stanford).
papercode - YueFan1014/VideoAgent — Memory-augmented multimodal agent with structured temporal + object memory for video understanding (ECCV 2024).
papercode - z-x-yang/DoraemonGPT — Dynamic-scene understanding agent with symbolic task memory, sub-task tools, and MCTS planning (ICML 2024).
papercode - Ziyang412/VideoTree — Query-adaptive hierarchical tree representation for long-video LLM reasoning (CVPR 2025).
papercode - HKUDS/VideoRAG — Retrieval-augmented generation over extreme long-context video corpora; powers the Vimo desktop chat-with-video app (KDD 2026).
code - yiwengxie/Chat-Video — Tracklet-centric versatile video understanding system enabling instance-level chat with videos.
papercode - VCA: Video Curious Agent — Curiosity-driven self-exploration agent with tree-search over video segments for efficient long-video QA.
paper
Agents and MCP servers that let LLMs drive professional editors (Premiere, DaVinci) or compose FFmpeg pipelines.
- mikechambers/adb-mcp — Reference MCP interface that exposes Adobe Photoshop and Premiere to LLM clients.
code - ayushozha/AdobePremiereProMCP — Premiere Pro MCP server with 1,000+ tools across timeline, color, audio, effects, and export.
code - hetpatel-11/Adobe_Premiere_Pro_MCP — Premiere Pro MCP covering project ops, ingest, sequence creation, transitions, effects, exports.
code - leancoderkavy/premiere-pro-mcp — MCP server for Adobe Premiere Pro via CEP/ExtendScript with 269 tools across 28 modules.
code - samuelgursky/davinci-resolve-mcp — MCP server integration for DaVinci Resolve Studio scripting API.
code - lordhoell/davinci-resolve-mcp — Claude Code skill + MCP exposing 440+ DaVinci Resolve tools for AI-assisted editing, color, and rendering.
code - wizenheimer/vibestudio — Headless zero-runtime FFmpeg MCP server in pure Bash for agent-driven editing pipelines.
code - burningion/video-editing-mcp — MCP interface for Video Jungle that produces OpenTimelineIO projects for DaVinci Resolve.
code - remyxai/FFMPerative — LLM-powered chat copilot that composes FFmpeg edits from natural language.
code - BAAI-Agents/Cradle — Generalist computer-control agent (screen-in, keyboard/mouse-out) that can drive CapCut, Meitu, and other editors.
papercode - showlab/Kiwi-Edit — Unified open-source framework for instruction-guided and reference-guided video editing in natural language.
papercode
- VideoGen-Eval — Agent-based system for video generation evaluation; useful prior art for agent-as-judge in production pipelines.
papercode - yunlong10/Awesome-LLMs-for-Video-Understanding — Companion survey of video-LMM literature (IEEE TCSVT).
- yunlong10/Awesome-Video-LMM-Post-Training — Tracks reasoning-oriented post-training of video LMMs (relevant when fine-tuning agent backbones).
- wentianli/awesome-video-editing — Paper list on cinematographic video editing and related CV tasks; broader than the agentic slice tracked here.
PRs welcome. See CONTRIBUTING.md for inclusion criteria and entry format. The bar is: agentic (planning / tool-use / multi-role / controller), video-focused, and substantive (paper, working code, or production system).
To the extent possible under law, contributors have waived all copyright and related rights to this work.
