Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Awesome Video Agents Awesome

A curated list of AI agents and agentic frameworks for video editing, video production, and video understanding-for-production.

Video creation is rapidly shifting from single-shot model inference toward agentic systems that plan, decompose tasks, use tools, orchestrate specialized roles (director, screenwriter, editor, cinematographer), and self-correct. This list tracks the projects driving that shift.

Scope & Inclusion Criteria

A project qualifies if it is an agent or agentic framework whose primary purpose is video editing, video production, or video understanding-in-service-of-production. It must satisfy at least one of:

  • Multi-step planning / tool-use for video tasks (storyboarding, shot selection, editing decisions, post-production)
  • Multi-agent orchestration for video pipelines (director / writer / editor / producer roles)
  • Agents that interact with video editing software, video APIs, or composition primitives
  • LLM/VLM-driven controllers wrapping video generation, editing, or understanding models

Out of scope (and curated elsewhere):

  • Pure video-generation foundation models with no agent layer (Sora, Veo, Kling, Runway core models) — see awesome-any2any-models for unified multimodal models
  • Benchmarks and evaluation suites for multimodal agents — see awesome-multimodal-agent-benchmarks
  • Generic NLE software, video diffusion architectures, or tracking/segmentation libraries

Contents


All-in-One Agentic Frameworks

End-to-end systems that span understanding, editing, and generation under a single agentic controller.

  • HKUDS/VideoAgent — All-in-one agentic framework for video understanding, editing, and remaking with intent-to-agent workflow routing. code
  • video-db/Director — Open-source framework for building video agents that reason across search, edit, compile, and generate over a VideoDB backend. code
  • HKUDS/ViMax — 12 specialized agents (director, screenwriter, producer, etc.) for end-to-end multi-shot video generation with RAG long-script design. code
  • diffusionstudio/agent — Agentic video editing framework built on a browser-based WebCodecs compositing engine. code

Multi-Agent Pipelines (Director / Writer / Editor)

Systems that explicitly simulate film-crew roles via multi-agent collaboration.

  • showlab/MovieAgent — Automated movie generation via multi-agent CoT planning across director, screenwriter, storyboard, and location agents. paper code
  • HITsz-TMG/FilmAgent — LLM multi-agent collaboration for end-to-end film automation in virtual 3D spaces (SIGGRAPH Asia 2024). paper code
  • HITsz-TMG/Anim-Director — Large multimodal model agent that autonomously directs controllable animation video generation (SIGGRAPH Asia 2024). paper code
  • Anim-Director/AniMaker — Multi-agent animated storytelling with MCTS-driven candidate clip generation (SIGGRAPH Asia 2025). paper code
  • Vchitect/Vlogger — LLM-as-director decomposes vlog generation into Script, Actor, ShowMaker, and Voicer roles (CVPR 2024). paper code
  • HL-hanlin/VideoDirectorGPT — LLM-guided planning for consistent multi-scene video generation with a layout-grounded video module (COLM 2024). paper code
  • X-PLUG/MM_StoryAgent — Open-source multi-agent paradigm for immersive narrated storybook video across text, image, and audio. paper code
  • DreamFactory — Multi-agent framework with director / art-director / screenwriter / artist roles for multi-scene long video generation. paper
  • AesopAgent — Agent-driven evolutionary RAG system from DAMO that turns story proposals into scripted, scored, voiced videos. paper
  • multimodal-art-projection/AutoMV — Multi-agent music video generation from raw audio plus lyrics with screenwriter, director, and verifier agents. paper code
  • Co-Director — Hierarchical multi-agent framework for generative video storytelling along creative-strategy / narrative-mode / aesthetic axes. paper

Video Editing Agents

Agents focused on editing existing footage — cut planning, trimming, montage, color, captions.

  • browser-use/video-use — Edit videos with Claude Code: word-boundary cuts, color grading, subtitles, self-evaluating output. Designed as a Claude Code skill. code
  • GVCLab/CutClaw — Autonomous multi-agent framework for hours-long montage editing with music synchronization (Playwriter / Editor / Reviewer agents). paper code
  • EditDuet — Editor + Critic multi-agent system for non-linear video editing from natural language (SIGGRAPH 2025, Adobe Research). paper
  • LAVE — LLM-powered plan-and-execute agent for video editing with language-augmented UI (IUI 2024). paper
  • GLANCE — Global-local coordination multi-agent framework for music-grounded non-linear mashup editing. paper
  • Prompt-Driven Agentic Video Editing — Modular pipeline using hierarchical semantic indexing for long-form, story-driven editing. paper

Video Generation / Production Agents

Agents that wrap or compose generative video models for controllable, multi-shot, or long-form output.

  • lichao-sun/Mora — Generalist video generation via a multi-agent framework that composes specialized visual agents to cover T2V / I2V / editing tasks. paper code
  • DuNGEOnmassster/VideoGen-of-Thought — Step-by-step multi-shot video synthesis from one sentence via dynamic storyline modeling (NeurIPS 2025 Workshop). paper code
  • GenMAC — Iterative DESIGN / GENERATION / REDESIGN multi-agent loop for compositional text-to-video (AAAI 2025). paper
  • Video-as-Agent/VideoAgent — Self-improving video generation that refines plans with VLM and execution feedback (for embodied planning). paper code
  • Vibe AIGC — Paradigm paper on agentic orchestration as a bridge between high-level creator intent and stochastic generative models. paper

Video Understanding Agents (for Production)

Understanding agents whose outputs feed editing, search, indexing, or selection — included when used agentically (planning, tool-use, memory), not as static VLM inference.

  • wxh1996/VideoAgent — LLM as central agent that iteratively calls VLM tools to answer queries about long-form video (ECCV 2024, Stanford). paper code
  • YueFan1014/VideoAgent — Memory-augmented multimodal agent with structured temporal + object memory for video understanding (ECCV 2024). paper code
  • z-x-yang/DoraemonGPT — Dynamic-scene understanding agent with symbolic task memory, sub-task tools, and MCTS planning (ICML 2024). paper code
  • Ziyang412/VideoTree — Query-adaptive hierarchical tree representation for long-video LLM reasoning (CVPR 2025). paper code
  • HKUDS/VideoRAG — Retrieval-augmented generation over extreme long-context video corpora; powers the Vimo desktop chat-with-video app (KDD 2026). code
  • yiwengxie/Chat-Video — Tracklet-centric versatile video understanding system enabling instance-level chat with videos. paper code
  • VCA: Video Curious Agent — Curiosity-driven self-exploration agent with tree-search over video segments for efficient long-video QA. paper

NLE & Software-Control Integrations (MCP & Tools)

Agents and MCP servers that let LLMs drive professional editors (Premiere, DaVinci) or compose FFmpeg pipelines.

Related Surveys & Papers


Contributing

PRs welcome. See CONTRIBUTING.md for inclusion criteria and entry format. The bar is: agentic (planning / tool-use / multi-role / controller), video-focused, and substantive (paper, working code, or production system).

License

CC0

To the extent possible under law, contributors have waived all copyright and related rights to this work.

About

A curated list of AI agents and agentic frameworks for video editing, video production, and video understanding-for-production.

Resources

Contributing

Stars

14 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors