Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Awesome Any-to-Any Models Awesome

A curated list of true any-to-any multimodal models: all-in, all-out.

Most "unified multimodal" and "omni" lists conflate three very different kinds of model: vision-language models that only output text, text-to-X generators that only emit a single media modality, and the rare model that is genuinely bidirectional across every canonical modality. This list keeps only the last kind.

Inclusion criteria

A model qualifies for the main list only if it satisfies the full any-to-any definition pioneered by NExT-GPT:

"A general-purpose multimodal LLM that accepts inputs and generates outputs in arbitrary combinations of text, image, video, and audio." — NExT-GPT (Wu et al., ICML 2024)

Concretely: the model must accept all of {text, image, audio, video} as input and generate all of {text, image, audio, video} as output. No "≥2 in / ≥2 out" loophole.

Tier definitions

  • Tier 1 — Gold Standard Any-to-Any. Text ⇄ image ⇄ audio ⇄ video, all directions, code and weights released. NExT-GPT, CoDi, CoDi-2 are the canonical members.
  • Tier 2 — Partial / Claimed Any-to-Any. Paper claims most of the quad but one output modality is missing, low quality, or unreleased; the gap is named in the one-liner. Marked ◐ in the matrix.
  • Tier 3 — Closed / Proprietary. API-only / partial release. GPT-4o, Gemini 2.5, Gemini Omni.

Everything else has been removed. In particular:

  • Qwen2.5-Omni, Qwen3-Omni, MiniCPM-o, VITA-1.5, Baichuan-Omni-1.5, InteractiveOmni generate text+speech only — no image or video out. They sit in the "Omni-In / Speech-Out" family, not here.
  • AnyGPT, OmniFlow, Ming-Omni, Ming-Flash-Omni, M2-omni, Omni-Diffusion, Dynin-Omni, AR-Omni, FlowBind cover three of the four canonical modalities but miss video out. They live in Tier 2.
  • LWM has video but no audio. Tier 2.
  • Chameleon, Janus / Janus-Pro / Janus-Flow, Show-o, Show-o2, Transfusion, Emu3 / Emu3.5, BAGEL, OmniGen2, Lumina-mGPT, VILA-U, Nexus-Gen, UniToken are image+text only. See showlab/Awesome-Unified-Multimodal-Models.
  • 4M / 4M-21 is visual-derivatives only (no audio, no real video).
  • M3GPT, Point-Bind / Point-LLM target niche modality pairs, not the canonical quad.

Quality over quantity: the value of this list is its strictness.

Contents

Modality matrix

Legend: Y = supported, - = not supported, ◐ = claimed but partial / weakly evidenced / unreleased head. Text in/out is assumed present for every entry. Sorted within each tier by recency.

Tier 1 — Full quad, released

Model (year) Image In Image Out Audio In Audio Out Video In Video Out
NExT-OMNI (ICLR 2026) Y Y Y Y Y Y ◐
MIO (EMNLP 2025) Y Y Y Y Y Y ◐
Spider (2024) Y Y Y Y Y Y
X-VILA (2024) Y Y Y Y Y Y
NExT-GPT (ICML 2024) Y Y Y Y Y Y
CoDi-2 (CVPR 2024) Y Y Y Y Y Y
CoDi (NeurIPS 2023) Y Y Y Y Y Y

Notes on ◐ markers above:

  • NExT-OMNI explicitly notes "preliminary" short-video generation; long-form / high-resolution video out is acknowledged as future work.
  • MIO generates video as keyframe sequences only (no audio track, not full-frame-rate).

Tier 2 — Partial / claimed (one modality missing or weak)

Model (year) Image In Image Out Audio In Audio Out Video In Video Out Gap
Omni-Diffusion (2026) Y Y Y Y ◐ - no video out
AR-Omni (2026) Y Y Y Y - - no video in/out
Dynin-Omni (2026) Y Y Y Y Y - no video out
FlowBind (NeurIPS 2025) Y Y Y Y - - text/image/audio only
Ming-Flash-Omni (2025) Y Y Y Y Y - no video out
Ming-Omni (2025) Y Y Y Y Y - no video out
M2-omni (2025) Y Y Y Y Y - no video out
OmniFlow (CVPR 2025) Y Y Y Y - - no video
Unified-IO 2 (CVPR 2024) Y Y Y Y Y ◐ video out unreliable
AnyGPT (ACL 2024) Y Y Y Y - - no video
LWM (2024) Y Y - - Y Y no audio

Tier 3 — Closed / proprietary

Model (year) Image In Image Out Audio In Audio Out Video In Video Out
Gemini Omni Flash (2026) Y ◐ Y ◐ Y Y
Gemini 2.5 (2025) Y Y Y Y Y ◐
GPT-4o (2024/25) Y Y Y Y ◐ -

Gemini Omni Flash launched May 2026 with full quad input and video output; image/audio output are on the announced roadmap but not yet shipped. Gemini 2.5 video out is via the Veo-coupled API. GPT-4o video understanding is frame-sampling, not native, and there is no video out.

Tier 1: Gold Standard Any-to-Any

Full bidirectional coverage of text, image, audio, and video, with released code and weights. Sorted newest first.

  • NExT-OMNI — 7B discrete-flow-matching omnimodal foundation model on Qwen2.5; unified any-to-any across text, image, audio, video via metric-induced probability paths; short-video generation marked preliminary. [paper](https://arxiv.org/abs/2510.13721) [openreview](https://openreview.net/forum?id=odatOcBi61) (ICLR 2026)
  • MIO — First open-source any-to-any foundation model on discrete multimodal tokens across text, image, speech-with-voice, and video; four-stage training; video out limited to keyframe sequences. [paper](https://arxiv.org/abs/2409.17692) [code](https://github.com/MIO-Team/MIO) (EMNLP 2025)
  • Spider — Any-to-Many MLLM that emits multiple non-text modalities in a single response (Text + {Image, Audio, Video}); decoders are SD 1.5 / AudioLDM / Zeroscope-v2. [paper](https://arxiv.org/abs/2411.09439) [code](https://github.com/Layjins/Spider) (2024)
  • X-VILA — Cross-modal alignment LLM with VideoCrafter2 + SD 1.5 + AudioLDM decoders; introduces visual-embedding highway to reduce info loss. [paper](https://arxiv.org/abs/2405.19335) (2024)
  • NExT-GPT — The canonical end-to-end any-to-any MM-LLM; image/video/audio/text in and out via diffusion decoders and modality-switching instruction tuning. [paper](https://arxiv.org/abs/2309.05519) [code](https://github.com/NExT-GPT/NExT-GPT) (ICML 2024, Oral)
  • CoDi-2 — In-context, interleaved, interactive any-to-any generation; zero-shot reasoning and compositionality across text/vision/audio/video. [paper](https://arxiv.org/abs/2311.18775) [code](https://github.com/microsoft/i-Code/tree/main/CoDi-2) (CVPR 2024)
  • CoDi — Original composable diffusion; arbitrary input combinations to arbitrary output combinations across image/video/audio/text via aligned diffusion latents. [paper](https://arxiv.org/abs/2305.11846) [code](https://github.com/microsoft/i-Code/tree/main/i-Code-V3) (NeurIPS 2023)

Tier 2: Partial / Claimed Any-to-Any

Almost-there models with an explicit gap in the strict quad. Listed for completeness; gap is in the matrix above and the one-liner below.

  • Omni-Diffusion — First mask-based discrete-diffusion any-to-any across text, speech, image; no video output. [paper](https://arxiv.org/abs/2603.06577) [code](https://github.com/VITA-MLLM/Omni-Diffusion) (2026)
  • AR-Omni — Unified AR model for text + image + streaming speech under one decoder; no external diffusion, but no video in or out. [paper](https://arxiv.org/abs/2601.17761) [project](https://modalitydance.github.io/AR-Omni/) (2026)
  • Dynin-Omni — 8B masked-diffusion unifying text, image, video understanding and text+image+speech generation; explicitly does not emit video. [paper](https://arxiv.org/abs/2604.00007) [project](https://dynin.ai/omni/) (2026)
  • FlowBind — Bidirectional-flow any-to-any across text, image, audio with shared latent and invertible per-modality flows; no video. [paper](https://arxiv.org/abs/2512.15420) [code](https://github.com/yeonwoo378/flowbind) (NeurIPS 2025)
  • Ming-Flash-Omni — Sparse MoE (100B / 6.1B active); image gen, speech gen, generative segmentation; video understanding but no video generation. [paper](https://arxiv.org/abs/2510.24821) [blog](https://www.inclusion-ai.org/blog/ming-flash-omni-preview/) (2025)
  • Ming-Omni — MoE with modality-specific routers; image+speech generation from any-modality input; no video generation. [paper](https://arxiv.org/abs/2506.09344) (2025)
  • M2-omni — Omni-MLLM with step-balance pretraining; outputs interleaved audio, image, text — video input only. [paper](https://arxiv.org/abs/2502.18778) (2025)
  • OmniFlow — Modular multi-modal rectified flows; text/image/audio joint distribution; no video. [paper](https://arxiv.org/abs/2412.01169) [code](https://github.com/jacklishufan/OmniFlows) (CVPR 2025)
  • Unified-IO 2 — Single AR model for vision+language+audio+action; image/audio generation strong, video generation acknowledged as unreliable. [paper](https://arxiv.org/abs/2312.17172) [code](https://github.com/allenai/unified-io-2) (CVPR 2024)
  • AnyGPT — Discrete tokens across speech, text, image, and music with an unchanged LLM backbone; no video. [paper](https://arxiv.org/abs/2402.12226) [code](https://github.com/OpenMOSS/AnyGPT) (ACL 2024)
  • LWM — Million-context AR model; image+video in and out via RingAttention; no audio. [paper](https://arxiv.org/abs/2402.08268) [code](https://github.com/LargeWorldModel/LWM) (2024)

Tier 3: Closed / Proprietary

API-only or partial-release systems whose vendor docs claim any-to-any but cannot be independently audited. Capability claims match vendor documentation as of 2026-05.

  • Gemini Omni / Gemini Omni Flash — Announced at I/O 2026 as Google's first "any-to-any" model family; full quad input with video output today, image and audio output positioned as forthcoming. (Google DeepMind, May 2026)
  • Gemini 2.5 — Native multi-speaker audio in/out, native image generation (Nano Banana), video in. Video out only via Veo-coupled API. (Google DeepMind, 2025)
  • GPT-4o — Voice-to-voice and native autoregressive image generation (rolled out March 2025); video is frame-sampling on input only; no video out. (OpenAI, 2024–25)

Surveys and position papers

Datasets for any-to-any training

  • AnyInstruct — Multimodal instruction dataset (text/image/speech/music interleaved) released with AnyGPT.
  • MIO training mix — Four-stage interleaved corpus across text, image, video, and speech.
  • Unified-IO 2 dataset — Large multi-task multi-modal mixture covering 35+ benchmarks.
  • TMM (Text-formatted Many-Modal) — First X-to-Xs dataset released with Spider for many-modal-output training.

Related lists

Strict any-to-any is narrow by design. For the adjacent ecosystems:

Contributing

See CONTRIBUTING.md. PRs must demonstrate that the proposed model truly accepts all of {text, image, audio, video} as input and emits all of {text, image, audio, video} as output (or transparently mark the gap and the appropriate tier). A working video-generation demo or paper figure is required for Tier 1.

License

CC0

To the extent possible under law, the maintainers have waived all copyright and related or neighboring rights to this work.

About

A curated list of true any-to-any multimodal models: all-in, all-out.

Resources

Contributing

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors