A curated list of true any-to-any multimodal models: all-in, all-out.
Most "unified multimodal" and "omni" lists conflate three very different kinds of model: vision-language models that only output text, text-to-X generators that only emit a single media modality, and the rare model that is genuinely bidirectional across every canonical modality. This list keeps only the last kind.
A model qualifies for the main list only if it satisfies the full any-to-any definition pioneered by NExT-GPT:
"A general-purpose multimodal LLM that accepts inputs and generates outputs in arbitrary combinations of text, image, video, and audio." — NExT-GPT (Wu et al., ICML 2024)
Concretely: the model must accept all of {text, image, audio, video} as input and generate all of {text, image, audio, video} as output. No "≥2 in / ≥2 out" loophole.
- Tier 1 — Gold Standard Any-to-Any. Text ⇄ image ⇄ audio ⇄ video, all directions, code and weights released. NExT-GPT, CoDi, CoDi-2 are the canonical members.
- Tier 2 — Partial / Claimed Any-to-Any. Paper claims most of the quad but one output modality is missing, low quality, or unreleased; the gap is named in the one-liner. Marked
◐in the matrix. - Tier 3 — Closed / Proprietary. API-only / partial release. GPT-4o, Gemini 2.5, Gemini Omni.
Everything else has been removed. In particular:
- Qwen2.5-Omni, Qwen3-Omni, MiniCPM-o, VITA-1.5, Baichuan-Omni-1.5, InteractiveOmni generate text+speech only — no image or video out. They sit in the "Omni-In / Speech-Out" family, not here.
- AnyGPT, OmniFlow, Ming-Omni, Ming-Flash-Omni, M2-omni, Omni-Diffusion, Dynin-Omni, AR-Omni, FlowBind cover three of the four canonical modalities but miss video out. They live in Tier 2.
- LWM has video but no audio. Tier 2.
- Chameleon, Janus / Janus-Pro / Janus-Flow, Show-o, Show-o2, Transfusion, Emu3 / Emu3.5, BAGEL, OmniGen2, Lumina-mGPT, VILA-U, Nexus-Gen, UniToken are image+text only. See showlab/Awesome-Unified-Multimodal-Models.
- 4M / 4M-21 is visual-derivatives only (no audio, no real video).
- M3GPT, Point-Bind / Point-LLM target niche modality pairs, not the canonical quad.
Quality over quantity: the value of this list is its strictness.
- Modality matrix
- Tier 1: Gold Standard Any-to-Any
- Tier 2: Partial / Claimed Any-to-Any
- Tier 3: Closed / Proprietary
- Surveys and position papers
- Datasets for any-to-any training
- Related lists
- Contributing
- License
Legend: Y = supported, - = not supported, ◐ = claimed but partial / weakly evidenced / unreleased head.
Text in/out is assumed present for every entry. Sorted within each tier by recency.
| Model (year) | Image In | Image Out | Audio In | Audio Out | Video In | Video Out |
|---|---|---|---|---|---|---|
| NExT-OMNI (ICLR 2026) | Y | Y | Y | Y | Y | Y ◐ |
| MIO (EMNLP 2025) | Y | Y | Y | Y | Y | Y ◐ |
| Spider (2024) | Y | Y | Y | Y | Y | Y |
| X-VILA (2024) | Y | Y | Y | Y | Y | Y |
| NExT-GPT (ICML 2024) | Y | Y | Y | Y | Y | Y |
| CoDi-2 (CVPR 2024) | Y | Y | Y | Y | Y | Y |
| CoDi (NeurIPS 2023) | Y | Y | Y | Y | Y | Y |
Notes on ◐ markers above:
- NExT-OMNI explicitly notes "preliminary" short-video generation; long-form / high-resolution video out is acknowledged as future work.
- MIO generates video as keyframe sequences only (no audio track, not full-frame-rate).
| Model (year) | Image In | Image Out | Audio In | Audio Out | Video In | Video Out | Gap |
|---|---|---|---|---|---|---|---|
| Omni-Diffusion (2026) | Y | Y | Y | Y | ◐ | - | no video out |
| AR-Omni (2026) | Y | Y | Y | Y | - | - | no video in/out |
| Dynin-Omni (2026) | Y | Y | Y | Y | Y | - | no video out |
| FlowBind (NeurIPS 2025) | Y | Y | Y | Y | - | - | text/image/audio only |
| Ming-Flash-Omni (2025) | Y | Y | Y | Y | Y | - | no video out |
| Ming-Omni (2025) | Y | Y | Y | Y | Y | - | no video out |
| M2-omni (2025) | Y | Y | Y | Y | Y | - | no video out |
| OmniFlow (CVPR 2025) | Y | Y | Y | Y | - | - | no video |
| Unified-IO 2 (CVPR 2024) | Y | Y | Y | Y | Y | ◐ | video out unreliable |
| AnyGPT (ACL 2024) | Y | Y | Y | Y | - | - | no video |
| LWM (2024) | Y | Y | - | - | Y | Y | no audio |
| Model (year) | Image In | Image Out | Audio In | Audio Out | Video In | Video Out |
|---|---|---|---|---|---|---|
| Gemini Omni Flash (2026) | Y | ◐ | Y | ◐ | Y | Y |
| Gemini 2.5 (2025) | Y | Y | Y | Y | Y | ◐ |
| GPT-4o (2024/25) | Y | Y | Y | Y | ◐ | - |
Gemini Omni Flash launched May 2026 with full quad input and video output; image/audio output are on the announced roadmap but not yet shipped. Gemini 2.5 video out is via the Veo-coupled API. GPT-4o video understanding is frame-sampling, not native, and there is no video out.
Full bidirectional coverage of text, image, audio, and video, with released code and weights. Sorted newest first.
- NExT-OMNI — 7B discrete-flow-matching omnimodal foundation model on Qwen2.5; unified any-to-any across text, image, audio, video via metric-induced probability paths; short-video generation marked preliminary.
[paper](https://arxiv.org/abs/2510.13721)[openreview](https://openreview.net/forum?id=odatOcBi61)(ICLR 2026) - MIO — First open-source any-to-any foundation model on discrete multimodal tokens across text, image, speech-with-voice, and video; four-stage training; video out limited to keyframe sequences.
[paper](https://arxiv.org/abs/2409.17692)[code](https://github.com/MIO-Team/MIO)(EMNLP 2025) - Spider — Any-to-Many MLLM that emits multiple non-text modalities in a single response (Text + {Image, Audio, Video}); decoders are SD 1.5 / AudioLDM / Zeroscope-v2.
[paper](https://arxiv.org/abs/2411.09439)[code](https://github.com/Layjins/Spider)(2024) - X-VILA — Cross-modal alignment LLM with VideoCrafter2 + SD 1.5 + AudioLDM decoders; introduces visual-embedding highway to reduce info loss.
[paper](https://arxiv.org/abs/2405.19335)(2024) - NExT-GPT — The canonical end-to-end any-to-any MM-LLM; image/video/audio/text in and out via diffusion decoders and modality-switching instruction tuning.
[paper](https://arxiv.org/abs/2309.05519)[code](https://github.com/NExT-GPT/NExT-GPT)(ICML 2024, Oral) - CoDi-2 — In-context, interleaved, interactive any-to-any generation; zero-shot reasoning and compositionality across text/vision/audio/video.
[paper](https://arxiv.org/abs/2311.18775)[code](https://github.com/microsoft/i-Code/tree/main/CoDi-2)(CVPR 2024) - CoDi — Original composable diffusion; arbitrary input combinations to arbitrary output combinations across image/video/audio/text via aligned diffusion latents.
[paper](https://arxiv.org/abs/2305.11846)[code](https://github.com/microsoft/i-Code/tree/main/i-Code-V3)(NeurIPS 2023)
Almost-there models with an explicit gap in the strict quad. Listed for completeness; gap is in the matrix above and the one-liner below.
- Omni-Diffusion — First mask-based discrete-diffusion any-to-any across text, speech, image; no video output.
[paper](https://arxiv.org/abs/2603.06577)[code](https://github.com/VITA-MLLM/Omni-Diffusion)(2026) - AR-Omni — Unified AR model for text + image + streaming speech under one decoder; no external diffusion, but no video in or out.
[paper](https://arxiv.org/abs/2601.17761)[project](https://modalitydance.github.io/AR-Omni/)(2026) - Dynin-Omni — 8B masked-diffusion unifying text, image, video understanding and text+image+speech generation; explicitly does not emit video.
[paper](https://arxiv.org/abs/2604.00007)[project](https://dynin.ai/omni/)(2026) - FlowBind — Bidirectional-flow any-to-any across text, image, audio with shared latent and invertible per-modality flows; no video.
[paper](https://arxiv.org/abs/2512.15420)[code](https://github.com/yeonwoo378/flowbind)(NeurIPS 2025) - Ming-Flash-Omni — Sparse MoE (100B / 6.1B active); image gen, speech gen, generative segmentation; video understanding but no video generation.
[paper](https://arxiv.org/abs/2510.24821)[blog](https://www.inclusion-ai.org/blog/ming-flash-omni-preview/)(2025) - Ming-Omni — MoE with modality-specific routers; image+speech generation from any-modality input; no video generation.
[paper](https://arxiv.org/abs/2506.09344)(2025) - M2-omni — Omni-MLLM with step-balance pretraining; outputs interleaved audio, image, text — video input only.
[paper](https://arxiv.org/abs/2502.18778)(2025) - OmniFlow — Modular multi-modal rectified flows; text/image/audio joint distribution; no video.
[paper](https://arxiv.org/abs/2412.01169)[code](https://github.com/jacklishufan/OmniFlows)(CVPR 2025) - Unified-IO 2 — Single AR model for vision+language+audio+action; image/audio generation strong, video generation acknowledged as unreliable.
[paper](https://arxiv.org/abs/2312.17172)[code](https://github.com/allenai/unified-io-2)(CVPR 2024) - AnyGPT — Discrete tokens across speech, text, image, and music with an unchanged LLM backbone; no video.
[paper](https://arxiv.org/abs/2402.12226)[code](https://github.com/OpenMOSS/AnyGPT)(ACL 2024) - LWM — Million-context AR model; image+video in and out via RingAttention; no audio.
[paper](https://arxiv.org/abs/2402.08268)[code](https://github.com/LargeWorldModel/LWM)(2024)
API-only or partial-release systems whose vendor docs claim any-to-any but cannot be independently audited. Capability claims match vendor documentation as of 2026-05.
- Gemini Omni / Gemini Omni Flash — Announced at I/O 2026 as Google's first "any-to-any" model family; full quad input with video output today, image and audio output positioned as forthcoming. (Google DeepMind, May 2026)
- Gemini 2.5 — Native multi-speaker audio in/out, native image generation (Nano Banana), video in. Video out only via Veo-coupled API. (Google DeepMind, 2025)
- GPT-4o — Voice-to-voice and native autoregressive image generation (rolled out March 2025); video is frame-sampling on input only; no video out. (OpenAI, 2024–25)
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities — Comprehensive 2025 survey of unified-MLLM space; covers fusion, tokenization, training across diffusion / MLLM-AR / MLLM-AR+Diffusion families. (2025, updated 2026)
- From Specific-MLLM to Omni-MLLM: A Survey — Maps the transition from single-modality MLLMs to omni-MLLMs. (2024)
- Toward Native Multimodal Modeling: A Roadmap — Position paper on native multimodal training versus encoder-LLM-decoder stitching. (2026)
- MAGUS: Multi-Agent Framework for Universal Multimodal Understanding and Generation — Argues for modular multi-agent decomposition as an alternative to monolithic any-to-any. (2025)
- AnyInstruct — Multimodal instruction dataset (text/image/speech/music interleaved) released with AnyGPT.
- MIO training mix — Four-stage interleaved corpus across text, image, video, and speech.
- Unified-IO 2 dataset — Large multi-task multi-modal mixture covering 35+ benchmarks.
- TMM (Text-formatted Many-Modal) — First X-to-Xs dataset released with Spider for many-modal-output training.
Strict any-to-any is narrow by design. For the adjacent ecosystems:
- Image + Text unified models (Chameleon, Janus, Show-o, Transfusion, Emu3, BAGEL, OmniGen2, Lumina-mGPT, VILA-U, etc.): see showlab/Awesome-Unified-Multimodal-Models and AIDC-AI/Awesome-Unified-Multimodal-Models.
- Omni-In / VLMs (perception, text-only output): Qwen2.5-Omni, Qwen3-Omni, MiniCPM-o, VITA-1.5, Baichuan-Omni-1.5, InteractiveOmni, Ola, OneLLM, PandaGPT, ImageBind-LLM, Video-LLaVA, InternVL, LLaVA-OneVision. Broader resource: Awesome-MLLM. Survey: From Specific-MLLM to Omni-MLLM.
- Omni-Out / T2X (single-direction generation): Stable Diffusion 3, Flux, Sora, Veo 3.1, Suno, AudioLDM, MusicGen, Stable Audio, Mochi, LTX-2, Wan 2.5. Broader resource: Awesome-LLMs-meet-Multimodal-Generation.
- Visual derivatives any-to-any (RGB / depth / normals / semantic / etc., no audio or video): see Apple 4M and 4M-21.
See CONTRIBUTING.md. PRs must demonstrate that the proposed model truly accepts all of {text, image, audio, video} as input and emits all of {text, image, audio, video} as output (or transparently mark the gap and the appropriate tier). A working video-generation demo or paper figure is required for Tier 1.
To the extent possible under law, the maintainers have waived all copyright and related or neighboring rights to this work.
