MiniMax H3: Open Omni-Modal Video Generation Model
MiniMax H3 is an open omni-modal video generation model with native 32 kHz stereo audio, 768p output, and 2K in-context regeneration. Native ComfyUI support.
MiniMax H3
Video GenerationNative AudioOmni-ModalText-to-VideoImage-to-VideoMiniMax H3 is an open general-purpose omni-modal generation model that jointly understands text, images, video, and audio, and generates up to 15 seconds of video with native 32 kHz stereo audio in a single pass. It is the latest model in MiniMax's Hailuo video line, powered by a 33.1B dense single-stream omni transformer with a Qwen3-VL-32B text encoder.
| Developer | MiniMax |
| Release Date | 2026-08 |
| Architecture | 33.1B dense single-stream omni transformer (13B adaLN branch) |
| Text Encoder | Qwen3-VL-32B (layer 50 hidden states) |
| Video VAE | Temporal causal VAE, f16t4d24 |
| Audio VAE | 32 kHz stereo to 40 Hz latent tokens |
| License | MiniMax H3 Community License |
MiniMax H3 generates up to 15 seconds of video at 24 FPS with native 32 kHz stereo audio, in any of 11 languages. Output defaults to a 768px short edge; the H3-Regenerate-2K module re-generates results at 2K resolution in-context. The model was officially open-sourced on August 3, 2026 under the MiniMax H3 Community License Agreement, with native ComfyUI support merged the same day (Comfy-Org/ComfyUI #15224).
Checkpoints
| Checkpoint | Modes |
|---|---|
| FL2VA | Text-to-video, first-frame, last-frame, first+last-frame |
| Ref2VA | Reference-based: up to 9 images, 3 videos, 3 audio clips (12 files max mixed) |
Resources
Guides and workflows related to this model series.
Comments
Sign in with GitHub to join the discussion.