MiniMax H3: Open Omni-Modal Video Generation Model

ComfyUI Wiki

MiniMax H3 is an open omni-modal video generation model with native 32 kHz stereo audio, 768p output, and 2K in-context regeneration. Native ComfyUI support.

M

MiniMax H3

Video GenerationNative AudioOmni-ModalText-to-VideoImage-to-Video

MiniMax H3 is an open general-purpose omni-modal generation model that jointly understands text, images, video, and audio, and generates up to 15 seconds of video with native 32 kHz stereo audio in a single pass. It is the latest model in MiniMax's Hailuo video line, powered by a 33.1B dense single-stream omni transformer with a Qwen3-VL-32B text encoder.

DeveloperMiniMax
Release Date2026-08
Architecture33.1B dense single-stream omni transformer (13B adaLN branch)
Text EncoderQwen3-VL-32B (layer 50 hidden states)
Video VAETemporal causal VAE, f16t4d24
Audio VAE32 kHz stereo to 40 Hz latent tokens
LicenseMiniMax H3 Community License

MiniMax H3 generates up to 15 seconds of video at 24 FPS with native 32 kHz stereo audio, in any of 11 languages. Output defaults to a 768px short edge; the H3-Regenerate-2K module re-generates results at 2K resolution in-context. The model was officially open-sourced on August 3, 2026 under the MiniMax H3 Community License Agreement, with native ComfyUI support merged the same day (Comfy-Org/ComfyUI #15224).

Checkpoints

CheckpointModes
FL2VAText-to-video, first-frame, last-frame, first+last-frame
Ref2VAReference-based: up to 9 images, 3 videos, 3 audio clips (12 files max mixed)

Resources

Guides and workflows related to this model series.

No articles found.

Comments

Sign in with GitHub to join the discussion.

Loading comments…