MiniMax H3: Open Omni-Modal Video Generation Model

ComfyUI Wiki

Run MiniMax H3 in ComfyUI: official T2V, I2V, R2V workflows, native 32 kHz stereo audio, FL2VA pruned INT8 checkpoint, Qwen3-VL-32B text encoder, GGUF and FP8 quant options.

M

MiniMax H3

Video GenerationNative AudioOmni-ModalText-to-VideoImage-to-Video

MiniMax H3 is an open general-purpose omni-modal generation model that jointly understands text, images, video, and audio, and generates up to 15 seconds of video with native 32 kHz stereo audio in a single pass. It is the latest model in MiniMax's Hailuo video line, powered by a 33.1B dense single-stream omni transformer with a Qwen3-VL-32B text encoder.

DeveloperMiniMax
Release Date2026-08
Architecture33.1B dense single-stream omni transformer (13B adaLN branch)
Text EncoderQwen3-VL-32B (layer 50 hidden states)
Video VAETemporal causal VAE, f16t4d24
Audio VAE32 kHz stereo to 40 Hz latent tokens
LicenseMiniMax H3 Community License

MiniMax H3 generates up to 15 seconds of video at 24 FPS with native 32 kHz stereo audio, in any of 11 languages. Output defaults to a 768px short edge; the H3-Regenerate-2K module re-generates results at 2K resolution in-context. The model was officially open-sourced on August 3, 2026 under the MiniMax H3 Community License Agreement, with native ComfyUI support merged the same day (Comfy-Org/ComfyUI #15224).

Checkpoints

CheckpointModes
FL2VAText-to-video, first-frame, last-frame, first+last-frame
Ref2VAReference-based: up to 9 images, 3 videos, 3 audio clips (12 files max mixed)

Resources

Guides and workflows related to this model series.

MiniMax H3 in ComfyUI: Complete Video Generation Guide
MiniMax H3 in ComfyUI: Complete Video Generation Guide

Download and run MiniMax H3 in ComfyUI: T2V, I2V and R2V workflows, official model and VAE files, GGUF/INT8/NVFP4 quant options, performance tips, and troubleshooting.

Comments

Sign in with GitHub to join the discussion.

Loading comments…