MiniMax H3: Open Omni-Modal Video Generation Model

ComfyUI Wiki

Run MiniMax H3 in ComfyUI: official T2V, I2V, R2V workflows, native 32 kHz stereo audio, FL2VA pruned INT8 checkpoint, Qwen3-VL-32B text encoder, GGUF and FP8 quant options.

M

MiniMax H3

Video GenerationNative AudioOmni-ModalText-to-VideoImage-to-Video

MiniMax H3 is an open general-purpose omni-modal generation model that jointly understands text, images, video, and audio, and generates up to 15 seconds of video with native 32 kHz stereo audio in a single pass. It is the latest model in MiniMax's Hailuo video line, powered by a 33.1B dense single-stream omni transformer with a Qwen3-VL-32B text encoder.

DeveloperMiniMax
Release Date2026-08
Architecture33.1B dense single-stream omni transformer (13B adaLN branch)
Text EncoderQwen3-VL-32B (layer 50 hidden states)
Video VAETemporal causal VAE, f16t4d24
Audio VAE32 kHz stereo to 40 Hz latent tokens
LicenseMiniMax H3 Community License

MiniMax H3 generates up to 15 seconds of video at 24 FPS with native 32 kHz stereo audio, in any of 11 languages. Output defaults to a 768px short edge; the H3-Regenerate-2K module re-generates results at 2K resolution in-context. The model was officially open-sourced on August 3, 2026 under the MiniMax H3 Community License Agreement, with native ComfyUI support merged the same day (Comfy-Org/ComfyUI #15224).

Checkpoints

CheckpointModes
FL2VAText-to-video, first-frame, last-frame, first+last-frame
Ref2VAReference-based: up to 9 images, 3 videos, 3 audio clips (12 files max mixed)

Resources

Guides and workflows related to this model series.

MiniMax H3 in ComfyUI: Complete Video Generation Guide
MiniMax H3 in ComfyUI: Complete Video Generation Guide

Download MiniMax H3 for ComfyUI: minimax_h3_fl2va_pruned_int8_convrot.safetensors (19.5 GB), qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors (14.6 GB), T2V/I2V/R2V workflows.

Comments

Sign in with GitHub to join the discussion.

Loading comments…