MiniMax H3: Open Omni-Modal Video Model With Native Audio
MiniMax launches H3, an omni-modal video model generating 2K 15-second clips with native stereo audio, multimodal context understanding, and instruction-based editing.
MiniMax has officially launched MiniMax H3, a general-purpose omni-modal generation model that jointly understands text, images, video, and audio. H3 generates video with native stereo audio at up to 2K resolution and 15 seconds in length, and the company plans to release the model weights in the coming days.
H3 succeeds MiniMax's Hailuo 01/02 video model line. It is a video-generation model, not one of the M-series language models MiniMax open-sourced earlier this year.
Multimodal context understanding
Real creative work means blending information across modalities. H3 accepts any combination of text, images, video, and audio as input, and lets you describe the relationship between them in plain language. MiniMax's official demo prompts a shot as: "Reference the Hitchcock camera movement from Video 1, have the character in Image 2 sing, with the vocals matching Audio 3." The model handles the full-modality understanding itself.
An example image input used in MiniMax's official H3 multimodal context demo
H3-generated video combining a reference clip, an image, and an audio track in a single prompt (source: MiniMax blog)
Key capabilities
- Native 2K video - up to 15 seconds per clip at 2K resolution, achieved through in-context regeneration instead of a separate super-resolution module.
- Native stereo audio - dialogue, sound effects, and music are generated in the same pass as the picture, with no separation between voice, SFX, and music domains.
- Omni-reference generation - up to 9 reference images, 3 reference videos, and 3 reference audio clips can steer a single generation.
- Instruction-based editing - edit existing images, videos, or audio by describing the change in natural language.
- V2V motion transfer - carry the motion of a source video into a new scene.
- Strong text and brand rendering - built for advertising, branding, e-commerce, product design, and UI/UX use cases.
Native stereo audio demo
Native stereo sound generation demo (source: MiniMax blog)
Pricing
MiniMax positions H3's price-performance as industry-leading: at 2K, the per-second price is less than a third of mainstream models, and at 768p it is less than half the price of rivals' 720p tiers. Pay-as-you-go pricing starts at $0.13 per second for 2K output.
Availability
MiniMax H3 is available now through MiniMax's platform API and third-party providers such as fal.ai. MiniMax says it plans to open up the model weights in the coming days, subject to applicable laws and regulations; as of this writing, no Hugging Face repository has been published yet, and a full technical report is expected soon.
In ComfyUI, H3 is available through the official MiniMax partner node (API-based), merged into ComfyUI on July 31 (Comfy-Org/ComfyUI #15167). Native local support is expected once the open weights ship.
Comments
Sign in with GitHub to join the discussion.