MiniMax Music 3: ComfyUI Setup and Model Guide
Run MiniMax Music 3 in ComfyUI: official Comfy-Org checkpoints, FP16, FP32 and INT8 DiT variants, pruned text encoders, the Flow-VAE decoder, and text-to-music workflows.
MiniMax Music 3
Music GenerationText to MusicFlow MatchingAudioMiniMax Music 3 is MiniMax's open music generation model for complete songs up to five minutes long. An 8B Global LLM and a 0.6B Local LLM drive a Flow-Matching synthesizer and a Flow-VAE decoder to produce 32 kHz stereo WAV audio from structured captions and lyrics. It is officially supported in ComfyUI.
| Developer | MiniMax |
| Release Date | 2026-09 |
| Architecture | Hybrid-LM (8B Global + 0.6B Local LLM) with Flow Matching synthesis (2.4B) |
| Audio Output | 32 kHz, 16-bit stereo WAV |
| Max Length | Up to 5 minutes per song |
| License | MiniMax Music 3 Community License |
MiniMax Music 3 generates complete songs from two complementary inputs: lyrics (with optional section tags such as [Intro], [Verse], [Chorus], [Bridge] and [Outro]) and a music description covering genre, mood, instrumentation, vocal performance and production. It maintains musical themes, rhythm, vocal identity and arrangement across long sequences, so structures like intro, verse, pre-chorus, chorus, bridge and outro stay coherent.
How it works
MiniMax Music 3 is a hierarchical autoregressive model that separates global musical structure from local acoustic detail:
- The 8B Global LLM, initialized from Qwen3-8B, predicts the first RVQ codebook frame by frame and models the song's long-range structure.
- The 0.6B Local LLM predicts the remaining acoustic codebooks within each frame and restores fine-grained detail.
Instead of decoding only from discrete tokens, the synthesis module fuses the final hidden states of both LLMs and feeds them through a 2.4B Flow-Matching stage into the Flow-VAE latent, then a 123M Flow-VAE decoder adapted from MiniMax Speech. The training tokenizer uses eight RVQ layers: one semantic codebook of 16,384 entries plus seven acoustic codebooks of 1,024 entries each.
Installation
MiniMax Music 3 is natively supported in ComfyUI through the official Comfy-Org repackaged repository. Place the files in their respective folders:
| File | Destination |
|---|---|
minimax_music3_dit_fp16.safetensors / minimax_music3_dit_fp32.safetensors / minimax_music3_dit_int8_convrot.safetensors | ComfyUI/models/diffusion_models/ |
minimax_music3_text_encoder_bf16.safetensors / minimax_music3_text_encoder_pruned_bf16.safetensors / minimax_music3_text_encoder_pruned_int8_convrot.safetensors | ComfyUI/models/text_encoders/ |
minimax_music3_dav.safetensors | ComfyUI/models/vae/ |
The DiT ships in FP16 (about 4.9 GB), FP32 and INT8 convrot variants; the text encoder combos an 8B Global and 0.6B Local LLM, with pruned BF16 and pruned INT8 convrot options for a smaller footprint.
Resources
- MiniMax Music 3 ComfyUI tutorial (official docs)
- MiniMax Music 3 text-to-music workflow (official template)
- Music Production Toolkit news
- MiniMax Music 3 demo
Guides and workflows related to this model series.
Comments
Sign in with GitHub to join the discussion.