MiniMax Music 3: ComfyUI Setup and Model Guide

ComfyUI Wiki

Run MiniMax Music 3 in ComfyUI: official Comfy-Org checkpoints, FP16, FP32 and INT8 DiT variants, pruned text encoders, the Flow-VAE decoder, and text-to-music workflows.

M

MiniMax Music 3

Music GenerationText to MusicFlow MatchingAudio

MiniMax Music 3 is MiniMax's open music generation model for complete songs up to five minutes long. An 8B Global LLM and a 0.6B Local LLM drive a Flow-Matching synthesizer and a Flow-VAE decoder to produce 32 kHz stereo WAV audio from structured captions and lyrics. It is officially supported in ComfyUI.

DeveloperMiniMax
Release Date2026-09
ArchitectureHybrid-LM (8B Global + 0.6B Local LLM) with Flow Matching synthesis (2.4B)
Audio Output32 kHz, 16-bit stereo WAV
Max LengthUp to 5 minutes per song
LicenseMiniMax Music 3 Community License

MiniMax Music 3 generates complete songs from two complementary inputs: lyrics (with optional section tags such as [Intro], [Verse], [Chorus], [Bridge] and [Outro]) and a music description covering genre, mood, instrumentation, vocal performance and production. It maintains musical themes, rhythm, vocal identity and arrangement across long sequences, so structures like intro, verse, pre-chorus, chorus, bridge and outro stay coherent.

How it works

MiniMax Music 3 is a hierarchical autoregressive model that separates global musical structure from local acoustic detail:

  • The 8B Global LLM, initialized from Qwen3-8B, predicts the first RVQ codebook frame by frame and models the song's long-range structure.
  • The 0.6B Local LLM predicts the remaining acoustic codebooks within each frame and restores fine-grained detail.

Instead of decoding only from discrete tokens, the synthesis module fuses the final hidden states of both LLMs and feeds them through a 2.4B Flow-Matching stage into the Flow-VAE latent, then a 123M Flow-VAE decoder adapted from MiniMax Speech. The training tokenizer uses eight RVQ layers: one semantic codebook of 16,384 entries plus seven acoustic codebooks of 1,024 entries each.

Installation

MiniMax Music 3 is natively supported in ComfyUI through the official Comfy-Org repackaged repository. Place the files in their respective folders:

FileDestination
minimax_music3_dit_fp16.safetensors / minimax_music3_dit_fp32.safetensors / minimax_music3_dit_int8_convrot.safetensorsComfyUI/models/diffusion_models/
minimax_music3_text_encoder_bf16.safetensors / minimax_music3_text_encoder_pruned_bf16.safetensors / minimax_music3_text_encoder_pruned_int8_convrot.safetensorsComfyUI/models/text_encoders/
minimax_music3_dav.safetensorsComfyUI/models/vae/

The DiT ships in FP16 (about 4.9 GB), FP32 and INT8 convrot variants; the text encoder combos an 8B Global and 0.6B Local LLM, with pruned BF16 and pruned INT8 convrot options for a smaller footprint.

Resources

Guides and workflows related to this model series.

No articles found.

Comments

Sign in with GitHub to join the discussion.

Loading comments…