YuE2-3B in ComfyUI: Open Music Generation From MAP

ComfyUI Wikinews

YuE2-3B ComfyUI workflow: the MAP team's open music generation model turns lyrics into songs with editable scores, beating Suno v5 on WildSongBench. Weights now in master.

The Multimodal Art Projection (m-a-p) team has open-sourced YuE2-3B, a 3B music generation model that turns lyrics and a style prompt into complete songs with vocals and accompaniment. Unlike typical lyrics-to-song systems, YuE2 exposes an editable score layer: melody and chords are written down as symbolic plans that you can revise before rendering. On the team's WildSongBench evaluation it outscores Suno v5. ComfyUI's native YuE2 support has now merged into master (PR #16250), so a normal update is all you need.

Overview

YuE2-3B is the successor to the original YuE lyrics-to-song model. The team positions it as frontier-quality open music generation: on 192 WildSongBench prompts, YuE2 with best-of-8 sampling reaches a SongBench average of 6.9632, compared with 6.8721 for Suno v5, the highest among all evaluated open and proprietary models.

YuE2 song quality and text alignment on WildSongBench

Frontier song quality and text alignment on 192 WildSongBench prompts. YuE2 uses symbolic planning; Bo8 means best-of-8.

The headline feature is symbolic planning with an editable score. Instead of jumping straight from text to audio, YuE2 first writes a plan containing melody and chords (in ABC notation), then renders audio from it. Because the plan is explicit, you can:

  • Compose and edit: full melody plus chords, melody-only, or direct generation without a score
  • Bring your own score: feed an existing ABC score and have YuE2 sing it
  • Edit with an agent: convert musical feedback into score, style, and lyric revisions, then let the model render the next version. The team demonstrates a 14-version edit chain that moves one song from Mandarin pop to English jazz.

Architecture

YuE2 architecture

One AR–NAR Mixture-of-Transformers backbone writes the score and semantic tokens, then generates acoustic latents through flow matching. The VAE turns them into stereo audio.

YuE2 uses an AR–NAR Mixture-of-Transformers backbone. The autoregressive stage plans the score and semantic tokens; the non-autoregressive stage produces acoustic latents via flow matching, and a dedicated VAE decodes them into 48 kHz stereo audio. Both the planning and synthesis APIs are exposed separately, so developers can inspect or replace the symbolic plan.

Example output

The model card ships full-length demos covering original compositions and covers:

ExampleStyleLength
Cyber MetalEnglish cyber metal5:00
今晚不眠Mandarin funk / nu-disco3:24
PassionEnglish rock3:55
Auld Lang SyneJazz-funk cover3:10
最炫民族风Ballad cover4:45
Jingle BellsHeavy metal cover1:09

The cover examples show the model reimagining an existing song in a completely different genre, and the agentic editing demo walks through 9 planning steps and 14 rendered versions of a single song.

Running it

YuE2-3B runs locally on a 24 GB NVIDIA GPU (Linux, Python 3.10+) without quantization. The team distributes an inference wheel alongside the 7.3 GB checkpoint on Hugging Face:

python -m pip install huggingface-hub==0.36.2
hf download m-a-p/YuE2-3B yue2_infer-0.1.5-py3-none-any.whl --local-dir .
python -m pip install ./yue2_infer-0.1.5-py3-none-any.whl

The pipeline exposes Hugging Face loading, text guidance (CFG), and separate planning and synthesis APIs. Full notebooks and the cover/editing recipes are in the README.

Availability

There is no native ComfyUI support yet. The long-standing community node ComfyUI_YuE wraps the YuE series in ComfyUI; it has not been updated for YuE2 at the time of writing, so running YuE2-3B today requires the official Python inference package. A hosted demo is available on the project page.

Native ComfyUI support has merged into master

On September 11, ComfyUI core landed native YuE2 support in PR #16250, and it has since merged into master alongside follow-up fixes: longer max song duration (PR #16292, up to 900 seconds), AMD fixes plus more controls on the Generate ABC node (PR #16293), and official README acknowledgment (PR #16303). If you are on a recent master build (or the next stable release after v0.35.0), YuE2 is included and no branch checkout is needed. The integration ships three nodes: YuE2 Generate ABC for the symbolic score plan, YuE2 Generate Music for rendering, and Empty YuE2 Latent Audio. A ready-to-use all-in-one checkpoint (yue2.safetensors, 7.8 GB) is published on the Comfy-Org Hugging Face account, and a test workflow is attached to the pull request. The core two-node flow matches how the community uses it: keep the ABC text from Generate ABC to reuse the same melody across renders.

Early renders sound reasonable, and the rough edges from the first days are already smoothed over:

  • YuE2 pins exact versions of PyTorch, Transformers, and other shared packages, so installing the standalone YuE2 requirements can overwrite packages ComfyUI needs. The native integration avoids the multi-venv problem, but check your environment after setup.
  • Community node packs such as EmeraldApple-AI/ComfyUI-YuE2 and a ScryptHunter fork that isolates the YuE2 wheel also appeared within a day; they are experimental one-person projects for now.
  • YuE2's own VAE only converts audio to and from continuous acoustic latents, which means LoRA training for YuE2 is not feasible yet: the discrete semantic music tokens and a YuE2-specific training pipeline have not been released.

Comments

Sign in with GitHub to join the discussion.

Loading comments…
YuE2-3B in ComfyUI: Open Music Generation From MAP | ComfyUI Wiki