Mage-Flow-4B: RL-Aligned Text-to-Image Generation
Mage-Flow-4B is Microsoft Asia's RL-aligned 4B text-to-image model with native resolution up to 2048 and 20-step inference.
Mage-Flow-4B (RL-aligned)
Text-to-Image4B ParametersRL-Aligned20 StepsThe RL-aligned text-to-image variant of the Mage-Flow family. Uses rectified flow matching in the Mage-VAE latent space with Qwen3-VL as the text encoder. Achieves **GenEval 0.90** — the highest among all open-source models — surpassing FLUX.2 (0.87), Qwen-Image (0.87), and Z-Image (0.84). On DPG-Bench, scores 86.26, competitive with much larger models.
| Developer | Microsoft Asia |
| Release Date | 2026-07-22 |
| Architecture | 4B NR-MMDiT + Mage-VAE |
| Text Encoder | Qwen3-VL (4B) |
| License | MIT |
| Output Resolution | 512×512 to 2048×2048 (any aspect ratio) |
| Inference Steps | 20 |
| VRAM | ~18-20 GB at 1024×1024 |
| GenEval | 0.90 |
ComfyUI Support
Mage-Flow-4B has native ComfyUI support via the integration PR (#15026). Download the repackaged weights from Comfy-Org/Mage-Flow on Hugging Face, including int8 ConvRot quantized variants.
See the Mage-Flow news article for more details on the full model family.
Guides and workflows related to this model series.
Comments
Sign in with GitHub to join the discussion.