Z-Image: 6B S3-DiT Foundation Model by Tongyi-MAI — Apache-2.0
ComfyUI Wiki
Z-Image is Tongyi-MAI's 6B parameter single-stream DiT foundation model for text-to-image generation. Supports full CFG, fine-tuning, and 28-50 step inference.
Z
Z-Image
Text-to-Image6B ParametersS3-DiTFine-tunableOpen SourceZ-Image (造相) is the foundation model of the Z-Image family by Alibaba's Tongyi-MAI team. It features a 6B parameter Scalable Single-Stream DiT (S3-DiT) architecture — an undistilled base model that supports full Classifier-Free Guidance (CFG), fine-tuning, and negative prompting. Engineered for high-quality generation, rich aesthetics, strong diversity, and controllability.
| Developer | Tongyi-MAI (Alibaba) |
| Release Date | 2025-12 |
| Architecture | 6B Scalable Single-Stream DiT (S3-DiT) |
| License | Apache-2.0 |
| Output Resolution | 512×512 to 2048×2048 (any aspect ratio) |
| Inference Steps | 28-50 |
| Guidance Scale | 3.0-5.0 |
| Text Encoder | Qwen 3 4B |
Guides and workflows related to this model series.
No articles found.
Comments
Sign in with GitHub to join the discussion.