Z-Image: 6B S3-DiT Foundation Model by Tongyi-MAI — Apache-2.0

ComfyUI Wiki

Z-Image is Tongyi-MAI's 6B parameter single-stream DiT foundation model for text-to-image generation. Supports full CFG, fine-tuning, and 28-50 step inference.

Z

Z-Image

Text-to-Image6B ParametersS3-DiTFine-tunableOpen Source

Z-Image (造相) is the foundation model of the Z-Image family by Alibaba's Tongyi-MAI team. It features a 6B parameter Scalable Single-Stream DiT (S3-DiT) architecture — an undistilled base model that supports full Classifier-Free Guidance (CFG), fine-tuning, and negative prompting. Engineered for high-quality generation, rich aesthetics, strong diversity, and controllability.

DeveloperTongyi-MAI (Alibaba)
Release Date2025-12
Architecture6B Scalable Single-Stream DiT (S3-DiT)
LicenseApache-2.0
Output Resolution512×512 to 2048×2048 (any aspect ratio)
Inference Steps28-50
Guidance Scale3.0-5.0
Text EncoderQwen 3 4B

Guides and workflows related to this model series.

No articles found.

Comments

Sign in with GitHub to join the discussion.

Loading comments…