LLaDA-Image: InclusionAI's 6B Image Generation and Editing Model
LLaDA-Image is InclusionAI's open 6B model for text-to-image generation and native instruction editing, with a 4-step Turbo variant and a community ComfyUI node pack.
LLaDA-Image
Text-to-ImageImage Editing6BText RenderingOpen 6B unified image model from InclusionAI (Ant Group). One checkpoint handles text-to-image generation and native instruction-guided editing, with bilingual Chinese and English text rendering and an open training recipe. LLaDA-Image Turbo distills the pipeline to 2 to 4 sampling steps.
LLaDA-Image is a unified diffusion model where both the backbone and the DiT are diffusion models trained in a single framework. Instead of bolting an editing add-on onto a text-to-image model, a single checkpoint covers generation, instruction-guided editing with reference preservation, and VQ-conditioned generation through the SigVQ conditioning module.
On Qwen-Image-Bench it scores 53.53 in English and 53.38 in Chinese, state-of-the-art overall results for an open 6B model at release. The full training recipe is documented in the accompanying arXiv report.
Models
| Model | Steps | Description | License | Released |
|---|---|---|---|---|
| LLaDA-Image Turbo | 2-4 | Twin-DMD distilled variant packaged for ComfyUI (BF16 / INT8) | Apache-2.0 | 2026-09-04 |
Select a model above for curated weight downloads, tutorials, and related information.
Comments
Sign in with GitHub to join the discussion.