LLaDA-Image: InclusionAI's 6B Image Generation and Editing Model

ComfyUI Wiki

LLaDA-Image is InclusionAI's open 6B model for text-to-image generation and native instruction editing, with a 4-step Turbo variant and a community ComfyUI node pack.

L

LLaDA-Image

Text-to-ImageImage Editing6BText Rendering

Open 6B unified image model from InclusionAI (Ant Group). One checkpoint handles text-to-image generation and native instruction-guided editing, with bilingual Chinese and English text rendering and an open training recipe. LLaDA-Image Turbo distills the pipeline to 2 to 4 sampling steps.

LLaDA-Image is a unified diffusion model where both the backbone and the DiT are diffusion models trained in a single framework. Instead of bolting an editing add-on onto a text-to-image model, a single checkpoint covers generation, instruction-guided editing with reference preservation, and VQ-conditioned generation through the SigVQ conditioning module.

On Qwen-Image-Bench it scores 53.53 in English and 53.38 in Chinese, state-of-the-art overall results for an open 6B model at release. The full training recipe is documented in the accompanying arXiv report.

Models

ModelStepsDescriptionLicenseReleased
LLaDA-Image Turbo2-4Twin-DMD distilled variant packaged for ComfyUI (BF16 / INT8)Apache-2.02026-09-04

Select a model above for curated weight downloads, tutorials, and related information.

Comments

Sign in with GitHub to join the discussion.

Loading comments…