Qwen-Image-Flash: NVIDIA's 4-Step DMD2 Distillation of Qwen-Image

ComfyUI Wiki

Qwen-Image-Flash is NVIDIA's four-step DMD2 distillation of Qwen-Image that keeps the base MMDiT transformer and generates images in 4 function evaluations.

Q

Qwen-Image-Flash

Text-to-ImageFew-StepDistillationOpen Source

NVIDIA's distilled Qwen-Image: the 20.43B MMDiT denoising transformer is replaced by DMD2 student weights and sampled in four function evaluations on a packaged shift-3 FlowMatch trajectory, with the Qwen2.5-VL text encoder and Qwen-Image VAE unchanged.

DeveloperNVIDIA
Release Date2026-07-23
Base ModelQwen/Qwen-Image
ArchitectureDiffusion Transformer (MMDiT), distilled
Parameters20.43B in the denoising transformer (28.85B in the full pipeline)
Steps4 function evaluations, static shift-3 FlowMatch Euler scheduler
LicenseNVIDIA Open Model License

What is Qwen-Image-Flash?

Qwen-Image-Flash is a four-step, DMD2-distilled build of Qwen/Qwen-Image from NVIDIA. Improved Distribution Matching Distillation was run through NVIDIA FastGen, NVIDIA Model Optimizer and NVIDIA AutoModel; only the transformer weights change, so the rest of the Qwen-Image pipeline (Qwen2.5-VL text encoder, Qwen tokenizer, Qwen-Image VAE, 60-layer transformer shape) is inherited unchanged.

  • 20.43B parameters in the denoising transformer, 28.85B across the full learned pipeline
  • 4 function evaluations instead of the base model's full schedule
  • Packaged FlowMatch Euler scheduler with a static shift-3 trajectory, producing sigmas [1.0, 0.9, 0.75, 0.5, 0.0]
  • Distilled with English captions; inherited Chinese-language capability is not evaluated in the model card

Running the distilled checkpoint

The official release is a Diffusers checkpoint, and the distillation internalizes the teacher's guidance: sampling needs num_inference_steps=4 together with true_cfg_scale=1.0. Dropping the weights into a Qwen-Image workflow that still applies CFG 4.0 would double the guidance.

from diffusers import QwenImagePipeline
import torch

pipe = QwenImagePipeline.from_pretrained(
    "nvidia/Qwen-Image-Flash", torch_dtype=torch.bfloat16
).to("cuda")

image = pipe(
    prompt="A red fox in a snowy pine forest at golden hour, photorealistic, sharp focus",
    width=1024, height=1024,
    num_inference_steps=4,
    true_cfg_scale=1.0,
).images[0]

The model card lists Diffusers 0.38.0 with Transformers 5.12.1, SGLang Diffusion, vLLM-Omni and TensorRT-LLM VisualGen as the supported runtimes, with H100 and B200 as the validated hardware. ComfyUI is not part of the official engine list: treat the single-file transformer weights as architecture-compatible with the Qwen-Image family rather than as an officially supported runtime.

Guides and workflows related to this model series.

No articles found.

Comments

Sign in with GitHub to join the discussion.

Loading comments…