- Home
- Models
- Qwen Image
- Qwen-Image-Flash: NVIDIA's 4-Step DMD2 Distillation of Qwen-Image
Qwen-Image-Flash: NVIDIA's 4-Step DMD2 Distillation of Qwen-Image
Qwen-Image-Flash is NVIDIA's four-step DMD2 distillation of Qwen-Image that keeps the base MMDiT transformer and generates images in 4 function evaluations.
Qwen-Image-Flash
Text-to-ImageFew-StepDistillationOpen SourceNVIDIA's distilled Qwen-Image: the 20.43B MMDiT denoising transformer is replaced by DMD2 student weights and sampled in four function evaluations on a packaged shift-3 FlowMatch trajectory, with the Qwen2.5-VL text encoder and Qwen-Image VAE unchanged.
| Developer | NVIDIA |
| Release Date | 2026-07-23 |
| Base Model | Qwen/Qwen-Image |
| Architecture | Diffusion Transformer (MMDiT), distilled |
| Parameters | 20.43B in the denoising transformer (28.85B in the full pipeline) |
| Steps | 4 function evaluations, static shift-3 FlowMatch Euler scheduler |
| License | NVIDIA Open Model License |
What is Qwen-Image-Flash?
Qwen-Image-Flash is a four-step, DMD2-distilled build of Qwen/Qwen-Image from NVIDIA. Improved Distribution Matching Distillation was run through NVIDIA FastGen, NVIDIA Model Optimizer and NVIDIA AutoModel; only the transformer weights change, so the rest of the Qwen-Image pipeline (Qwen2.5-VL text encoder, Qwen tokenizer, Qwen-Image VAE, 60-layer transformer shape) is inherited unchanged.
- 20.43B parameters in the denoising transformer, 28.85B across the full learned pipeline
- 4 function evaluations instead of the base model's full schedule
- Packaged FlowMatch Euler scheduler with a static shift-3 trajectory, producing sigmas
[1.0, 0.9, 0.75, 0.5, 0.0] - Distilled with English captions; inherited Chinese-language capability is not evaluated in the model card
Running the distilled checkpoint
The official release is a Diffusers checkpoint, and the distillation internalizes the teacher's guidance: sampling needs num_inference_steps=4 together with true_cfg_scale=1.0. Dropping the weights into a Qwen-Image workflow that still applies CFG 4.0 would double the guidance.
from diffusers import QwenImagePipeline
import torch
pipe = QwenImagePipeline.from_pretrained(
"nvidia/Qwen-Image-Flash", torch_dtype=torch.bfloat16
).to("cuda")
image = pipe(
prompt="A red fox in a snowy pine forest at golden hour, photorealistic, sharp focus",
width=1024, height=1024,
num_inference_steps=4,
true_cfg_scale=1.0,
).images[0]The model card lists Diffusers 0.38.0 with Transformers 5.12.1, SGLang Diffusion, vLLM-Omni and TensorRT-LLM VisualGen as the supported runtimes, with H100 and B200 as the validated hardware. ComfyUI is not part of the official engine list: treat the single-file transformer weights as architecture-compatible with the Qwen-Image family rather than as an officially supported runtime.
Guides and workflows related to this model series.
Comments
Sign in with GitHub to join the discussion.