- Home
- Models
- Qwen Image
- Qwen-Image 2.1: 7B Unified Generation and Editing with RGBA Output
Qwen-Image 2.1: 7B Unified Generation and Editing with RGBA Output
Qwen-Image 2.1 is Alibaba's 7B single-stream DiT for unified text-to-image generation and editing, writing real RGBA alpha from up to 10 reference images.
Qwen-Image 2.1
Text-to-ImageImage EditingTransparencyOpen SourceAlibaba's open-weight successor to Qwen-Image: a 7B single-stream diffusion transformer (32 layers) that generates and edits with the same weights, writes real RGBA alpha channels, accepts up to 10 reference images, and outputs natively from 2048x2048 up to 2752x1536.
| Developer | Alibaba Cloud (Qwen Team) |
| Release Date | 2026-09 |
| Architecture | Single-stream DiT (7B visual generation component, 32 layers) |
| Text Encoder | Qwen3-VL-8B (Comfy-Org repackaging) |
| License | Qwen Research License |
| Generation Modes | Text-to-image, instruction editing, transparent (RGBA) output, subject extraction |
| Native Resolutions | 2048x2048 (1:1) up to 2752x1536 (16:9), with matching portrait ratios |
What is Qwen-Image 2.1?
Qwen-Image 2.1 is the second generation of Alibaba's open-weight Qwen-Image line, and it replaces the 20B MMDiT backbone of the original release with a 7B visual generation component built from 32 single-stream DiT layers. The model card describes four changes:
- Compact and efficient. Mixed-granularity attention plus prefix KV cache reuse keeps image quality high at a much lower compute cost than the 20B generation.
- Native transparency with unified creation and editing. One model writes regular or transparent RGBA images from text, edits transparent layers, and extracts a subject out of a photograph.
- Versatile editing. Up to 10 reference images per request, local edits specified with circles, painted annotations or a separate mask, and identity preservation for people and products.
- Realistic textures and refined aesthetics. Improved typography, portrait lighting and fine detail.
Native output sizes start at 2048x2048 for 1:1 and go to 2400x1792 (4:3), 2528x1696 (3:2) and 2752x1536 (16:9), with the matching portrait ratios.
Transparent images with a real alpha channel
The headline change is transparency that survives into the file. Instead of generating on a flat backdrop and cutting the subject out afterwards with a matting model, Qwen-Image 2.1 writes the alpha channel itself, so stickers, product renders and UI assets come out ready to composite. The model card recommends an explicit prompt format to trigger it:
This is an RGBA image with transparency. <your subject description>.
The image has alpha channel and the background is transparent.ComfyUI support
Qwen-Image 2.1 landed in ComfyUI on day one with the integration PR (Comfy-Org/ComfyUI #16400), and the official templates require ComfyUI 0.37.0 or newer. Repackaged single-file weights live in Comfy-Org/Qwen-Image-2.1, including BF16 and INT8 convrot transformers plus the Qwen3-VL-8B text encoder and optional prompt enhancer.
Guides and workflows related to this model series.
Comments
Sign in with GitHub to join the discussion.