Zen Image Edit: Qwen-Image 2.1 on a 0.8B Encoder in ComfyUI

ComfyUI Wikinews

Zen Image Edit runs Qwen-Image 2.1 on a Qwen3.5-0.8B text encoder plus a 158M fusion adapter, cutting the encoder from 17.5 GB to 1.7 GB, with two ComfyUI routes.

Zen Image Edit (Hugging Face) keeps the Qwen-Image 2.1 DiT and VAE but swaps the 17.5 GB Qwen3-VL-8B text encoder for Qwen3.5-0.8B (1.7 GB) plus a 158M text-fusion adapter trained to reproduce the original encoder's conditioning. One pipeline covers text-to-image, character and scene editing, and transparent RGBA output, and the model card ships both Diffusers code and a ComfyUI route.
Text-to-image sample from Zen Image Edit at 1024 px

Text-to-image at 1024 px, 30 steps, generated by the pipeline. Source: AiArtLab/zen-image-edit.

What Zen Image Edit changes

Qwen-Image 2.1's conditioning stack is what makes it expensive to hold in memory: the base model's text encoder is Qwen3-VL-8B at roughly 17.5 GB in fp16, and it has to sit next to a 14.5 GB DiT. Zen Image Edit replaces it with a small multimodal encoder and an adapter that imitates what the large one produced, both from plain text and from text read together with reference images.

ComponentZen Image Edit
TransformerQwen-Image 2.1 DiT, 32 layers, 14.5 GB fp16, with the 158M text-fusion adapter merged in
Text encoderQwen3.5-0.8B, 1.7 GB fp16, tokenizer and processor unchanged
Conditioning fidelitycosine 0.95 on text, 0.97 on the vision positions of edit prompts, against the native Qwen3-VL-8B encoder
VAEfinetuned Qwen-Image 2.1 decoder, 16x spatial, fp32
SchedulerFlowMatchEulerDiscreteScheduler, plain static shift 5.0
Peak VRAMabout 17.5 GB resident, less with enable_model_cpu_offload()

The adapter lives inside the DiT as its text-fusion block, so the whole model is one self-contained Diffusers folder and no separate 17.5 GB encoder is needed anywhere. The sampler runs a plain static shift of 5.0 instead of the base model's dynamic shifting.

The bundled adapter is revision v12. Its attention-branch position table covers 2,304 slots and it was finetuned at the real 1024 px edit geometry, so long reference sequences keep their positions instead of falling into a zero-padded tail. That moved the vision cosine against the native encoder from 0.93 to 0.97, while text stayed at 0.95.

A finetuned VAE

The bundled decoder is not the stock Qwen-Image 2.1 one. It is a decoder-only finetune built on madebyollin/texture-fix-vae-for-qwen-image-2.1, trained to remove the 2 px lattice that nearest-exact upsampling locks into the output stride. Only the decoder was trained, over 5,300 steps, with one extra loss term: the peak-to-peak difference between the reconstruction's and the target's 2x2 sublattice means. Faithfully reproduced content contributes zero to that term, so only the added lattice is penalised. The encoder is bit-identical to the original, which leaves the latent space unchanged.

DecoderPSNRLPIPSLattice (/255)
original Qwen-Image 2.133.0730.053162.631
texture-fix32.8400.053880.125
this VAE33.3970.052910.000

Reconstruction on 32 held-out images at 512 px. The model card notes the lattice is not visible at 100% zoom, which is why it is measured rather than judged by eye.

Examples

Single-image editing: background change, subject kept

Editing with one condition image: the background changes, the subject is kept.

Character swap with two condition imagesCharacter swap with two condition images
Two condition images: <image1> keeps its pose, clothing and sceneThe identity is copied from <image2>

Editing follows the base model's convention: the first image is the one being edited (<image1>) and everything after it is a reference. The model card points out that feeding the reference first is the usual reason a swap "does not happen", because the model then edits the reference instead.

Editing with three condition images

Three condition images: target and composition from <image1>, the person from <image2>, colour and lighting from <image3>.

Transparent RGBA output

Transparent (RGBA) generation runs through the same pipeline.

Running it in ComfyUI

There are two independent ComfyUI routes, and neither one needs the 17.5 GB encoder.

Node pack: recoilme/zen-image-edit-comfyui. The DiT, VAE, sampler and prefix KV cache stay stock ComfyUI (a build with Qwen-Image 2.1 support is required), and the node only supplies the small encoder and the adapter:

  1. Download adapter_v12.safetensors (0.64 GB) into ComfyUI/models/.
  2. Put qwen_image_2.1_bf16.safetensors in models/diffusion_models/ and qwen_image_2.1_vae_bf16.safetensors in models/vae/, both from Comfy-Org/Qwen-Image-2.1.
  3. Fetch the text encoder: hf download Qwen/Qwen3.5-0.8B --local-dir ComfyUI/models/text_encoders/qwen3.5_0.8b.
  4. Restart ComfyUI. It needs transformers >= 5.17.

The author reports cosine 1.0000 against the Diffusers build of the same folder, and both example graphs are verified end to end.

Single-file checkpoint: T8mars/comfyui-Zen-Image-Edit-T8. Installable from ComfyUI Manager, this route packs the DiT, the VAE, the Zen encoder, the fusion adapter and the tokenizer assets into one zen_image_edit_qwen21_single.safetensors checkpoint, published as t8star/Zen-Image-Edit-Comfy. Its loader returns MODEL, CLIP and VAE, so no separate downloads are needed, and a CLIP-only loader works with the native diffusion and VAE loaders. The conversion was validated on ComfyUI 0.37.0 with an RTX 5090 Laptop at 24 GB, on both the text-to-image and the image-edit graphs. The same repository also republishes the Viggle turbo adapters as LoRAs that load in ComfyUI's built-in nodes, with workflows.

Examples generated through the ComfyUI node pack

Two character references in, three scenes out, through the ComfyUI node at 768x1280. Source: recoilme/zen-image-edit-comfyui.

Documented limitations

The model card lists the gaps rather than leaving them to be found:

  • English only. That is all the adapter was trained and tested on, and other languages drift.
  • Numerals on signage. "OPEN 24 HOURS" renders as "OPEN 26 HOURS" on every seed tried. Words are fine.
  • Non-photographic references transfer less faithfully than photographic ones, since the adapter imitates the native encoder and inherits its ceiling.
  • Batch size above 1 at 1024 px peaks near 28 GB, so one prompt per call is the safe mode.
A text rendering failure case

Words render correctly, numerals on signage do not. Source: AiArtLab/zen-image-edit.

Availability

The weights are published as AiArtLab/zen-image-edit. For 16 GB cards there is a GGUF Q5_K build at recoilme/zen-image-edit-gguf that takes the transformer from 13.55 GiB to 4.52 GiB during a run. Running the Diffusers pipeline needs a diffusers built with Qwen-Image-2.1 support and trust_remote_code=True.

Comments

Sign in with GitHub to join the discussion.

Loading comments…
Zen Image Edit: Qwen-Image 2.1 on a 0.8B Encoder in ComfyUI | ComfyUI Wiki