Qwen-Image 2.1 Consistency LoRA: Edits That Stay on Frame
A community LoRA for Qwen-Image 2.1 keeps edits on the original frame: the same prompt and seed come back without drift or stray repainting, with two ComfyUI workflows behind it.
The dashed line marks where a feature sits in the original and the arrow shows how far it moved. Same prompt and seed in every column, rendered in ComfyUI on the Comfy-Org INT8 Qwen-Image 2.1 weights.
What it fixes
Every Qwen-Image 2.1 edit is a fresh generation, so the model re-composes the whole picture instead of changing only the part you asked about. On a global restyle that shows up as drift: a waistband 40 px lower, a horizon 25 px higher, a face that no longer lines up with the original. On a local edit it shows up as repainting, where hair strands, signs and texture outside the edit get redrawn along with it.
The LoRA makes the edit happen on the original's frame. There is no trigger word. Load it, write the edit instruction as usual, and the pixels outside the change stay where they were.
Measured drift
The model card reports the numbers from held-out edits that were never in the training set. Offsets are the worst corner of the picture measured against the original, and the style column compares each render with plain Qwen-Image 2.1's look, where a lower distance means a closer match:
| held-out edits, plain Qwen 2.1 vs the LoRA | no LoRA | step 1500 | step 2000 |
|---|---|---|---|
| restyles (36): worst corner off, median | 24.3 px | 1.6 px | 0.9 px |
| restyles: worst corner off, worst case | 61.1 px | 12.9 px | 3.5 px |
| restyles under 3 px off | 3 % | 75 % | 97 % |
| restyles (15): colour distance vs plain Qwen's look | 7.0 | 6.3 | 10.9 |
| relights and local edits (12): worst corner off, median | 0.7 px | 0.1 px | 0.1 px |
Two seeds of plain Qwen-Image 2.1 differ from each other by 7.0 on that colour scale, so the step 1500 restyles look as much like plain Qwen as plain Qwen looks like itself. Step 2000 lines up more tightly but starts to change the look: whiter paper and less colour on ink and gouache.
A second set measures repainting rather than drift, on 18 recolour and remove edits of held-out photos. Each render is first lined up with its original, so what is counted is real repainting, not offset:
| outside the edited object | no LoRA | step 1500 | step 2000 |
|---|---|---|---|
| pixels that changed noticeably | 12.8 % | 7.5 % | 7.5 % |
| PSNR against the original | 27.7 dB | 31.6 dB | 31.7 dB |
For reference, encoding and decoding a picture through the Qwen-Image 2.1 VAE alone changes about 2 % of the pixels at 37 dB, so that is the floor. The edit itself still happened in all 18 cases, with and without the LoRA.
A side-by-side run of the same edits with and without the LoRA, from the model card.
Step 1500 or step 2000
| file | notes |
|---|---|
qwen-image-2.1-consistency.safetensors | step 1500, the recommended starting point. Keeps Qwen's own look. |
qwen-image-2.1-consistency-2000.safetensors | step 2000: the tightest alignment, but paintings come out a little paler. |
Both are rank 32 and use ComfyUI key naming (diffusion_model.transformer_blocks.*), and all 384 tensors load onto the Comfy-Org Qwen-Image 2.1 weights.
Running it in ComfyUI
The adapter is a model-only LoRA, so no custom nodes are needed. To add it to any Qwen-Image 2.1 edit graph:
- Put the
.safetensorsfile inComfyUI/models/loras/and add a LoraLoaderModelOnly node right after the model loader at strength 1.0. Lower strengths let some of the drift back in. - Use a Text Encode Qwen Image 2.1 node with your picture as
image_1,resolution0, and the edit instruction as the prompt. - Sample on the encoder's
latentoutput with KSampler at 25 steps, CFG 1,eulerandsimple, denoise 1. A latent of any other size makes Qwen zoom by the size ratio, which no LoRA can undo. - Decode with VAE Decode and then Split Image with Alpha, since the Qwen-Image 2.1 VAE decodes RGBA.
The pack's Qwen-Image 2.1 Edit example graph is the one these numbers come from. Two ready-made workflows ship with the repo:
The second workflow turns the LoRA off and instead uses a Realign to Source node to line the finished edit back up with the original, which is the better path for edits that need to move something. Both were built with the author's ComfyUI-AusBoss nodes, but the card notes any equivalent nodes will do.
How it was trained
The dataset is built from real Qwen-Image 2.1 edits, with the drift measured and taken out of each target so that every training example sits on its source's frame:
- 950 edit pairs from 257 pictures: 110 portraits rendered with Qwen-Image 2.1 for this dataset, 133 Unsplash photos and 14 hand-picked pictures from the author's reference folders.
- Pair types: forward pairs (picture to edit), reverse pairs (edit back to picture), and local edits that keep the original's pixels everywhere outside the edited object.
- Training: ostris/ai-toolkit with
arch: qwen_image_2on the Comfy-Org INT8 ConvRot base, the reference kept at the target's size. Rank 32, alpha 32, AdamW8bit at lr 1e-4, batch 1,shifttimesteps, target resolution 1408. 3,000 steps on one H100 took about 1.9 s per step, with checkpoints every 250.
Limits
The card is direct about how narrow the release is. In a September 28 update the author writes that the dataset may have overfit the LoRA on certain kinds of edits, and that it can ignore requests that need real movement such as a new pose or a turned head. A v2 is planned once a new dataset is put together, and until then the realign workflow is the recommended path for those edits.
The other stated limits:
- Not yet tested with 4- or 8-step turbo LoRAs, at 2 MP, or at CFG above 1.
- Comic and anime restyles redraw every outline, so some shapes can still shift a few pixels on their own. The worst restyle at step 1500 was 12.9 px off at a corner.
- It keeps the picture in place. It does not make a weak edit stronger: if plain Qwen will not do the edit at all, the LoRA will not change that.
Availability
- LoRA weights: ausboss/Qwen-Image 2.1 Consistency LoRA
- Base model repack: Comfy-Org/Qwen-Image-2.1
- Author's node pack: ausboss/ComfyUI-AusBoss
Comments
Sign in with GitHub to join the discussion.