Turbo8: 8-Step Distillation LoRA for Qwen-Image 2.1
Turbo8 distills Qwen-Image 2.1 to 8 steps with no CFG, covering text-to-image, editing, RGBA output and subject extraction, with two ready ComfyUI workflows.
Qwen-Image 2.1 has attracted several few-step routes since it was released on September 20, 2026: the Pruna 5-step and 8-step LoRAs, the Viggle turbo LoRA, and Alibaba PAI's official 4-step Fun Acc distillation. Turbo8 is a community entry in the same race, and its distinguishing point is breadth: instead of accelerating text-to-image alone, it is trained to hold the whole base model's behaviour, including the RGBA and multi-reference editing paths that the earlier adapters handled less well.
All samples use Turbo8 with 8 steps and CFG 1 at 1024², first seed, with no cherry-picking within a prompt. Prompts were chosen for display from the held-out evaluation set. Source: Turbo8-LoRA-Qwen-Image-2.1.
What it does
The adapter is a rank-128 LoRA on the attention, MLP, modulation and timestep-embedding layers of Qwen-Image 2.1, distilled so the base weights stay untouched. At inference it runs with 8 steps, CFG 1, sampler euler, scheduler simple, and the author notes that going above 8 steps or past CFG 1 does not improve results.
The claim is coverage rather than one task: text-to-image up to the model's native 2K, English and Chinese text rendering, instruction editing with up to 10 reference images, transparent generation, and subject extraction (passing a photo and asking for an RGBA cutout). The model card also documents the training route, which is a three-stage recipe: 3,100 fixed-noise 40-step teacher renders recorded at the student's 8 noise levels, 3,000 steps of trajectory regression, then 2,500 steps of DMD2 distribution matching with a fake-score critic and a GAN head. The critic is a second LoRA on the same frozen transformer.
Running it in ComfyUI
Turbo8 targets ComfyUI's native Qwen-Image 2.1 support directly:
- Put
turbo8_lora_step2500.safetensorsinComfyUI/models/loras/. - Load
comfyui/turbo8_t2i.json(text-to-image and RGBA) orcomfyui/turbo8_edit.json(editing and extraction). - Keep steps 8, CFG 1, sampler
euler, schedulersimple.
The T2I graph includes a ModelSamplingFlux node (max_shift 0.6935, base_shift 0.5) wired to width and height, which reproduces the resolution-dependent noise schedule the LoRA was distilled with, including at 2K. According to the model card, all 227 LoRA layers map in ComfyUI and a 1024² image takes about 4 seconds on an RTX PRO 6000.
How it compares to the teacher
The model card evaluates Turbo8 against its own teacher, the base model at 40 steps, over 192 held-out prompts: DrawBench plus extra seeds, text-rendering prompts, RGBA subjects, OmniEdit edits and extractions.
| Metric | Teacher (40 steps) | Turbo8 (8 steps) |
|---|---|---|
| PickScore, T2I | 22.15 | 21.94 |
| CLIPScore, T2I | 26.89 | 26.89 |
| PickScore, RGBA | 20.51 | 20.07 |
| Text exact match (OCR) | 95.0% | 75.0% |
| Text character accuracy | 99.6% | 95.3% |
| Transparency rate (RGBA + extraction) | 0.875 | 0.925 |
| Seed diversity (lower is more varied) | 0.705 | 0.670 |
| Edit similarity to teacher (DINOv2) | n/a | 0.967 |
| Extraction alpha IoU vs teacher | n/a | 0.85 |
Teacher at 40 steps (upper row of each pair) against Turbo8 at 8 steps (lower row), same seed. Samples cover text-to-image, text rendering, edits, RGBA and extraction.
Documented limitations
The model card lists the gaps rather than leaving them to be found:
- Dense or long text (paragraphs, small print) degrades much more than at 40 steps, while short headlines, signs and labels stay reliable.
- On some busy scenes the composition can differ from the teacher's at the same seed.
- Extraction inherits the base model's failure cases: in the evaluation, 3 of 16 extractions came out empty for both the teacher and Turbo8, and Turbo8 sometimes keeps more surrounding context in the cutout.
In the community
In Banodoco's #qwen-image channel, users tested Turbo8 against the earlier few-step routes over the following days. The common finding was that it needs reduced strength and more than 8 steps to hold up on editing: reports settled around strength 0.65 and about 16 steps before body artefacts appear, and several users ranked it ahead of the 6-step Viggle turbo for edits while still preferring Pruna's 8-step schedule for some work.
Comments
Sign in with GitHub to join the discussion.