Viggle Turbo v0.3: Qwen-Image 2.1 at 6 and 9 Steps in ComfyUI

ComfyUI Wikinews

Viggle's v0.3 distillation of Qwen-Image 2.1 adds a 9-step mode and ships official ComfyUI nodes, workflows and merged single-file weights for 6-step generation and editing.

Viggle Turbo v0.3 (Hugging Face) is the next cut of Viggle's few-step distillation of Qwen-Image 2.1: 6 steps with no classifier-free guidance instead of 40, now with a second 9-step mode that hands the last two steps back to the base model. The bigger change for ComfyUI users is packaging: v0.3 ships its own custom nodes, text-to-image and edit workflows, and merged single-file transformers in int8, fp8 and GGUF, instead of leaving the port to the community.
A 360 degree equirectangular panorama generated from a single photo with Viggle Turbo v0.3

A 360 degree equirectangular panorama built from one perspective photo, generated by v0.3 in 9 steps at 2176x1088. Source: Viggle's demo Space comparison tab.

What v0.3 changes

The distillation itself is unchanged: a rank-256 LoRA on the base Qwen-Image 2.1 transformer, sampled on a fixed sigma schedule with true_cfg_scale=1.0 and no negative prompt. v0.3 moves the balance rather than the method.

  • 6 steps, retuned. Against v0.2.1 the results are less grainy and cleaner on flat surfaces, with fine texture a little softer. Viggle is explicit that this is not a strict upgrade: if you preferred the crisper v0.2.1 look, that adapter stays in the repository.
  • A new 9-step mode. Seven turbo steps, then the LoRA is switched off and the unmodified base model finishes the last two. Detail is finer and small rendered text comes out right more often. It takes roughly 1.4 to 1.5 times as long as 6 steps, which is still about 3.5 times faster than the 40-step base.
  • A stated ceiling. Viggle writes that 6 steps is close to what this student can do: every gain found since v0.2.1 traded something away, sharper coming with more grain and less grain coming with a softer look. Beyond that point quality has to be paid for with steps, which is what the 9-step mode does.

The 9-step mode

The 9-step path is not just a longer schedule. The pipeline computes the text and reference K/V once and reuses them; the 7 turbo steps fill that cache, so the first base-model step has to recompute it before the LoRA is disabled and the base takes over. In diffusers that means a step callback and a wrapper around the transformer forward, both shown in the model card:

SIGMAS_9 = [1.0, 0.9583, 0.9167, 0.875, 0.75, 0.5, 0.25, 1 / 6, 1 / 12]

The 6-step schedule keeps the low-noise nodes 0.875, 0.75, 0.5, 0.25 and adds steps at the high-noise end only: 5 steps is [1, 0.875, 0.75, 0.5, 0.25], 7 steps is [1, 0.9583, 0.9167, 0.875, 0.75, 0.5, 0.25]. Moving the low-noise nodes instead makes results softer, and passing a plain step count without the sigma list does not work.

9 steps currently runs in diffusers and the demo Space only. The ComfyUI workflows implement the 6-step schedule.

Official ComfyUI support

This is the part v0.2 did not have. The release now carries a ComfyUI port in the comfyui/ folder of the model repo, tested against ComfyUI 0.37.0, which has native Qwen-Image 2.1 support:

  • viggle_turbo.py adds two nodes. Copy it into ComfyUI/custom_nodes/ and restart.
  • Viggle Turbo Sigmas supplies the 6-step schedule with the pipeline's resolution-dependent shift, to be used with euler and BasicGuider instead of a KSampler scheduler, no CFG and no negative prompt.
  • Viggle Turbo LoRA (unmerged) applies the adapter at runtime the way diffusers does. The stock LoRA loaders merge it into the weights, which Viggle measures as dropping about 30 percent of this adapter's update in bf16 and adding noise on int8. The unmerged node costs 10 to 25 percent more time per step.
  • Workflows for text-to-image and editing, plus variants for the merged single-file weights and for GGUF.

With the int8 defaults, the r128 LoRA and the prompt enhancer on, the workflows peak at about 26 GB of VRAM at 1248x832. The model card also notes that the ComfyUI port was mostly written with an AI coding assistant, so rough edges are expected and fixes are welcome.

Merged single-file transformers

If you would rather not load a LoRA at all, v0.3 also ships the base transformer with the rank-256 adapter already merged in fp32 and quantized once. The files go into models/diffusion_models/; GGUF versions need ComfyUI-GGUF, and every path still needs Viggle Turbo Sigmas from the custom node. Only the 6-step mode is available this way.

FormatSizeComfyUI loaderLPIPS v0.3LPIPS v0.2.1
GGUF Q8_07.7 GBUnet Loader (GGUF)0.0510.054
int8, Comfy-Org convrot recipe7.3 GBLoad Diffusion Model0.0570.060
GGUF Q6_K6.0 GBUnet Loader (GGUF)0.0550.067
fp8 e4m3fn, weight-only7.3 GBLoad Diffusion Model0.0680.070
GGUF Q5_K_M5.1 GBUnet Loader (GGUF)0.0760.083
GGUF Q4_K_M4.3 GBUnet Loader (GGUF)0.1000.118
int8 plus the r128 LoRA (the LoRA workflows)8.0 GBLoad Diffusion Model0.0410.044

LPIPS is Viggle's mean VGG distance against diffusers running the rank-256 LoRA, over 96 held-out requests with the same prompt, inputs, seed and noise; lower is closer. For scale, ComfyUI and diffusers differ by about 0.03 to 0.04 with no LoRA loaded.

Merged weights are close to, but not identical to, the LoRA path. On roughly 8 of those 96 requests the merged int8 or Q8_0 build lands on a different composition or outfit, against 3 to 5 for the LoRA workflow. Q4_K_M drifts visibly more, at 23 to 29 of 96, and Viggle only recommends it when nothing larger fits in memory.

40-step base against 6 and 9 steps

The examples below come from the Comparison tab of Viggle's own demo Space. Every column uses the same prompt, inputs, seed and noise.

40-step base modelViggle Turbo v0.3 at 6 stepsViggle Turbo v0.3 at 9 steps
Base model, 40 stepsv0.3, 6 stepsv0.3, 9 steps

A character reference carried into a mangrove boardwalk scene at 1344x1760. The 6-step result is the cleanest of the two turbo columns; 9 steps brings back some of the fine texture.

The 9-step mode is aimed at the two places the 6-step student still trails the base model: dense rendered text and small detail. A vertical app screenshot with a lot of small interface text at 1536x2720 is a reasonable stress test, and the comparison tab runs one.

40-step base model on a dense app screenshotViggle Turbo v0.3 at 9 steps on the same screenshot
Base model, 40 stepsv0.3, 9 steps

The distillation also runs with up to 3 reference images in diffusers, and the six-image group portrait below exercises that path: six identities, one prompt, one bar interior.

40-step base model on a six-reference group portraitViggle Turbo v0.3 at 6 steps on the same group portrait
Base model, 40 stepsv0.3, 6 steps

Speed

Same comparisons, per-image seconds as reported in the demo Space:

CaseResolutionBase, 40 stepsv0.3, 6 stepsv0.3, 9 steps
Portrait from a reference1344x176014.0 s2.9 s4.0 s
Six-reference group portrait1248x188825.8 s5.9 s8.8 s
Dense app screenshot1536x272026.9 s4.9 s6.8 s
360 degree panorama2176x108814.3 s2.9 s4.1 s

What it still does not do

The model card is unusually direct about the remaining gap. Complicated edits are still where the student trails: multi-reference composition, face swaps and identity-preserving instructions can produce duplicated or ghosted figures and identity drift. Small or long rendered text garbles more often than with the base model, and 9 steps helps without fixing it. Colours come out a few percent less saturated. 2K output, RGBA output, mask-guided edits and edits with more than 3 references are checked by eye on the comparison tab only, with no benchmark claimed.

Comments

Sign in with GitHub to join the discussion.

Loading comments…
Viggle Turbo v0.3: Qwen-Image 2.1 at 6 and 9 Steps in ComfyUI | ComfyUI Wiki