Looped-DiT: A 260M Model Beats One 6.5x Bigger
SenseTime's OpenSenseNova released Looped-DiT, which reruns shared transformer blocks inside each denoising step so a small model can beat a much larger one.
The method: a shared block group loops N times per denoising step, with self-modulating attention regulating the loop and deep supervision training the state after every pass.
Why naive looping fails
Running the same blocks twice does not reliably improve a diffusion transformer. The paper traces that to two problems: weak supervision across intermediate loops, and attention updates that progressively erode local information as the loop count grows. Looped-DiT adds two components to fix them:
- Deep Supervision decodes the state after every loop through the post-loop blocks and trains each of those predictions against the same clean-image target.
- Self-Modulating Attention regulates the attention updates inside the loop, using exclusive self-attention (XSA) or a head-wise attention gate.
The backbone is the pixel-space MMDiT from MiniT2I, conditioned on a frozen FLAN-T5-Large. The repository covers training for B/32, B/16 and L/16, inference at any loop depth, evaluation across six benchmarks, and dataset preparation scripts for each training set.
Results
Scores use EMA weights, 100 Euler steps, guidance 6.0 and loop depth 4:
| Model | Patch | GenEval | DPG | PRISM | CoRe | Spatial | TIIF | Avg |
|---|---|---|---|---|---|---|---|---|
| Looped-DiT B/32 | 32 | 85.1 | 85.3 | 54.4 | 44.5 | 52.3 | 76.1 | 66.3 |
| Looped-DiT B/16 | 16 | 87.4 | 87.0 | 67.0 | 53.5 | 54.6 | 79.7 | 71.5 |
The claim that matters is in the paper's abstract: under matched-parameter and matched-compute settings the looped design consistently outperforms non-looped baselines, and deeper loops buy more than extra denoising steps do under a fixed inference budget. The authors also report that deeper loops progressively correct mistakes made in earlier loops, which they describe as behavior suggestive of latent reasoning.
Running it
Both checkpoints are published as PyTorch files, sensenova/Looped-DiT-B16 and sensenova/Looped-DiT-B32, and inference is a single module call. --loops sets the loop depth, and passing several depths gives one row each so the depth can be compared on a single prompt:
hf download sensenova/Looped-DiT-B16 looped-dit-b16.pt --local-dir checkpoints
python -m looped_dit.sample --checkpoint checkpoints/looped-dit-b16.pt \
--prompt "a red cube on top of a blue sphere" --loops 1 2 3 4 --out loops.pngThe paper's main results use Euler with 100 steps, guidance 6.0, loop depth 4 at 512x512 on EMA weights in bfloat16. Depth 4 is the trained depth, but the authors note other depths work without retraining, so the loop count is a knob at inference time rather than a fixed property of the checkpoint.
Availability
This is a research release, not a ComfyUI integration. There is currently no ComfyUI node, no Comfy-Org repackaging and no diffusers pipeline, so running it means the repository's own Python environment: a CUDA build of PyTorch 2.1 or newer, requirements.txt, and the evaluation stacks (mmdet, vLLM, modelscope) installed separately since the benchmark data and Mask2Former weights are not bundled. Given the release date and the 51 GitHub stars it has gathered so far, a community port would not be surprising, but nothing exists at the time of writing.
Comments
Sign in with GitHub to join the discussion.