LongLive-Plug: NVIDIA Distills Few-Step LoRAs for H3 and Wan

ComfyUI Wikinews

NVIDIA's LongLive-Plug distills few-step, CFG and long-context LoRAs once per backbone and reuses them across 54 downstream models, with weights for MiniMax H3 and Wan.

LongLive-Plug is an NVIDIA once-for-all distillation framework from the SANA team: instead of distilling every downstream video model separately, it learns reusable capabilities as LoRAs on a base model and attaches them to compatible downstream models with no retraining. Weights for MiniMax H3, Wan2.1-T2V-14B and Wan2.2-TI2V-5B landed on Hugging Face on 2026-09-29 alongside the paper (arXiv 2609.38154).
Matched transfer comparisons: native, naive four-step, per-target distillation, and LongLive-Plug

Matched frames from the paper's Figure 3. Naive four-step sampling (second column) blurs the frame; the transferred adapters (fourth column) hold up at four steps. Boxes mark identical coordinates across methods.

One distillation, many downstream models

Building a specialised video model often ends with a distillation stage, to cut sampling steps or to make long rollouts hold together. That stage is normally repeated for every model: collect task data, run the teacher, optimise again. A depth-conditioned Wan fine-tune and a robotics world model trained on the same backbone each paid for their own few-step LoRA.

LongLive-Plug splits that work into three capabilities that are distilled once per backbone family and then carried over as ordinary LoRAs:

  1. Few-step LoRA: fewer sampling steps, trained with distribution matching and a CFG-guided teacher.
  2. CFG LoRA: folds classifier-free guidance into a single conditional forward pass.
  3. Long-context LoRA: corrects error accumulation in causal autoregressive rollouts.

The adapters are designed to survive the changes a downstream model makes. Adding conditioning branches or expanding output channels does not invalidate them, so the same file set attaches to full fine-tunes, task LoRAs and models with extra control modules.

The CFG weight works like a dial

Guidance is normally paid for twice per step: one conditional pass, one unconditional pass. The CFG LoRA removes the second pass, and because guidance is distilled into the adapter, the LoRA weight itself becomes the guidance control even though the adapter was trained at one fixed scale. Raising it strengthens prompt attributes without re-running the unconditional branch.

The important detail is that scaling the whole coupled LoRA does not do the same thing: it moves the update and the step count together and degrades the output. That is why the few-step and CFG adapters are separate files, and why they are weighted separately. The paper's recommended ratio is few-step : CFG = 1 : 0.5, and the model cards repeat that these numbers are adapter weights, not the model's native CFG scale.

Independent CFG control with a CFG-only LoRA

The paper's Figure 5. Scaling the coupled LoRA alone barely responds to the prompt; adding a separately weighted CFG LoRA strengthens the boxed attributes while the coupled LoRA stays fixed.

What was released

The Hugging Face collection holds six adapters, one few-step and one CFG file per backbone.

Base modelFew-step adapterCFG adapterNotes
MiniMax H3LongLive-Plug-MiniMax-H3-few-step (2.6 GB)LongLive-Plug-MiniMax-H3-cfg (2.6 GB)PEFT format. The cards say to use the two separately for now, combined use is not recommended yet
Wan2.1-T2V-14BLongLive-Plug-Wan2.1-T2V-14B-few-step (1.2 GB)LongLive-Plug-Wan2.1-T2V-14B-cfgShips a generator_lora_lightx2v.safetensors alongside the PEFT export
Wan2.2-TI2V-5BLongLive-Plug-Wan2.2-TI2V-5B-few-step (1.3 GB)LongLive-Plug-Wan2.2-TI2V-5B-cfgPEFT adapter_model.safetensors

The Wan2.1-14B few-step file is the one that drops straight into an existing Wan workflow: it is exported in the lightx2v naming that ComfyUI's Wan LoRA loaders already read, so it goes in ComfyUI/models/loras/ like any other Wan speed LoRA. The MiniMax H3 adapters are PEFT exports and need conversion before ComfyUI will load them; Kijai has published a converted bf16 version at Kijai/MiniMax-H3-experimental as minimax_h3_ELM_longlive_plug_4step_lora_bf16.safetensors (1.96 GB).

Coverage and cost

The paper verifies training-free deployment on 54 downstream models across three backbone families (24 on Wan2.1-14B, 24 on Wan2.2-TI2V-5B, six on MiniMax H3) and eight task categories: world modeling, robotics, structure-conditioned generation, camera and trajectory control, video editing and restoration, subject and avatar generation, audio and RGBA generation, and style adaptation. Against native 20 to 50 step schedules, attaching the base-distilled adapters gives four-step, CFG-free inference, a 5x to 12.5x reduction in denoising steps.

The cost argument is the other half of the paper. Base distillation is about 80 H100 GPU-hours (700 iterations across 32 GPUs, roughly 2.5 hours). Distilling the four downstream tasks individually instead costs a further 83.9 (depth), 150.0 (world modeling), 86.8 (pose) and 56.1 (robotics) GPU-hours, about 456.8 in total. Reusing the base adapters stays at the fixed ~80 GPU-hours, with no task-specific data collection.

Additional Wan2.1-14B transfer cases

The paper's Figure 9: LongVie 2, ABot-PhysWorld, MagicTryOn and Wan-Alpha, covering world modeling, robotics, subject conditioning and RGBA output, each comparing two native frames with the same timestamps after four-step transfer.

MiniMax H3

The H3 side of the release is the narrowest. Six of the 54 verified models are H3, and the paper states its ten H3 comparisons are single-seed cases with no repeated-run confidence intervals, that the native default is a reference operating point rather than ground truth, and that static frames do not evaluate audio quality or synchronisation. The examples show four-step H3 holding character outlines better than naive four-step Euler sampling while appearance can still differ from the undistilled multi-step result.

H3 task transfer: action control and line-art colouring

The paper's Figure 11: H3-World under a forward-action condition and LineartAnime, comparing undistilled multi-step inference, naive four-step sampling, and four-step inference with LongLive-Plug.

Limits

  • The transfer idea works on the backbones it was distilled for. Wan2.1, Wan2.2 and MiniMax H3 each have their own adapter set, and the paper does not claim cross-family reuse.
  • Effect on H3 is measured on ten single-seed cases with no confidence intervals, and the audio track is not part of the evaluation.
  • For MiniMax H3 the two adapters should be used one at a time for now, so few-step speed and CFG control cannot be composed there yet.
  • Guidance control is a weight dial, not a new sampler: the few-step and CFG files have to be loaded and weighted separately to get both.

Availability

Project page: nvlabs.github.io/LongLive/LongLive-Plug
Paper: arXiv 2609.38154
Demo film: YouTube
Weights: Efficient-Large-Model/longlive-plug
ComfyUI conversion of the H3 few-step adapter: Kijai/MiniMax-H3-experimental

Comments

Sign in with GitHub to join the discussion.

Loading comments…
LongLive-Plug: NVIDIA Distills Few-Step LoRAs for H3 and Wan | ComfyUI Wiki