Qwen-Image-2.1 Next-Scene LoRA: Chained Storyboards
akhaliq's Next-Scene LoRA teaches Qwen-Image 2.1 to render the next directed shot from a single frame, and to chain into storyboards through a ComfyUI LoRA node.
Four directed shots from one photograph. Each frame is the previous output fed back with a new instruction, at the recommended step-2,500 checkpoint and strength 0.7.
What it adds to Qwen-Image 2.1
The next-scene format comes from lovis93/next-scene-qwen-image-lora-2509, which has passed 313,000 downloads on the Qwen-Image 2509 edit model. This is the port to the current base: retrained natively on Qwen-Image 2.1, where the capability does not exist yet. It pairs with akhaliq's earlier Multiple-Angles LoRA, which controls where the camera is, while Next-Scene controls where the story goes next.
Prompts start with Next Scene: and lead with the camera direction, for example "Next Scene: The camera pulls back to a sweeping aerial view, revealing the fleet behind the cliffs":
Next Scene: The camera tracks forward and tilts down as rain begins to streak the lens.
Next Scene: Cut to a low-angle shot; sunlight breaks through the clouds behind her.
Next Scene: The camera pans right, revealing the fleet massing behind the ridge.Chaining, with a measured recipe
Feeding each output back in as the next input is what turns the adapter into a storyboard tool, but the author measured a failure mode instead of guessing at one: naive chaining at strength 0.9 accumulates film grain until it becomes visible color noise by hop 3 or 4. The verified recipe for clean chains is strength 0.7 or lower plus a 0.5 pixel gaussian blur on every chained input, which breaks the grain-amplification feedback loop. The four-hop example above was rendered exactly that way.
![]() | ![]() |
|---|---|
| Hop 1: the camera pulls back to a wide shot | Hop 2: pans right, the astronaut bounds across the regolith |
![]() | ![]() |
|---|---|
| Hop 3: cut to a low-angle shot as the astronaut salutes | Hop 4: push in on the gold visor, the lunar module reflected in it |
Which checkpoint to use
The repository ships three converted checkpoints plus the raw training outputs:
| Checkpoint | Verdict |
|---|---|
next_scene_step2500.safetensors | Recommended. Final step, converted to the diffusers key layout. Strongest scene transformation while holding composition and identity. |
next_scene_step1500.safetensors | Softer alternative, gentler on identity. Use it if 2,500 pushes too hard. |
next_scene_step1000.safetensors | Weakest transformation, mostly useful for studying the training curve. |
checkpoints/qwen21_next_scene_v1_*.safetensors | Raw ai-toolkit format. The MLP keys use ai-toolkit's fused naming, which diffusers silently skips. Only usable through ai-toolkit or its own loader. |
The author also published a checkpoint sweep on one storm prompt across step 1,000, 1,500 and 2,500, with a no-LoRA control. The base model already attempts the storm, since Qwen-Image 2.1 edits well on its own, but step 2,500 is the cell that rolls a full storm front over the valley while keeping the explorer, the ruins and the palette intact.
One cinematic starter, the same storm instruction at each checkpoint, plus a no-LoRA control.
How it was trained
- Base: Qwen-Image 2.1,
arch: qwen_image_2. - Data: 1,144 scene-progression pairs, consecutive frames from the Blender open films Sintel and Tears of Steel (CC-BY), captioned as director-style "Next Scene: ..." instructions by Qwen3-VL-4B over each frame-to-next-frame pair. Both the raw pairs and the captioned set are published.
- Training: ostris/ai-toolkit, rank 64 / alpha 64, adamw8bit, learning rate 1e-4, weighted timesteps, caption dropout 0.05, 512 pixels, 2,500 steps, on a uint3-quantized base in bf16.
Scope and limits
The adapter was trained on cinematic film frames, one animated and one live action. It generalizes to photographs, but it is strongest on filmic scenes, landscapes and establishing shots, and static portraits are out of scope by design. Training was done at 512 pixels; larger canvases inherit the behavior but were not trained on. Per-250-step sample grids are kept in checkpoints/samples/, and a Trackio dashboard plus a demo Space are linked from the model card.
ComfyUI availability
The recommended checkpoint is already converted, so it loads in any Qwen-Image 2.1 workflow through a LoRA node at strength 0.7 to 0.9, with the image passed as the edit or reference input. There is no separate workflow file for the adapter, and the root-level next_scene_step*.safetensors files are the ones to use in ComfyUI; the checkpoints/ copies are ai-toolkit format.




Comments
Sign in with GitHub to join the discussion.