Qwen-Image-2.1 Next-Scene LoRA: Chained Storyboards

ComfyUI Wikinews

akhaliq's Next-Scene LoRA teaches Qwen-Image 2.1 to render the next directed shot from a single frame, and to chain into storyboards through a ComfyUI LoRA node.

Qwen-Image-2.1-Next-Scene-LoRA is a rank-64 adapter that teaches Qwen-Image 2.1 one specific edit: given the current frame and a director-style instruction, generate the next shot instead of editing the one you have. Camera moves, reveals and lighting shifts are applied while the world, the characters and the style stay put, and the output can be fed back in as the next input to build a storyboard.
Four hops from a single photograph, each frame generated from the previous one

Four directed shots from one photograph. Each frame is the previous output fed back with a new instruction, at the recommended step-2,500 checkpoint and strength 0.7.

What it adds to Qwen-Image 2.1

The next-scene format comes from lovis93/next-scene-qwen-image-lora-2509, which has passed 313,000 downloads on the Qwen-Image 2509 edit model. This is the port to the current base: retrained natively on Qwen-Image 2.1, where the capability does not exist yet. It pairs with akhaliq's earlier Multiple-Angles LoRA, which controls where the camera is, while Next-Scene controls where the story goes next.

Prompts start with Next Scene: and lead with the camera direction, for example "Next Scene: The camera pulls back to a sweeping aerial view, revealing the fleet behind the cliffs":

Next Scene: The camera tracks forward and tilts down as rain begins to streak the lens.
Next Scene: Cut to a low-angle shot; sunlight breaks through the clouds behind her.
Next Scene: The camera pans right, revealing the fleet massing behind the ridge.

Chaining, with a measured recipe

Feeding each output back in as the next input is what turns the adapter into a storyboard tool, but the author measured a failure mode instead of guessing at one: naive chaining at strength 0.9 accumulates film grain until it becomes visible color noise by hop 3 or 4. The verified recipe for clean chains is strength 0.7 or lower plus a 0.5 pixel gaussian blur on every chained input, which breaks the grain-amplification feedback loop. The four-hop example above was rendered exactly that way.

First hopSecond hop
Hop 1: the camera pulls back to a wide shotHop 2: pans right, the astronaut bounds across the regolith
Third hopFourth hop
Hop 3: cut to a low-angle shot as the astronaut salutesHop 4: push in on the gold visor, the lunar module reflected in it

Which checkpoint to use

The repository ships three converted checkpoints plus the raw training outputs:

CheckpointVerdict
next_scene_step2500.safetensorsRecommended. Final step, converted to the diffusers key layout. Strongest scene transformation while holding composition and identity.
next_scene_step1500.safetensorsSofter alternative, gentler on identity. Use it if 2,500 pushes too hard.
next_scene_step1000.safetensorsWeakest transformation, mostly useful for studying the training curve.
checkpoints/qwen21_next_scene_v1_*.safetensorsRaw ai-toolkit format. The MLP keys use ai-toolkit's fused naming, which diffusers silently skips. Only usable through ai-toolkit or its own loader.

The author also published a checkpoint sweep on one storm prompt across step 1,000, 1,500 and 2,500, with a no-LoRA control. The base model already attempts the storm, since Qwen-Image 2.1 edits well on its own, but step 2,500 is the cell that rolls a full storm front over the valley while keeping the explorer, the ruins and the palette intact.

Checkpoint sweep across training steps with a no-LoRA control

One cinematic starter, the same storm instruction at each checkpoint, plus a no-LoRA control.

How it was trained

  • Base: Qwen-Image 2.1, arch: qwen_image_2.
  • Data: 1,144 scene-progression pairs, consecutive frames from the Blender open films Sintel and Tears of Steel (CC-BY), captioned as director-style "Next Scene: ..." instructions by Qwen3-VL-4B over each frame-to-next-frame pair. Both the raw pairs and the captioned set are published.
  • Training: ostris/ai-toolkit, rank 64 / alpha 64, adamw8bit, learning rate 1e-4, weighted timesteps, caption dropout 0.05, 512 pixels, 2,500 steps, on a uint3-quantized base in bf16.

Scope and limits

The adapter was trained on cinematic film frames, one animated and one live action. It generalizes to photographs, but it is strongest on filmic scenes, landscapes and establishing shots, and static portraits are out of scope by design. Training was done at 512 pixels; larger canvases inherit the behavior but were not trained on. Per-250-step sample grids are kept in checkpoints/samples/, and a Trackio dashboard plus a demo Space are linked from the model card.

ComfyUI availability

The recommended checkpoint is already converted, so it loads in any Qwen-Image 2.1 workflow through a LoRA node at strength 0.7 to 0.9, with the image passed as the edit or reference input. There is no separate workflow file for the adapter, and the root-level next_scene_step*.safetensors files are the ones to use in ComfyUI; the checkpoints/ copies are ai-toolkit format.

Comments

Sign in with GitHub to join the discussion.

Loading comments…
Qwen-Image-2.1 Next-Scene LoRA: Chained Storyboards | ComfyUI Wiki