MiniMax H3 360 Orbit LoRA for ComfyUI: One Photo, Full Orbit

ComfyUI Wikinews

A community LoRA for MiniMax H3 FL2VA turns one photo into a full 360 camera orbit that lands back on its own first frame, with ideal ComfyUI settings.

MiniMax-H3-360-Orbit-LoRA is a community adapter for MiniMax H3 FL2VA that walks the camera all the way around a subject while the scene stays frozen in a single instant, then lands on the exact frame it started from. Feed the same image as the first and the last keyframe and the clip loops back onto itself, so orbits can be chained without a visible seam. Pablo Dawson published the weights on September 27 together with the settings, the prompt and the training recipe behind them.
Four 360 degree orbits generated from a single photo each with the MiniMax H3 360 Orbit LoRA

Four separate orbits, each generated from one photo as the first and last keyframe. Source: MiniMax-H3-360-Orbit-LoRA.

What the LoRA does

MiniMax H3 has two video paths, and the gap between them is what this adapter fills. Reference-to-video treats its input images as loose appearance references, so a clip does not end on a known frame and two clips cannot be joined cleanly. First-and-last-frame generation pins both ends, which makes the joints exact, but out of the box it barely moves when the first and last frames are the same picture, because the model reads the request as a still.

The adapter teaches the model a real, geometry-consistent orbit instead. The scene is declared frozen, only the camera travels, and because the end frame is pinned to the start frame the motion closes into a loop.

An orbit rendered from a single photo, with the same image used as the first and last keyframe.

Compared with the base model

Each comparison below renders the same input three times with an identical seed, prompt, resolution and step count. From left to right: base model with the first frame only, base model with first and last frame, and the base plus this LoRA with first and last frame.

Comparison of the base model with one keyframe, the base model with two identical keyframes, and the model plus the orbit LoRA

A skate scene: the base model either drifts away from the opening view or freezes, while the LoRA completes the orbit and returns to the first frame.

Second comparison of base model and LoRA orbits around a standing figure

The same three-way comparison around a standing figure.

Full-resolution MP4 of the three-way comparison above.

The LoRA on its own, without the base-model panels next to it.

The prompt

The author asks for this prompt to be used verbatim:

One frozen instant. Only the camera moves. In a continuous 360 orbit. Preserve every person and object in exactly the same world position, orientation, shape and pose throughout the shot. Airborne objects remain suspended at the captured height and angle: no wobbling, shaking, spinning, drifting, falling or continued action. Keep faces, hands, clothing, liquids and the background motionless while retaining their natural appearance. Camera parallax is the only source of apparent movement. No cuts, zoom, morphing or added objects.
settingvalue
base weightsMiniMax H3 FL2VA pruned, the INT8 ConvRot repack from Comfy-Org/MiniMax-H3
keyframesthe same image as the first and the last frame, for a full 360 degree loop
resolution768 × 768
frames73 (about 3 s at 24 fps)
steps28
guidancenone. MiniMax H3 is guidance-distilled, so there is no CFG and no negative prompt
LoRA strength1.0
audiooff

Running it in ComfyUI

The adapter is a single model-only LoRA with ComfyUI naming (diffusion_model.* keys), so no custom nodes are required:

  1. Download minimax_h3_flf2v_lora_v1.safetensors into ComfyUI/models/loras/.
  2. Open a MiniMax H3 first-and-last-frame workflow and point it at the pruned INT8 FL2VA base with its text encoder and the H3 video and audio VAEs.
  3. Apply the LoRA at strength 1.0 through a model-only LoRA loader.
  4. Load one image into both keyframe slots and paste the prompt above unchanged.

The card notes that the LoRA was trained and tested with ostris/ai-toolkit and its minimax_h3 extension, and that the training adapter used during training is not needed at inference time.

How it was trained

The dataset is deliberately one-note. Every clip shows the motion the LoRA should learn and nothing else, so the model never sees a subject move:

  • 28 orbit renders around hand-picked human Gaussian splats. Because the splats are static 3D scenes, every frame is geometrically consistent by construction, and the camera is the only thing that changes between frames.
  • Clip format: 768 × 768, 73 frames, 24 fps, one caption for all clips (the prompt above), caption dropout 0.05, no audio.
  • Adaptor: rank 16, alpha 16, applied to all transformer linear layers except adaln_proj, 208 modules in total.
  • Training: 3,000 steps (about 107 epochs), batch size 1, AdamW 8-bit at lr 1e-4, bf16 with gradient checkpointing, flow-matching shifted timesteps, MSE loss. One NVIDIA A100 80 GB took 10 h 24 min, roughly 12.5 s per step.

Limitations

The author is explicit about how narrow the release is:

  • The domain is narrow. Training used 28 human-centric square clips of 73 frames. Other subjects, other aspect ratios and other clip lengths are untested.
  • Small movements can survive. Some subtle motion, blinks in particular, may still appear even though the scene is meant to be frozen.

Availability

Download minimax_h3_flf2v_lora_v1.safetensors (148 MB)

Comments

Sign in with GitHub to join the discussion.

Loading comments…
MiniMax H3 360 Orbit LoRA for ComfyUI: One Photo, Full Orbit | ComfyUI Wiki