MiniMax H3 360 Orbit LoRA for ComfyUI: One Photo, Full Orbit
A community LoRA for MiniMax H3 FL2VA turns one photo into a full 360 camera orbit that lands back on its own first frame, with ideal ComfyUI settings.
Four separate orbits, each generated from one photo as the first and last keyframe. Source: MiniMax-H3-360-Orbit-LoRA.
What the LoRA does
MiniMax H3 has two video paths, and the gap between them is what this adapter fills. Reference-to-video treats its input images as loose appearance references, so a clip does not end on a known frame and two clips cannot be joined cleanly. First-and-last-frame generation pins both ends, which makes the joints exact, but out of the box it barely moves when the first and last frames are the same picture, because the model reads the request as a still.
The adapter teaches the model a real, geometry-consistent orbit instead. The scene is declared frozen, only the camera travels, and because the end frame is pinned to the start frame the motion closes into a loop.
An orbit rendered from a single photo, with the same image used as the first and last keyframe.
Compared with the base model
Each comparison below renders the same input three times with an identical seed, prompt, resolution and step count. From left to right: base model with the first frame only, base model with first and last frame, and the base plus this LoRA with first and last frame.
A skate scene: the base model either drifts away from the opening view or freezes, while the LoRA completes the orbit and returns to the first frame.
The same three-way comparison around a standing figure.
Full-resolution MP4 of the three-way comparison above.
The LoRA on its own, without the base-model panels next to it.
The prompt
The author asks for this prompt to be used verbatim:
One frozen instant. Only the camera moves. In a continuous 360 orbit. Preserve every person and object in exactly the same world position, orientation, shape and pose throughout the shot. Airborne objects remain suspended at the captured height and angle: no wobbling, shaking, spinning, drifting, falling or continued action. Keep faces, hands, clothing, liquids and the background motionless while retaining their natural appearance. Camera parallax is the only source of apparent movement. No cuts, zoom, morphing or added objects.Recommended settings
| setting | value |
|---|---|
| base weights | MiniMax H3 FL2VA pruned, the INT8 ConvRot repack from Comfy-Org/MiniMax-H3 |
| keyframes | the same image as the first and the last frame, for a full 360 degree loop |
| resolution | 768 × 768 |
| frames | 73 (about 3 s at 24 fps) |
| steps | 28 |
| guidance | none. MiniMax H3 is guidance-distilled, so there is no CFG and no negative prompt |
| LoRA strength | 1.0 |
| audio | off |
Running it in ComfyUI
The adapter is a single model-only LoRA with ComfyUI naming (diffusion_model.* keys), so no custom nodes are required:
- Download
minimax_h3_flf2v_lora_v1.safetensorsintoComfyUI/models/loras/. - Open a MiniMax H3 first-and-last-frame workflow and point it at the pruned INT8 FL2VA base with its text encoder and the H3 video and audio VAEs.
- Apply the LoRA at strength 1.0 through a model-only LoRA loader.
- Load one image into both keyframe slots and paste the prompt above unchanged.
The card notes that the LoRA was trained and tested with ostris/ai-toolkit and its minimax_h3 extension, and that the training adapter used during training is not needed at inference time.
How it was trained
The dataset is deliberately one-note. Every clip shows the motion the LoRA should learn and nothing else, so the model never sees a subject move:
- 28 orbit renders around hand-picked human Gaussian splats. Because the splats are static 3D scenes, every frame is geometrically consistent by construction, and the camera is the only thing that changes between frames.
- Clip format: 768 × 768, 73 frames, 24 fps, one caption for all clips (the prompt above), caption dropout 0.05, no audio.
- Adaptor: rank 16, alpha 16, applied to all transformer linear layers except
adaln_proj, 208 modules in total. - Training: 3,000 steps (about 107 epochs), batch size 1, AdamW 8-bit at lr 1e-4, bf16 with gradient checkpointing, flow-matching shifted timesteps, MSE loss. One NVIDIA A100 80 GB took 10 h 24 min, roughly 12.5 s per step.
Limitations
The author is explicit about how narrow the release is:
- The domain is narrow. Training used 28 human-centric square clips of 73 frames. Other subjects, other aspect ratios and other clip lengths are untested.
- Small movements can survive. Some subtle motion, blinks in particular, may still appear even though the scene is meant to be frozen.
Availability
- LoRA weights: pablodawson/MiniMax-H3-360-Orbit-LoRA
- Base model repack: Comfy-Org/MiniMax-H3
- Training tool: ostris/ai-toolkit
Comments
Sign in with GitHub to join the discussion.