H3 Person Remover LoRA: Delete a Person From Video
Akatz Labs' Person Remover LoRA tracks a person with SAM 3.1, fills the mask green and rebuilds the background in overlapping windows, with a ready ComfyUI workflow.
From the eight-example showcase: original clip, green mask, reconstructed background.
The author's before/after showcase, re-encoded at 1280 pixels wide.
What the adapter does, and what it does not
The LoRA does not segment people and does not invent the clean background on its own. SAM 3.1 supplies the tracking mask from a text description such as man in gray shirt, and you still have to supply a clean version of the video's first frame with the person already gone, either from an image editor or an image-editing model. From there the window reroll workflow does the rest: green-masked frames go in, H3 regenerates the covered region, and the reconstructed boundary frame carries the continuation into the next window.
Default removal prompt:
Remove the green-masked person and reconstruct the background. Preserve the rest of the video, including its camera motion and frame timing.
Window reroll, by the numbers
The workflow is tuned rather than generic, and the recorded settings are worth copying:
| Setting | Value |
|---|---|
| Scheduler | beta / simple |
| CFG | 1 |
| Video / audio sigma shift | 12 / 3 |
| Mask expansion | 5 pixels |
| Frame rate | 24 fps |
| Seed | 904234 |
No Turbo or VFX LoRA is required. The workflow uses H3 frame lengths of the form 17n + 5 (22, 39, 56, 73 ...) and recommends starting at 22 frames, since a longer window costs more memory without automatically improving the result. Each continuation uses the generated boundary frame as its next reference and carries 18 frames of generated video and audio history, and the relay removes the overlap and trims the result back to the source frame count. Output is silent by default; you connect the original audio from Get Video Components to the final Create Video node to keep the soundtrack.
Input requirements are specific: 24 fps video, width and height divisible by 32, and ideally a short continuous shot of about five seconds.
Training record
V1 is a rank-16 adapter on the pruned Ref2VA model, trained for 2,000 optimizer updates after an initial 250 updates of static five-frame pairs and preservation regularization.
| Setting | Recorded value |
|---|---|
| Architecture | minimax_h3_ref2va, pruned |
| Rank / alpha | 16 / 16, excluding adaln_proj |
| Tensors | BF16, 416 tensors across 208 adapter targets |
| Optimizer / learning rate | AdamW8bit / 5e-5 |
| Text encoder | NVFP4 Qwen3VL |
| Task / regularization resolution | 1152 / 256, aspect-ratio buckets |
| Task frame lengths | 5, 56, 73, 90, 124 at 24 fps |
| Regularization length | 107 frames at 24 fps, with audio |
| Runtime | AI Toolkit 0.13.23 plus a recorded aligned-video-guide patch |
The mixed phase's training pool holds 128 static pairs, 80 moving-mask excerpts from 20 scenes, and six preservation clips. The dataset card documents the construction, split and review status, and training/ carries the configs, schedules, model hashes and runtime patch.
Limitations
Akatz Labs is direct about what still fails:
- H3 regenerates the whole scene, so unmasked areas are not pixel-locked to the original even though the training pairs preserve pixels outside their masks.
- A missed hand, hair edge, shadow, reflection or occluded body part can survive, and a larger mask asks the model to invent more background.
- A poor clean first frame propagates through later windows, and strong camera motion or hard cuts can break continuity.
- Rerolling a window changes its continuation, so the rebuilt suffix has to be inspected too.
- The recorded dataset-acceptance flag remains false: passing technical checks is not the same as human approval of every example.
Availability
The workflow needs a current ComfyUI with native MiniMax H3 and SAM 3.1 support, plus H3-Person-Remover-V1.safetensors in models/loras/. The published workflow pins the H3 ref2va pruned INT8 ConvRot transformer, the NVFP4 Qwen3VL text encoder, the H3 video and audio VAEs, and SAM 3.1 multiplex; prior validation ran on an RTX 4090 with ComfyUI 0.37.0, which the author describes as a tested configuration rather than a minimum specification. H3 Relay provides the window preview and reroll controls that the workflow uses.
Comments
Sign in with GitHub to join the discussion.