VOID: Video Object and Interaction Deletion with 5B CogVideoX Architecture by Netflix

ComfyUI Wiki

VOID by Netflix is a video object removal model that deletes objects from videos along with all induced physical interactions, built on CogVideoX-Fun 5B.

V

VOID

Video EditingObject RemovalVideo InpaintingCogVideoX5B

VOID (Video Object and Interaction Deletion) by Netflix removes objects from videos along with all interactions they induce on the scene — not just secondary effects like shadows and reflections, but physical interactions like objects falling when a person is removed. Built on CogVideoX-Fun-V1.5-5b-InP, it uses a novel quadmask conditioning (4-value mask encoding primary object, overlap, affected regions, and background). Features a two-pass architecture with optical flow-warped refinement for temporal consistency.

DeveloperNetflix
Release Date2026-04
ArchitectureCogVideoX 3D Transformer 5B + Quadmask Conditioning
LicenseApache-2.0
Resolution384×672 (default)
Max Frames197
VRAM40GB+ (A100 recommended)
PrecisionBF16 (FP8 quantization for memory efficiency)

Guides and workflows related to this model series.

No articles found.

Comments

Sign in with GitHub to join the discussion.

Loading comments…