VOID: Video Object and Interaction Deletion with 5B CogVideoX Architecture by Netflix
ComfyUI Wiki
VOID by Netflix is a video object removal model that deletes objects from videos along with all induced physical interactions, built on CogVideoX-Fun 5B.
V
VOID
Video EditingObject RemovalVideo InpaintingCogVideoX5BVOID (Video Object and Interaction Deletion) by Netflix removes objects from videos along with all interactions they induce on the scene — not just secondary effects like shadows and reflections, but physical interactions like objects falling when a person is removed. Built on CogVideoX-Fun-V1.5-5b-InP, it uses a novel quadmask conditioning (4-value mask encoding primary object, overlap, affected regions, and background). Features a two-pass architecture with optical flow-warped refinement for temporal consistency.
| Developer | Netflix |
| Release Date | 2026-04 |
| Architecture | CogVideoX 3D Transformer 5B + Quadmask Conditioning |
| License | Apache-2.0 |
| Resolution | 384×672 (default) |
| Max Frames | 197 |
| VRAM | 40GB+ (A100 recommended) |
| Precision | BF16 (FP8 quantization for memory efficiency) |
Guides and workflows related to this model series.
No articles found.
Comments
Sign in with GitHub to join the discussion.