Bernini v2: Semantic Planning Video Model by ByteDance

ComfyUI Wiki

Bernini-Diffusers-v2 is ByteDance's semantic-planning video model with a Wan2.2 dual DiT renderer, packaged for ComfyUI with FP8 and NVFP4 weights.

B

Bernini v2

Video GenerationDiTByteDanceComfyUI

ByteDance's Bernini-Diffusers-v2 — the second diffusers release of the Bernini framework. A Qwen2.5-VL semantic planner produces plan tokens that a Wan2.2 dual DiT renderer turns into video, supporting text-to-video, reference-to-video, video editing, and aesthetic transfer in a single workflow.

DeveloperByteDance
Release Date2026-08-13
ArchitectureQwen2.5-VL semantic planner + Wan2.2 dual DiT renderer (high/low noise, switch at 0.875)
LicenseApache-2.0
VariantsFP8 scaled · NVFP4 (both DiTs) via the ComfyUI pack
CapabilitiesT2V · R2V · V2V editing · reference-guided editing · aesthetic transfer

Guides and workflows related to this model series.

No articles found.

Comments

Sign in with GitHub to join the discussion.

Loading comments…