Kandinsky 6.0: ComfyUI Setup and Model Guide
Run Kandinsky 6.0 in ComfyUI: 29B Pro and 3B Lite generate 5-second video with synchronized 44 kHz audio, plus a super-resolution model for Full HD.
Kandinsky 6.0
Video GenerationAudio-VideoText-to-VideoImage-to-VideoComfyUIKandinsky 6.0 Video is a family of foundation diffusion models for synchronized text-to-audio-video generation. Pro (29B) and Lite (3B) both produce 5-second clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video and image-to-audio-video modes, and a separate super-resolution model raises the output to Full HD.
| Developer | Kandinsky Lab |
| Release Date | 2026-10-06 (open weights) |
| Architecture | Multimodal video and audio diffusion transformer with flow matching (29B Pro, 3B Lite) |
| Text Encoder | Qwen2.5-VL with a CLIP pooled embedding |
| Audio | MMAudio VAE and BigVGAN-style vocoder, synchronized 44 kHz output |
| Output | 5-second clips at 24 fps, up to Full HD after super resolution |
| Generation Modes | Text-to-audio-video, image-to-audio-video |
| License | MIT |
Kandinsky 6.0 Video generates image and soundtrack together in one pass rather than dubbing audio on afterwards. A multimodal video and audio diffusion transformer models both streams, with a Qwen2.5-VL text encoder supplying token-level text embeddings and a CLIP model supplying a pooled embedding. The visual side decodes through a Hunyuan Video VAE, and the audio side runs an MMAudio VAE plus a BigVGAN-style vocoder for 44 kHz waveform output. Sampling uses flow matching with shift = 5.0.
Models
The family ships in two sizes, each with a base, a distilled 10-step and a pretrained checkpoint, plus a dedicated super-resolution pair:
| Checkpoint | Size | Notes |
|---|---|---|
| Kandinsky 6.0 Pro | 29B | Flagship generation checkpoint |
| Kandinsky 6.0 Pro distill | 29B | Distilled to 10 steps with PiFlow, used by default in the ComfyUI templates |
| Kandinsky 6.0 Pro pretrain | 29B | Pretrained base for fine-tuning |
| Kandinsky 6.0 Lite | 3B | Compact generation checkpoint |
| Kandinsky 6.0 Lite distill | 3B | Distilled 10-step variant |
| Kandinsky 6.0 Lite pretrain | 3B | Pretrained base for fine-tuning |
| Kandinsky 6.0 VSR | - | Super-resolution pipeline, x2, x2.25 and x4, 4 steps per tile |
| Kandinsky 6.0 VSR distill | - | Distilled super-resolution, 2 steps per tile |
Super resolution
Kandinsky6SRPipeline upscales generated frames with tiled diffusion in K-VAE latent space and blends overlapping tiles with Hann windows. resolution_scale accepts 2, 2.25 (the default route after generation) or 4, and min_overlap and tiles_batch_size trade tile blending and VRAM for speed.
ComfyUI
The official ComfyUI extension is published by kandinskylab on the ComfyUI Registry as kandinsky6 and kandinsky6-sr. Install both through ComfyUI Manager, restart, and open Workflow → Browse Templates → kandinsky6 for the bundled Text to Video+Audio and Image to Video+Audio templates. Model files, including the companion Diffusers JSON configs, are fetched by the extension's Download models button into the correct ComfyUI folders.
Resources
Guides and workflows related to this model series.
Comments
Sign in with GitHub to join the discussion.