ComfyUI News & Open-Source AI Releases

ComfyUI ecosystem updates — open-source model releases, custom nodes, workflows, and tools for image, video, and audio generation.

Multimodal

OpenMOSS Releases MOVA - Open-Source Synchronized Video and Audio Generation Model

Multimodal
Base model
OpenMOSS Releases MOVA - Open-Source Synchronized Video and Audio Generation Model

OpenMOSS team releases MOVA (MOSS Video and Audio), an end-to-end synchronized video and audio generation foundation model that generates video and audio in a single inference pass, achieving precise lip-sync and environment-aware sound effects, fully open-sourcing model weights, training and inference code

Open-Weights

Speech recognition

Microsoft Releases VibeVoice-ASR - Speech Recognition Model Supporting 60-Minute Long Audio Single-Pass Processing

Speech recognition
Base model
Microsoft Releases VibeVoice-ASR - Speech Recognition Model Supporting 60-Minute Long Audio Single-Pass Processing

Microsoft open-sources VibeVoice-ASR, a 9B parameter unified speech recognition model capable of processing up to 60 minutes of audio in a single pass, jointly completing recognition, speaker diarization, and timestamping in one inference process

Open-Weights

Text to image

Tencent Open Sources Hunyuan Image 3.0 - World's Largest Open-Source Text-to-Image Model

Text to image
Base model
Tencent Open Sources Hunyuan Image 3.0 - World's Largest Open-Source Text-to-Image Model

Tencent releases Hunyuan Image 3.0, the first open-source commercial-grade native multimodal image generation model with a total of 80B parameters, featuring world knowledge reasoning capabilities and supporting complex semantic understanding of thousands of characters

Open-Weights