Text to speechHiggs TTS 3 is a 4B parameter text-to-speech model supporting 100+ languages with zero-shot voice cloning, expressive emotional control, and inline prosody/sound effects for voice agent applications.
Open-Weights
ComfyUI ecosystem updates — open-source model releases, custom nodes, workflows, and tools for image, video, and audio generation.
Text to speechHiggs TTS 3 is a 4B parameter text-to-speech model supporting 100+ languages with zero-shot voice cloning, expressive emotional control, and inline prosody/sound effects for voice agent applications.
Open-Weights
Text to videoSulphur 2 is a community fine-tune of LTX 2.3 offering text-to-video and image-to-video generation with a built-in prompt enhancer and distill LoRA, trained on 125K+ curated clips.
Open-Weights
OpenMOSS team releases MOVA (MOSS Video and Audio), an end-to-end synchronized video and audio generation foundation model that generates video and audio in a single inference pass, achieving precise lip-sync and environment-aware sound effects, fully open-sourcing model weights, training and inference code
Open-Weights
Alibaba Tongyi Lab releases Z-Image-Base, the non-distilled foundation model of the Z-Image series, preserving full generative potential and providing an ideal base for community fine-tuning and custom development
Open-Weights
DeepSeek open-sources DeepSeek-OCR-2, introducing the new DeepEncoder V2 vision encoder that mimics human reading logic through visual causal flow mechanism, achieving 91.09% accuracy on OmniDocBench v1.5
Open-Weights
Moonshot AI officially releases Kimi K2.5, a 1T parameter native multimodal agent model, continually pre-trained on 15 trillion mixed visual and text tokens, supporting image and video understanding with Agent Swarm autonomous collaboration mechanism
Open-Weights
Alibaba Qwen open-sources Qwen3-TTS series voice generation models, achieving 97ms ultra-low latency through Dual-Track modeling, supporting 3-second voice cloning and natural language voice design, covering 10 major languages
Open-Weights
Microsoft open-sources VibeVoice-ASR, a 9B parameter unified speech recognition model capable of processing up to 60 minutes of audio in a single pass, jointly completing recognition, speaker diarization, and timestamping in one inference process
Open-Weights
NVIDIA open-sources PersonaPlex-7B-v1, a 7 billion parameter full-duplex speech-to-speech dialogue model based on Moshi architecture, supporting simultaneous listening and speaking, natural interruptions, and role customization with interruption response latency as low as 240 milliseconds
Open-Weights
Text to videoStable Video Infinity 2.0 Pro has been officially released, adding support for the Wan 2.2 base model, providing HIGH and LOW versions of LoRA models for generating infinite-length video content in ComfyUI
Image editingQwen-Image-Layered can decompose images into multiple RGBA layers, with each layer independently editable without affecting other content, supporting operations like recoloring, replacement, deletion, resizing, and repositioning
Open-Weights
Image to 3DMicrosoft introduces TRELLIS.2, a 4 billion parameter large 3D generative model that converts images to high-quality 3D assets in seconds, with full PBR materials and complex topology support
Open-Weights
Tsinghua University team open sources TurboDiffusion, a video generation acceleration framework that achieves 100-205x end-to-end diffusion generation acceleration on a single RTX 5090 while maintaining video quality
Alibaba Cloud PAI team releases Z-Image-Turbo-Fun-Controlnet-Union, a ControlNet model supporting multiple control conditions including Canny, HED, Depth, Pose, and MLSD, trained at 1328 resolution
Open-Weights
Text to imageAlibaba AIDC-AI team releases Ovis-Image, a 7B parameter text-to-image model focused on high-quality text rendering, achieving excellent results on multiple text rendering benchmarks while running efficiently on a single high-end GPU
Open-Weights
Alibaba Tongyi Lab releases Z-Image-Turbo, an efficient 6B parameter image generation model that produces high-quality images in just 8 sampling steps and runs smoothly on consumer GPUs with 16GB VRAM
Open-Weights
MultimodalByteDance introduces the Sa2VA multimodal model, combining SAM2 and LLaVA technologies to achieve dense segmentation and visual question answering for both images and videos, attaining top performance on multiple benchmarks
Open-Weights
Tencent releases Hunyuan Image 3.0, the first open-source commercial-grade native multimodal image generation model with a total of 80B parameters, featuring world knowledge reasoning capabilities and supporting complex semantic understanding of thousands of characters
Open-Weights
Nunchaku team releases optimized version of 4-Bit Lightning Qwen-Image-Edit-2509 model, supporting 4/8 step fast inference, running smoothly on 8GB VRAM environment
Alibaba Tongyi Lab launches Wan-Animate, a unified character animation framework based on Wan2.2, supporting character animation generation and video character replacement, with open-sourced model weights and inference code
Open-Weights