Image to 3DTencent launches Hunyuan3D 2.0 system with a two-stage process for generating high-quality 3D models, featuring multiple open-source model series that support creating high-resolution 3D assets from text and images
Open-Weights
ComfyUI ecosystem updates — open-source model releases, custom nodes, workflows, and tools for image, video, and audio generation.
Image to 3DTencent launches Hunyuan3D 2.0 system with a two-stage process for generating high-quality 3D models, featuring multiple open-source model series that support creating high-resolution 3D assets from text and images
Open-Weights
Kuaishou Technology has launched ReCamMaster, a generative video technology that allows users to create new camera perspectives and motion paths from a single video.
Open-Weights
Text to videoLuChen Technology releases Open-Sora 2.0 open-source video generation model, achieving performance close to top commercial models with just $200,000 in training costs
Open-Weights
Ali Tongyi Lab launches multifunctional video creation and editing model VACE, integrating various video processing tasks into a single framework to lower the barrier of video creation
Open-Weights
Microsoft Research introduces intelligent layered generation based on global text prompts, supporting creation of transparent images with 50+ independent layers
Open-Weights
Image to videoTencent's Hunyuan team releases open-source model for generating 5-second videos from single images, featuring smart motion generation and custom effects
Open-Weights
MultimodalAlibaba introduces a document analysis system capable of understanding both text and images, improving processing efficiency for complex documents by over 10%
Open-Weights
Text to imageTHUDM releases CogView4 open-source image generation model with native Chinese support, leading in multiple benchmark tests
Open-Weights
MultimodalSesame Research introduces dual-Transformer conversational voice model CSM, achieving human-like interaction with open-source core architecture
Open-Weights
Alibaba has officially open-sourced its latest video generation model, Wan2.1, which can run with only 8GB of video memory, supporting high-definition video generation, dynamic subtitles, and multi-language dubbing, surpassing models like Sora with a total score of 86.22% on the VBench leaderboard
Open-Weights
Alibaba International Digital Commerce Group (AIDC-AI) releases the ComfyUI Copilot plugin, simplifying the user experience of ComfyUI through natural language interaction and AI-driven functionality, supporting Chinese interaction, and offering intelligent node recommendations and other features
Open-Weights
Alibaba has announced that its latest video generation model, WanX 2.1, will be open-sourced in the second quarter of 2025, supporting high-definition video generation, dynamic subtitles, and multi-language dubbing, ranking first on the VBench leaderboard with a total score of 84.7%
Open-Weights
MultimodalGoogle introduces the new PaliGemma 2 mix model, supporting various visual tasks including image description, OCR, object detection, and providing 3B, 10B, and 28B scale versions
Open-Weights
Skywork has open-sourced its latest video generation model, SkyReels-V1, which supports text-to-video and image-to-video generation, featuring cinematic lighting effects and natural motion representation, and is now available for commercial use.
Open-Weights
Researchers have proposed a new video relighting method, Light-A-Video, which achieves temporally smooth video relighting effects through Consistent Light Attention (CLA) and Progressive Light Fusion (PLF).
Open-Weights
StepFun has released the open-source text-to-video model Step-Video-T2V, which has 300 billion parameters, supports the generation of high-quality videos up to 204 frames, and provides an online experience platform
Open-Weights
Image to 3DKuaishou officially releases CineMaster text-to-video generation framework, enabling high-quality video content creation through 3D-aware technology
Open-Weights
Alibaba's latest open-source project InspireMusic, a unified audio generation framework based on FunAudioLLM, supporting music creation, song generation and various audio synthesis tasks.
Open-Weights
Text to imageAlibaba Research Institute open sources image generation tool ACE++, supporting character-consistent image generation from single input through context-aware content filling technology, offering online experience and three specialized models.
Open-Weights
ByteDance research team releases OmniHuman-1 human animation framework, capable of generating high-quality human video animations from a single image and motion signals.
Open-Weights