NUS team introduces OmniConsistency project, achieving GPT-4o level image stylization consistency with only 2,600 image pairs and 500 hours of GPU training, now supports ComfyUI nodes
Open-Weights
ComfyUI ecosystem updates — open-source model releases, custom nodes, workflows, and tools for image, video, and audio generation.
NUS team introduces OmniConsistency project, achieving GPT-4o level image stylization consistency with only 2,600 image pairs and 500 hours of GPU training, now supports ComfyUI nodes
Open-Weights
Black Forest Labs launches FLUX.1 Kontext series models, the first to achieve context-aware editing based on text and image inputs, supporting character consistency preservation and localized editing features
Open-Weights
Text to imageByteDance and Michigan State University jointly propose ID-Patch method, achieving multi-identity personalized image generation with improved identity preservation and generation efficiency, suitable for group photos, advertisements and other multi-character scenarios.
Open-Weights
Character animationTencent releases HunyuanVideo-Avatar, enabling high-fidelity, emotion-controllable digital human videos from images and audio, suitable for short videos, e-commerce ads, and more.
Open-Weights
Pixel-Reasoner, based on Qwen2, offers global and local pixel-level visual understanding and reasoning, supports detailed zoom-in analysis, and advances visual language models.
MIT
IndexTTS releases version 1.5 with significantly improved model stability and English performance, supporting pinyin pronunciation correction and punctuation-controlled pauses, outperforming mainstream TTS systems in multiple tests
Open-Weights
Image to 3DStepFun releases Step1X-3D open source framework that generates high-quality 3D geometry and textures from single images, including 2M high-quality dataset, complete training code and model weights
Open-Weights
MultimodalByteDance releases BAGEL, an open-source multimodal foundation model with 7B active parameters, supporting understanding and generation across text, image, and video, and achieving strong results on public benchmarks.
Open-Weights
Text to imageAfter recent ComfyUI updates, some custom nodes may cause frontend issues. This article summarizes affected plugins and solutions, recommending timely updates or uninstallation of relevant plugins.
Open-Weights
The Step1X-3D team has released a high-fidelity 3D asset generation solution, making models, datasets, and code open source, supporting various 3D asset and texture generation tasks.
Open-Weights
MultimodalTencent introduces HunyuanCustom, a multimodal video customization framework that supports text, image, audio, and video conditions, enabling highly consistent video generation across multiple scenarios
Open-Weights
Image editingInsert Anything is an open-source unified framework that can seamlessly insert elements such as people, objects, and clothing from reference images into target scenes, supporting various application scenarios
Open-Weights
Text to videoFlexiAct, developed jointly by Tsinghua University and Tencent ARC Lab, can transfer actions from a reference video to any target image while maintaining identity consistency
Open-Weights
MultimodalStep1X-Edit is a powerful open-source image editing framework that supports image editing via natural language instructions, offering capabilities similar to GPT-4o and Gemini2 Flash。
Open-Weights
InpaintingFlex.2-preview is an open-source text-to-image diffusion model with universal control and built-in inpainting capabilities, providing creators with more powerful image generation tools
Open-Weights
Image to videoSand AI has open-sourced MAGI-1, an autoregressive model that generates videos chunk by chunk, offering 24B and 4.5B parameter versions and supporting multiple video generation modes
Open-Weights
Image to videoSkyworkAI releases new open-source video generation model SkyReels-V2, supporting infinite-length videos with both text-to-video and image-to-video capabilities
Open-Weights
Text to speechNari Labs introduces open-source text-to-speech model Dia 1.6B, capable of generating multi-character dialogues directly from text with emotion and non-verbal communication
Open-Weights
MultimodalKunlun Wanwei's SkyReels team releases and open-sources SkyReels-V2, the world's first infinite-length film generative model using Diffusion Forcing framework, capable of producing high-quality, long-duration video content
Open-Weights
Text to imageShakker Labs launches the new FLUX.1-dev-ControlNet-Union-Pro-2.0 model with optimized control effects, support for multiple control modes, and smaller model size
Open-Weights