FLUX 3: Black Forest Labs' Multimodal Video, Audio and Image Model
Black Forest Labs unveils FLUX 3, a unified multimodal foundation model generating 20-second video with native audio, images, and action prediction for robotics.
Black Forest Labs has announced FLUX 3, a new multimodal foundation model trained jointly on images, video, and audio within a unified architecture. The model can generate 20-second video clips with native audio, synthesize and edit images across multiple styles, and extends to action prediction for robotics applications.
FLUX 3 demo: multimodal video generation with native audio (source: BFL blog)
How FLUX 3 Works
FLUX 3 builds on BFL's Self-Flow approach, which aligns multimodal generation and understanding within the same underlying architecture. Instead of treating each media type as a separate task, the model learns from images, video, and audio simultaneously — the sound has to match the impact, the motion has to obey physics, and the future has to follow from the past.
Video Capabilities
FLUX 3 Video can generate clips with native audio lasting up to 20 seconds from text prompts, images, or existing videos. Core capabilities include:
- Text-to-video generation with native audio
- Image-to-video — animating from a starting frame or using images as visual references
- Video-to-video — carrying central elements from a source clip into a new scene
- Generative continuation from input video and audio
- Keyframe-to-video for controlled transitions between defined moments
- Multilingual dialogue generation
- Agentic chaining of individual clips into longer, multi-shot sequences
- Strong typography generation and animated designs
In preliminary evaluations, FLUX 3 Video was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Runway Gen-4.5 in 77%, and Luma Ray 3.2 in 93% of comparisons. The model is particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities.
Image Capabilities
FLUX 3 can generate and edit images across a wide variety of styles, aspect ratios, and resolutions. Early evaluations show significant improvement over earlier FLUX versions in handling complex prompts and multilingual text rendering. FLUX 3 Image early access is expected in the coming weeks.
Action Prediction
FLUX 3's world understanding extends to action prediction. BFL has developed FLUX-mimic with mimic robotics, combining the FLUX 3 video backbone with robot learning for dexterous manipulation. Audi is testing the system for production tasks, with some tasks fine-tuned using as little as 30 minutes of robot data.
Availability
FLUX 3 Video and FLUX 3 Action are available through Early Access. Black Forest Labs plans to release API access, private model weights, and an open-weight FLUX 3 Dev version later this year.
At the time of writing, there is no native ComfyUI support for FLUX 3 yet — it requires the official BFL API or early access program.
Comments
Sign in with GitHub to join the discussion.