FLUX 3: Multimodal AI Model for Video, Image, Audio and Action
FLUX 3 is Black Forest Labs' first multimodal foundation model, jointly trained on video, image, audio, and action prediction via Self-Flow.
FLUX 3
MultimodalVideo GenerationAudio GenerationText-to-ImageAction PredictionBlack Forest Labs' first multimodal foundation model — jointly learns from images, video, audio, and action prediction in a unified architecture. Built on the Self-Flow approach for efficient multimodal alignment. Capable of text-to-video with native audio, image-to-video, video-to-video, generative audio continuation, and high-quality text-to-image generation.
| Developer | Black Forest Labs |
| Announced | 2026-07-23 |
| Architecture | Multimodal Flow Matching (Self-Flow) |
| Status | Early Access (Video), Coming Soon (Image, Dev) |
| Capabilities | Video + Audio, Image Generation, Action Prediction |
| Key Research | Self-Flow |
What is FLUX 3?
FLUX 3 is Black Forest Labs' next-generation multimodal model that jointly learns from images, videos, and audio within a single unified architecture. Unlike previous FLUX models that focused solely on image generation, FLUX 3 builds a shared representation of the physical world: how objects hold together, how things move, and how events sound.
No single modality provides a complete description of reality. Images capture spatial structures at a point in time, videos restore temporal dynamics, audio reveals causal acoustic relationships, and language links perception to instructions. FLUX 3 learns from all of them simultaneously, using mutual constraints across modalities to produce more coherent and physically grounded outputs.
Built on the Self-Flow approach for efficiently aligning multimodal generation and understanding, FLUX 3 significantly scales up compute and data resources to train across video, images, and audio at the same time.
What Can FLUX 3 Do?
Video + Audio (FLUX 3 Video — Early Access Now)
FLUX 3 Video generates highly diverse video clips with native audio up to 20 seconds in length at 720p resolution in a single generation:
- Text-to-video generation with synchronized audio
- Image-to-video (animation from a starting frame or visual references)
- Video-to-video carrying central elements (e.g. the same character) into new scenes
- Generative video-audio continuation from input video and audio
- Keyframe-to-video for controlled transitions between defined moments
- Multilingual dialogue generation
- Agentic chaining of individual clips into longer multi-shot sequences
- Broad range of visual styles and aspect ratios beyond conventional cinematic output
- Strong human facial expressions, physically grounded sound-event association, and multilingual capabilities
Image (FLUX 3 Image — Early Access Coming Soon)
FLUX 3 Image will offer text-to-image synthesis and image editing across a wide variety of styles, aspect ratios, and resolutions. Preliminary evaluations show significant improvements over earlier FLUX versions in complex prompt handling and text generation accuracy, including high-accuracy text rendering in multiple languages.
Action Prediction (FLUX 3 Action / FLUX-mimic — Partner Access)
FLUX 3's world understanding extends to action prediction through native integration (building on Self-Flow) and via fine-tuning the pretrained video backbone as a dynamics-aware foundation. In partnership with mimic robotics, the FLUX-mimic video-action model is being tested on real production tasks at Audi for dexterous manipulation and production deployment.
Launch Plan
| Capability | Access | Status |
|---|---|---|
| FLUX 3 Video | API + Private Weight Access | Early Access Now |
| FLUX 3 Image | API + Private Weight Access | Early Access Coming Weeks |
| FLUX 3 Action | Research & Commercial Partners (mimic) | Partners Now |
| FLUX 3 Dev | Open-Weight Multimodal Backbone | Coming Later |
Early Evaluations
In preliminary human preference evaluations for 10-second text-to-video clips at 720p with audio:
| Comparison | Preference |
|---|---|
| vs Runway Gen-4.5 | 77% preferred FLUX 3 |
| vs Grok Imagine Video | 69% preferred FLUX 3 |
| vs Kling v3 Pro | 60% preferred FLUX 3 |
| vs Happy Horse v1 | 59% preferred FLUX 3 |
| vs Happy Horse 1.1 | 57% preferred FLUX 3 |
| vs Seedance 2.0 | 52% preferred FLUX 3 |
| vs Luma Ray 3.2 | 93% preferred FLUX 3 |
|> These results are preliminary and the team expects further improvements during the Early Access phase. Source: Black Forest Labs — FLUX 3 Blog Post.
Links
Guides and workflows related to this model series.
Comments
Sign in with GitHub to join the discussion.