FLUX 3: Multimodal AI Model for Video, Image, Audio and Action

ComfyUI Wiki

FLUX 3 is Black Forest Labs' first multimodal foundation model, jointly trained on video, image, audio, and action prediction via Self-Flow.

F

FLUX 3

MultimodalVideo GenerationAudio GenerationText-to-ImageAction Prediction

Black Forest Labs' first multimodal foundation model — jointly learns from images, video, audio, and action prediction in a unified architecture. Built on the Self-Flow approach for efficient multimodal alignment. Capable of text-to-video with native audio, image-to-video, video-to-video, generative audio continuation, and high-quality text-to-image generation.

DeveloperBlack Forest Labs
Announced2026-07-23
ArchitectureMultimodal Flow Matching (Self-Flow)
StatusEarly Access (Video), Coming Soon (Image, Dev)
CapabilitiesVideo + Audio, Image Generation, Action Prediction
Key ResearchSelf-Flow

What is FLUX 3?

FLUX 3 is Black Forest Labs' next-generation multimodal model that jointly learns from images, videos, and audio within a single unified architecture. Unlike previous FLUX models that focused solely on image generation, FLUX 3 builds a shared representation of the physical world: how objects hold together, how things move, and how events sound.

No single modality provides a complete description of reality. Images capture spatial structures at a point in time, videos restore temporal dynamics, audio reveals causal acoustic relationships, and language links perception to instructions. FLUX 3 learns from all of them simultaneously, using mutual constraints across modalities to produce more coherent and physically grounded outputs.

Built on the Self-Flow approach for efficiently aligning multimodal generation and understanding, FLUX 3 significantly scales up compute and data resources to train across video, images, and audio at the same time.

What Can FLUX 3 Do?

Video + Audio (FLUX 3 Video — Early Access Now)

FLUX 3 Video generates highly diverse video clips with native audio up to 20 seconds in length at 720p resolution in a single generation:

  • Text-to-video generation with synchronized audio
  • Image-to-video (animation from a starting frame or visual references)
  • Video-to-video carrying central elements (e.g. the same character) into new scenes
  • Generative video-audio continuation from input video and audio
  • Keyframe-to-video for controlled transitions between defined moments
  • Multilingual dialogue generation
  • Agentic chaining of individual clips into longer multi-shot sequences
  • Broad range of visual styles and aspect ratios beyond conventional cinematic output
  • Strong human facial expressions, physically grounded sound-event association, and multilingual capabilities

Image (FLUX 3 Image — Early Access Coming Soon)

FLUX 3 Image will offer text-to-image synthesis and image editing across a wide variety of styles, aspect ratios, and resolutions. Preliminary evaluations show significant improvements over earlier FLUX versions in complex prompt handling and text generation accuracy, including high-accuracy text rendering in multiple languages.

Action Prediction (FLUX 3 Action / FLUX-mimic — Partner Access)

FLUX 3's world understanding extends to action prediction through native integration (building on Self-Flow) and via fine-tuning the pretrained video backbone as a dynamics-aware foundation. In partnership with mimic robotics, the FLUX-mimic video-action model is being tested on real production tasks at Audi for dexterous manipulation and production deployment.

Launch Plan

CapabilityAccessStatus
FLUX 3 VideoAPI + Private Weight AccessEarly Access Now
FLUX 3 ImageAPI + Private Weight AccessEarly Access Coming Weeks
FLUX 3 ActionResearch & Commercial Partners (mimic)Partners Now
FLUX 3 DevOpen-Weight Multimodal BackboneComing Later

Early Evaluations

In preliminary human preference evaluations for 10-second text-to-video clips at 720p with audio:

ComparisonPreference
vs Runway Gen-4.577% preferred FLUX 3
vs Grok Imagine Video69% preferred FLUX 3
vs Kling v3 Pro60% preferred FLUX 3
vs Happy Horse v159% preferred FLUX 3
vs Happy Horse 1.157% preferred FLUX 3
vs Seedance 2.052% preferred FLUX 3
vs Luma Ray 3.293% preferred FLUX 3

|> These results are preliminary and the team expects further improvements during the Early Access phase. Source: Black Forest Labs — FLUX 3 Blog Post.

Guides and workflows related to this model series.

No articles found.

Comments

Sign in with GitHub to join the discussion.

Loading comments…