presentofai

Black Forest Labs releases FLUX 3 multimodal foundation model

TL;DR

Black Forest Labs releases FLUX 3, a single open-weights model that jointly generates images, video, and audio, and drives a real robot policy in under 80 ms, marking the first time one flow-matching backbone spans all three modalities plus physical action.

What happened

  • Black Forest Labs (BFL) released FLUX 3 on September 23, 2026, its first model to unify image, video, audio, and action prediction in one set of weights.
  • The architecture scales Self-Flow, BFL's flow-matching plus self-supervised reconstruction method introduced in March 2026, with significantly more compute and data across all modalities simultaneously.
  • FLUX 3 Video generates clips up to 20 seconds with native audio in a single pass, supporting text-to-video, image-to-video, video-to-video, keyframe-to-video, and generative audio-video continuation.
  • Video prediction consumes over 95% of training compute; audio accounts for under 0.5% of tokens, yet the model associates sounds with physical events as a core capability.
  • The same backbone powers FLUX-mimic, a robot policy that runs in under 80 ms on a single RTX 5090, connecting generative modeling directly to physical control.

Why it matters

  • One backbone for perception and action collapses the traditional stack of separate vision, audio, and control models, lowering integration cost for robotics and embodied AI developers.
  • Human preference results are aggressive: FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%, positioning it at or near the top of the video generation market.
  • The only near-parity result is against Seedance 2.0 and Gemini Omni Flash at 52%, signaling that Google and ByteDance remain credible competition.
  • The open-weights release path (Video and Action in early access, Image to follow, open weights last) gives BFL commercial leverage while preserving the open-source narrative that built the FLUX brand.
  • Training modalities as mutual constraints (sound must match impact, motion must obey mass) is a design bet that physical plausibility scales better than modality-specific fine-tuning, with direct implications for sim-to-real robotics pipelines.

What to watch next

  • Whether open weights ship on schedule and whether the robot policy weights are included, which would determine how quickly the robotics research community can build on FLUX-mimic.
  • Benchmark responses from Seedance 2.0 and Gemini Omni Flash, the two models that matched FLUX 3 in preference tests, and whether they close the gap on audio-visual coherence.
  • Adoption of Self-Flow as a training paradigm by third parties: the Apache-2.0 reference implementation is already public, and scaled uptake would validate BFL's architectural thesis beyond its own models.

Originally published on Present of AI, a daily source-linked AI news timeline. Read the full timeline or browse the open dataset.