Black Forest Labs (BFL) has launched FLUX 3, a multimodal basis mannequin that learns from photographs, movies and audio inside a single structure. It is usually the primary FLUX mannequin to ship video, audio and motion prediction from one set of weights.
The Black Forest Labs (BFL) analysis workforce argues that no single modality provides an entire description of the world. Photos seize spatial construction at one immediate. Video restores time and exposes bodily dynamics. Audio reveals causal relationships between mechanical occasions and sound. Every is handled as a lossy projection of the identical underlying actuality.
Coaching on all of them without delay means the modalities constrain one another. The sound has to match the affect. The movement has to obey the mass. The analysis workforce calls FLUX 3 its first mannequin constructed solely on that precept.
The strategy beneath: Self-Circulation
FLUX 3 builds on Self-Circulation, BFL’s technique for aligning multimodal era and understanding in a single structure. Self-Circulation combines the circulation matching goal with a self-supervised function reconstruction goal. The reference implementation on GitHub is Apache-2.0 and makes use of SiT-XL/2 with per-token timestep conditioning. It trains with a 25% per-token masks ratio and self-distillation from an EMA trainer at layer 20 to a pupil at layer 8.
That launched checkpoint is an ImageNet 256×256 analysis mannequin, not FLUX 3. BFL states that it ‘considerably scaled up compute and information sources’ on the identical strategy to coach FLUX 3 throughout video, photographs and audio concurrently. Self-Circulation itself was launched in March 2026, so it isn’t new to this launch. What’s new is the size.
What FLUX 3 Video does
FLUX 3 Video generates clips as much as 20 seconds lengthy in a single era, with native audio. The supported modes cowl text-to-video, image-to-video, video-to-video from a reference clip, keyframe-to-video for managed transitions, and generative video-audio continuation from enter video and audio.
BFL additionally lists multilingual dialogue, agentic chaining of clips into multi-shot sequences, and robust typography era with animated designs. The BFL workforce experiences specific energy in human facial expressions and in associating sounds with bodily occasions.
Efficiency
BFL workforce revealed preliminary human choice outcomes. The setup was 10-second text-to-video clips at 720p with audio. FLUX 3 was most well-liked over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%. In opposition to Grok Think about Video the determine is as much as 69%, then Kling v3 Professional at 60%, Pleased Horse v1 at 59% and Pleased Horse 1.1 at 57%. In opposition to Seedance 2.0 and Gemini Omni Flash the result’s 52%, near a coin flip.
Interactive Explorer
enter
Key Takeaways
- FLUX 3 is one circulation matching spine educated collectively on picture, video and audio.
- FLUX 3 Video generates as much as 20 seconds with native audio in a single era.
- Video prediction consumes over 95% of the coaching compute; audio is underneath 0.5% of tokens.
- The identical spine drives FLUX-mimic, a robotic coverage operating underneath 80 ms on one RTX 5090.
- Entry is gated: Video and Motion are in early entry, Picture follows, open weights come final.
Try the FLUX 3 announcement, the FLUX 3 x mimic technical put up and the Self-Circulation paper. All credit score for this analysis goes to the researchers of this venture.
