Source: ai-research/bfl-flux-3-announcement-2026-07-29.md (primary — BFL’s own announcement, https://bfl.ai/blog/flux-3), ai-research/bfl-flux-3-mimic-2026-07-29.md (primary companion — FLUX-mimic robotics post, https://bfl.ai/blog/flux-3-mimic), raw/newsletter-theneurondaily-com-064df62297.md (secondary — The Neuron, 2026-07-27).
Product: FLUX 3 | URL: https://bfl.ai/blog/flux-3 | Date: the announcement page is undated; secondary coverage dates the launch to 2026-07-23.
FLUX 3 is Black Forest Labs’ new multimodal foundation model — one architecture jointly trained on images, video, and audio (built on BFL’s Self-Flow approach, scaled up), on the argument that each modality is a lossy projection of the same underlying reality and training them together produces a model of how the world looks, moves, and sounds. The headline creative capability is video with native audio up to 20 seconds in a single generation; the headline strategic claim is that the same backbone drives FLUX-mimic, a video-action robotics model tested and deployed at Audi. FLUX 3 Video is in gated Early Access now, with Image and further capabilities staged over the coming weeks and months.
Key Takeaways
- One backbone, four modalities. FLUX 3 jointly learns images, video, and audio in a unified architecture, with action prediction as the fourth output — both native in FLUX 3 and via finetuned specialist models on the pretrained video backbone. It scales up BFL’s Self-Flow work (unifying generation and representation learning).
- Video with native audio, up to 20 seconds per generation. All video outputs come with native audio. Inputs: pure text prompts, or references such as images and video (per the primary); the capability list adds keyframe-to-video for controlled transitions, video-to-video, and generative video-audio continuation (per secondary quotes of the announcement — see Open Questions on the extraction gap).
- Chained clips reach several minutes. BFL says capabilities can be combined into “sequences lasting several minutes,” with visual references keeping characters consistent across scenes. Multilingual dialogue, synced sound effects, animated typography, and multiple aspect ratios are part of the pitch.
- Company-run evals, not independent — treat accordingly. In BFL’s own preliminary preference tests (10-second 720p text-to-video with audio), FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and Luma Ray 3.2 in 93% — but only 52% against Seedance 2.0 and Gemini Omni Flash, 60% against Kling v3 Pro, and up to 69% against Grok Imagine Video. BFL itself labels these preliminary and expects them to change.
- The same backbone runs robots. FLUX-mimic (built with mimic robotics) decodes actions from the FLUX 3 backbone’s internal world representation; Audi is testing and deploying it on production tasks — kitting, ECU insertion into tight fixtures, assembly, and soft materials (seals, cables) conventional automation couldn’t handle. Full-system robot reaction time: 101 ms.
- Video prediction is the whole game economically. Over 95% of FLUX 3’s training compute goes to video prediction; audio is under 0.5% of tokens in a 720p clip. BFL’s thesis: rendering the world accurately forces learning how it behaves — content creation and physical AI are two applications of one foundation.
- Access is gated and staged. Early Access for FLUX 3 Video now (request form; rollout in phases with safety testing), Image early access “in the following weeks,” further capabilities over weeks and months. The Neuron summarizes this as availability rolling out “in batches.”
What FLUX 3 Is
- BFL frames FLUX 3 as “a checkpoint on our mission to develop real-world visual intelligence: models that perceive, predict, and act across physical and digital environments” — a world-model play, not just a bigger generator.
- The multimodal argument: images capture spatial structure at an instant, video restores time and physical dynamics, audio reveals causal links between mechanical events and sound. Trained jointly, “the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past.”
- Foundation: Self-Flow, BFL’s method for aligning multimodal generation and understanding in one architecture, “significantly scaled up” in compute and data — tens of millions of hours of general video plus hundreds of thousands of hours of human/robot manipulation video (per the mimic post).
- Next-generation goal stated in the announcement: unify perceptual, action, and language prediction in the same model; application frontier named as interactive image/video editing, simulation, computer use, and physical AI.
Video
- Up to 20 seconds per single generation, always with native audio; strong points BFL claims while still in development: human facial expressions, associating sounds with physical events, and multilingual capability.
- Longer sequences: chain clips into multi-shot, multi-minute pieces with character consistency held by visual references.
- Eval setup disclosed: 10-second text-to-video clips at 720p with audio, human preference comparisons. Full company-run table: Luma Ray 3.2 93%, Runway Gen-4.5 77%, Grok Imagine Video up to 69%, Kling v3 Pro 60%, Happy Horse v1 59%, Happy Horse 1.1 57%, Seedance 2.0 52%, Gemini Omni Flash 52%. These are BFL’s own numbers on its own harness, explicitly preliminary — no independent replication exists yet.
Image
- FLUX 3 Image synthesizes and edits across styles, aspect ratios, and resolutions; midtraining evals show “significant improvement over earlier versions of FLUX,” notably complex-prompt handling and high-accuracy multilingual text rendering.
- Early access for Image opens “in the following weeks” — as of the announcement only Video is requestable.
Action and FLUX-mimic
- Two routes to action prediction: native integration in FLUX 3 itself, and finetuning specialist action models from the pretrained video backbone with limited task-specific data.
- FLUX-mimic = FLUX 3 backbone + lightweight action decoder over intermediate features of the video-prediction path (approach pioneered in mimic-video). With a completely frozen backbone it outperforms previous vision-language-action models; finetuned, BFL claims state-of-the-art success rates.
- Adding action prediction to the training curriculum dipped video quality by up to 10%, fully recovered after 3,500 steps — BFL’s evidence that one backbone carries both content creation and robotics without permanent cost.
- Deployment: backbone optimized to under 80 ms on a single RTX 5090; the full mimic robot system reacts in 101 ms. Audi Production Lab (Christoph Schneider, quoted) confirms testing and deployment on soft-body manipulation “impossible with conventional robotics.”
What It Means for This Wiki’s Video Stack
- Native synced audio is the differentiator, not raw visual quality. The current stack pattern is silent generative clips (Seedance, Higgsfield) plus separate voice/avatar layers (HeyGen) and post-hoc sound. FLUX 3 collapses dialogue, SFX, and picture into one generation for clips up to 20s — if the multilingual dialogue holds up outside BFL’s demos, that removes an entire sync step for short-form work.
- No quality-driven urgency to switch off Seedance. BFL’s own eval has FLUX 3 at just 52% preference over Seedance 2.0 — a coin flip on the vendor’s home turf. The reasons to watch it are native audio, keyframe control, and multi-minute character-consistent chaining, which map directly onto the storyboard-driven pipelines in animated-shortfilm-seedance-pipeline and fable5-seedance-4k-film-workflow.
- Not usable in the stack yet. Access is a gated early-access form, no API pricing or aggregator availability announced — nothing to wire into Higgsfield MCP-style tooling today. Watch for it landing on aggregators the way FLUX.2 did.
- HyperFrames is orthogonal but downstream. HyperFrames renders video from HTML rather than generating footage; FLUX 3 clips would enter that workflow as media assets, same as Seedance or Higgsfield output.
- FLUX 3 Image is a coming competitor to the image layer. Multilingual in-image text rendering and complex-prompt handling target exactly the strengths this wiki tracks for Nano Banana and GPT Image 2 — worth a head-to-head when Image early access opens.
Open Questions
- Extraction gap on the primary: the announcement page’s JS-rendered bullet lists (the video core-capability list and the launch-plan list) did not survive Tavily extraction. The capability list (text-to-video, image-to-video, video-to-video, keyframe-to-video, video-audio continuation, multilingual dialogue, aspect ratios, agentic clip chaining, typography) is reconstructed from secondary sources quoting the announcement verbatim; verify against the live page.
- No pricing, parameter count, or license terms disclosed anywhere in the primary. Secondary coverage (MarkTechPost, Hugging Face community post) says the launch plan ends with an open-weight release of the backbone, but that could not be confirmed from the extracted primary text.
- Eval methodology undisclosed beyond “10-second 720p text-to-video with audio”: no sample size, prompt set, or rater details. All preference numbers are company-run; no independent benchmarks exist yet.
- Exact launch date is not printed on the announcement page; secondary coverage consistently dates it 2026-07-23.
- Early-access timeline — “batches” (The Neuron’s wording) with no stated criteria or dates; unknown when/whether FLUX 3 Video reaches the aggregators (fal, Replicate, Krea, Higgsfield) where this stack could actually use it.
Related
- HyperFrames — the wiki’s default video-authoring layer; FLUX 3 clips would feed it as media assets
- Higgsfield Overview — current generative-video aggregator in the stack
- Seedance Claude Skill — Free Filmmaking — Seedance 2.0 is the closest current-stack rival (52% in BFL’s own evals)
- Animated Shortfilm Seedance Pipeline — the storyboard/keyframe workflow FLUX 3’s keyframe-to-video and chaining target
- Fable 5 + Seedance 4K Film Workflow — multi-scene character-consistency workflow FLUX 3 claims to natively solve
- HeyGen + Seedance 2 Avatar Shots — the silent-clip-plus-voice-layer pattern FLUX 3’s native dialogue would collapse
- Nano Banana — image-generation rival for the coming FLUX 3 Image (multilingual in-image text)
- ChatGPT Image (GPT Image 2) — the other in-image-text leader FLUX 3 Image will be measured against