Key takeaways
- German AI company Black Forest Labs (BFL) has released Flux 3, a multimodal foundation model that learns from images, video, and audio…
- Images show spatial structure, video captures how it changes over time, and audio can reveal links between mechanical events and the sounds…
- It supports text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven links…
What happened
German AI company Black Forest Labs (BFL) has released Flux 3, a multimodal foundation model that learns from images, video, and audio together. In BFL's early tests, it beat several rivals in video generation. " The company is part of a broader push to build so-called world models. No single modality captures reality in full, BFL argues.
The company plans to release Flux 3 Image in early access within the next few weeks. BFL says the model can also predict actions based on its understanding of the world. The company worked with Mimic Robotics to develop Flux-mimic, a video-action model now being tested on production tasks at Audi.
Flux 3 is based on Self-Flow, BFL's approach for teaching one model to generate and understand content at the same time. A multimodal transformer uses dedicated components to convert images, video, and audio into a shared internal representation and then turn it back into outputs. A component for actions provides the foundation for robotics applications.
BFL says this unified learning process delivers better results than the previously standard flow-matching method, both in generation quality and in the model's grasp of the physical world. BFL is rolling out all capabilities in stages, with early-access phases for feedback and safety testing. Flux 3 Video is already available, and Flux 3 Image is set to follow in the coming weeks.
Why it matters
Images show spatial structure, video captures how it changes over time, and audio can reveal links between mechanical events and the sounds they produce. Training on all three together lets them fill gaps for one another, giving the model more information than training on each modality separately. Flux 3 can now generate videos with native audio for the first time, with clips up to 20 seconds long.
It supports text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven links between clips for longer multi-shot sequences. BFL says the model is especially good at human facial expressions and matching sounds to physical events. 5 in 77 percent, and over Grok Imagine Video in 69 percent. The margins narrow against stronger competitors. 0 and Gemini Omni Flash at 52 percent each.
BFL says the results are preliminary, and no independent tests are available yet. Matching leading systems such as Seedance, which has already reached Hollywood, and Gemini Omni Flash would put Flux 3 among the top video models. BFL also expects Flux 3 to improve image generation, especially for complex prompts and accurate text rendering in multiple languages.
What to watch
Action prediction will initially be offered through select partners. " Longer term, the company is working on next-generation models that aim to combine perception, action, and language prediction in a single model.




