German AI startup Black Forest Labs has officially launched the multimodal foundation model Flux3. The model is built on the Self-Flow architecture, equipped with dedicated image, video, audio, and motion codecs, achieving unified understanding and generation of physical and digital environments.

image.png

Outperforming Luma and Runway, competing with top models

As the first model to support native audio generation, Flux3 can output synchronized audio-video clips up to 20 seconds in one go, covering functions such as text/image/video to video, transition based on keyframes, and multilingual dialogue. In early testing at 720p resolution and 10-second clips, Flux3 outperformed Luma Ray3.2 with a 93% win rate, crushed Runway Gen-4.5 with a 77% win rate, and maintained a slight advantage over top models such as Seedance2.0 and Gemini Omni Flash.

Entering the robotics field, phased release strategy

The company collaborated with Mimic Robotics to develop the video action model Flux-mimic for the robotics field. It has already started production task testing in an Audi factory, extending AI capabilities from the pure digital world into physical manufacturing scenarios. BFL adopts a phased release strategy, with Flux3 Video now available, and Flux3 Image and the open-source weight "Flux3 Dev" will be released in the near future.