On July 23, the German AI startup Black Forest Labs officially released the multimodal foundation model Flux3, marking a major breakthrough in generative video technology. The model is built on the Self-Flow architecture, featuring dedicated image, video, audio, and motion codecs, enabling unified understanding and generation of physical and digital environments.

As the first model to support native audio generation, Flux3 can output synchronized audio-video clips up to 20 seconds long, covering functions such as text/image/video-to-video conversion, transition based on keyframes, and multilingual dialogue. In early tests at 720p resolution and 10-second clips, Flux3 demonstrated strong competitiveness, outperforming Luma Ray3.2 (93% win rate) and Runway Gen-4.5 (77% win rate), and maintaining a slight advantage over top models like Seedance2.0 and Gemini Omni Flash.
Additionally, the company developed the video action model Flux-mimic for the robotics field in collaboration with Mimic Robotics, which has already begun production task testing in an Audi factory. BFL adopts a phased release strategy, with Flux3Video now available, and Flux3Image and the open-source weight "Flux3Dev" will be released in the near future. By breaking through the limitations of single-modal information, Flux3 is accelerating the advancement of AI toward a "world model" with perception and action capabilities in the physical world.





