On July 23, the German AI startup Black Forest Labs officially released the multimodal foundation model Flux3, marking a major breakthrough in generative video technology. The model is built on the Self-Flow architecture, featuring dedicated image, video, audio, and motion codecs, enabling unified understanding and generation of physical and digital environments.

QQ20260724-100544.jpg

As the first model to support native audio generation, Flux3 can output synchronized audio-video clips up to 20 seconds long, covering functions such as text/image/video-to-video conversion, transition based on keyframes, and multilingual dialogue. In early tests at 720p resolution and 10-second clips, Flux3 demonstrated strong competitiveness, outperforming Luma Ray3.2 (93% win rate) and Runway Gen-4.5 (77% win rate), and maintaining a slight advantage over top models like Seedance2.0 and Gemini Omni Flash.

Additionally, the company developed the video action model Flux-mimic for the robotics field in collaboration with Mimic Robotics, which has already begun production task testing in an Audi factory. BFL adopts a phased release strategy, with Flux3Video now available, and Flux3Image and the open-source weight "Flux3Dev" will be released in the near future. By breaking through the limitations of single-modal information, Flux3 is accelerating the advancement of AI toward a "world model" with perception and action capabilities in the physical world.