Black Forest Lab FLUX3 Multimodal Model Makes Its Debut: Generate 20-Second Audio and Video in One Go, Significantly Outperforming Grok and Seedance
Black Forest Labs released FLUX3, a unified multimodal model for joint image, video, and audio learning. Built on the Self-Flow self-supervised flow matching framework, it extends FLUX series for multimodal generation and understanding. Supports text-to-video, image-to-video, generating up to 20s videos with native synchronized audio, outperforming predecessors.....