Black Forest Labs launches FLUX 3: one multimodal model for image, video, audio, and robot action prediction
[AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model
FLUX 3 is Black Forest Labs' new multimodal model that handles image, video, audio, and action prediction under one unified architecture. The official blog claims it beats Seedance 2.0, Gemini Omni, and Grok Imagine in preference tests for video generation, with native audio on all outputs. Capabilities include text-to-video, image-to-video, video-to-video, keyframe transitions, multilingual dialogue, and agentic chaining of clips into multi-shot sequences. The team also announced FLUX3-mimic, which uses the same backbone to drive dexterous robot manipulation in real factory settings, developed with mimic. The post doesn't disclose parameter count, training data size, or a release timeline beyond mentioning an open-weights Dev version is coming.
Why it matters: BFL launches FLUX 3, a unified architecture spanning video, audio, and robot action prediction, claiming wins over Seedance 2.0 and others in preference tests. This is a substantial expansion from image gen into multimodal and embodied AI. Score stays below 85 because only blo...