Skip to content
AI HOT (Curated Pool)

Black Forest Labs launches FLUX 3, a multimodal model generating 20s video with native audio in one pass

Black Forest Labs 发布 FLUX 3 多模态模型,支持单次生成 20 秒视频与原生音频

Black Forest Labs released FLUX 3 in Early Access, using a unified architecture to jointly learn images, video, and audio. It generates up to 20 seconds of video with native audio in one pass, covering text-to-video, image-to-video, video-to-video, keyframe-to-video, and multi-shot sequences. In human evaluations on 10s 720p clips with sound, FLUX 3 beats Grok Imagine Video 69% of the time, and Seedance 2.0 and Gemini Omni Flash 52% each. The lab is also working with Mimic Robotics to use FLUX 3 as a robot behavior prediction model. The post doesn't disclose parameter count, inference latency, or the scope of Early Access.

Why it matters: Black Forest Labs drops FLUX 3 — a unified multimodal model that outputs 20-second video with native audio in one shot, not the old image-model-plus-audio-plugin approach. The 69% win rate vs Grok Imagine Video gives it teeth. Not an 85 because there's no public access or thir...

Read the original ↗Export Markdown