Skip to content
Hacker News front page

BFL launches FLUX 3: a single model that generates video, images, and audio together

Flux 3

FLUX 3 is BFL's new multimodal foundation model that jointly trains on images, video, and audio in a unified architecture. Instead of treating each modality separately, it learns from their mutual constraints—sound matching impact, motion obeying mass—to build a representation of the physical world. Video generation is now in early access: text-to-video, image-to-video, and video-to-video, up to 20 seconds with native audio. In early evals, FLUX 3 beats Grok Imagine Video in 69% of comparisons, Runway Gen-4.5 in 77%, and Luma Ray 3.2 in 93%; it leads Kling v3 Pro 60% of the time. Image generation early access opens in the coming weeks. For action prediction, BFL partnered with mimic robotics to test a finetuned version on real production tasks at Audi. The post does not disclose pricing or a general release date.

Why it matters: BFL released FLUX 3, a multimodal foundation model jointly trained on images, video, and audio. Video generation is in early access with up to 20-second clips and native audio, claiming a 69% win rate in initial comparisons. This is a substantive product update with concrete n...

Read the original ↗Export Markdown