FLUX 3 now controls robots: one model generates video and predicts actions
Flux 3 X Mimic: The Next Generation of Video-Action Models
Black Forest Labs put an early FLUX 3 onto mimic's robots, tested on Audi production lines. FLUX 3 jointly generates images, video, and audio; video prediction alone accounts for over 95% of training compute. To avoid visual artifacts, the model had to learn contact, motion, and causality. Adding action prediction as a low-dimensional modality caused a temporary 10% drop in video quality, fully recovered after 3,500 steps. The same backbone now handles both video generation and robot actions. FLUX-mimic attaches a lightweight action decoder to FLUX 3's video prediction features, reading actions from the learned world representation. The post doesn't disclose success rates, latency, or deployment scale—treat this as an architecture proof point, not a production-ready system.
Why it matters: Black Forest Labs put an early FLUX 3 on mimic robots tested at Audi — one model doing video generation and action prediction, with a 10% quality dip that recovered in 3,500 steps. Cross-modal + physical deployment + concrete numbers hit all three HKR axes. Not scoring higher ...