Ant Lingbo open-sources LingBot-Video, a MoE video base model for embodied AI
蚂蚁灵波开源全球首个面向具身智能的MoE视频基模LingBot-Video
Ant Lingbo open-sourced LingBot-Video, the first MoE-based video generation model built for embodied AI. It has 30B total parameters but activates only ~3B during inference, roughly 3× faster than a dense model of similar size. Training used 70,000 hours of robot-related video—dexterous manipulation, navigation, egocentric interaction. On the RBench benchmark for robot manipulation videos it scored 0.620, ahead of Wan2.6 (0.607) and Seedance 1.5 Pro (0.584). Internal tests also place it above NVIDIA Cosmos 3 and Hunyuan Video 1.5 on physical plausibility and motion consistency. The model targets robot action prediction, simulation data generation, and world-model research. Code is public.
Why it matters: Ant Lingbo open-sourced the first MoE video foundation model for embodied AI — 30B total params, ~3B activated during inference, 3x faster than dense models of similar scale, trained on 70k hours of real robot video. HKR all hit, but it's a fresh release with no external repro...