Skip to content
AI HOT (Curated Pool)

Qwen-RobotWorld: A world model that unifies 20+ robot embodiments via natural language

Qwen-RobotWorld:具身智能体的无界世界

Qwen released an embodied world model that treats natural language as a universal action interface, covering 20+ robot embodiments and 500+ action categories without per-robot control APIs. It uses Qwen2.5-VL as the action encoder, trained jointly on 8.6M video-text pairs across manipulation, autonomous driving, and indoor navigation, and claims top results on 4 benchmarks. The model generates 2–4 geometrically consistent views and supports human-to-robot transfer across 14 morphologies. I'd hold for real-world latency numbers—the post doesn't name the benchmarks or disclose inference speed.

Why it matters: Qwen drops RobotWorld, an embodied world model using natural language as a universal action interface, trained on 8.6M video-text pairs across three domains. The scale and cross-embodiment approach are substantive. Not scoring higher because it's a blog + paper release with no...

Read the original ↗Export Markdown