Gemini Robotics 2 merges perception and planning, but keeps action control separate
Gemini Robotics 2 的架构逻辑:机器人的感知与规划正在走向合并
Google launched the Gemini Robotics 2 family in August 2026 with three tiers: cloud-based ER 2 for reasoning, a VLA model for action, and On-Device 2 running on hardware. ER 2 handles vision and video but outputs only text and tool calls—no direct motor control. This split isn't a step back; physics forces it. Large models can't run at 200 Hz, and small models lack general knowledge. On-Device 2 reached 53.3% task success on SO101 with just 0.25–1.7 hours of demo data per task, up from 6.7%. But screwing in a lightbulb hit only 36%, and sweeping into a dustpan 32%—soft manipulation remains tough. The article introduces the concept of an 'embodiment tax': data, compute, latency, and safety costs to onboard new hardware. The post doesn't disclose ER 2 inference latency or Gemini Robotics 2 control frequency.
Why it matters: A clear architectural breakdown of Google's newly released Gemini Robotics 2, explaining why embodied models are layered rather than monolithic, with concrete numbers and comparisons. The robotics focus limits broader resonance, placing it at the 78-point featured threshold.