Gemini gets agentic video understanding that can watch and act on screen
Google DeepMind 为 Gemini 推出 agentic 视频理解功能
Google DeepMind added agentic video understanding to Gemini: it can watch a video of a UI and then perform the same clicks, typing, and scrolling itself. Instead of just describing what it sees, Gemini executes multi-step tasks like filling web forms or completing an order in a mobile app. The feature is now available for testing in the Gemini app and Google AI Studio. The post doesn't disclose latency or success rates—real-world UI agent reliability is still a big open question.
Why it matters: Google DeepMind added agentic video understanding to Gemini — it learns UI workflows from screen recordings and executes multi-step tasks, now available in the Gemini app and AI Studio. Hits all three HKR axes, but the post doesn't disclose latency or success rate, the two num...