Skip to content
AI HOT (Curated Pool)

Meituan LongCat open-sources VitaBench 2.0, a long-horizon dynamic agent benchmark

美团 LongCat 开源 VitaBench 2.0:长期动态智能体基准新标杆

Meituan's LongCat team open-sourced VitaBench 2.0, a benchmark for testing how well agents model users over long, dynamic real-life scenarios. It includes 56 simulated users, 819 complex tasks, over 2,000 shifting preferences, and 66 executable tools—averaging 2,093 interaction events per user across roughly 1,580 days. Even the top model, Claude-Opus-4.6, barely scored above 0.5 in open-book mode. Thinking mode didn't consistently help on personalization tasks, and all models saw a sharp drop on tasks requiring proactive questions. The benchmark and tools are open-sourced.

Why it matters: Meituan LongCat open-sourced a large-scale long-horizon agent benchmark with concrete numbers on data volume, task design, and results — not a vague leaderboard. Score isn't higher because there's only one WeChat post so far, no cross-source confirmation yet, and the benchmark...

Read the original ↗Export Markdown