Qwen-AgentWorld open-sourced: an agent that predicts before it acts
Qwen-AgentWorld 开源:让 Agent 学会"先预测,再行动"
Qwen released Qwen-AgentWorld, a native language world model covering seven domains: MCP, Search, Terminal, SWE, Web, OS, and Android. Trained on over 10 million real interaction trajectories through CPT→SFT→RL, it scored 58.71 on AgentWorldBench, edging out GPT-5.4 (58.25) and Claude Opus 4.8. As a decoupled environment simulator, it hit 50.3% F1 on WideSearch via Sim RL, beating real-environment RL at 45.6%. When used as an agent foundation model with LWM warm-up, it transfers to seven benchmarks—three of which never appeared in training. Both model and benchmark are open-sourced.
Why it matters: Qwen dropped an agent model with a clear methodology and benchmark — not concept hype. The 7-environment coverage and 10M+ training traces make it substantive, but it just went open-source and the community hasn't reproduced it yet, so the score stays below 85.