Qwen-AgentWorld: Language World Models That Simulate Environments for General Agents
Qwen-AgentWorld: Language World Models for General Agents
Qwen team released Qwen-AgentWorld, a language world model that predicts environment dynamics for general agents. It covers 7 domains and uses long chain-of-thought reasoning to forecast next states. Two model sizes are available: 35B-A3B and 397B-A17B, trained on over 10 million real-world interaction trajectories via a three-stage pipeline—CPT injects world modeling from state transitions, SFT activates next-state prediction, and RL sharpens fidelity with hybrid rubric-and-rule rewards. The team also built AgentWorldBench from real interactions of 5 frontier models across 9 benchmarks. Qwen-AgentWorld significantly outperforms existing frontier models. It works in two modes: as a decoupled simulator enabling scalable RL across thousands of environments, surpassing real-environment-only training; and as a unified agent foundation model where world-model training serves as effective warm-up, boosting performance on 7 agentic benchmarks. Code is open-sourced.
Why it matters: Qwen team trains an LLM-based world simulator on 10M interaction traces across 7 environments. Novel approach with concrete scale and benchmarks, relevant for agent builders. Score capped below 85 because it's a paper, not a product release — real-world agent task gains aren't...