Same Model, 20-Point Gap: DeepSeek's Harness Dependency and the Hidden Ceiling of Synthetic Data
DeepSeek V4 Flash scored 82.7 on its official harness but dropped to a 46.7% pass rate on third-party setups—a 20-point gap from the same model. A joint paper from Stanford, UC Berkeley, and others explains why: training an agent with a single LLM as the user simulator causes the policy to exploit the simulator's narrow response patterns, with policy entropy collapsing from 1.9 to 0.4 nats. DeepSeek lacked a first-party product to collect real interaction data, so its training relied entirely on synthetic environments with limited behavioral diversity. DSH, released on August 13, is their answer—it makes the agent loop a hot-swappable plugin so the training environment can co-evolve with the policy, an engineering implementation of the paper's Co-Training approach.
Why it matters: Hits all three HKR axes: the 20-point gap is intriguing, the evidence chain from official footnotes to third-party repros is solid, and it directly resonates with agent developers. Capped below 85 because this is a benchmarking methodology exposé, not a model or product launch...