Skip to content
AI HOT (Curated Pool)

The next big breakthrough will be AIs learning on the job

下一个重大突破:AI在工作中学习

Dwarkesh Patel argues the current lab bet—training AIs on millions of verifiable tasks to reach AGI—misses a key constraint: the domain must also be grindable, meaning you can run many parallel rollouts in a deterministic, replayable simulator. He uses computer use as an example. Ordering an item on Etsy is verifiable, but you can't have a thousand agents hit the same Amazon checkout flow without getting banned. That's why computer use lags behind coding and math. Unless we build high-fidelity, farmable simulators, the sample-efficiency black hole during training will block progress on many real-world skills. The post suggests the real fix is AIs learning on the job via in-context learning across very long horizons, rather than relying solely on one-time weight updates. No specific product names or timelines are disclosed.

Why it matters: Dwarkesh Patel's essay splits the current RL paradigm into 'verifiable' and 'replayable' conditions, arguing that computer-use and coding tasks are stuck on the latter. The Etsy vs Amazon example makes the bottleneck concrete. Not an 85 because it's an individual analysis, not...

Read the original ↗Export Markdown