Dwarkesh Patel interviews Noam Brown on 10,000-agent swarms, alignment, and recursive self-improvement
Dwarkesh 对谈 Noam Brown:智能体集群、对齐与递归自我改进
Noam Brown, a core contributor to OpenAI's o1 reasoning models, now works on multi-agent systems. His team just solved a Millennium Prize Problem using 10,000 agents, 130 billion tokens, and 88 hours of compute. Brown frames multi-agent as parallel test-time compute: a single agent hits a latency wall, so you throw more agents at the problem to go faster, at the cost of some efficiency. In the 5.6 release's Ultra Mode, 4 agents cut solve time in half; 16 agents push it further, especially on parallel-friendly tasks like math. The conversation also covers what math progress signals for recursive self-improvement, degrading chain-of-thought quality, and how to verify alignment before kicking off RSI.
Why it matters: Noam Brown is a core contributor to the o1 reasoning line, and this interview comes with a concrete result (Millennium Prize problem) and real numbers, not just speculation. The multi-agent-as-parallel-inference frame and the alignment preconditions for RSI are directly useful...