Skip to content
Trending storyPast story

Noam Brown on 10,000-agent swarms solving math problems and recursive self-improvement

1 report1 sourceupdated 9 days ago

What happened

From the coverage

Noam Brown 是 OpenAI o1 推理模型的核心贡献者,现在在做多智能体系统。他的团队上周用一万个 AI 智能体、花了 1300 亿个 token、跑了 88 小时,解出了一个千禧年大奖难题。Brown 把多智能体看成一种并行的“推理时计算”:单个模型推理久了延迟太高,就多扔几个智能体一起干,用效率换速度。在 5.6 版本的 Ultra 模...

From AI HOT 精选

Coverage

Follow the reports to see the story from different sides.

Sep 17
  1. AI HOT (Curated Pool)Pick
    Dwarkesh Patel interviews Noam Brown on 10,000-agent swarms, alignment, and recursive self-improvement

    Noam Brown, a core contributor to OpenAI's o1 reasoning models, now works on multi-agent systems. His team just solved a Millennium Prize Problem using 10,000 agents, 130 billion tokens, and 88 hours of compute. Brown frames multi-agent as parallel test-time compute: a single agent hits a latency wall, so you throw more agents at the problem to go faster, at the cost of some efficiency. In the 5.6 release's Ultra Mode, 4 agents cut solve time in half; 16 agents push it further, especially on parallel-friendly tasks like math. The conversation also covers what math progress signals for recursive self-improvement, degrading chain-of-thought quality, and how to verify alignment before kicking off RSI.