This one's worth opening because Anthropic laid out the engineering ledger of a math search campaign. Two Claude Code sessions burned 31M output tokens. Round one produced 650 ideas—all failed—but left a ledger of 106 partial survivors, each tagged with where it died and what new evidence would reopen it. Round two coordinated ~60 subagents to recheck that ledger and stitch a final route via stepping-stone transfers.
I'd discount the headline number a bit. The proven lower bound of Riemann zeta zeros on the critical line went from 41.6% to 67.2%—still nowhere near proving the Riemann Hypothesis, and Anthropic admits this technique won't get there. Baluyot et al. already reached 67.25% under a narrow-box assumption in 2024-2026. Claude's real contribution was dropping that assumption, not inventing 25.6 percentage points from scratch.
The search architecture is the useful bit. Each route got kill criteria, and subagents returned death reports, not half-baked proofs: what broke, which local derivations still hold. One hostile review subagent caught a hidden math error and supplied a fix. Failure information became a do-not-repeat list, and it even enabled cross-route stepping stones—the post documents one case where a seemingly useless local judgment from round one resurfaced 12 hours later to fill a critical gap in a new route.
The verification side is less rosy. Lean code is sorry-free and hooks directly into Mathlib's riemannZeta definition—solid L2 machine checking. But the effective form mentioned in the paper isn't in the headline statements, and the reference script for symbolic checks didn't make it into the frozen repo. L3 intent alignment has backing from two Anthropic mathematicians, but no independent third-party review. L4 peer understanding amounts to Brian Conrey and Dan Goldston reading on short notice—no public referee report a day after release.
Stack this against OpenAI's ten math advances and GPT-5 on Erdős, and the pattern is clear: generation and formalization are accelerating fast, while human understanding and absorption stay flat. The bottleneck is shifting from discovery to comprehension.