Skip to content
Hacker News front page

Prime Intellect ran 153 autonomous NanoGPT optimizer runs across 18 frontier models

NanoGPT Speedrun Frontier

Prime Intellect ran a speedrun where 18 frontier models, each paired with a coding agent harness, autonomously optimized a NanoGPT training script to minimize validation loss. Fable 5 with Claude Code reached a loss of 2,726, closing 81.7% of the gap from the human baseline to the optimal record. Opus 5 and Kimi K3 followed at 53.6% and 52.2%. DeepSeek V4 Pro closed only 12.3%, ranking 13th. The experiment consumed hundreds of millions of tokens, with the longest run lasting 8.7 days. Caveat: this benchmark tests a narrow skill—autonomous optimization of known code—not general capability, but it reveals how different models handle stability and exploration over long-horizon tasks. The post does not disclose the specific optimization strategies each model used or why certain runs failed.

Why it matters: Prime Intellect's speedrun pits 18 models against each other on autonomous code optimization, with Fable 5 + Claude Code posting a dominant lead. The concrete numbers, leaderboard, and methodology make this genuinely useful for agent builders and benchmark watchers. Not scorin...

Read the original ↗Export Markdown