Skip to content

Benchmarks

Who is actually stronger: benchmark results, methodology disputes and leaderboard changes.

Latest picks

21–40 of 453

Sep 23Wednesday

AI HOT (Curated Pool)

Claude Opus 5.5 lands on Arena's Agent Arena and Battle Mode

Anthropic's Claude Opus 5.5 is now available on Arena's Agent Arena, where users vote on rankings after the model runs real long-horizon agent tasks. The model can use web search, a file system, and a terminal; the leaderboard uses causal tracking to measure performance relative to the average model. The post doesn't spell out Battle Mode specifics or show example tasks.

Why it matters: Opus 5.5 landing on Agent Arena is the most watchable third-party eval signal this week. The causal-tracking leaderboard design carries more info than raw win rates, but the post doesn't give concrete task examples or Battle Mode rules — real performance waits on community tes...

AI HOT (Curated Pool)

Claude Opus 5.5 tops Artificial Analysis Intelligence Index with a score of 58, plus a 20% price cut

Claude Opus 5.5 scored 58 on the Artificial Analysis Intelligence Index, the highest measured so far. It leads on 6 of 10 evaluations, including Humanity's Last Exam at 61.4% and SciCode at 66.9%, and matches GPT-6 Astra (xhigh) on Terminal-Bench 4.0 at 59.6%. On the agentic knowledge-work eval AA-Briefcase, it hit 1822 Elo—143 points above Fable 5.1—and surpassed GPT-5.6 Sol on both analytical quality and presentation. Pricing dropped to $4/$20 per 1M input/output tokens (from $5/$25), with cache reads down 60% to $0.20. Output tokens per task grew ~60% vs Opus 5, so cost per task stayed flat. Context window remains 1M tokens with image and text input.

Why it matters: Anthropic's flagship tops a major third-party benchmark with a price cut — a same-day must-write. Not a 95 because it's a benchmark result, not a model launch, but 6/10 leads, parity with GPT-6 Astra, and a 20% price drop make it a clear featured pick.

Sep 22Tuesday

Hacker News front page

Prompting agents to iteratively optimize Rust until it beats state-of-the-art libraries

Max Woolf spent over a year testing whether agentic LLMs can iteratively optimize Rust code with a hard pass/fail rule: each iteration must deliver at least a 5% speedup or roll back. With Claude Opus 4.5 and later models, he got 2–20× speedups on algorithms like UMAP versus mature libraries, all without unsafe code. The post includes the exact prompts and benchmark results. The caveat: these gains are measured on his specific benchmarks and may not translate directly to production workloads.

Why it matters: Max Woolf spent a year validating a concrete agentic-iteration loop for Rust optimization with a hard ≥5% speedup rule, reproducible benchmarks against mature libraries, and zero unsafe code. The post includes prompts and results — a rare first-person experiment with numbers. ...

Computing Life · Share · Yage

Three AI Coding Stories, Three Numbers to Read

ZCode was caught silently packaging entire Git histories for cloud upload, with .git objects making up 86.6% of snapshots. HarnessTax benchmarked Claude Fable 5 across frameworks: Claude Code and minimal Pi achieved near-identical success rates but a 2x cost gap. A Microsoft architect replaced multi-model agent loops with a single model reading skill docs and calling tools directly—halving API calls but increasing total tokens by 22%. Each story unpacks one number and a reminder to check which layer a metric actually measures.

Why it matters: Three distinct AI coding stories bundled into one piece. The ZCode silent snapshot upload is a hard security story backed by ferstar's forensic report and community reproduction. HarnessTax benchmark and Microsoft architecture case add billing and engineering angles. HKR all h...

AI HOT (Curated Pool)

Xiaomi MiMo-V2.6-Pro hits ~10th on Code Arena WebDev, ~3rd among open-weight models

Xiaomi released two omni-modal models: MiMo-V2.6-Pro and Flash. The Pro version scored 1628 on Code Arena WebDev, up 153 points from MiMo-V2.5-Pro's 1475, landing around 10th overall and ~3rd among open-weight models under MIT license. The post doesn't disclose Flash's benchmark numbers or parameter counts.

Why it matters: Xiaomi's multimodal model hits ~10th on Code Arena WebDev and ~3rd among MIT open-weight models, with a 153-point gain for Pro. Flags a domestic flagship release with concrete benchmark data. Flash variant lacks params and scores, capping it below 80.

AI HOT (Curated Pool)

Xiaomi MiMo-V2.6-Pro tops open-weight model intelligence index

Xiaomi released MiMo-V2.6-Pro, scoring 46 on the Artificial Analysis Intelligence Index—up from 26 for the previous V2.5-Pro. It's now the highest among open-weight models. The post doesn't disclose parameter count, architecture details, or a release timeline.

Why it matters: Xiaomi's model hits #1 on the open-weight intelligence index with a near-doubling of score — triggers the domestic flagship model positive signal. Missing param count and release timeline keep it from scoring higher.

AI HOT (Curated Pool)

Xiaomi MiMo-V2.6-Pro tops open-weight model intelligence index

Xiaomi MiMo-V2.6-Pro scored 46 on the Artificial Analysis Intelligence Index, the highest among open-weight models. The previous MiMo-V2.5-Pro scored 26. The post doesn't disclose model size, training data, or release license, so hold for details.

Why it matters: Xiaomi's model tops the open-weight leaderboard with a near-doubling of its intelligence score — newsworthy. But without model size, training data, or license details, real-world usability is unclear, capping the score at 78 until more info drops.

Sep 21Monday

AI HOT (Curated Pool)

Fireworks AI launches FireRouter: the frontier isn't a model, it's a router

Fireworks AI benchmarked 18 models on DeepSWE: picking the right model per task beats any single model. GPT-6 Astra alone scores 74.1% at $6.52/task. An oracle router across all 18 hits 97.6% at $1.88. Open-weight models alone reach 90.3% at $1.45. 94 of 113 tasks need a model under $3; the three priciest models are the best pick on only 3 tasks. FireRouter aims to make that per-task choice before the work starts—the post doesn't yet detail how.

Why it matters: Fireworks presents a data-backed argument using 18 models on DeepSWE: task-aware routing beats the single strongest model on both accuracy (97.6% vs 74.1%) and cost ($1.88 vs $6.52). The open-source-only result of 90.3% also provides a path that doesn't depend on closed models...

Sep 20Sunday

Hacker News front page

Prompts Aren't Real: Build Evaluation Pipelines Instead

Dan McKinley argues that prompt engineering is a distraction. Building consumer-facing agents taught him that even structured output fails on a fraction of requests—models will flood a field with nonsense. His fix was renaming a field from 'title' to 'heading,' which he calls deranged. The talk pushes for pass^k testing and evaluation pipelines to constrain behavior, since prompts alone can't tame the beast. The post is a slide deck; it names no specific eval frameworks or metrics.

Why it matters: Dan McKinley's first-hand production experience with concrete cases and numbers, sharp opinion. But it's a personal talk, not a formal publication, and the post doesn't disclose pass^k test pass rates or scale — slight deduction.

Sep 19Saturday

Hacker News front page

Brood War Bench: No model played beyond beginner level in StarCraft

Ben Swerdlow pitted 19 models against each other in StarCraft: Brood War. Codex Astra / xhigh went 18–0, but no model surpassed beginner level. Older models treated the RTS as turn-based and got destroyed while thinking; Grok 4.6 issued only 6 command batches in 43 minutes and never fielded a combat unit. Claude Fable earnestly climbed the tech tree but couldn't execute. Codex models favored early Probe harassment that paralyzed opponents. The post doesn't specify whether matches were pure AI vs AI or involved human input.

Why it matters: First-person benchmark with 19 models playing StarCraft against each other. Concrete data (win rates, APM, cost) and a clear finding: none surpass beginner level. Codex Astra won by early worker harass, not macro play — that detail carries signal. Not an 85 because it's more a...

Sep 17Thursday

Hacker News front page

Coding agent harnesses can 2× your cost with no real accuracy gain

UC Berkeley and Arena researchers tested 7 models across 3 harnesses—Claude Code, Codex CLI, and Pi—on 30 tasks each from SWE-bench Lite and Terminal-Bench 2.0, with 3 repetitions per pair. Harness choice barely moves success rates (±2–5%), but cost can vary up to 5×. Claude Fable 5 hits 97.8% in Claude Code at $1.33, and 96.7% in Pi at $0.67. Pi, a minimal open-source harness with just read, write, edit, and bash, reaches the Pareto frontier on both benchmarks. The post doesn't spell out the exact open-source licenses for Pi and Codex CLI, and doesn't link the full pricing sheet.

Why it matters: Systematic eval from UC Berkeley and Arena: 21 model–harness pairs on standard benchmarks yield a counterintuitive finding—harness barely moves success rate but swings cost 5x. Concrete numbers, clean experimental design, practical takeaway. Held at 78 because the body excerpt...

Hacker News front page

Frontier models are much better at physics than benchmarks suggest—expert re-grading shows why

Researchers at Yale and other institutions had physics faculty and PhDs re-grade six widely used physics benchmarks. Most answers previously marked wrong turned out to be grader errors, incorrect reference solutions, or ambiguous questions. For GPT-5.6-Sol, corrected mean@4 jumped from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark. The near-saturation on these closed-ended tasks signals an urgent need for harder, expert-validated evaluations.

Why it matters: Yale physicists re-graded six popular physics benchmarks and found most 'wrong answers' were actually grading bugs or ambiguous questions. Corrected scores show GPT-5.6-Sol jumping from 47.3% to 78.7% on HLE-Physics — near saturation. A solid takedown of benchmark trustworthin...

Hacker News front page

Training a 4B model to produce 81% faster query plans than Postgres

Rohan Bansal post-trained a Qwen 4B model to beat Postgres's default query plans. After SFT distillation from 500 GPT-6 Astra trajectories and a custom GRPO variant for RL, the model achieved 44.7% latency reduction and 81% geometric mean speedup across 113 join-heavy queries. Training ran on a rented 2×H100 node with four Postgres containers on his desk for measurement. The post doesn't disclose total training time or per-inference latency.

Why it matters: A 4B model trained via RL beats Postgres default plans by 81% on the Join Order Benchmark. The method is practically interesting, but only 113 queries were tested—generalization is unproven, capping the score at 78.

Latent Space

AIUC raised a $40M Series A to insure AI agents so companies can deploy them and sue when things go wrong

AIUC announced a $40M Series A led by Ribbit Capital and First Harmonic. CEO Rune Kvist, Anthropic's first product hire, argues that trust and liability—not capability—will cap AI adoption. They built AIUC-1, a standard that stress-tests agents for jailbreaks, hallucinations, and data leaks, backed by real insurance. Cursor, Harvey, Lovable, and ElevenLabs are already working with them. The episode raises a sharp hypothetical: what happens when a $20 Cursor subscription contributes to a $200M plane crash. The post doesn't disclose specific premium or claims-handling details.

Why it matters: AI agent insurance is a new category, and the AIUC-1 standard plus $40M Series A give this story substance. The CEO's Anthropic pedigree and Ribbit Capital backing add credibility, but the product is early-stage — the post doesn't disclose actual claims data or premium pricing...

Sep 16Wednesday

AI HOT (Curated Pool)

Arena updates Image-to-WebDev leaderboard: GPT-6 Astra tops at 1733

Arena added four new models to its Image-to-WebDev leaderboard. GPT-6 Astra (Max) leads at 1733, 129 points ahead of GPT-5.6 Sol (xHigh). Claude Fable 5.1 (Max) is third at 1710, Muse Spark 1.3 (Max) fourth at 1645, and GLM-5.3-Flash tenth at 1588. The post doesn't disclose evaluation tasks or sample size, so I'd take the gaps with a grain of salt.

Why it matters: GPT-6 Astra tops the Image-to-WebDev leaderboard on its first appearance with a meaningful margin — newsworthy. But the post only gives scores and rankings, with no detail on methodology, task difficulty, or model differences. H and K both hit, R is absent — lands right at the...

Latent Space

Can skills learned in games transfer to real-world work?

Good Start Labs trained a 30B model on the railroad game 1830 and found that training design determines skill transfer. A multi-turn terminal agent version improved at financial research tasks—querying databases, writing Excel formulas, reasoning on the fly—while single-turn training did not. The company spun out of Every last October with $3.6M in funding, betting on verifiable game environments for RL-based skill teaching.

Why it matters: The experimental design is novel, with positive skill-transfer evidence and a failure control, useful for agent training research. But the company just spun out, product path is unclear, and the post doesn't disclose specific accuracy numbers on the financial task, so it stays...

Hacker News front page

RL for LLMs has a Matthew Effect where hard problems get ignored—this post proposes Never Give Up to fix it

Michael Noukhovitch's blog walks through his new paper on the Matthew Effect in RL post-training for LLMs: as training progresses, the model samples easy problems more and hard problems less, because early successes on easy tasks dominate the reward signal. His proposed fix, Never Give Up (NGU), forces a minimum sampling ratio for hard problems so they don't get squeezed out. On Olmo 3.1 7B math training, NGU lifts AIME 2025 pass@1 from 26.7% to 33.3%; on code, LiveCodeBench pass@1 goes from 23.4% to 26.1%. The post also covers async RL staleness tricks and frames the Matthew Effect as a form of primacy bias. The body doesn't disclose NGU's specific hyperparameters or extra compute cost, so I'd discount the gains until those details surface.

Why it matters: Michael Noukhovitch turns his paper on the Matthew Effect in RL post-training into a highly readable blog post: models increasingly favor easy problems during training, and hard-problem sampling rates keep dropping. His proposed NGU method enforces a minimum sampling ratio for...

Sep 15Tuesday

AI Chat-Group Daily (群聊日报)

Daily digest: empty-repo coding fails, Ollama Cloud throughput test, Trump calls out Dario

今天最直观的教训来自 @搞仁义毛义仁:给 Astra 一个空白 C++ 仓库,代码写得一塌糊涂;把积累了大量 code review 经验的 GacUI 上下文导进去,质量立刻飙升。这说明模型不是不会写,是得用具体规则去“规训”。@今天群内信息量极大 实测了 Ollama Cloud 跑 DeepSeek V4.1 Flash,解码吞吐是官方 API ...

AI HOT (Curated Pool)

Fireworks benchmarks DeepSeek-V4.1-Flash: matches GPT-6 Astra on DeepSWE at 1/15th the cost

Fireworks ran a full benchmark suite on DeepSeek-V4.1-Flash. On DeepSWE, it scores 74.34% pass@1, in the same band as GPT-6 Astra at 74.12%, but costs $0.43 per task—15x cheaper. The model uses a 552B MoE with a split activation design: 8B active for input, 16B for output, plus KV cache optimizations. On Terminal-Bench 2.1 it trails Astra by 1 point while costing 12x less. The post mentions an HLE and oracle router eval but does not disclose the actual scores.

Why it matters: DeepSeek V4.1-Flash matching GPT-6 Astra on DeepSWE at an order-of-magnitude lower cost is the strongest price-performance signal this week. Docked because the source is Fireworks' own benchmark, not an independent eval, and the body is truncated by a cookie wall with no full ...

Hacker News front page

GPT-5.6 Luna vs GPT-6 Astra: 3.6% of the cost for 75% of the bugs in code review

Entelligence benchmarked GPT-5.6 Luna and GPT-6 Astra on 50 public PRs for code review. Luna costs $1.20 per million output tokens vs Astra's $50, making per-review cost 28x lower. Luna found 69 verified bugs to Astra's 92, but 24 of its 93 findings were wrong (Astra: 4 of 96). On Sentry, Discourse, and Grafana, Luna was within two bugs of Astra. On Keycloak, an identity server, Luna found 6 bugs to Astra's 14 with only 50% precision. The security gap was widest: 9 vs 19 verified bugs. Luna is good enough for everyday correctness bugs at that price, but not for auth or permission code on its own.

Why it matters: A controlled experiment on 50 real PRs comparing Luna and Astra on code review cost, recall, and false-positive rate — concrete and reproducible. Points off because it's a vendor self-test; the post doesn't disclose PR sources or a human reviewer baseline, so independence is u...