Skip to content

Benchmarks

Who is actually stronger: benchmark results, methodology disputes and leaderboard changes.

Latest picks

1–20 of 453

Today · Sep 30Wednesday

AI HOT picks · Models

GPT-6.1 Sol replaces GPT-6 Sol 7 days after launch, 1 point behind GPT-6 Astra on intelligence index

GPT-6.1 Sol replaced GPT-6 Sol 7 days after launch. Its intelligence index is 1 point below GPT-6 Astra, and pricing stays at $2 per million input tokens and $10 per million output tokens. The cached-read discount rises from 90% to 95%.

Why it matters: Side-by-side intelligence index, cost and token-efficiency figures let readers judge the trade-off GPT-6.1 Sol makes between price and capability.

Yesterday · Sep 29Tuesday

Hacker News front page

Pac-Bench: One-shot Pac-Man benchmark, Claude Opus 5.5 scores 99/100

Jon Clegg built a Pac-Man benchmark: one prompt, one HTML page, scored automatically by Opus 5.5. Claude Opus 5.5 hit 99/100 via Claude Code at $1.99, generating a 10.8 KB page in 9 minutes with near-arcade audio. Claude Fable 5.1 scored 96 but cost $5.87. Grok 4.7 and GPT-5.6-sol scored 94 and 90; the latter cost just $0.72 in under 5 minutes. Scoring covers controls, ghost behavior, stuck detection, maze layout, and sound. The post doesn't explain why some models ran Phase 2 or how much the harness affects scores. Worth noting: this measures model-plus-toolchain combos, not bare model capability.

Why it matters: A 30-model Pac-Man benchmark with Claude Opus 5.5 hitting 99/100 via Claude Code at $1.99 is solid signal. Capped at 78 because it's an individual project, not an official release, so authority is limited despite strong HKR.

AI HOT (Curated Pool)

GPT-6 Luna (Max) ranks #23 on Agent Arena at $0.05 per task

Arena evaluated GPT-6 Luna (Max) on 8K real agent conversations. It ranked #23, up 6 spots from GPT-5.6 Luna (xHigh), with a net gain of +1.6%. Cost is $0.05 per task. The post doesn't disclose latency or task-type breakdown.

Why it matters: First Agent Arena run for GPT-6 Luna with a real ranking and cost figure. But the post doesn't disclose latency or task breakdown, so real-world usability is unclear — score sits at the featured threshold.

AI HOT (Curated Pool)

Claude Sonnet 5.5 enters Arena's Agent Arena and Battle Mode

Anthropic's Claude Sonnet 5.5 is now available for voting in Arena's Agent Arena. The leaderboard evaluates models on millions of real-world long-horizon agent tasks where models can use web search, filesystem, and terminal tools. Rankings use a causal tracking method to measure how much a model outperforms the average.

Why it matters: Claude Sonnet 5.5 hitting Arena's Agent leaderboard is a direct user-facing eval signal, hitting all three HKR axes. Score capped at 74 because the post only describes the methodology — no specific win rates or rankings disclosed. Adjust upward once concrete numbers drop.

Hacker News front page

Anthropic launches Claude Sonnet 5.5: 30%+ faster, up to 30% cheaper than Sonnet 5

Claude Sonnet 5.5 is the second model in the 5.5 family, aimed at everyday coding, bug fixes, and polished docs. It scores 70.6% on Terminal-Bench 4.0 vs. Sonnet 5's 10.3%. Pricing stays at $2/$10 per million input/output tokens, but it uses fewer tokens per task, cutting per-task cost by up to 30%. Speed is up 30%+. For the first time, a Sonnet model ships with cyber safeguards because its cybersecurity capabilities now match Opus 5. Haiku 5.5 is coming in a few weeks.

Why it matters: Anthropic officially released Claude Sonnet 5.5, the second model in the 5.5 family. Terminal-Bench jumped from 10.3% to 70.6%, 30% faster with 30% lower per-task cost at unchanged pricing. A same-day must-write model update. Not 95 because it's a complement to Opus 5.5, not a...

AI HOT (Curated Pool)

Claude Opus 5.5 (High) hits #2 on Agent Arena and reshapes the Pareto frontier

Anthropic's Claude Opus 5.5 (High) landed at #2 on Agent Arena with a +12.15% net improvement, behind only Fable 5.1 (Max). Median cost is $1.31 per task—40% cheaper than Opus 5 (High) and 56% cheaper than Opus 5 (Max). It ranked #1 on Steerability at +14.50%. The post doesn't disclose a release date or other model comparisons.

Why it matters: Anthropic model hitting #2 on Agent Arena with a significant price drop is a same-day must-write product signal. The +12.15% net improvement and $1.31 median cost provide hard data, and steerability gains are a bonus. Not scoring higher because this is still a benchmark — real...

AI HOT (Curated Pool)

Claude Opus 5.5 (High) hits #2 on Agent Arena, costs 56% less than Opus 5 (Max)

Anthropic's Claude Opus 5.5 (High) reached #2 on Agent Arena with a +12.15% net improvement, behind only Fable 5.1 (Max). It also costs 56% less than Opus 5 (Max). The post doesn't disclose exact pricing or latency—I'd discount the cost claim until we see real usage numbers.

Why it matters: Opus 5.5 landing #2 on Agent Arena with a claimed 56% cost cut makes it a notable Anthropic update today. Score capped below 85 because the post omits pricing and latency — the cost advantage needs real-world confirmation.

Sep 28Monday

Computing Life · Share · Yage

Agents mentioned PayPal 139 times, picked it 0 times: software distribution is changing

Armature ran 5,300 sandbox sessions where Claude Code, Codex, and Cursor integrated payment and email tools into real engineering repos. PayPal got 139 mentions and zero code commits; Stripe won 88% of payment tests. Tool choice follows language stack: the same email task picked Resend in TypeScript, SendGrid in Python, Postmark in Go, and Azure ACS in Java. Big markets concentrate heavily; long-tail categories split three ways across agents. Armature sells agent-adoption optimization starting at $5,000/month and disclosed the conflict upfront. The post doesn't spell out how the 42% selection-consistency figure was calculated; our own recomputation across reasonable definitions landed at 41.5%–48.2%, consistent with the published number. Agent tool selection is a context-dependent function, not a global ranking—traditional SEO logic breaks here.

Why it matters: Armature's 5,300 sandbox runs deliver first-hand data on how agents pick tools — the PayPal 0 vs Stripe 88% contrast is solid. All three HKR axes hit, but the testing org is a commercial entity and only 31% of data is released so far, capping it below 85.

Sep 26Saturday

Hacker News front page

LLM watermarking degrades AI agent performance and speed

Lasso Security tested LLM watermarking on AI agents and found it hurts tool-calling accuracy by 3–5 points on BFCL V3, adds 11 seconds of latency, and increases output length by 20–30%. The post doesn't name the specific models or agent frameworks tested.

Why it matters: Has concrete benchmark data answering a production-relevant question: what's the performance cost of watermarking on agents. But the post doesn't disclose which models/frameworks were tested, so the numbers are directional only — hence the score cap.

Sep 25Friday

Hacker News front page

NSA is spending billions this year to test frontier AI models, far above prior estimates

Two sources say the NSA told lawmakers in a classified briefing that it is spending billions in taxpayer money this year to evaluate and test advanced AI models. The figure is far higher than previously known, leading lawmakers to estimate a full federal AI regulatory system could cost tens of billions per year. Trump has mostly resisted stronger federal AI oversight. The article does not name which models are being tested, whose compute is used, or how the money breaks down.

Why it matters: Exclusive disclosure of a classified budget figure with solid information density; hits all three HKR axes. Held below 85 because the body doesn't disclose which models are tested, whose compute is used, or how the money breaks down — major factual gaps mean it's a policy sign...

Sep 24Thursday

Hacker News front page

A daily-updated LLM value chart that plots price against intelligence to find the frontier

The site plots 420 models from the Artificial Analysis Intelligence Index against blended API price, drawing a value frontier where no cheaper model is smarter. Claude Opus 5.5 leads at $8/1M tokens with a 57.6 intelligence score. Meta's Muse Spark 1.3 tops the $2–$8 band at 48.1, Xiaomi's MiMo-V2.6-Pro wins $0.54–$2 at 46.3, and Z AI's GLM 5.3 Flash takes the under-$0.24 tier at 41.8. The post doesn't disclose how the intelligence index is built, and Coding/Math sub-scores are listed as empty for many models, so I'd hold off on those comparisons.

Why it matters: A daily-updated price-performance leaderboard using Artificial Analysis data — genuinely useful for model selection. Hits H and K, but lacks the controversy or identity hook for R, so it lands at the featured threshold of 72.

Hacker News front page

Paper Instruments releases Paper Office, a Python suite for agents to safely edit Word, PowerPoint, and Excel files

Paper Office is a suite of Python packages that wrap python-docx, python-pptx, and OpenPyxl with safety checks and broader editing capabilities. Across 5 models and 61 tasks, Paper packages plus guidance passed 92.5% of trials, vs 80.7% for the upstream packages alone and 69.5% for Anthropic's Office skills. Agents resorted to raw OOXML editing in only 1.6% of Paper runs, compared to 78.7% without skills and 50.5% with Anthropic skills. The team argues that low agent adoption in consulting, law, and banking stems from tools that silently corrupt formatting, break references, or produce client-unready output. Paper Office keeps the familiar imports and adds cross-run text search, native Word redlines, comment threads, content controls, cross-document composition, and package-level diff saves, refusing unsafe operations instead of quietly breaking files.

Why it matters: Paper Instruments open-sourced a suite of Office file-editing libraries for agents, adding safety checks and broader editing capabilities on top of python-docx and friends. Across 5 models and 61 tasks, they hit 92.5% pass rate — 11.8 points above bare upstream libs and well a...

Hacker News front page

Open-source prompt-injection detectors catch 0–1% of realistic AI agent attacks buried in tool output

This benchmark hides 629 AgentDojo injection attacks inside tool outputs and tests Regex vs. Meta Prompt Guard 2. Regex catches 0%, Prompt Guard 2 catches 1%. The attacks aren't sent directly to the model—they're buried in search results, email bodies, and similar tool responses, so current detectors are effectively blind. Code and reproduction steps are public; the post doesn't include comparisons with commercial detectors.

Why it matters: 629 AgentDojo attacks buried in tool output, Regex catches 0%, Prompt Guard 2 catches 1%. Cleanly exposes the blind spot in indirect injection detection. Code and repro steps are public, which adds practical value. Held at 78 because it's a single benchmark without cross-detec...

AI HOT (Curated Pool)

Claude Opus 5.5 tops Coding Agent Index, but per-task cost rises to $13.04

Artificial Analysis tested Claude Opus 5.5 under Claude Code max effort and it scored 66 on the Coding Agent Index, up from Opus 5's 60. All three subtests improved: Terminal-Bench 4.0 63.1%, DeepSWE v1.1 68.4%, SWE-Atlas-QnA 66.4%. The trade-off: per-task cost jumped from $3 to $13.04. The post doesn't break down how max effort drove the cost increase.

Why it matters: Claude Opus 5.5 tops the Coding Agent Index with a 6-point jump to 66, but $13.04 per task is the hard number. Anthropic substantive update + independent third-party benchmark + concrete data — all three HKR axes hit. Not scoring higher because this is a single benchmark, not ...

AI HOT (Curated Pool)

MiMo-V2.6-Pro, Claude Opus 5.5, GPT-6 Luna/Sol launch, shifting the intelligence–cost Pareto frontier

Artificial Analysis reports that four new models this week added 11 points on the Intelligence Index vs. cost-per-task Pareto frontier. GPT-6 Luna contributed 5 points, Claude Opus 5.5 contributed 4, and the remaining two came from MiMo-V2.6-Pro and GPT-6 Sol. The post doesn't disclose specific scores, pricing, or latency—hold off on conclusions until full benchmarks drop.

Why it matters: Artificial Analysis's Pareto frontier chart is a hard reference for model selection — four new models landing 11 points at once pushes the boundary out meaningfully. GPT-6 Luna taking 5 points suggests competitiveness across cost tiers; Claude Opus 5.5's 4 points aren't far be...

Sep 23Wednesday

OpenAI News

OpenAI releases MentalHealthBench, an open benchmark co-developed with 80+ licensed clinicians to evaluate AI in realistic mental health conversations

OpenAI open-sourced MentalHealthBench, a benchmark built with over 80 licensed psychologists and psychiatrists across 22 countries. It tests AI on realistic mental health conversations ranging from everyday stress to emergencies, covering adults, teens, and caregivers. The eval goes beyond safety filters: it checks whether models seek context, preserve user agency, and offer actionable guidance when appropriate. OpenAI stresses ChatGPT isn't a substitute for therapy, but the benchmark tracks progress on empathy and steering people toward real-world support. The paper and benchmark are publicly available.

Why it matters: OpenAI released an open mental health benchmark built with 80+ licensed clinicians, covering a wide range of scenarios with finer evaluation dimensions than typical safety tests. It's directly useful for AI safety and product teams. Not scoring higher because it's an eval tool...

Latent Space

Claude Opus 5.5 launches with Fable 5.1-level performance at 40% lower cost, plus a rare focus on writing quality

Anthropic released Claude Opus 5.5, the first model in the new 5.5 family. It matches Claude Fable 5.1 on most tasks, costs 40% less to run than Opus 5, and is about 30% faster. The launch unusually highlights writing improvements: the model puts key info up front and follows user style rules. Artificial Analysis notes that token usage on frontier tasks jumped ~80%, so per-task cost remains around $6—similar to Opus 5. OpenAI shipped GPT-6 Sol and Luna an hour later at 50% lower prices than GPT-5.6, but Opus 5.5's launch post hit 17M views and dominated the day. Anthropic's system card also reports multi-agent scaling with up to 100 parallel agents for the first time. Latent Space tested both and switched to Opus 5.5 as the default model immediately, calling the writing quality a night-and-day difference over Sol 6.

Why it matters: Anthropic drops the first model in a new flagship family, claiming Fable 5.1 parity at 40% lower cost, with writing improvements front and center — a directly actionable upgrade signal for heavy Claude users. Held below 90 because we only have the official claim and Latent Spa...

Computing Life · Share · Yage

Same tool toggle: Nemotron-3 550B gained, Mistral-Medium-3.5 crashed

A new paper breaks down coding agent harnesses into three independent toggles and measures each one. The most striking result: switching from dedicated file tools to a pure CLI made Nemotron-3 550B's SWE-Bench Verified score jump 3.6 pp while cutting per-task cost from $2.33 to $1.11, but Mistral-Medium-3.5-128B dropped from 68.60% to 45.40%. Trajectory analysis shows 550B composing dense shell one-liners, while Mistral failed to locate files in 32.80% of tasks and submitted no edits. On Terminal-Bench 2.1, both models improved under CLI mode. Planning boosted the 30B model from 13.60% to 25.20% but only saved ~30% cost for larger models without accuracy gains. Context management mainly prevents window overflow; at 128k the gap shrinks to 2.7 pp, and complex read-back mechanisms were almost never invoked. The takeaway: no universal best harness design—it depends on the model's CLI fluency and the task type.

Why it matters: A controlled experiment that isolates three harness design switches and shows Nemotron-3 and Mistral-Medium-3.5 reacting in opposite directions, with concrete numbers and engineering takeaways. Not an 85 because it's a single preprint without cross-source cluster yet, but HKR ...

AI HOT (Curated Pool)

Arena launches GPT-6 Sol and GPT-6 Luna testing, scores coming soon

Arena is now testing two new OpenAI models, GPT-6 Sol and GPT-6 Luna, with scores not yet released. You can try them on real agent tasks and vote to feed the leaderboard. The post doesn't disclose model size, release date, or pricing.

Why it matters: GPT-6's first public appearance, two variants live on Arena running agent tasks — strong suspense and signal. Deduction for thin info: no scale, pricing, or release date disclosed, just a test entry point.

AI HOT (Curated Pool)

OpenAI launches GPT-6 Sol and Luna, API pricing cut 50% vs GPT-5.6

OpenAI added two cheaper models to the GPT-6 family: Sol and Luna, with API prices halved across input and output. Sol costs $2/$10 per 1M tokens, Luna $0.10/$0.50. Sol scored 33.2% on AutomationBench at xhigh effort at 9% of Claude Opus 5's cost per task, and 56.4% on Agents' Last Exam at max effort at 60% lower cost. On internal factuality evals, Sol makes about half as many mistakes as its predecessor. The post does not specify a launch date beyond 'available now.'

Why it matters: Official OpenAI release of new GPT-6 models with a 50% API price cut and Sol's agent benchmark cost at 9% of a competitor — industry-shaking. HKR all hit, with solid pricing and benchmark data. Minus 3 points because the post doesn't fully detail the capability gap between Sol...