Skip to content

Benchmarks

Who is actually stronger: benchmark results, methodology disputes and leaderboard changes.

Latest picks

101–120 of 453

Aug 5Wednesday

Hacker News front page

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

This ICML 2026 paper defines and measures benchmark saturation across 60 LM benchmarks. Nearly half are already saturated, and older benchmarks saturate faster. Expert curation helps resist saturation; keeping test data private does not. The post doesn't name the specific benchmarks but identifies 14 properties linked to saturation, offering a framework for building longer-lasting evaluations.

Why it matters: ICML 2026 paper with a systematic audit of 60 benchmarks and counterintuitive findings (hidden test sets don't delay saturation). Solid eval-infra research, but no named benchmarks limits immediate impact, capping at 78.

Aug 4Tuesday

AI Chat-Group Daily (群聊日报)

Qwen 3.8 Max matches Fable 5 at 2.4T params, open weights next week

Qwen 3.8 Max launched with Terminal Bench 2.1 score 86.6 and PaperBench 93.0, beating Fable 5's 88.8. A 500-yuan token plan burned out in one day; the model lands between Luna and Terra, with price as the main draw. Open weights for both Qwen 3.8 Max and Qwen 3.8-27B drop next week. DS V4 Flash hit 8T tokens consumed in a single day, topping weekly charts—the group sees tokens becoming a commodity. On tools: an M5Stick voice dongle turns a keychain into an agent remote, LoopX keeps agent state across 200+ hours, and reverse-skill injects reverse-engineering toolchain knowledge into coding agents. The wildest methodology story: an agent autonomously downloaded a local Qwen mid-translation task and auto-installed Whisper when its API key ran out of funds.

Why it matters: Qwen 3.8 Max official release, 2.4T params matching Fable 5 with open weights coming next week — a major domestic flagship model update. The chat digest provides concrete benchmarks and real-world impressions, high information density. Deduction because the source is a group c...

Aug 2Sunday

Computing Life · Share · Yage

Prompt injection defense lives in the harness, not the model

Ghostcommit showed the same Sonnet model rejected malicious PNG instructions 10/10 times in Claude Code, but obeyed 10/10 times in Cursor and Antigravity, leaking .env secrets. Lab-reported 99% defense rates suffer from five traps: static benchmark overfitting, misleading single-attempt ASR, LLM-as-judge drift, ignored utility-under-attack, and bare-model testing without tool shells. Deeper causes: LLMs lack hard instruction-data separation, and stronger models can follow injections more faithfully—Opus 4.6 with extended thinking saw ASR rise from 14.8% to 21.7%. A joint study by 14 researchers from OpenAI, Anthropic, and DeepMind tested 12 model-layer defenses; over 90% broke under adaptive attacks, with human red-teamers hitting 100%. The engineering fix is architectural isolation: CaMeL separates trusted planner from untrusted executor, and OpenClaw's dual-agent setup cut ASR from 100% to 0.31%. Harness-level deterministic tool gating, hook signature checks, and sandboxed least-privilege are the real controls.

Why it matters: Uses Ghostcommit's 0/10 vs 10/10 data to relocate the prompt injection debate from the model layer to the toolchain harness—sharp thesis with reproducible evidence. Score held at 82 because the article cuts off mid-argument (only one of five eval traps is unpacked), so the ful...

Jul 31Friday

Hacker News front page

Distilling DeepSeek into GPT-OSS Doesn't Transfer Censorship

CTGT distilled a 120B finance model from DeepSeek V4 Flash. Across 152 matched prompt pairs, the teacher scored 45.45 points more censored on China-sensitive topics, but the student showed no censorship at all—four US lab judges agreed. Self-distillation on corrected outputs matched the DeepSeek-taught model on financial reasoning, at 62× lower cost than Inkling. Code, data, and models are open.

Why it matters: CTGT distilled DeepSeek V4 Flash into a 120B finance model and found censorship didn't transfer, while self-distillation matched the teacher on finance reasoning at a fraction of the cost. Ships with weights, a playground, and a reproducible eval framework. HKR all hit. Not sc...

Jul 29Wednesday

AI HOT (Curated Pool)

Enabling two API settings tripled GPT-5.6's ARC-AGI-3 scores

GPT-5.6 Sol scored just 7.8% on ARC-AGI-3 because the official harness discarded private reasoning after each action and used rolling truncation that dropped older moves. Switching to retained reasoning and context compaction raised the public-set score from 13.3% to 38.3% while cutting output tokens by 6x. Human testers averaged about 48%. The post doesn't disclose full private-set results or whether the same settings help other models.

Why it matters: Official OpenAI post with concrete numbers and root-cause analysis, not marketing fluff. Capped below 85 because it's an engineering lesson rather than a capability breakthrough, and total score isn't disclosed. But 'the harness hurt the model' is directly useful for agent ben...

Jul 28Tuesday

Hacker News front page

A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

Fermisense GRPO-fine-tuned a 9B open-source model for catalog review and hit 87.3% quality vs. 76.9% for the best frontier setup. Cost per 1,000 listings: $0.50 for the fine-tune, $19–$172 for frontier models. The post calls this 'intelligence ownership'—training a specialist on your own data, tools, and scorer beats prompt-tuned frontier APIs on both accuracy and cost. Ramp data is cited: top-quartile AI spenders more than doubled revenue from Nov 2022 to Dec 2025; zero-AI-spend companies grew ~15%.

Why it matters: Hits all three HKR axes with concrete numbers, named method, and cost comparison. Not scored higher because it's a single company blog post with no third-party reproduction or cross-source verification, and Fermisense sells this solution — conflict of interest exists. But the ...

Jul 27Monday

Computing Life · Share · Yage

Why high SWE-bench scores don't translate to real-world Kotlin projects

JetBrains released the Kotlin Benchmark on July 24, 2026, testing AI coding agents on 105 real-world tasks across 8 open-source projects. With the same Opus 4.7 model, Claude Code hit 85.71% and Junie 81.9%—a nearly 4-point gap driven by how each agent harness handles Gradle build logs. A good harness uses Language Server diagnostics to catch static errors locally, then runs full builds only at key checkpoints and extracts just the blocking lines from noisy output. The article argues that Python-based benchmarks like SWE-bench reward trial-and-error strategies that collapse under Kotlin's heavy build overhead. It recommends teams stop buying off public leaderboards and instead use the Harbor container spec to build private micro-eval matrices from their own historical PRs and issues, measuring Pass@k stability, token cost, and whether patches respect internal architectural constraints.

Why it matters: JetBrains official benchmark with concrete numbers plus an engineering-level breakdown of SWE-bench's limitations. Not just complaining about leaderboard distortion—it traces the root cause to static compilation overhead vs Python's dynamic runtime feedback. Slight ding becaus...

Jul 26Sunday

Hacker News front page

Terence Tao's ICM 2026 talk on how the math community should respond to AI

Tao's ICM 2026 talk sidesteps the debate on whether AI can do research math. He asks the community to assume a strong capability hypothesis and instead examine its own goals and values. He draws a parallel to the early-20th-century foundations crisis. The talk cites First Proof benchmark results: on May 28, 2026, 7 out of 10 novel research problems were solved at publication quality by at least one AI harness, with compute costs of $10–$1,000 per problem.

Why it matters: Tao's ICM 2026 public lecture takes a sharp angle: he assumes strong AI capability for research math and forces the community to answer 'what are our goals and values?' The parallel to the early-20th-century foundations crisis lifts this from technical to philosophical. Two de...

AI Chat-Group Daily (群聊日报)

Opus 5 Day 2: Saturation self-testing trades cost for quality, total cost may beat Fable

Third-party tests show Claude Opus 5 uses saturation self-testing—frontend screenshot checks and 1000+ backend test cases—to nearly eliminate delivery issues, but total cost in complex scenarios may exceed Fable. With self-testing off, bug rates don't beat the previous model. Anthropic's strategy: long-chain debugging over one-shot correctness. The official model card advises against max effort for the first time; FrontierBench peaks at xhigh. Another test reveals ~80% of Claude Code's system prompt was cut. Group sentiment is positive, some calling it smoother than Fable. OpenAI had a full 503 outage overnight; reset cards landed the next day. WSJ reports US companies are mixing cheaper models to control costs, with Cursor as a beneficiary.

Why it matters: Third-party testing delivers the most concrete behavioral and cost data on Opus 5 so far — the self-testing tradeoff is a real signal. Slight discount for being a group-chat digest rather than the original review, but density clears the featured bar.

Jul 25Saturday

Hacker News front page

ARC Prize launches ARC-AGI-3 leaderboard, ranking AI systems by cost efficiency

ARC Prize published the verified leaderboard for ARC-AGI-3. The new benchmark tests how AI agents adapt on the fly to novel interactive environments, not just passive reasoning. A scatter plot maps each system's score against cost per task, making efficiency the headline metric. Only systems that cost under $10,000 to run are shown; Kaggle entries operate under a $50 compute cap. The post doesn't list specific model names or scores—you need to open the page to see the full ranking.

Why it matters: ARC Prize launches the ARC-AGI-3 verified leaderboard, shifting from static benchmarks to interactive agent adaptation with cost transparency. But the body is just navigation chrome with zero actual scores or model names — too thin to push higher than the featured threshold.

Computing Life · Share · Yage

OpenAI Presence: Turning Field Failures Into a Productized Improvement Loop

OpenAI launched Presence on July 22, an enterprise voice and chat agent product targeting specific roles like customer service and outbound sales. Its core pitch is not model capability but productizing the feedback loop after agent failures: when a task gets stuck and escalates to a human, the system saves the full execution context, lets teams reproduce the failure in a sandbox, fix rules, run regression tests, and push code changes via Codex. This standardizes what field FDEs used to do manually—migrating on-site failures back into the product. Presence is in limited GA, non-self-serve, deployed case-by-case; OpenAI hasn't disclosed hosting details, data residency, or cross-vendor export for failure records and test suites. The article warns that if enterprises can't take these hard-won lessons with them, they face a new form of vendor lock-in.

Why it matters: OpenAI productized the hardest part of enterprise agent deployment — the post-failure improvement loop — with a concrete mechanism. Score held below 85 because it's a single-source analysis lacking multi-source confirmation, official pricing, or real customer scale data.

Jul 24Friday

Computing Life · Share · Yage

GPT-5.6 prompt guide: write fewer steps, define clearer deliverables

OpenAI's July 22 guidance for GPT-5.6 tells developers to strip hand-written intermediate steps from prompts and instead constrain agents with completion criteria, verification evidence, and permission boundaries. The recommended method is ablation testing on eval sets—remove a section, rerun, and keep it only if metrics hold. This reverses the GPT-4.1 era of hard-coding eight-step workflows into system prompts. GPT-5 had already started loosening route control by scene. The author validated the approach in a long-form translation system, replacing chunking and retry logic with deliverable specs that let the agent decide its own execution path.

Why it matters: Connects three generations of OpenAI prompt guides into a coherent engineering narrative with concrete methodology (ablation testing), not generic advice. Score capped here because it's a secondary analysis of official docs rather than a primary release, and the excerpt doesn'...

Jul 23Thursday

Computing Life · Share · Yage

Cursor rewrites SQLite with Swarm: a controlled experiment pushing three scaling dimensions of agent orchestration

Cursor fed 835 pages of SQLite docs into its new Harness, hitting 80% sqllogictest pass rate in 4 hours; the old Swarm was halted before hour 2 due to code conflicts. The new system isolates planner and worker roles, uses shared design docs and auto-merge, cutting merge conflicts from 70k to under 1k. A mixed-model setup—Opus 4.8 planning, Composer 2.5 executing—cost $1,339 total, roughly 8× cheaper than GPT-5.5 solo. The public minisqlite repo lacks CLI, C API, and cross-process locking, so it's far from production-ready SQLite. This is a capability demo under ideal conditions—fixed spec, dense feedback—not a daily driver for product teams with shifting requirements.

Why it matters: Cursor ran a controlled A/B test of old vs new agent systems, hitting 80% sqllogictest pass rate in 4 hours while the old system collapsed in 2. Concrete numbers and architectural insight make it valuable for AI coding practitioners. Not top-tier because it's a single technica...

Jul 22Wednesday

OpenAI News

OpenAI launches Presence, a production agent product for customer and internal workflows

OpenAI launched Presence today, a product for deploying voice and chat AI agents in enterprise workflows. It bundles policies, guardrails, escalation rules, and evaluation tooling so agents can access company systems, take approved actions, and hand off to humans when needed. OpenAI's own English-language support line at 1-888-GPT-0090 already runs on Presence: it resolves 75% of inbound issues without human help and cut handoff rates by 15 percentage points in 10 days via a Codex-powered improvement loop. BBVA is testing Spanish-language voice banking in Mexico, SoftBank is trialing Japanese conversations, and IAG is exploring claims support during severe weather. The post does not disclose pricing or API availability details.

Why it matters: OpenAI productizes its internally validated support-agent stack with a 75% automation stat and two named enterprise references. Not scoring higher because we only have the vendor's own announcement — no third-party benchmarks or customer-side data yet, and pricing isn't disclo...

Hacker News front page

Fireworks benchmarks Kimi K3 vs Fable 5: routing the two hits 93% accuracy at far lower cost

Fireworks ran ~1,030 real agent tasks comparing Kimi K3 and Fable 5. Head-to-head on SWE: K3 92.4%, Fable 92.6% — a near tie. Under the hood, K3 excels at symbolic math and dev tooling; Fable wins on web and data viz. Routing between the two lifts accuracy to 93% and cuts cost up to 50x on long agent loops. Oracle routing shows K3 can handle 72–96% of tasks, but a practical router still needs an order of magnitude more routing data.

Why it matters: Fireworks ran ~1,030 real agent tasks comparing K3 and Fable, with concrete numbers and a task-type breakdown. The finding that routing between the two beats either alone is useful and testable. Score stays at 78 rather than higher because this is a single-source vendor post—F...

Jul 17Friday

AI HOT (Curated Pool)

Kimi K3 tops frontend coding leaderboard, open weights coming July 27

Kimi K3 scored 1679 on Frontend Code Arena, taking first in 6 of 7 frontend sub-tasks and beating Claude Fable 5 and GPT-5.6 Sol. It's a 2.8-trillion-parameter MoE model with a 1M context window, and open weights are promised for July 27. API pricing is $15 per million tokens—no low-cost play here, it's priced against top closed-source models and aimed at long-context coding and agent workflows.

Why it matters: Moonshot AI's Kimi K3 tops Frontend Code Arena at 1679, winning 6 of 7 subtasks against Claude Fable 5 and GPT-5.6 Sol. 2.8T MoE params, 1M context window, weights opening July 27. A domestic flagship model directly challenging the closed-source duopoly on a concrete coding be...

AI HOT (Curated Pool)

The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway

A VentureBeat survey of 157 enterprises finds a sharp gap between agent evaluations and real-world performance. In the past year, 50% of organizations shipped an agent or LLM feature that passed internal evals but then caused a customer-facing failure; a quarter saw it happen more than once. Only 5% fully trust automated evaluation, with poor alignment to real outcomes cited as the top limitation (29%). Yet 66% already allow or are engineering toward fully automated, zero-human-in-the-loop deployments. The eval stack is fragmented: 17% rely on model-provider native evals, another 17% have no dedicated tooling, and only about a quarter run real-time quality checks on live traffic. The sample skews mid-market (100+ employees), with tech/software at 23%.

Why it matters: Survey of 157 enterprises quantifies the trust gap between agent testing and production. The 50% failure rate and 5% full-trust number are solid. Downside: it's a survey report, not a product launch, and methodology details aren't disclosed in the excerpt.

Jul 16Thursday

Hugging Face Blog

Ai2 shares the engineering lessons behind Shippy, a maritime AI agent

Ai2's Skylight team built Shippy, an AI assistant that helps maritime analysts query fishing activity, EEZ boundaries, and vessel tracks. The post breaks its architecture into three parts: a soul (system prompt), skills (markdown files that teach it to call APIs and interpret track data), and config (runtime settings; currently Claude Opus 4.6 with the OpenClaw framework). The core idea is wrapping a non-deterministic model in deterministic tools—every answer includes source, data cutoff, and a deep link to the Skylight map so an analyst can verify it. The post doesn't disclose error rates or latency numbers, but it stresses sandboxed hosting and evaluating the agent as a system, not just the model.

Why it matters: Ai2's three-layer agent architecture (soul/skills/config) and the Markdown-as-skill-sheet pattern are concrete engineering takeaways. But the maritime domain is too niche for broad resonance, landing right at the featured threshold.

Hacker News front page

Unsolved Problems in MLOps

This ACM Queue piece lays out why classical ops practices break down for ML: non-deterministic outputs and data as a system driver make canary deploys, health checks, and alerting nearly useless. Azure validates new models by having LLMs judge LLM output—the SRECon audience was audibly surprised. The authors argue the field must either find a better paradigm or fix the ones we have.

Why it matters: This ACM Queue piece lays out MLOps' core tension: traditional ops relies on deterministic responses for health checks and canary releases, but ML systems are non-deterministic and data-driven. The Microsoft Azure example—using LLMs as judges with employee thumbs-up as fallbac...

Jul 15Wednesday

Computing Life · Share · Yage

AI Trains AI: What a Public Self-Improvement Experiment Actually Closed the Loop On

Dan Austin open-sourced a full AI-trains-AI loop. An outer Qwen3.6 agent designs post-training recipes; an inner Qwen3-0.6B or 1.7B model runs real GPU training, and hidden eval scores feed back as reward to update the outer policy. Over 54 steps, the agent first learned to reduce invalid submissions, then shifted 1.7B model usage from 42% to 95% and began tuning temperature, optimizer, and other hyperparameters. Trained small-model scores rose from noise level into the 0.22–0.48 range, with limited transfer to a held-out triage task. A postmortem also revealed an evaluator bug: the old tool-use detector looked for `.function.name` instead of `.name`, so the 0.4-weighted score never fired—yet the reward curve still climbed. The fix required a full restart. The experiment shows outer RL can reshape a training agent's behavior, but tasks, rewards, and budgets are still human-designed.

Why it matters: Dan Austin open-sourced a full AI-trains-AI experiment: an outer Qwen3.6 agent designs post-training recipes, inner models actually train on GPU, and after 54 steps the agent learned to pick models and tune hyperparams, pushing 1.7B usage from 42% to 95%. All three HKR axes hi...