Skip to content

Benchmarks

Who is actually stronger: benchmark results, methodology disputes and leaderboard changes.

Latest picks

41–60 of 453

Sep 14Monday

Hacker News front page

The AI job market in 2026: who gets hired, what they earn, and which roles are fading

Maksim Ilin pulls together LinkedIn, WEF, Stanford, PwC, and Bain data for a September 2026 snapshot of the AI labor market. AI Engineer is the most-hired role; Research Scientist at a frontier lab is the most prestigious, with median pay around $746K at Anthropic and $1.15M at OpenAI L5. The fastest-growing niches are agentic systems and Forward Deployed Engineering—postings for the latter jumped over 1,000% YoY. Prompt engineer has faded as a job title; the skill remains but dissolved into other roles. AI skills now command a 62% wage premium in the US. Bain projects more than 1.3M AI jobs in the US by 2027 against roughly 645K available workers. In Europe, over half of AI postings sit outside tech departments, and Germany shows seven AI-user roles for every AI-developer role. Entry-level hiring got harder: employment for 22–25-year-olds in AI-exposed occupations now trails the rest by 19%.

Why it matters: A multi-source synthesis of the 2026 AI job market with concrete salary figures and role trends — high reference value for practitioners. Capped at 72 because it's a personal blog aggregating secondary data, not an original institutional report with primary research.

Hacker News front page

30 SVG prompts benchmark 2025–2026 LLMs on pelican-bicycle-style drawing tests

Tom Gally built a site with Claude Fable 5.1 that extends Simon Willison's pelican-riding-a-bicycle test into 30 SVG drawing prompts. The 2026 run covers six models—GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash, DeepSeek V4 Pro, Qwen3.8 Max, and Fugu Ultra v2—while the 2025 run includes ten models like Claude Sonnet 4.5 and GPT-5.1. Each image shows generation time and cost: DeepSeek V4 Pro finished in 1 min 36 s at $0.10, Qwen3.8 Max took over 12 minutes, and Fugu Ultra v2 cost $1.02. The post presents raw SVG outputs without subjective ratings, so you compare the drawings directly.

Why it matters: Simon Willison's pelican test is a community staple, and this expands it to 30 prompts across 6 models with timing and cost — dense, useful signal. The deliberate lack of subjective scoring means readers have to flip through images themselves, which costs it a bit of immediate...

Computing Life · Share · Yage

The AI Benchmark Yardstick Moved Faster Than the Models

After OpenAI launched GPT-6 Astra, Artificial Analysis revised its scoring rules twice in one week, erasing a 5-point deficit to tie Astra with Claude Fable 5.1—without any model update. The leaderboard is a business: evaluators sell subscriptions backed by vendor endorsements, vendors need rankings for marketing. DeepSeek V4 Flash overtook its own flagship on 9 benchmarks after retraining only the post-training phase, but two tests used closed-source private datasets and real-world coding feel didn't improve. The same model scored 62.7% vs 99.9% on ARC-AGI-3 depending on the execution harness. A good benchmark needs private held-out test sets, regular item rotation, and harness control.

Why it matters: A well-sourced industry commentary with concrete version numbers and score shifts, exposing how a benchmark vendor rewrote its scoring rules twice in one week after GPT-6 Astra's release, flipping the ranking from a 5-point deficit to a tie for first. Hits all three HKR axes a...

Sep 11Friday

Hacker News front page

When code is correct but sloppy: measuring LLM-generated bloat

Sebastian at Earendil applied SlopCodeBench metrics to measure AI-generated code bloat. Agent code averaged 0.33 verbosity vs. 0.15 for human repos, and 0.68 erosion vs. 0.31. In multi-round, context-cleared iterations, even SOTA models hit 0% strict pass rate—bad decisions compound. The simplest effective metric is LOC change, but it breaks under Goodhart's law. The post does not spell out which directions he plans to explore next.

Why it matters: Earendil's post quantifies AI code bloat with two novel metrics—verbosity and erosion—using their SlopCodeBench. Concrete data, fresh angle. Downside: it's a single blog post, not peer-reviewed, and the benchmark isn't open-sourced, so reproducibility is unclear. But the topic...

Sep 10Thursday

Hacker News front page

LRU is harder to beat than KV-cache papers suggest, tested on 393 Claude Code sessions

This repo replays 68k requests from 393 real Claude Code sessions to test agentic KV-cache eviction policies. LRU is harder to beat than papers claim—many new policies look good on paper benchmarks but fall apart on real agent traces. The post doesn't give exact hit-rate numbers, but the core finding is that real access patterns differ sharply from academic benchmarks. Don't rush to replace LRU in production.

Why it matters: Replays real agent traces to stress-test KV-cache eviction policies, directly pushing back on papers that only cite academic benchmarks. 393 sessions and 68k requests is solid scale, but the repo doesn't disclose specific hit-rate numbers, so the score stays at the featured th...

AI HOT (Curated Pool)

OpenRouter launches Fusion: a compound model that debates across models before synthesizing a final answer

OpenRouter Fusion is a compound inference pipeline, not a new model. It fans out one prompt to up to 8 panelist models in parallel, has a judge compare their answers for consensus and blind spots, then lets the calling model write a final synthesis. A default three-model panel costs roughly 4–5× a single completion and takes 2–3× longer. On the DRACO deep-research benchmark, a budget panel of Gemini 3 Flash, Kimi K2.6, and DeepSeek V4 Pro scored ~64.7%, close to Claude Fable 5’s solo 65.3%. OpenRouter’s own test paired two Claude Opus 4.8 runs and saw a 6.7-point gain over a single run. The team positions Fusion as an escalation path for complex research and high-stakes decisions, not for simple chat. The post does not disclose per-token pricing, only the cost multiplier.

Why it matters: Fusion is a multi-model debate-and-synthesize workflow, not a new model. Concrete cost/latency numbers and DRACO benchmark data give it substance beyond marketing. But it's a routing-layer product update, not a foundation-model breakthrough — capped at the low end of featured,...

AI HOT (Curated Pool)

Cognition launches SWE-2 coding model, pushing the cost-performance frontier

Cognition's SWE-2 hits 50.0% on FrontierCode 1.1 Main, 8 points above SWE-1.7, at 64% lower cost than Fable 5.1. It's post-trained from the 2.8T-parameter Kimi K3 using a single-run RL method that trains all reasoning-effort levels together. The medium effort level solves tasks in 53 steps vs. 127 for SWE-1.7. The post doesn't disclose exact API pricing.

Why it matters: SWE-2 hits 50.0% on FrontierCode 1.1 Main, +8 pts over its predecessor, and costs 64% less than Fable 5.1 — Cognition's closest model to the frontier yet. Not an 85 because it still trails GPT-6 Astra by a few points and uses Kimi K3 as the base rather than an in-house model, ...

Sep 7Monday

Computing Life · Share · Yage

AI raised the floor, but grading rubrics still penalize the ceiling

Two large-scale RCTs show the same pattern: AI lifts the floor of student work while present, but once removed, performance drops, and traditional rubrics actively penalize deeper reasoning. In a Turkish high school math experiment, ChatGPT-assisted practice scores jumped 48%, yet closed-book exam scores fell 17% below the control group. In a Milan business writing study, students who spelled out failure conditions and causal mechanisms received systematically lower grades. The floor is borrowed from external compute; the ceiling only grows when rubrics reward it.

Why it matters: Two large-scale RCTs with hard numbers expose the illusion of AI-assisted learning: practice scores soar but closed-book tests drop, and students copy answers without reasoning. Strong HKR, but it's a synthesis piece rather than a primary research release, so it stays below 85.

r/LocalLLaMA

Qwen 3.8-27B NVFP4 beats Q5_K_M and nears BF16 after tweaking temp and min_p

A user compared Qwen 3.8-27B NVFP4 (NInfer) against Q5_K_M (llama.cpp) on a 5090. With default sampling, Q5 led on IFBench strict (76% vs 74%), but NVFP4's loose score was 78%, meaning its failures were mostly minor formatting drift. After lowering temp to 0.9 and setting min_p to 0.05, NVFP4 hit 80% strict while Q5 stayed at 76%. The official BF16 baseline is 79.5% on 300 samples; NVFP4's 80% on 50 samples has variance but the trend held across reruns. NVFP4 also ran nearly 3x faster and used far less VRAM. The takeaway: default sampling hides NVFP4's real quality—tighten it and you get near-BF16 instruction following.

Why it matters: First-person benchmark with concrete numbers and a reproducible recipe — not armchair theory. Directly useful for local LLM users. Score capped because it's a single-GPU, single-model case relying on the niche NInfer engine, limiting generalizability.

Sep 6Sunday

AI HOT (Curated Pool)

OpenAI repeatedly revised GPT-6 Astra benchmarks after launch, hallucination rate briefly halved from 4.2% to 2%

Fortune reported that OpenAI changed multiple benchmark scores for GPT-6 Astra after the September 3 launch. Astra's hallucination rate dropped from 4.2% to 2% then reverted; Anthropic Fable 5.1's math score was briefly cut by nearly 10 points. OpenAI called it normal pre-release validation, but Stanford researchers noted the system card lacks details on the hallucination eval—not even the number of test items. Worth flagging: these are best-case scores under any compute budget, not what a typical ChatGPT user would see.

Why it matters: GPT-6 Astra's launch is already a top-tier event; Fortune catching post-launch benchmark revisions — including competitor score changes — hits all three HKR axes. Held below 95 because it's a single-source report so far and OpenAI's response is vague.

AI HOT (Curated Pool)

OpenAI GPT-6 Astra tops Code Arena WebDev leaderboard, 35 points ahead of Claude Fable 5.1

GPT-6 Astra (Max) hit 1797 points on the Code Arena WebDev leaderboard, 35 points ahead of Claude Fable 5.1 (Max) in second place. Claude Opus 5 (Max) scored 1688 in third. The post is a single tweet—no sample size, task breakdown, or latency info disclosed, so I'd discount the claim for now.

Why it matters: GPT-6 tops Claude on Code Arena WebDev for the first time — a 35-point lead is conversation-worthy. But the source is a single tweet with no sample size, task breakdown, or latency data, so the score stays below 80. Wait for more sources before adjusting.

Sep 5Saturday

AI Chat-Group Daily (群聊日报)

GPT-6 Astra opens to all: faster but pricier, with a concurrent rate-limit war

GPT-6 Astra rolled out to all Pro users, landing in Codex CLI and Copilot. Early tests show a task that took 12 minutes now finishes in 6, but per-task cost is ~75% higher than Sol—one API code review burned $100. Tibo and Anthropic both reset all user quotas the same day, while Codex patched an infinite-usage exploit. A detailed Cerebras benchmark reveals real-world agentic throughput is only ~357 tps vs. the advertised 1,500 tps; the same task cost $1.57 in 3 minutes versus ~$0.017 locally. Zhipu GLM-5.3-Flash hit just 20 tps on domestic inference cards, while the same weights on Ollama Cloud reached 70 tps. In industry news, the US is drafting rules to block Chinese access to overseas AI servers, DeepSeek plans to buy over 160,000 Huawei chips for inference, and Saudi Arabia's Humain M3 was exposed as a rebranded MiniMax M3.

Why it matters: GPT-6 Astra's full rollout is the week's biggest product move, and this chat digest delivers first-day speed and cost data with real numbers. The cap at 78 reflects the source being an anonymized group-chat compilation rather than a primary official post, and some details (e.g...

Hacker News front page

Artificial Analysis launches Intelligence Index v4.2 with private test sets to prevent gaming

Artificial Analysis updated its model benchmark to v4.2, adding two new evaluations: AA-Briefcase and GDP.pdf. AA-Briefcase uses a private test set to simulate multi-week knowledge work projects and assess holistic agentic capability. GDP.pdf requires models to synthesize evidence across 4,592 pages of professional documents, graded against 1,275 atomic criteria where a task passes only if every criterion is met. Claude Fable 5.1 leads the index, followed by GPT-6 Astra, which shows an ~85 Elo gain over GPT-5.6 Sol. Private test sets now account for 40% of the weighting, double the v4.1 figure, specifically to reduce gaming by labs.

Why it matters: AA's leaderboard refresh matters for model selection workflows — the private test sets and 4,592-page document eval are more grounded than saturated public benchmarks. Not scoring higher because this is methodology iteration, not a capability breakthrough, and the post only gi...

Hacker News front page

EEBench benchmarks whether AI can design real circuit boards that actually work

EEBench released a circuit design benchmark that uses declarative code (atopile) instead of GUI clicking, so models work directly on components and constraints. It runs SPICE simulations to check voltages, tolerances, cost, and real part availability. Claude Opus 5 leads at 61.6%, with Grok 4.6 at 57.1%. In one energy-meter task, a design picked a 22µF nominal cap that delivered only 11.4µF at 4.7V bias—far below the 545µF requirement—and failed. xAI already included EEBench in the Grok 4.6 model card under engineering acceleration.

Why it matters: EEBench benchmarked frontier models on circuit design using declarative code and SPICE simulation — Claude Opus 5 leads at 61.6%. Timing is sharp, landing right after GPT-6 Astra's KiCad demo. Score sits at the featured threshold because PCB design is niche for the general AI ...

Sep 4Friday

r/LocalLLaMA

GPT-6 Astra hit 98.6% on ARC AGI-3 — don't fall for the hype

OpenAI reported GPT-6 Astra at 98.6% on ARC AGI-3, but used a proprietary harness instead of the standard one. Nvidia already hit 100% with its AVO harness, and earlier systems like Arc-Skill and VISTA also neared perfect scores. Under the standard harness, Astra drops to 66%. That's still solid, but it's not AGI. The post doesn't spell out what OpenAI's custom harness changed, so I'd discount the 98.6% figure for now.

Why it matters: This post dismantles OpenAI's 98.6% narrative with two numbers — Nvidia's 100% and a 66% on the standard harness — high information density and strong conflict. Not scoring higher because the source is a Reddit individual post, not an institutional review, and the body doesn't...

AI Chat-Group Daily (群聊日报)

Flash models hit SOTA: Gemini 3.8 Flash and Muse Spark 1.3 launch, cheap models now cover 90% of tasks

Google launched Gemini 3.8 Flash at $0.75/M tokens input, scoring 71% on DeepSWE and beating Sol and Opus 5 on multiple agent benchmarks. Meta released Muse Spark 1.3 the same day, hitting 61–62 on AA Intelligence Index, matching Grok 4.6; Contributor tier costs just $0.10/$0.20 but trains on user data by default. A group member shared two-week usage stats: 1.28B tokens on GLM 5.3, with over 90% of tasks handled by cheap models. Uncle Bob proposed a multi-agent pipeline completing tasks in about one hour, insisting deterministic tools like tests and linters won't go away. GPT-6 confirmed for September 3 morning launch. LatePost exposed China's embodied AI funding bubble: among 22 companies valued over 10B RMB, one at 20B spent under 40M on R&D last year. NYC will ban student-facing generative AI tools for K-8.

Why it matters: Gemini 3.8 Flash launch with Flash-tier pricing beating Sol and Opus 5 on agent benchmarks. The source is a curated group chat digest, not a first-party announcement, which caps the score slightly, but the signal density and real-world testing notes are solid.

AI HOT (Curated Pool)

OpenAI launches GPT-6 Astra, focused on computer use and alignment

GPT-6 Astra can operate across apps, build test software, and tackle open science problems. OSWorld real-desktop task time dropped from 75 to 40 minutes, and workplace automation rose from 18% to 41%. On alignment, unguarded jailbreak rate fell from 48% to 0%. The author says $2,000 in compute solved 10 decade-old math and theoretical CS problems, but tool-augmented benchmarks still trail Claude.

Why it matters: GPT-6 Astra launch is an industry-shaking event. The computer-use and 0% jailbreak numbers are concrete, hitting all three HKR axes. Score not at 98-100 only because we currently have a tweet summary without an official blog or third-party verification; can bump higher once mo...

Hacker News front page

Which tools Claude Code, Codex, and Cursor pick in 16,893 real coding sessions

Armature ran nearly 17k experiments across 75 repos and 1,163 prompt variants to see which services Claude Code, Codex, and Cursor actually install. They simulated four personas—vibe coder, junior, senior, and enterprise engineer—and had agents go from analysis to implementation. The post discloses partial findings: in object storage, Cloudflare R2 started beating Amazon S3 once a simulated human was added to the loop; in databases, Neon was repeatedly recommended. Full leaderboards and raw traces are published, but the article body cuts off before covering more categories.

Why it matters: Armature ran 16,893 simulated sessions to surface tool selection preferences across Claude Code, Codex, and Cursor—solid sample size, useful signal. The caveat: Armature sells growth services to dev-tool companies, and while they disclose it upfront, that stake puts a question...

AI HOT (Curated Pool)

Perplexity to integrate GPT-6 Astra, CEO says it tops WANDR benchmark

Perplexity CEO Aravind Srinivas says the company will integrate OpenAI's newly released GPT-6 Astra, claiming it far outperforms other models on deep and broad research tasks at lower cost. It will roll out to Perplexity Computer Pro and Max users first. The post does not disclose a launch date, WANDR scores, or cost figures.

Why it matters: A top AI search product quickly adopting the latest flagship model is newsworthy. But the post lacks WANDR scores, cost figures, and a launch timeline — the actual improvement is still unclear, so it doesn't push past 85.

AI HOT (Curated Pool)

GPT-6 Astra scores 66% on ARC-AGI-3, nears 100% with persistent conversation

François Chollet reports GPT-6 Astra's ARC-AGI-3 results. Standard harness yields 66%; persistent conversation with custom compaction pushes it near 100%. Cost is roughly $360 per task. The post doesn't disclose task count or latency. Worth noting: near-perfect scores rely on a conversational harness, not raw model output, so it's not yet general reasoning out of the box.

Why it matters: Chollet himself posted GPT-6 Astra's ARC-AGI-3 scores: 66% standard, near-100% with multi-turn dialogue and custom context compression, at ~$360 per task. The contrast and the cost make it a must-read for the reasoning-eval crowd, but it's not single-pass reasoning, so it does...