Skip to content

#评测/基准

2 today

Sep 19Saturday

Hacker News front page

Brood War Bench: No model played beyond beginner level in StarCraft

Ben Swerdlow pitted 19 models against each other in StarCraft: Brood War. Codex Astra / xhigh went 18–0, but no model surpassed beginner level. Older models treated the RTS as turn-based and got destroyed while thinking; Grok 4.6 issued only 6 command batches in 43 minutes and never fielded a combat unit. Claude Fable earnestly climbed the tech tree but couldn't execute. Codex models favored early Probe harassment that paralyzed opponents. The post doesn't specify whether matches were pure AI vs AI or involved human input.

Why it matters: First-person benchmark with 19 models playing StarCraft against each other. Concrete data (win rates, APM, cost) and a clear finding: none surpass beginner level. Codex Astra won by early worker harass, not macro play — that detail carries signal. Not an 85 because it's more a...

Hacker News front page

PlanetScale releases TIN: a full-text search extension for Postgres

PlanetScale released TIN, a GA full-text search extension for Postgres. It handles boolean, phrase, fuzzy, and regex queries with correct MVCC visibility under concurrent writes. Benchmarks on 150M Stack Exchange documents show index build time and mixed-query latency; I'd want to see direct comparisons with existing Postgres options before drawing conclusions.

TechCrunch · AI

a16z-backed Vals aims to become the gold standard for AI benchmarking

Vals wants to be a neutral third-party benchmark for AI models, backed by Andreessen Horowitz. With model makers publishing their own scores, Vals aims to be the trusted referee. The post doesn't disclose its evaluation methodology or early customers yet.

Hacker News front page

Apple M6 Pro tops Geekbench 7 single-core chart

Apple M6 Pro scored 4150 in Geekbench 7 single-core, the highest public result so far. The 18-core chip runs at 4.78 GHz base, split into 6+12 clusters, with 48 GB RAM. Multi-core hit 37565. But this is just a benchmark—real-world power, thermals, and shipping devices aren't covered here.

Sep 18Friday

Hacker News front page

ByteShape releases full ShapeLearn quant for Qwen 3.8 27B, hitting 99.63% accuracy at 13.1 GB VRAM

ByteShape released full ShapeLearn quantized models for Qwen 3.8 27B. All five models sit on the quality-speed frontier across six GPUs. Default pick GPU-5 (IQ4_XS, 3.84 bpw, 13.1 GB) scores 99.63% of BF16 at 93.66 tok/s on an RTX 5090. If VRAM is tight, GPU-4 (IQ3_S, 11.0 GB) still delivers 98.72% accuracy and runs faster. Supports MTP and DFlash2 speculative decoding; DFlash2 is faster for text-only but needs an extra 1.1 GB VRAM and doesn't handle images. The post doesn't disclose training data or optimization budget details.

Sep 17Thursday

Hacker News front page

AI now beats some of the best human forecasters

The Economist reports that AI has outperformed some top human forecasters in predicting geopolitical events. The article references specific competitions and models, but the body doesn't disclose model names, dataset size, or error margins. I'd hold off on the exact lead until the full evaluation is available.

Hacker News front page

Coding agent harnesses can 2× your cost with no real accuracy gain

UC Berkeley and Arena researchers tested 7 models across 3 harnesses—Claude Code, Codex CLI, and Pi—on 30 tasks each from SWE-bench Lite and Terminal-Bench 2.0, with 3 repetitions per pair. Harness choice barely moves success rates (±2–5%), but cost can vary up to 5×. Claude Fable 5 hits 97.8% in Claude Code at $1.33, and 96.7% in Pi at $0.67. Pi, a minimal open-source harness with just read, write, edit, and bash, reaches the Pareto frontier on both benchmarks. The post doesn't spell out the exact open-source licenses for Pi and Codex CLI, and doesn't link the full pricing sheet.

Why it matters: Systematic eval from UC Berkeley and Arena: 21 model–harness pairs on standard benchmarks yield a counterintuitive finding—harness barely moves success rate but swings cost 5x. Concrete numbers, clean experimental design, practical takeaway. Held at 78 because the body excerpt...

Hacker News front page

Frontier models are much better at physics than benchmarks suggest—expert re-grading shows why

Researchers at Yale and other institutions had physics faculty and PhDs re-grade six widely used physics benchmarks. Most answers previously marked wrong turned out to be grader errors, incorrect reference solutions, or ambiguous questions. For GPT-5.6-Sol, corrected mean@4 jumped from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark. The near-saturation on these closed-ended tasks signals an urgent need for harder, expert-validated evaluations.

Why it matters: Yale physicists re-graded six popular physics benchmarks and found most 'wrong answers' were actually grading bugs or ambiguous questions. Corrected scores show GPT-5.6-Sol jumping from 47.3% to 78.7% on HLE-Physics — near saturation. A solid takedown of benchmark trustworthin...

Hacker News front page

Training a 4B model to produce 81% faster query plans than Postgres

Rohan Bansal post-trained a Qwen 4B model to beat Postgres's default query plans. After SFT distillation from 500 GPT-6 Astra trajectories and a custom GRPO variant for RL, the model achieved 44.7% latency reduction and 81% geometric mean speedup across 113 join-heavy queries. Training ran on a rented 2×H100 node with four Postgres containers on his desk for measurement. The post doesn't disclose total training time or per-inference latency.

Why it matters: A 4B model trained via RL beats Postgres default plans by 81% on the Join Order Benchmark. The method is practically interesting, but only 113 queries were tested—generalization is unproven, capping the score at 78.

Latent Space

AIUC raised a $40M Series A to insure AI agents so companies can deploy them and sue when things go wrong

AIUC announced a $40M Series A led by Ribbit Capital and First Harmonic. CEO Rune Kvist, Anthropic's first product hire, argues that trust and liability—not capability—will cap AI adoption. They built AIUC-1, a standard that stress-tests agents for jailbreaks, hallucinations, and data leaks, backed by real insurance. Cursor, Harvey, Lovable, and ElevenLabs are already working with them. The episode raises a sharp hypothetical: what happens when a $20 Cursor subscription contributes to a $200M plane crash. The post doesn't disclose specific premium or claims-handling details.

Why it matters: AI agent insurance is a new category, and the AIUC-1 standard plus $40M Series A give this story substance. The CEO's Anthropic pedigree and Ribbit Capital backing add credibility, but the product is early-stage — the post doesn't disclose actual claims data or premium pricing...

Sep 16Wednesday

NVIDIA Blog

NVIDIA Vera Rubin NVL72 tops MLPerf Inference v6.1 in debut

NVIDIA's Vera Rubin NVL72 topped MLPerf Inference v6.1 in its first run. It's the post-Blackwell flagship with 72 GPUs linked via NVLink, built for large-scale inference. The post doesn't disclose exact scores or comparison models—only claims "leading performance." For buyers, this suggests inference throughput and latency improvements over H100/B200, but detailed numbers are needed to calculate ROI.

AI HOT (Curated Pool)

Arena updates Image-to-WebDev leaderboard: GPT-6 Astra tops at 1733

Arena added four new models to its Image-to-WebDev leaderboard. GPT-6 Astra (Max) leads at 1733, 129 points ahead of GPT-5.6 Sol (xHigh). Claude Fable 5.1 (Max) is third at 1710, Muse Spark 1.3 (Max) fourth at 1645, and GLM-5.3-Flash tenth at 1588. The post doesn't disclose evaluation tasks or sample size, so I'd take the gaps with a grain of salt.

Why it matters: GPT-6 Astra tops the Image-to-WebDev leaderboard on its first appearance with a meaningful margin — newsworthy. But the post only gives scores and rankings, with no detail on methodology, task difficulty, or model differences. H and K both hit, R is absent — lands right at the...

Latent Space

Can skills learned in games transfer to real-world work?

Good Start Labs trained a 30B model on the railroad game 1830 and found that training design determines skill transfer. A multi-turn terminal agent version improved at financial research tasks—querying databases, writing Excel formulas, reasoning on the fly—while single-turn training did not. The company spun out of Every last October with $3.6M in funding, betting on verifiable game environments for RL-based skill teaching.

Why it matters: The experimental design is novel, with positive skill-transfer evidence and a failure control, useful for agent training research. But the company just spun out, product path is unclear, and the post doesn't disclose specific accuracy numbers on the financial task, so it stays...

Hacker News front page

RL for LLMs has a Matthew Effect where hard problems get ignored—this post proposes Never Give Up to fix it

Michael Noukhovitch's blog walks through his new paper on the Matthew Effect in RL post-training for LLMs: as training progresses, the model samples easy problems more and hard problems less, because early successes on easy tasks dominate the reward signal. His proposed fix, Never Give Up (NGU), forces a minimum sampling ratio for hard problems so they don't get squeezed out. On Olmo 3.1 7B math training, NGU lifts AIME 2025 pass@1 from 26.7% to 33.3%; on code, LiveCodeBench pass@1 goes from 23.4% to 26.1%. The post also covers async RL staleness tricks and frames the Matthew Effect as a form of primacy bias. The body doesn't disclose NGU's specific hyperparameters or extra compute cost, so I'd discount the gains until those details surface.

Why it matters: Michael Noukhovitch turns his paper on the Matthew Effect in RL post-training into a highly readable blog post: models increasingly favor easy problems during training, and hard-problem sampling rates keep dropping. His proposed NGU method enforces a minimum sampling ratio for...

Sep 15Tuesday

AI Chat-Group Daily (群聊日报)

Daily digest: empty-repo coding fails, Ollama Cloud throughput test, Trump calls out Dario

今天最直观的教训来自 @搞仁义毛义仁:给 Astra 一个空白 C++ 仓库,代码写得一塌糊涂;把积累了大量 code review 经验的 GacUI 上下文导进去,质量立刻飙升。这说明模型不是不会写,是得用具体规则去“规训”。@今天群内信息量极大 实测了 Ollama Cloud 跑 DeepSeek V4.1 Flash,解码吞吐是官方 API ...

AI HOT (Curated Pool)

Fireworks benchmarks DeepSeek-V4.1-Flash: matches GPT-6 Astra on DeepSWE at 1/15th the cost

Fireworks ran a full benchmark suite on DeepSeek-V4.1-Flash. On DeepSWE, it scores 74.34% pass@1, in the same band as GPT-6 Astra at 74.12%, but costs $0.43 per task—15x cheaper. The model uses a 552B MoE with a split activation design: 8B active for input, 16B for output, plus KV cache optimizations. On Terminal-Bench 2.1 it trails Astra by 1 point while costing 12x less. The post mentions an HLE and oracle router eval but does not disclose the actual scores.

Why it matters: DeepSeek V4.1-Flash matching GPT-6 Astra on DeepSWE at an order-of-magnitude lower cost is the strongest price-performance signal this week. Docked because the source is Fireworks' own benchmark, not an independent eval, and the body is truncated by a cookie wall with no full ...

Hacker News front page

GPT-5.6 Luna vs GPT-6 Astra: 3.6% of the cost for 75% of the bugs in code review

Entelligence benchmarked GPT-5.6 Luna and GPT-6 Astra on 50 public PRs for code review. Luna costs $1.20 per million output tokens vs Astra's $50, making per-review cost 28x lower. Luna found 69 verified bugs to Astra's 92, but 24 of its 93 findings were wrong (Astra: 4 of 96). On Sentry, Discourse, and Grafana, Luna was within two bugs of Astra. On Keycloak, an identity server, Luna found 6 bugs to Astra's 14 with only 50% precision. The security gap was widest: 9 vs 19 verified bugs. Luna is good enough for everyday correctness bugs at that price, but not for auth or permission code on its own.

Why it matters: A controlled experiment on 50 real PRs comparing Luna and Astra on code review cost, recall, and false-positive rate — concrete and reproducible. Points off because it's a vendor self-test; the post doesn't disclose PR sources or a human reviewer baseline, so independence is u...

Hacker News front page

Nari Labs tops Coval voice AI benchmarks on latency and accuracy for both STT and TTS

Nari Labs placed both its Qwen3-ASR and Qwen3-TTS 1.7B models on the quality-latency Pareto frontier in Coval's voice AI benchmarks. STT hits 44 ms median time-to-final-segment with 3.6% WER, second only to AssemblyAI's 3.5% but at less than a quarter of the cost. TTS achieves 63 ms median time-to-first-audio and 3.8% WER, ranking first, while tying for the cheapest public price at $10 per 1M characters. The official Qwen3 TTS Flash Realtime endpoint scores 8.8% WER and 692 ms latency on the same benchmark, so Nari's serving stack makes a big difference. The post doesn't disclose the audio dataset makeup or p95/p99 tail latencies.

Sep 14Monday

Hacker News front page

The AI job market in 2026: who gets hired, what they earn, and which roles are fading

Maksim Ilin pulls together LinkedIn, WEF, Stanford, PwC, and Bain data for a September 2026 snapshot of the AI labor market. AI Engineer is the most-hired role; Research Scientist at a frontier lab is the most prestigious, with median pay around $746K at Anthropic and $1.15M at OpenAI L5. The fastest-growing niches are agentic systems and Forward Deployed Engineering—postings for the latter jumped over 1,000% YoY. Prompt engineer has faded as a job title; the skill remains but dissolved into other roles. AI skills now command a 62% wage premium in the US. Bain projects more than 1.3M AI jobs in the US by 2027 against roughly 645K available workers. In Europe, over half of AI postings sit outside tech departments, and Germany shows seven AI-user roles for every AI-developer role. Entry-level hiring got harder: employment for 22–25-year-olds in AI-exposed occupations now trails the rest by 19%.

Why it matters: A multi-source synthesis of the 2026 AI job market with concrete salary figures and role trends — high reference value for practitioners. Capped at 72 because it's a personal blog aggregating secondary data, not an original institutional report with primary research.

Hacker News front page

30 SVG prompts benchmark 2025–2026 LLMs on pelican-bicycle-style drawing tests

Tom Gally built a site with Claude Fable 5.1 that extends Simon Willison's pelican-riding-a-bicycle test into 30 SVG drawing prompts. The 2026 run covers six models—GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash, DeepSeek V4 Pro, Qwen3.8 Max, and Fugu Ultra v2—while the 2025 run includes ten models like Claude Sonnet 4.5 and GPT-5.1. Each image shows generation time and cost: DeepSeek V4 Pro finished in 1 min 36 s at $0.10, Qwen3.8 Max took over 12 minutes, and Fugu Ultra v2 cost $1.02. The post presents raw SVG outputs without subjective ratings, so you compare the drawings directly.

Why it matters: Simon Willison's pelican test is a community staple, and this expands it to 30 prompts across 6 models with timing and cost — dense, useful signal. The deliberate lack of subjective scoring means readers have to flip through images themselves, which costs it a bit of immediate...

Computing Life · Share · Yage

The AI Benchmark Yardstick Moved Faster Than the Models

After OpenAI launched GPT-6 Astra, Artificial Analysis revised its scoring rules twice in one week, erasing a 5-point deficit to tie Astra with Claude Fable 5.1—without any model update. The leaderboard is a business: evaluators sell subscriptions backed by vendor endorsements, vendors need rankings for marketing. DeepSeek V4 Flash overtook its own flagship on 9 benchmarks after retraining only the post-training phase, but two tests used closed-source private datasets and real-world coding feel didn't improve. The same model scored 62.7% vs 99.9% on ARC-AGI-3 depending on the execution harness. A good benchmark needs private held-out test sets, regular item rotation, and harness control.

Why it matters: A well-sourced industry commentary with concrete version numbers and score shifts, exposing how a benchmark vendor rewrote its scoring rules twice in one week after GPT-6 Astra's release, flipping the ranking from a 5-point deficit to a tie for first. Hits all three HKR axes a...

Sep 13Sunday

Hacker News front page

CadQuery vs OpenSCAD for AI-generated 3D-printable parts: a benchmark

ModelRift ran six AI agents on three printable parts each in CadQuery and OpenSCAD. All six STLs passed. The difference is failure modes: OpenSCAD is faster but can't inspect part edges; CadQuery supports B-rep validity checks but is slower. The hardest task, an M24 threaded adapter, both solved in one shot. OpenSCAD averaged 2,061 seconds vs CadQuery's 2,406. The post doesn't spell out whether agents used identical prompt templates, only that task descriptions were the same.

Sep 11Friday

Hacker News front page

When code is correct but sloppy: measuring LLM-generated bloat

Sebastian at Earendil applied SlopCodeBench metrics to measure AI-generated code bloat. Agent code averaged 0.33 verbosity vs. 0.15 for human repos, and 0.68 erosion vs. 0.31. In multi-round, context-cleared iterations, even SOTA models hit 0% strict pass rate—bad decisions compound. The simplest effective metric is LOC change, but it breaks under Goodhart's law. The post does not spell out which directions he plans to explore next.

Why it matters: Earendil's post quantifies AI code bloat with two novel metrics—verbosity and erosion—using their SlopCodeBench. Concrete data, fresh angle. Downside: it's a single blog post, not peer-reviewed, and the benchmark isn't open-sourced, so reproducibility is unclear. But the topic...

Sep 10Thursday

Hacker News front page

LRU is harder to beat than KV-cache papers suggest, tested on 393 Claude Code sessions

This repo replays 68k requests from 393 real Claude Code sessions to test agentic KV-cache eviction policies. LRU is harder to beat than papers claim—many new policies look good on paper benchmarks but fall apart on real agent traces. The post doesn't give exact hit-rate numbers, but the core finding is that real access patterns differ sharply from academic benchmarks. Don't rush to replace LRU in production.

Why it matters: Replays real agent traces to stress-test KV-cache eviction policies, directly pushing back on papers that only cite academic benchmarks. 393 sessions and 68k requests is solid scale, but the repo doesn't disclose specific hit-rate numbers, so the score stays at the featured th...

AI Chat-Group Daily (群聊日报)

Chat digest: Astra capacity crunch, DeepSeek V4.1 Flash benchmarks, Codex quota bug, and why xHigh saves more credits than Medium

OpenAI's Tibo publicly admitted unprecedented Astra demand and may pause new Pro subscriptions; users report lag even during off-peak hours and frequent WebSocket disconnects. DeepSeek V4.1 Flash scored 81.2 on OpenDesign's design benchmark—98% of Astra's quality at 1.4% of the cost—but the API's mandatory training clause and not-so-cheap real pricing gave users pause. A Codex quota display bug caused panic today; Tibo promised compensation but most users never got it. A counterintuitive finding: xHigh mode actually consumes fewer total credits than Medium because it plans more accurately and loops less. Also: Jacob Coxon quit with a warning about AI arms-race risks, Apple announced the foldable iPhone Duo starting around $2,800, and the Navier–Stokes proof cost roughly $15M in API fees.

AI HOT (Curated Pool)

OpenRouter launches Fusion: a compound model that debates across models before synthesizing a final answer

OpenRouter Fusion is a compound inference pipeline, not a new model. It fans out one prompt to up to 8 panelist models in parallel, has a judge compare their answers for consensus and blind spots, then lets the calling model write a final synthesis. A default three-model panel costs roughly 4–5× a single completion and takes 2–3× longer. On the DRACO deep-research benchmark, a budget panel of Gemini 3 Flash, Kimi K2.6, and DeepSeek V4 Pro scored ~64.7%, close to Claude Fable 5’s solo 65.3%. OpenRouter’s own test paired two Claude Opus 4.8 runs and saw a 6.7-point gain over a single run. The team positions Fusion as an escalation path for complex research and high-stakes decisions, not for simple chat. The post does not disclose per-token pricing, only the cost multiplier.

Why it matters: Fusion is a multi-model debate-and-synthesize workflow, not a new model. Concrete cost/latency numbers and DRACO benchmark data give it substance beyond marketing. But it's a routing-layer product update, not a foundation-model breakthrough — capped at the low end of featured,...

AI HOT (Curated Pool)

Cognition launches SWE-2 coding model, pushing the cost-performance frontier

Cognition's SWE-2 hits 50.0% on FrontierCode 1.1 Main, 8 points above SWE-1.7, at 64% lower cost than Fable 5.1. It's post-trained from the 2.8T-parameter Kimi K3 using a single-run RL method that trains all reasoning-effort levels together. The medium effort level solves tasks in 53 steps vs. 127 for SWE-1.7. The post doesn't disclose exact API pricing.

Why it matters: SWE-2 hits 50.0% on FrontierCode 1.1 Main, +8 pts over its predecessor, and costs 64% less than Fable 5.1 — Cognition's closest model to the frontier yet. Not an 85 because it still trails GPT-6 Astra by a few points and uses Kimi K3 as the base rather than an in-house model, ...

Sep 9Wednesday

AI Chat-Group Daily (群聊日报)

Astra effort tier benchmarks: low beats Sol high, web 6 Pro is the cheapest entry point

Tibo calibrated Astra: low now outperforms Sol high. A Codex CLI speed benchmark shows Astra API Fast hits 124 tps at $3.01 per 3 tasks, while Pro Normal costs almost nothing at 36 tps. Web ChatGPT 6 Pro can read GitHub repos and write code without consuming Codex quota—currently the cheapest Astra entry. GPT-6 triggers Computer Use more aggressively than previous versions. On the industry side, Tencent Hy4 topped OpenRouter weekly usage, H100 rental prices rose to $3.28/hr, and ByteDance plans to release a real-time spatial video world model next month.

Sep 7Monday

Computing Life · Share · Yage

AI raised the floor, but grading rubrics still penalize the ceiling

Two large-scale RCTs show the same pattern: AI lifts the floor of student work while present, but once removed, performance drops, and traditional rubrics actively penalize deeper reasoning. In a Turkish high school math experiment, ChatGPT-assisted practice scores jumped 48%, yet closed-book exam scores fell 17% below the control group. In a Milan business writing study, students who spelled out failure conditions and causal mechanisms received systematically lower grades. The floor is borrowed from external compute; the ceiling only grows when rubrics reward it.

Why it matters: Two large-scale RCTs with hard numbers expose the illusion of AI-assisted learning: practice scores soar but closed-book tests drop, and students copy answers without reasoning. Strong HKR, but it's a synthesis piece rather than a primary research release, so it stays below 85.

r/LocalLLaMA

Qwen 3.8-27B NVFP4 beats Q5_K_M and nears BF16 after tweaking temp and min_p

A user compared Qwen 3.8-27B NVFP4 (NInfer) against Q5_K_M (llama.cpp) on a 5090. With default sampling, Q5 led on IFBench strict (76% vs 74%), but NVFP4's loose score was 78%, meaning its failures were mostly minor formatting drift. After lowering temp to 0.9 and setting min_p to 0.05, NVFP4 hit 80% strict while Q5 stayed at 76%. The official BF16 baseline is 79.5% on 300 samples; NVFP4's 80% on 50 samples has variance but the trend held across reruns. NVFP4 also ran nearly 3x faster and used far less VRAM. The takeaway: default sampling hides NVFP4's real quality—tighten it and you get near-BF16 instruction following.

Why it matters: First-person benchmark with concrete numbers and a reproducible recipe — not armchair theory. Directly useful for local LLM users. Score capped because it's a single-GPU, single-model case relying on the niche NInfer engine, limiting generalizability.

Sep 6Sunday

AI HOT (Curated Pool)

OpenAI repeatedly revised GPT-6 Astra benchmarks after launch, hallucination rate briefly halved from 4.2% to 2%

Fortune reported that OpenAI changed multiple benchmark scores for GPT-6 Astra after the September 3 launch. Astra's hallucination rate dropped from 4.2% to 2% then reverted; Anthropic Fable 5.1's math score was briefly cut by nearly 10 points. OpenAI called it normal pre-release validation, but Stanford researchers noted the system card lacks details on the hallucination eval—not even the number of test items. Worth flagging: these are best-case scores under any compute budget, not what a typical ChatGPT user would see.

Why it matters: GPT-6 Astra's launch is already a top-tier event; Fortune catching post-launch benchmark revisions — including competitor score changes — hits all three HKR axes. Held below 95 because it's a single-source report so far and OpenAI's response is vague.

AI HOT (Curated Pool)

OpenAI GPT-6 Astra tops Code Arena WebDev leaderboard, 35 points ahead of Claude Fable 5.1

GPT-6 Astra (Max) hit 1797 points on the Code Arena WebDev leaderboard, 35 points ahead of Claude Fable 5.1 (Max) in second place. Claude Opus 5 (Max) scored 1688 in third. The post is a single tweet—no sample size, task breakdown, or latency info disclosed, so I'd discount the claim for now.

Why it matters: GPT-6 tops Claude on Code Arena WebDev for the first time — a 35-point lead is conversation-worthy. But the source is a single tweet with no sample size, task breakdown, or latency data, so the score stays below 80. Wait for more sources before adjusting.

Sep 5Saturday

AI Chat-Group Daily (群聊日报)

GPT-6 Astra opens to all: faster but pricier, with a concurrent rate-limit war

GPT-6 Astra rolled out to all Pro users, landing in Codex CLI and Copilot. Early tests show a task that took 12 minutes now finishes in 6, but per-task cost is ~75% higher than Sol—one API code review burned $100. Tibo and Anthropic both reset all user quotas the same day, while Codex patched an infinite-usage exploit. A detailed Cerebras benchmark reveals real-world agentic throughput is only ~357 tps vs. the advertised 1,500 tps; the same task cost $1.57 in 3 minutes versus ~$0.017 locally. Zhipu GLM-5.3-Flash hit just 20 tps on domestic inference cards, while the same weights on Ollama Cloud reached 70 tps. In industry news, the US is drafting rules to block Chinese access to overseas AI servers, DeepSeek plans to buy over 160,000 Huawei chips for inference, and Saudi Arabia's Humain M3 was exposed as a rebranded MiniMax M3.

Why it matters: GPT-6 Astra's full rollout is the week's biggest product move, and this chat digest delivers first-day speed and cost data with real numbers. The cap at 78 reflects the source being an anonymized group-chat compilation rather than a primary official post, and some details (e.g...

Hacker News front page

Artificial Analysis launches Intelligence Index v4.2 with private test sets to prevent gaming

Artificial Analysis updated its model benchmark to v4.2, adding two new evaluations: AA-Briefcase and GDP.pdf. AA-Briefcase uses a private test set to simulate multi-week knowledge work projects and assess holistic agentic capability. GDP.pdf requires models to synthesize evidence across 4,592 pages of professional documents, graded against 1,275 atomic criteria where a task passes only if every criterion is met. Claude Fable 5.1 leads the index, followed by GPT-6 Astra, which shows an ~85 Elo gain over GPT-5.6 Sol. Private test sets now account for 40% of the weighting, double the v4.1 figure, specifically to reduce gaming by labs.

Why it matters: AA's leaderboard refresh matters for model selection workflows — the private test sets and 4,592-page document eval are more grounded than saturated public benchmarks. Not scoring higher because this is methodology iteration, not a capability breakthrough, and the post only gi...

r/LocalLLaMA

Qwen3.8 27B on RX 7900 XTX: Ollama ROCm vs llama.cpp Vulkan benchmarks

A user benchmarked Qwen3.8 27B Q4_K_M on an RX 7900 XTX. Plain decode speed: llama.cpp Vulkan is only ~4% faster than Ollama ROCm (36 vs 34.4 t/s), while Ollama leads in prompt processing at 64K context (215.8 vs 192 t/s). The real gain comes from MTP (multi-token prediction): average generation jumps from 36 t/s to ~69-70 t/s, peaking above 80 t/s. This explains why community reports of 50-80+ t/s are mostly from speculative decoding, not raw single-token decode. The post does not disclose MTP + ngram combination results.

Hacker News front page

EEBench benchmarks whether AI can design real circuit boards that actually work

EEBench released a circuit design benchmark that uses declarative code (atopile) instead of GUI clicking, so models work directly on components and constraints. It runs SPICE simulations to check voltages, tolerances, cost, and real part availability. Claude Opus 5 leads at 61.6%, with Grok 4.6 at 57.1%. In one energy-meter task, a design picked a 22µF nominal cap that delivered only 11.4µF at 4.7V bias—far below the 545µF requirement—and failed. xAI already included EEBench in the Grok 4.6 model card under engineering acceleration.

Why it matters: EEBench benchmarked frontier models on circuit design using declarative code and SPICE simulation — Claude Opus 5 leads at 61.6%. Timing is sharp, landing right after GPT-6 Astra's KiCad demo. Score sits at the featured threshold because PCB design is niche for the general AI ...

Sep 4Friday

r/LocalLLaMA

GPT-6 Astra hit 98.6% on ARC AGI-3 — don't fall for the hype

OpenAI reported GPT-6 Astra at 98.6% on ARC AGI-3, but used a proprietary harness instead of the standard one. Nvidia already hit 100% with its AVO harness, and earlier systems like Arc-Skill and VISTA also neared perfect scores. Under the standard harness, Astra drops to 66%. That's still solid, but it's not AGI. The post doesn't spell out what OpenAI's custom harness changed, so I'd discount the 98.6% figure for now.

Why it matters: This post dismantles OpenAI's 98.6% narrative with two numbers — Nvidia's 100% and a 66% on the standard harness — high information density and strong conflict. Not scoring higher because the source is a Reddit individual post, not an institutional review, and the body doesn't...

AI Chat-Group Daily (群聊日报)

Flash models hit SOTA: Gemini 3.8 Flash and Muse Spark 1.3 launch, cheap models now cover 90% of tasks

Google launched Gemini 3.8 Flash at $0.75/M tokens input, scoring 71% on DeepSWE and beating Sol and Opus 5 on multiple agent benchmarks. Meta released Muse Spark 1.3 the same day, hitting 61–62 on AA Intelligence Index, matching Grok 4.6; Contributor tier costs just $0.10/$0.20 but trains on user data by default. A group member shared two-week usage stats: 1.28B tokens on GLM 5.3, with over 90% of tasks handled by cheap models. Uncle Bob proposed a multi-agent pipeline completing tasks in about one hour, insisting deterministic tools like tests and linters won't go away. GPT-6 confirmed for September 3 morning launch. LatePost exposed China's embodied AI funding bubble: among 22 companies valued over 10B RMB, one at 20B spent under 40M on R&D last year. NYC will ban student-facing generative AI tools for K-8.

Why it matters: Gemini 3.8 Flash launch with Flash-tier pricing beating Sol and Opus 5 on agent benchmarks. The source is a curated group chat digest, not a first-party announcement, which caps the score slightly, but the signal density and real-world testing notes are solid.

AI HOT (Curated Pool)

OpenAI launches GPT-6 Astra, focused on computer use and alignment

GPT-6 Astra can operate across apps, build test software, and tackle open science problems. OSWorld real-desktop task time dropped from 75 to 40 minutes, and workplace automation rose from 18% to 41%. On alignment, unguarded jailbreak rate fell from 48% to 0%. The author says $2,000 in compute solved 10 decade-old math and theoretical CS problems, but tool-augmented benchmarks still trail Claude.

Why it matters: GPT-6 Astra launch is an industry-shaking event. The computer-use and 0% jailbreak numbers are concrete, hitting all three HKR axes. Score not at 98-100 only because we currently have a tweet summary without an official blog or third-party verification; can bump higher once mo...