Skip to content

#评测/基准

2 today

Today · Sep 30Wednesday · 2 items

AI HOT picks · Models

GPT-6.1 Sol replaces GPT-6 Sol 7 days after launch, 1 point behind GPT-6 Astra on intelligence index

GPT-6.1 Sol replaced GPT-6 Sol 7 days after launch. Its intelligence index is 1 point below GPT-6 Astra, and pricing stays at $2 per million input tokens and $10 per million output tokens. The cached-read discount rises from 90% to 95%.

Why it matters: Side-by-side intelligence index, cost and token-efficiency figures let readers judge the trade-off GPT-6.1 Sol makes between price and capability.

AI HOT picks · Models

Claude Sonnet 5.5 (High) 以 1699 分登 Code Arena: WebDev 第 4 名

Arena 公布 Claude Sonnet 5.5 (High) 在 Code Arena: WebDev 以 1699 分排第 4,混合价格约 $8 per Mtoken,比第 2、3 名便宜 80%。相较 Sonnet 5 (High) 的 1540 分提升 159 分,Reference-Based Design、Simulations、Gaming 均从第 30 多名升至第 4。

Yesterday · Sep 29Tuesday

Hacker News front page

Pac-Bench: One-shot Pac-Man benchmark, Claude Opus 5.5 scores 99/100

Jon Clegg built a Pac-Man benchmark: one prompt, one HTML page, scored automatically by Opus 5.5. Claude Opus 5.5 hit 99/100 via Claude Code at $1.99, generating a 10.8 KB page in 9 minutes with near-arcade audio. Claude Fable 5.1 scored 96 but cost $5.87. Grok 4.7 and GPT-5.6-sol scored 94 and 90; the latter cost just $0.72 in under 5 minutes. Scoring covers controls, ghost behavior, stuck detection, maze layout, and sound. The post doesn't explain why some models ran Phase 2 or how much the harness affects scores. Worth noting: this measures model-plus-toolchain combos, not bare model capability.

Why it matters: A 30-model Pac-Man benchmark with Claude Opus 5.5 hitting 99/100 via Claude Code at $1.99 is solid signal. Capped at 78 because it's an individual project, not an official release, so authority is limited despite strong HKR.

AI HOT (Curated Pool)

GPT-6 Luna (Max) ranks #23 on Agent Arena at $0.05 per task

Arena evaluated GPT-6 Luna (Max) on 8K real agent conversations. It ranked #23, up 6 spots from GPT-5.6 Luna (xHigh), with a net gain of +1.6%. Cost is $0.05 per task. The post doesn't disclose latency or task-type breakdown.

Why it matters: First Agent Arena run for GPT-6 Luna with a real ranking and cost figure. But the post doesn't disclose latency or task breakdown, so real-world usability is unclear — score sits at the featured threshold.

r/LocalLLaMA

95+ TPS through 100K tokens on Qwen 27B with a single 3090

A Reddit user reports running Qwen3.8 27B on a single RTX 3090, achieving 95+ tokens/sec throughput through 100K generated tokens with a 262K context window. This suggests local long-text generation is nearing practical speeds. However, the post body is blocked by Reddit, so the implementation details—quantization, inference framework, or whether this is a real benchmark—are not disclosed.

AI HOT (Curated Pool)

Claude Sonnet 5.5 enters Arena's Agent Arena and Battle Mode

Anthropic's Claude Sonnet 5.5 is now available for voting in Arena's Agent Arena. The leaderboard evaluates models on millions of real-world long-horizon agent tasks where models can use web search, filesystem, and terminal tools. Rankings use a causal tracking method to measure how much a model outperforms the average.

Why it matters: Claude Sonnet 5.5 hitting Arena's Agent leaderboard is a direct user-facing eval signal, hitting all three HKR axes. Score capped at 74 because the post only describes the methodology — no specific win rates or rankings disclosed. Adjust upward once concrete numbers drop.

Hacker News front page

Anthropic launches Claude Sonnet 5.5: 30%+ faster, up to 30% cheaper than Sonnet 5

Claude Sonnet 5.5 is the second model in the 5.5 family, aimed at everyday coding, bug fixes, and polished docs. It scores 70.6% on Terminal-Bench 4.0 vs. Sonnet 5's 10.3%. Pricing stays at $2/$10 per million input/output tokens, but it uses fewer tokens per task, cutting per-task cost by up to 30%. Speed is up 30%+. For the first time, a Sonnet model ships with cyber safeguards because its cybersecurity capabilities now match Opus 5. Haiku 5.5 is coming in a few weeks.

Why it matters: Anthropic officially released Claude Sonnet 5.5, the second model in the 5.5 family. Terminal-Bench jumped from 10.3% to 70.6%, 30% faster with 30% lower per-task cost at unchanged pricing. A same-day must-write model update. Not 95 because it's a complement to Opus 5.5, not a...

AI HOT (Curated Pool)

Claude Opus 5.5 (High) hits #2 on Agent Arena and reshapes the Pareto frontier

Anthropic's Claude Opus 5.5 (High) landed at #2 on Agent Arena with a +12.15% net improvement, behind only Fable 5.1 (Max). Median cost is $1.31 per task—40% cheaper than Opus 5 (High) and 56% cheaper than Opus 5 (Max). It ranked #1 on Steerability at +14.50%. The post doesn't disclose a release date or other model comparisons.

Why it matters: Anthropic model hitting #2 on Agent Arena with a significant price drop is a same-day must-write product signal. The +12.15% net improvement and $1.31 median cost provide hard data, and steerability gains are a bonus. Not scoring higher because this is still a benchmark — real...

AI HOT (Curated Pool)

Claude Opus 5.5 (High) hits #2 on Agent Arena, costs 56% less than Opus 5 (Max)

Anthropic's Claude Opus 5.5 (High) reached #2 on Agent Arena with a +12.15% net improvement, behind only Fable 5.1 (Max). It also costs 56% less than Opus 5 (Max). The post doesn't disclose exact pricing or latency—I'd discount the cost claim until we see real usage numbers.

Why it matters: Opus 5.5 landing #2 on Agent Arena with a claimed 56% cost cut makes it a notable Anthropic update today. Score capped below 85 because the post omits pricing and latency — the cost advantage needs real-world confirmation.

Sep 28Monday

Hacker News front page

TabPFN with zero training beats tuned XGBoost on 14 out of 14 tables

The author tested TabPFN and TabICL against tuned XGBoost on 14 datasets from the Grinsztajn benchmark. Both models skip gradient training on the target table, using training rows as context in a single forward pass, and won all 14 matchups. The advantage holds up to 32,000 rows. Inference latency ranges from 0.6 to 6 seconds per row, with wider tables costing more. The post also notes that the most-cited TabPFN version now requires an account to download. The XGBoost tuning budget and hyperparameter search space are not detailed in the article.

Computing Life · Share · Yage

Agents mentioned PayPal 139 times, picked it 0 times: software distribution is changing

Armature ran 5,300 sandbox sessions where Claude Code, Codex, and Cursor integrated payment and email tools into real engineering repos. PayPal got 139 mentions and zero code commits; Stripe won 88% of payment tests. Tool choice follows language stack: the same email task picked Resend in TypeScript, SendGrid in Python, Postmark in Go, and Azure ACS in Java. Big markets concentrate heavily; long-tail categories split three ways across agents. Armature sells agent-adoption optimization starting at $5,000/month and disclosed the conflict upfront. The post doesn't spell out how the 42% selection-consistency figure was calculated; our own recomputation across reasonable definitions landed at 41.5%–48.2%, consistent with the published number. Agent tool selection is a context-dependent function, not a global ranking—traditional SEO logic breaks here.

Why it matters: Armature's 5,300 sandbox runs deliver first-hand data on how agents pick tools — the PayPal 0 vs Stripe 88% contrast is solid. All three HKR axes hit, but the testing org is a commercial entity and only 31% of data is released so far, capping it below 85.

Sep 27Sunday

Sep 26Saturday

Hacker News front page

LLM watermarking degrades AI agent performance and speed

Lasso Security tested LLM watermarking on AI agents and found it hurts tool-calling accuracy by 3–5 points on BFCL V3, adds 11 seconds of latency, and increases output length by 20–30%. The post doesn't name the specific models or agent frameworks tested.

Why it matters: Has concrete benchmark data answering a production-relevant question: what's the performance cost of watermarking on agents. But the post doesn't disclose which models/frameworks were tested, so the numbers are directional only — hence the score cap.

Sep 25Friday

Hacker News front page

NSA is spending billions this year to test frontier AI models, far above prior estimates

Two sources say the NSA told lawmakers in a classified briefing that it is spending billions in taxpayer money this year to evaluate and test advanced AI models. The figure is far higher than previously known, leading lawmakers to estimate a full federal AI regulatory system could cost tens of billions per year. Trump has mostly resisted stronger federal AI oversight. The article does not name which models are being tested, whose compute is used, or how the money breaks down.

Why it matters: Exclusive disclosure of a classified budget figure with solid information density; hits all three HKR axes. Held below 85 because the body doesn't disclose which models are tested, whose compute is used, or how the money breaks down — major factual gaps mean it's a policy sign...

Computing Life · Share · Yage

Four report cards in one week: Step 5 Preview benchmarks, Jev's confidence blind spot, and what California's AI order actually requires

StepFun's Step 5 Preview posted four types of scores, but its self-reported Terminal-Bench result isn't on the public leaderboard yet, and weight license terms remain undisclosed. An independent dev tested TypeSafe's Jev classifier on 111 public hard questions: Jev's average confidence was 0.690 when wrong and 0.684 when right—confidence doesn't separate correct from incorrect. California's AI executive order only directs state agencies to draft legislative proposals; the only active law binding companies is 2025's SB 53.

Sep 24Thursday

Hacker News front page

A daily-updated LLM value chart that plots price against intelligence to find the frontier

The site plots 420 models from the Artificial Analysis Intelligence Index against blended API price, drawing a value frontier where no cheaper model is smarter. Claude Opus 5.5 leads at $8/1M tokens with a 57.6 intelligence score. Meta's Muse Spark 1.3 tops the $2–$8 band at 48.1, Xiaomi's MiMo-V2.6-Pro wins $0.54–$2 at 46.3, and Z AI's GLM 5.3 Flash takes the under-$0.24 tier at 41.8. The post doesn't disclose how the intelligence index is built, and Coding/Math sub-scores are listed as empty for many models, so I'd hold off on those comparisons.

Why it matters: A daily-updated price-performance leaderboard using Artificial Analysis data — genuinely useful for model selection. Hits H and K, but lacks the controversy or identity hook for R, so it lands at the featured threshold of 72.

Hacker News front page

Paper Instruments releases Paper Office, a Python suite for agents to safely edit Word, PowerPoint, and Excel files

Paper Office is a suite of Python packages that wrap python-docx, python-pptx, and OpenPyxl with safety checks and broader editing capabilities. Across 5 models and 61 tasks, Paper packages plus guidance passed 92.5% of trials, vs 80.7% for the upstream packages alone and 69.5% for Anthropic's Office skills. Agents resorted to raw OOXML editing in only 1.6% of Paper runs, compared to 78.7% without skills and 50.5% with Anthropic skills. The team argues that low agent adoption in consulting, law, and banking stems from tools that silently corrupt formatting, break references, or produce client-unready output. Paper Office keeps the familiar imports and adds cross-run text search, native Word redlines, comment threads, content controls, cross-document composition, and package-level diff saves, refusing unsafe operations instead of quietly breaking files.

Why it matters: Paper Instruments open-sourced a suite of Office file-editing libraries for agents, adding safety checks and broader editing capabilities on top of python-docx and friends. Across 5 models and 61 tasks, they hit 92.5% pass rate — 11.8 points above bare upstream libs and well a...

Hacker News front page

Open-source prompt-injection detectors catch 0–1% of realistic AI agent attacks buried in tool output

This benchmark hides 629 AgentDojo injection attacks inside tool outputs and tests Regex vs. Meta Prompt Guard 2. Regex catches 0%, Prompt Guard 2 catches 1%. The attacks aren't sent directly to the model—they're buried in search results, email bodies, and similar tool responses, so current detectors are effectively blind. Code and reproduction steps are public; the post doesn't include comparisons with commercial detectors.

Why it matters: 629 AgentDojo attacks buried in tool output, Regex catches 0%, Prompt Guard 2 catches 1%. Cleanly exposes the blind spot in indirect injection detection. Code and repro steps are public, which adds practical value. Held at 78 because it's a single benchmark without cross-detec...

AI Chat-Group Daily (群聊日报)

Opus 5.5 effort blind test: high mode costs 30% more tokens but catches real bugs tests miss

A double-blind test on real PRs shows Opus 5.5 high mode costs ~30% more tokens and 1.33× time vs medium, but wins 16 vs 7 in blind review by catching real bugs tests missed. Claude Code Cloud Sessions goes GA with $100 Pro / $250 Max trial credits. HLE-Diamond benchmark updated: GPT-6 Astra leads at 60.6%, Gemini 3.8 Flash surprises at 34.3% beating GPT-6 Sol. Muse phone calls were partly handled by human contractors; Meta rolled back the test. The newsletter's generation tool is now open source.

AI HOT (Curated Pool)

Claude Opus 5.5 tops Coding Agent Index, but per-task cost rises to $13.04

Artificial Analysis tested Claude Opus 5.5 under Claude Code max effort and it scored 66 on the Coding Agent Index, up from Opus 5's 60. All three subtests improved: Terminal-Bench 4.0 63.1%, DeepSWE v1.1 68.4%, SWE-Atlas-QnA 66.4%. The trade-off: per-task cost jumped from $3 to $13.04. The post doesn't break down how max effort drove the cost increase.

Why it matters: Claude Opus 5.5 tops the Coding Agent Index with a 6-point jump to 66, but $13.04 per task is the hard number. Anthropic substantive update + independent third-party benchmark + concrete data — all three HKR axes hit. Not scoring higher because this is a single benchmark, not ...

AI HOT (Curated Pool)

MiMo-V2.6-Pro, Claude Opus 5.5, GPT-6 Luna/Sol launch, shifting the intelligence–cost Pareto frontier

Artificial Analysis reports that four new models this week added 11 points on the Intelligence Index vs. cost-per-task Pareto frontier. GPT-6 Luna contributed 5 points, Claude Opus 5.5 contributed 4, and the remaining two came from MiMo-V2.6-Pro and GPT-6 Sol. The post doesn't disclose specific scores, pricing, or latency—hold off on conclusions until full benchmarks drop.

Why it matters: Artificial Analysis's Pareto frontier chart is a hard reference for model selection — four new models landing 11 points at once pushes the boundary out meaningfully. GPT-6 Luna taking 5 points suggests competitiveness across cost tiers; Claude Opus 5.5's 4 points aren't far be...

Sep 23Wednesday

OpenAI News

OpenAI releases MentalHealthBench, an open benchmark co-developed with 80+ licensed clinicians to evaluate AI in realistic mental health conversations

OpenAI open-sourced MentalHealthBench, a benchmark built with over 80 licensed psychologists and psychiatrists across 22 countries. It tests AI on realistic mental health conversations ranging from everyday stress to emergencies, covering adults, teens, and caregivers. The eval goes beyond safety filters: it checks whether models seek context, preserve user agency, and offer actionable guidance when appropriate. OpenAI stresses ChatGPT isn't a substitute for therapy, but the benchmark tracks progress on empathy and steering people toward real-world support. The paper and benchmark are publicly available.

Why it matters: OpenAI released an open mental health benchmark built with 80+ licensed clinicians, covering a wide range of scenarios with finer evaluation dimensions than typical safety tests. It's directly useful for AI safety and product teams. Not scoring higher because it's an eval tool...

Latent Space

Claude Opus 5.5 launches with Fable 5.1-level performance at 40% lower cost, plus a rare focus on writing quality

Anthropic released Claude Opus 5.5, the first model in the new 5.5 family. It matches Claude Fable 5.1 on most tasks, costs 40% less to run than Opus 5, and is about 30% faster. The launch unusually highlights writing improvements: the model puts key info up front and follows user style rules. Artificial Analysis notes that token usage on frontier tasks jumped ~80%, so per-task cost remains around $6—similar to Opus 5. OpenAI shipped GPT-6 Sol and Luna an hour later at 50% lower prices than GPT-5.6, but Opus 5.5's launch post hit 17M views and dominated the day. Anthropic's system card also reports multi-agent scaling with up to 100 parallel agents for the first time. Latent Space tested both and switched to Opus 5.5 as the default model immediately, calling the writing quality a night-and-day difference over Sol 6.

Why it matters: Anthropic drops the first model in a new flagship family, claiming Fable 5.1 parity at 40% lower cost, with writing improvements front and center — a directly actionable upgrade signal for heavy Claude users. Held below 90 because we only have the official claim and Latent Spa...

Computing Life · Share · Yage

Same tool toggle: Nemotron-3 550B gained, Mistral-Medium-3.5 crashed

A new paper breaks down coding agent harnesses into three independent toggles and measures each one. The most striking result: switching from dedicated file tools to a pure CLI made Nemotron-3 550B's SWE-Bench Verified score jump 3.6 pp while cutting per-task cost from $2.33 to $1.11, but Mistral-Medium-3.5-128B dropped from 68.60% to 45.40%. Trajectory analysis shows 550B composing dense shell one-liners, while Mistral failed to locate files in 32.80% of tasks and submitted no edits. On Terminal-Bench 2.1, both models improved under CLI mode. Planning boosted the 30B model from 13.60% to 25.20% but only saved ~30% cost for larger models without accuracy gains. Context management mainly prevents window overflow; at 128k the gap shrinks to 2.7 pp, and complex read-back mechanisms were almost never invoked. The takeaway: no universal best harness design—it depends on the model's CLI fluency and the task type.

Why it matters: A controlled experiment that isolates three harness design switches and shows Nemotron-3 and Mistral-Medium-3.5 reacting in opposite directions, with concrete numbers and engineering takeaways. Not an 85 because it's a single preprint without cross-source cluster yet, but HKR ...

AI HOT (Curated Pool)

Arena launches GPT-6 Sol and GPT-6 Luna testing, scores coming soon

Arena is now testing two new OpenAI models, GPT-6 Sol and GPT-6 Luna, with scores not yet released. You can try them on real agent tasks and vote to feed the leaderboard. The post doesn't disclose model size, release date, or pricing.

Why it matters: GPT-6's first public appearance, two variants live on Arena running agent tasks — strong suspense and signal. Deduction for thin info: no scale, pricing, or release date disclosed, just a test entry point.

Hacker News front page

Unreal Agent: async harness cuts agent costs by 40% on GPT-6 Astra

Unreal Labs open-sourced an agent harness that makes tool calls fully asynchronous: the model issues a call and moves on while the tool runs in the background, with results appended later. On Terminal-Bench, SWE-Atlas, DeepSWE, and ALE-CLI with GPT-6 Astra xhigh, it costs up to 40% less than Codex and up to 20% less than Pi, with pass rates roughly equal. The post doesn't report latency numbers or results with non-GPT-6 models.

AI HOT (Curated Pool)

OpenAI launches GPT-6 Sol and Luna, API pricing cut 50% vs GPT-5.6

OpenAI added two cheaper models to the GPT-6 family: Sol and Luna, with API prices halved across input and output. Sol costs $2/$10 per 1M tokens, Luna $0.10/$0.50. Sol scored 33.2% on AutomationBench at xhigh effort at 9% of Claude Opus 5's cost per task, and 56.4% on Agents' Last Exam at max effort at 60% lower cost. On internal factuality evals, Sol makes about half as many mistakes as its predecessor. The post does not specify a launch date beyond 'available now.'

Why it matters: Official OpenAI release of new GPT-6 models with a 50% API price cut and Sol's agent benchmark cost at 9% of a competitor — industry-shaking. HKR all hit, with solid pricing and benchmark data. Minus 3 points because the post doesn't fully detail the capability gap between Sol...

AI HOT (Curated Pool)

Claude Opus 5.5 lands on Arena's Agent Arena and Battle Mode

Anthropic's Claude Opus 5.5 is now available on Arena's Agent Arena, where users vote on rankings after the model runs real long-horizon agent tasks. The model can use web search, a file system, and a terminal; the leaderboard uses causal tracking to measure performance relative to the average model. The post doesn't spell out Battle Mode specifics or show example tasks.

Why it matters: Opus 5.5 landing on Agent Arena is the most watchable third-party eval signal this week. The causal-tracking leaderboard design carries more info than raw win rates, but the post doesn't give concrete task examples or Battle Mode rules — real performance waits on community tes...

AI HOT (Curated Pool)

Claude Opus 5.5 tops Artificial Analysis Intelligence Index with a score of 58, plus a 20% price cut

Claude Opus 5.5 scored 58 on the Artificial Analysis Intelligence Index, the highest measured so far. It leads on 6 of 10 evaluations, including Humanity's Last Exam at 61.4% and SciCode at 66.9%, and matches GPT-6 Astra (xhigh) on Terminal-Bench 4.0 at 59.6%. On the agentic knowledge-work eval AA-Briefcase, it hit 1822 Elo—143 points above Fable 5.1—and surpassed GPT-5.6 Sol on both analytical quality and presentation. Pricing dropped to $4/$20 per 1M input/output tokens (from $5/$25), with cache reads down 60% to $0.20. Output tokens per task grew ~60% vs Opus 5, so cost per task stayed flat. Context window remains 1M tokens with image and text input.

Why it matters: Anthropic's flagship tops a major third-party benchmark with a price cut — a same-day must-write. Not a 95 because it's a benchmark result, not a model launch, but 6/10 leads, parity with GPT-6 Astra, and a 20% price drop make it a clear featured pick.

Sep 22Tuesday

Hacker News front page

Prompting agents to iteratively optimize Rust until it beats state-of-the-art libraries

Max Woolf spent over a year testing whether agentic LLMs can iteratively optimize Rust code with a hard pass/fail rule: each iteration must deliver at least a 5% speedup or roll back. With Claude Opus 4.5 and later models, he got 2–20× speedups on algorithms like UMAP versus mature libraries, all without unsafe code. The post includes the exact prompts and benchmark results. The caveat: these gains are measured on his specific benchmarks and may not translate directly to production workloads.

Why it matters: Max Woolf spent a year validating a concrete agentic-iteration loop for Rust optimization with a hard ≥5% speedup rule, reproducible benchmarks against mature libraries, and zero unsafe code. The post includes prompts and results — a rare first-person experiment with numbers. ...

AI HOT (Curated Pool)

Kazike tests Grok 4.7 vs Xiaomi MiMo V2.6: the latter is the answer to the impossible triangle

The body does not disclose any test details. The title says Kazike compared Grok 4.7 with Xiaomi MiMo V2.6 and concluded that MiMo V2.6 is the answer to the 'impossible triangle'. However, the article was blocked by WeChat, showing only an environment anomaly and verification page, with no model parameters, test methodology, or specific results.

Computing Life · Share · Yage

Three AI Coding Stories, Three Numbers to Read

ZCode was caught silently packaging entire Git histories for cloud upload, with .git objects making up 86.6% of snapshots. HarnessTax benchmarked Claude Fable 5 across frameworks: Claude Code and minimal Pi achieved near-identical success rates but a 2x cost gap. A Microsoft architect replaced multi-model agent loops with a single model reading skill docs and calling tools directly—halving API calls but increasing total tokens by 22%. Each story unpacks one number and a reminder to check which layer a metric actually measures.

Why it matters: Three distinct AI coding stories bundled into one piece. The ZCode silent snapshot upload is a hard security story backed by ferstar's forensic report and community reproduction. HarnessTax benchmark and Microsoft architecture case add billing and engineering angles. HKR all h...

OpenAI News

OpenAI Publishes Priorities and Principles for Third-Party Safety Assessments

OpenAI outlines four priority areas for third-party safety assessments: safety case review, critical safeguard evaluation, capability evaluation, and deployment monitoring. The post stresses independence, scientific rigor, and security, and defines 'safety claim' and 'safety case.' It does not name specific assessors or timelines, but notes assessments may last weeks to months.

AI HOT (Curated Pool)

Xiaomi MiMo-V2.6-Pro hits ~10th on Code Arena WebDev, ~3rd among open-weight models

Xiaomi released two omni-modal models: MiMo-V2.6-Pro and Flash. The Pro version scored 1628 on Code Arena WebDev, up 153 points from MiMo-V2.5-Pro's 1475, landing around 10th overall and ~3rd among open-weight models under MIT license. The post doesn't disclose Flash's benchmark numbers or parameter counts.

Why it matters: Xiaomi's multimodal model hits ~10th on Code Arena WebDev and ~3rd among MIT open-weight models, with a 153-point gain for Pro. Flags a domestic flagship release with concrete benchmark data. Flash variant lacks params and scores, capping it below 80.

AI HOT (Curated Pool)

Xiaomi MiMo-V2.6-Pro tops open-weight model intelligence index

Xiaomi released MiMo-V2.6-Pro, scoring 46 on the Artificial Analysis Intelligence Index—up from 26 for the previous V2.5-Pro. It's now the highest among open-weight models. The post doesn't disclose parameter count, architecture details, or a release timeline.

Why it matters: Xiaomi's model hits #1 on the open-weight intelligence index with a near-doubling of score — triggers the domestic flagship model positive signal. Missing param count and release timeline keep it from scoring higher.

AI HOT (Curated Pool)

Xiaomi MiMo-V2.6-Pro tops open-weight model intelligence index

Xiaomi MiMo-V2.6-Pro scored 46 on the Artificial Analysis Intelligence Index, the highest among open-weight models. The previous MiMo-V2.5-Pro scored 26. The post doesn't disclose model size, training data, or release license, so hold for details.

Why it matters: Xiaomi's model tops the open-weight leaderboard with a near-doubling of its intelligence score — newsworthy. But without model size, training data, or license details, real-world usability is unclear, capping the score at 78 until more info drops.

AI HOT (Curated Pool)

Musk says Grok 4.7 puts xAI third in agentic coding

Elon Musk cites Artificial Analysis to claim Grok 4.7 ranks xAI third in agentic coding, behind only Anthropic and OpenAI. The post doesn't disclose the benchmark's metrics, scores, or version comparisons—only the ranking and competitors.

Sep 21Monday

AI HOT (Curated Pool)

Fireworks AI launches FireRouter: the frontier isn't a model, it's a router

Fireworks AI benchmarked 18 models on DeepSWE: picking the right model per task beats any single model. GPT-6 Astra alone scores 74.1% at $6.52/task. An oracle router across all 18 hits 97.6% at $1.88. Open-weight models alone reach 90.3% at $1.45. 94 of 113 tasks need a model under $3; the three priciest models are the best pick on only 3 tasks. FireRouter aims to make that per-task choice before the work starts—the post doesn't yet detail how.

Why it matters: Fireworks presents a data-backed argument using 18 models on DeepSWE: task-aware routing beats the single strongest model on both accuracy (97.6% vs 74.1%) and cost ($1.88 vs $6.52). The open-source-only result of 90.3% also provides a path that doesn't depend on closed models...

Sep 20Sunday

Hacker News front page

Prompts Aren't Real: Build Evaluation Pipelines Instead

Dan McKinley argues that prompt engineering is a distraction. Building consumer-facing agents taught him that even structured output fails on a fraction of requests—models will flood a field with nonsense. His fix was renaming a field from 'title' to 'heading,' which he calls deranged. The talk pushes for pass^k testing and evaluation pipelines to constrain behavior, since prompts alone can't tame the beast. The post is a slide deck; it names no specific eval frameworks or metrics.

Why it matters: Dan McKinley's first-hand production experience with concrete cases and numbers, sharp opinion. But it's a personal talk, not a formal publication, and the post doesn't disclose pass^k test pass rates or scale — slight deduction.

Hacker News front page

iPhone 18 Pro scores 172 on DXOMARK camera test, ranks second globally

DXOMARK tested the iPhone 18 Pro camera and gave it 172 points, placing it second globally. The variable aperture is the standout feature, keeping sharpness in complex scenes. Autofocus is solid, and flare is better controlled than last gen, though green spots can still appear when the iris is closed. Main: 48MP f/1.48–f/4.0 variable aperture, ultrawide: 48MP 120°, tele: 48MP 8x optical zoom.