Skip to content

#评测/基准

0 today

May 26Tuesday

Import AI (Jack Clark)

Import AI 458: Reckoning with the Future; and a Singularity Story

Jack Clark’s Import AI 458 excerpts his 2026 Cosmos HAI Lab Lecture, cites the Epoch Capabilities Index across 40-plus benchmarks, and argues that an AI system able to develop its own successor may arrive within two years or sooner.

Why it matters: HKR-H/K/R all pass: Jack Clark pairs ECI’s 40+ benchmarks with a two-year successor-system claim, giving this AGI-timeline essay both concrete detail and debate fuel.

May 24Sunday

Xinzhiyuan · WeChat

AI-generated articles now outnumber human-written ones: what is left for the brain?

Graphite sampled 43,000 CommonCrawl articles and found AI-generated English articles exceeded human-written ones from November 2024, with its detector reporting about a 4.2% false-positive rate and 0.6% false-negative rate.

Why it matters: HKR-H/K/R all pass: the article has a sharp web-content crossover claim, concrete sampling/error numbers, and clear data-quality resonance. Single-study sourcing and no platform-level impact keep it below the 78 band.

May 22Friday

AI HOT (Curated Pool)

Text Degeneration: A Production Failure Mode Most Benchmarks Do Not Track

Dharma-AI says in a Hugging Face post that large language models can produce repeated, incoherent, or logically confused text in production, and most mainstream benchmarks do not track this failure mode.

Why it matters: HKR-H/K/R all pass, but the post only discloses the failure pattern and benchmark blind spot, with no sample size, metric, or reproduction setup. This fits the lower featured threshold.

May 19Tuesday

AI Chat-Group Daily (群聊日报)

May 18, 2026 Chat Group Daily

The chat group daily says AI21 Labs cut 60% of staff and stopped selling model access, and cites a University of Waterloo paper where GPT-5.4 accuracy dropped from 100% to 23% after false peer-consensus injection; the snippet also mentions Meta layoff talk at 10%, but does not disclose source details or confirmation conditions.

Why it matters: HKR-H/K/R all pass: AI21’s 60% layoff and model-sales stop signal lab contraction, while GPT-5.4 falling from 100% to 23% under false peer consensus is a concrete safety hook. The chat-digest source keeps it at 78.

May 17Sunday

Dwarkesh Patel podcast

Notes on Pretraining Parallelisms and Failed Training Runs

Dwarkesh documents pretraining failure modes and parallelism tradeoffs: expert choice and token dropping can break causality in MoE routing, FP16 collectives can bias repeated additions after values exceed 1024, pretraining FLOPs are given as 6ND, B300 HBM is listed as 288GB, and FSDP communication can reach params × 3 with reduce-scatter.

Why it matters: HKR-H/K/R all pass: Dwarkesh’s notes expose concrete pretraining failure modes and numbers. The systems-training focus is specialized, so it sits in the high-quality band rather than same-day must-write.

May 15Friday

The Verge · AI

AI research papers are getting better, and it’s a big problem for scientists

The Verge describes Peter Degen investigating unusual citations to a 2017 paper: it rose from a few dozen citations over several years to being cited every few days, while the RSS snippet does not disclose the full sample size or review findings.

Why it matters: HKR-H/K/R all pass: the paradoxical angle, named investigation, and citation spike give it signal. The post lacks full sample size, so it stays in the lower featured band rather than becoming must-write.

May 11Monday

QbitAI · WeChat

Math Majors in Trouble: Fields Medalist Tests ChatGPT 5.5 Pro, Gets Paper-Level Result in 17 Minutes

Timothy Gowers tested ChatGPT 5.5 Pro on additive number theory problems, where it produced an optimal quadratic upper-bound construction in 17 minutes 5 seconds, then generated a LaTeX preprint in 47 minutes; the article says arXiv rejects AI-generated content, so the result remains on Gowers’s blog.

Why it matters: All three HKR axes pass: Gowers’ first-person test, 17m05s, and a 47-minute preprint are concrete and discussable. It is not a model release, but the named experiment and math-reasoning impact put it in the must-write band.

May 10Sunday

Xinzhiyuan · WeChat

Harsh Claim: Top Silicon Valley AI Is One Year Ahead of the World

Elad Gil claims top AI lab employees are 3-4 months ahead of Silicon Valley, while Silicon Valley is 3-6 months ahead of New York; the post cites Mythos’ 73% success rate in expert cyberattack simulations as evidence in a disputed “geographic time gap” argument.

Why it matters: HKR-H/K/R all pass: the lab-to-user lag hook is clickable, and the post cites 3–4 months, 3–6 months, and a 73% Mythos figure. It is secondhand commentary, not a model or product release, so it stays in the 72–77 threshold band.

May 8Friday

Ruan YiFeng's Weblog

Technology Enthusiast Weekly Issue 395: The Third Way of Software Development

Ruanyifeng Weekly issue 395 frames AI-assisted coding as a “mystery house” style of software development and cites HN SOTA, which ranks model popularity by scanning 200 top Hacker News topics each day and their programming or AI discussions.

Why it matters: HKR-H/K/R pass: the “third way/mystery house” framing, HN SOTA’s 200 daily HN topics, and developer workflow anxiety all land. It is commentary, not a model or product release, so it stays at 72.

May 5Tuesday

Synced · WeChat

Anthropic cofounder says AI self-improvement has a 60% chance by 2028

Anthropic cofounder Jack Clark says human-free AI R&D has over a 60% chance by end-2028. He cites SWE-Bench, CORE-Bench, MLE-Bench, and PostTrainBench: Claude Mythos Preview reaches 93.9% on SWE-Bench, and Opus 4.5 reaches 95.5% on CORE-Bench. The key signal is longer task horizons and post-training capability, not the “singularity” framing.

Why it matters: HKR-H/K/R all pass: a named Anthropic cofounder gives a 2028 timeline, backed by benchmark numbers. The headline is overheated, but the concrete claims and practitioner stakes justify P1.

May 4Monday

Import AI (Jack Clark)

Import AI 455: Automating AI Research

Jack Clark argues that no-human-involved AI R&D has a 60%+ chance of arriving by the end of 2028, citing SWE-Bench gains from Claude 2 at about 2% to Claude Mythos Preview at 93.9%, plus METR task horizons rising from 30 seconds in 2022 to 12 hours in 2026.

Why it matters: HKR-H/K/R all pass: Jack Clark anchors a >60% end-2028 automated-AI-R&D claim in SWE-Bench and METR numbers. This fits the 85–94 band for a notable figure’s AI-timeline essay, below model-release magnitude.

Apr 28Tuesday

The Verge · AI

Attack of the Killer Script Kiddies

The Verge discusses Claude Mythos and AI bug finding, citing DARPA AIxCC scans over 54 million code lines. Teams found most seeded flaws plus over a dozen unseeded bugs; the RSS snippet does not disclose Mythos benchmarks, pricing, or access terms.

Why it matters: HKR-H/K/R all pass: the hook is strong, DARPA AIxCC supplies concrete numbers, and the security angle resonates. No Claude Mythos benchmark, pricing, or access terms are disclosed, so it stays in the featured-threshold band.

Apr 24Friday

MIT Technology Review · AI

Health-care AI is here. We don’t know if it actually helps patients.

Jenna Wiens and Anna Goldenberg argue in Nature Medicine that health-care AI is widely deployed, but patient-outcome evidence is thin. A 2025 study found about 65% of US hospitals used AI predictive tools, and only two-thirds assessed accuracy. The key issue is post-deployment impact on clinical decisions.

Why it matters: HKR-H/K/R all pass: the story has a sharp evidence-gap hook, concrete 2025 hospital-use numbers, and clear safety resonance. It lacks a new model, regulation, or clinical trial result, so 76 fits the featured threshold.

Apr 22Wednesday

Hacker News front page

Show HN submissions tripled and are now mostly vibe-coded

Adrian Krebs scored 500 recent Show HN landing pages and says submissions have tripled, with 67% of pages triggering at least 2 AI design patterns. The method used Playwright plus an in-page script to check DOM and computed styles across 15 deterministic CSS/DOM signals; manual QA found about 5% to 10% false positives. The real signal is not model quality, but fast homogenization from AI default frontend templates.

Why it matters: This clears HKR-H/K/R: a sharp hook, a concrete 500-page method, and a real nerve for AI builders. I keep it at 78, not higher, because it is a single-author experiment rather than a product launch or a cross-source industry event.

Apr 20Monday

r/LocalLLaMA

Actually put Gemma 4 26B to work on something real: extract trading signals from 2,400 earnings calls

A Reddit user fine-tuned Gemma 4 26B on 800 labeled earnings-call transcripts and ran inference on 2,400 transcripts over 3 years on one RTX 4090 in about 14 hours. On 600 out-of-sample transcripts, one signal linked vaguer CFO guidance to about 1.8% sector-relative underperformance over 5 days with IC 0.04. A stronger signal showed 0.85 correlation with sector returns after checks and was discarded as a ghost factor; the key point is factor sanity checks, not the profit claim.

Why it matters: Strong HKR-H/K/R: this is a named first-person experiment with concrete setup, metrics, and a useful negative result. It stays at featured, not P1, because it is one Reddit test rather than a product release or industry-wide event.

Apr 16Thursday

Hacker News front page

AI cybersecurity is not proof of work

antirez argues AI bug finding is bounded by model intelligence level I, not by brute-force sampling alone; for the same code, execution paths eventually saturate. His concrete example is the OpenBSD SACK bug: weaker models fail even with unlimited tokens because they do not connect window validation, integer overflow, and the NULL branch. The key variable is model quality and access speed, not just more GPU.

Why it matters: High-quality commentary with HKR-H from the contrarian headline, HKR-K from the OpenBSD SACK mechanism and firsthand test, and HKR-R because it hits the 'more sampling vs better models' debate in AI security. Not a product, research release, or multi-source event, so it stays mid

Apr 7Tuesday

MIT Technology Review · AI

The one piece of data that could actually shed light on your job and AI

University of Chicago economist Alex Imas argues that AI job displacement depends less on task exposure and more on industry-level price elasticity data; the piece cites OpenAI estimating real estate agents as 28% exposed. It adds that the US task catalog started in 1998, and Anthropic compared it with millions of Claude chats in February. The key variable is whether lower prices raise demand enough, and the post does not disclose any economy-wide dataset yet.

Why it matters: Strong HKR-K: it reframes job impact around price elasticity, with concrete anchors like OpenAI's 28% exposure for real-estate agents and Anthropic's O*NET-to-Claude mapping. HKR-R is clear because it hits job displacement anxiety, but this is commentary, not a fresh dataset or a

Mar 13Friday

Ruan YiFeng's Weblog

Tech Enthusiast Weekly #388: Testing Is the New Moat

A Cloudflare engineer used AI to reimplement Next.js as vinext in 1 week, with $1,100 in token cost and 94% API coverage. The post cites early benchmarks: 4x faster builds and 57% smaller client bundles, with production Next.js apps already running on it. The sharper point is testing: SQLite has 156k lines of code, 92.05M lines of tests, and keeps its core TH3 suite closed.

Feb 27Friday

MIT Technology Review · AI

AI is rewiring how the world’s best Go players think

AI has become standard in pro Go training in South Korea, and the piece says competing professionally without it is now essentially impossible. It cites two figures: Shin Jin-seo matches AI moves 37.5% of the time versus a 28.5% player average, and AlphaGo Zero beat AlphaGo Lee 100-0 after three days of training. The shift to watch is training, not hype: KataGo is now a common tool, opening moves often mirror AI for the first 50 turns, and even top players still cannot fully explain its choices.

Why it matters: Strong HKR-H/K/R: the novelty is elite cognition shifting under AI, and the story brings concrete numbers plus a named tool. It is a reported commentary rather than a new model or product move, so it sits at the low end of featured.

Feb 5Thursday

MIT Technology Review · AI

This is the most misunderstood graph in AI

MIT Technology Review says METR’s plot shows frontier models’ software-task time horizon doubling about every seven months; Claude Opus 4.5 was estimated at about five hours in December 2025. The post stresses that five hours means human time for comparable tasks, not five autonomous model hours; METR gave Opus 4.5 a roughly 2-to-20-hour range. The key caveat: the plot mainly measures coding tasks and defines time horizon at 50% task success, not general AI ability.

Why it matters: HKR-H/K/R all land: the piece has a strong hook and clarifies the METR chart with concrete, testable details. It stays in the low featured band because this is authoritative explanatory commentary, not a new model, product, or research release.

Feb 1Sunday

Lex Fridman (YouTube RSS)

State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast #490

Lex Fridman, Sebastian Raschka, and Nathan Lambert discuss the 2026 AI race in podcast #490 and frame DeepSeek R1’s January 2025 release as a key inflection point. The episode names Claude Opus 4.5, Gemini 3, Z.ai GLM, Minimax, and Kimi Moonshot, but the post does not disclose a shared benchmark, cost table, or reproducible eval. The useful takeaway is the lens: gaps look more like compute, budget, and org culture than secret ideas.

Why it matters: High-quality commentary, not a news break. HKR-H and HKR-R pass because Lex Fridman, Sebastian Raschka, and Nathan Lambert frame China, agents, GPUs, and AGI for practitioners. HKR-K misses: the post names models and DeepSeek R1 but provides no shared benchmarks, cost table, or a