Skip to content

Benchmarks

Who is actually stronger: benchmark results, methodology disputes and leaderboard changes.

Latest picks

301–320 of 453

May 10Sunday

Xinzhiyuan · WeChat

Harsh Claim: Top Silicon Valley AI Is One Year Ahead of the World

Elad Gil claims top AI lab employees are 3-4 months ahead of Silicon Valley, while Silicon Valley is 3-6 months ahead of New York; the post cites Mythos’ 73% success rate in expert cyberattack simulations as evidence in a disputed “geographic time gap” argument.

Why it matters: HKR-H/K/R all pass: the lab-to-user lag hook is clickable, and the post cites 3–4 months, 3–6 months, and a 73% Mythos figure. It is secondhand commentary, not a model or product release, so it stays in the 72–77 threshold band.

QbitAI · WeChat

Zhejiang University introduces AdaMARP, an AI role-playing framework with scene direction

Zhejiang University and Tencent Youtu proposed AdaMARP for immersive role-playing, using a four-channel message format and a scene manager; its data pipeline includes 81 literary works, 20 synthetic themes, and AdaptiveBench with 100 evaluation seeds.

Why it matters: ACL 2026 role-play agent work brings four-channel messaging, a scene manager, and an 81-book dataset, clearing HKR-H/K. Narrow use cases and missing open-source or production evidence keep it at threshold featured.

r/LocalLLaMA

NVIDIA AI Releases Star Elastic: One Checkpoint Contains 30B, 23B, and 12B Reasoning Models

NVIDIA AI released Star Elastic, a single checkpoint that can zero-shot slice 30B, 23B, and 12B reasoning models in BF16, FP8, and NVFP4; when the 23B submodel handles thinking and the 30B model handles final answers, reported accuracy rises 16% and latency drops 1.9× on AIME-2025, GPQA, LiveCodeBench v5, and MMLU-Pro.

Why it matters: HKR-H/K/R all pass: Star Elastic has a concrete mechanism and testable numbers for inference deployment. Its reach is still narrower than a frontier-model release, so it sits in the high-quality featured band.

May 9Saturday

r/LocalLLaMA

80 tok/sec and 128K context on 12GB VRAM with Qwen3.6 35B A3B and llama.cpp MTP

Reddit user janvitos ran Qwen3.6-35B-A3B-MTP-GGUF with a llama.cpp MTP PR on an RTX 4070 Super. The posted benchmark shows 69.2-81.9 tok/s, 0.694-0.947 draft acceptance, 131072 context, and a -fitt 1536 setting that reserves 1536 MB for the draft model and KV cache.

Why it matters: HKR-H/K/R all pass with concrete single-user benchmark data and reproducible settings. Source is one Reddit post, so verification is thin; this lands above featured threshold, not in must-write range.

Synced · WeChat

StarVLA Open-Sources a Unified VLA Framework from HKUST and the Community

HKUST and the open-source community released StarVLA, a unified Vision-Language-Action framework that integrates backbones, action heads, training strategies, and evaluation interfaces; the repository has 2.2k GitHub stars and supports benchmarks including LIBERO, SimplerEnv, RoboTwin 2.0, RoboCasa-GR1, and BEHAVIOR-1K.

Why it matters: HKR-H/K/R all pass: StarVLA ships a concrete open-source VLA framework with unified interfaces, 2.2k stars, and named robotics benchmarks. The robotics scope keeps it in the 78–84 band, below model-release weight.

Synced · WeChat

DeepSeek Reportedly Raises RMB 50B, with Liang Wenfeng Funding 40%, Valuation Reaching RMB 350B

DeepSeek is negotiating a $7.3 billion funding round at an estimated $51.5 billion valuation; Liang Wenfeng reportedly plans to contribute 40%, while Tencent and China’s RMB 60 billion national AI fund are also in talks.

Why it matters: HKR-H/K/R all pass: the DeepSeek funding rumor has large numbers, a founder contribution ratio, and named backers. Because it is still reported as talks with no official confirmation, it stays at 84 and featured, not p1.

AI HOT (Curated Pool)

Claude Mythos Evaluation Shows 16-Hour Risk Horizon

METR evaluated an early Claude Mythos Preview build during a limited March 2026 window and estimated its 50% time horizon at at least 16 hours, with a 95% confidence interval of 8.5 to 55 hours.

Why it matters: HKR-H/K/R all pass: METR reports a concrete 16h risk-horizon estimate for Claude Mythos Preview. The single X-source and limited eval window keep it below P1, but it is strong featured safety signal.

May 8Friday

r/LocalLLaMA

Gemma 4 26B Hits 600 Tok/s on One RTX 5090

chain-77 benchmarked Gemma 4 26B with vLLM 0.19.2rc1, and DFlash raised output throughput on one RTX 5090 from 228 tok/s to 578 tok/s under 256 input tokens, 1024 output tokens, concurrency 1, and num_speculative_tokens=13.

Why it matters: HKR-H/K/R all pass: the single-GPU throughput hook is strong, and the post gives reproducible settings plus before/after speed. Reddit single-post evidence and one hardware setup keep it in the featured-threshold band.

AI HOT (Curated Pool)

Adaptive Parallel Reasoning: A New Paradigm for Efficient Reasoning Scaling

BAIR’s post describes adaptive parallel reasoning, where ThreadWeaver and Multiverse dynamically control parallel threads for math and code reasoning; the RSS snippet does not disclose benchmark scores, latency reductions, or reproducible settings.

Why it matters: BAIR authority supports the 72+ band, and HKR-H/K/R all pass. The post names mechanisms and dynamic thread control, but lacks scores, latency gains, and reproducible conditions, so it stays below 78.

Xinzhiyuan · WeChat

Token-Level Length Control: 3B Model Beats GPT 5.4 and Claude

UC Santa Barbara and Apple researchers introduced LenVM, which models remaining generation length as a token-level value function; Qwen2.5-3B with a 1.5B LenVM scored 62.6 on LIFEBench length control, above GPT-5.4 at 37.4 and Claude-Opus-4-6 at 35.5.

Why it matters: HKR-H/K/R all pass: the headline has a sharp small-model-vs-frontier hook, and the post gives LenVM's mechanism plus 62.6/37.4 benchmark numbers. The topic is narrow research, not a model or major product release, so it fits the 78-84 band.

QbitAI · WeChat

HIT and Huawei propose Dynamic-dLLM, a training-free acceleration framework with 4.48x speedup

HIT Shenzhen, Huawei, and Shenzhen Hetao College proposed Dynamic-dLLM, a training-free dLLM acceleration framework that raises LLaDA-8B-Instruct throughput on GSM8k from 8.32 TPS to 37.29 TPS with almost no accuracy loss.

Why it matters: HKR-H/K/R all pass: the 4.48x speedup is clickable, and GSM8k TPS figures add concrete substance. It is inference-optimization research, not a mainstream model launch, so it fits the 78–84 band.

Ruan YiFeng's Weblog

Technology Enthusiast Weekly Issue 395: The Third Way of Software Development

Ruanyifeng Weekly issue 395 frames AI-assisted coding as a “mystery house” style of software development and cites HN SOTA, which ranks model popularity by scanning 200 top Hacker News topics each day and their programming or AI discussions.

Why it matters: HKR-H/K/R pass: the “third way/mystery house” framing, HN SOTA’s 200 daily HN topics, and developer workflow anxiety all land. It is commentary, not a model or product release, so it stays at 72.

AI HOT (Curated Pool)

Donating the Open-Source Alignment Tool Petri

Anthropic transferred the open-source alignment testing tool Petri to Meridian Labs to preserve independence and credibility. Petri 3.0 separates auditor and target models, adds Dish for real prompts and deployment settings, and integrates Bloom.

Why it matters: HKR-H/K/R all pass: the independent donation is a real hook, Petri 3.0 and Dish add testable mechanisms, and audit credibility resonates. Anthropic open-source safety tooling is strong, but below a model-release-level event.

r/LocalLLaMA

11.67% ARC-AGI-2 Local Eval on a Single 4090: The TOPAS Recursive Architecture

Doug_Bitterbot says TOPAS scored 11.67% on ARC-AGI-2 using one RTX 4090 after about 14 days of training. The 100M-parameter checkpoint hit 36% locally, but recursive TTT caused null outputs on nearly half of Kaggle puzzles. The key detail is time management: the author expects 20% after threshold tuning and 3-5 more weeks of training.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post with unstable Kaggle submissions. It clears featured, not the higher research-release band.

AI HOT (Curated Pool)

Readable behavioral signals remain in frozen LLM hidden states, Cygnus boosts accuracy

Proprioceptive AI says Cygnus adds adapters to frozen LLMs and raises Qwen-32B on ARC-Challenge from 82.2% to 94.97%. It projects hidden states into a gl(4,R) Lie-algebra space to isolate “dark modes.” Watch replication; the post does not disclose full eval sets or controls.

Why it matters: HKR-H/K/R pass: the claim is novel, quantified, and practitioner-relevant. Kept at low featured because the source is an X post and full eval set, training details, and controls are not disclosed.

May 7Thursday

r/LocalLLaMA

Qwen/WebWorld 32B/14B/8B (Qwen3 finetune)

Qwen released WebWorld 32B/14B/8B, Qwen3 finetunes for training and evaluating web agents. It uses 1M+ real web trajectories and supports 30+ step simulation plus A11y Tree, HTML, XML, Markdown, and natural-language states. Agents trained on its synthetic trajectories gain 9.9% on MiniWob++ and 10.9% on WebArena.

Why it matters: HKR-H/K/R all pass: WebWorld has an agent hook, concrete scale, and benchmark gains. It is a useful Qwen research release for agent builders, but limited source detail keeps it below the 85 must-write band.

Xinzhiyuan · WeChat

Zhejiang University and Harvard open-source UniGeo for geometry-guided camera-controllable editing

Zhejiang University and Harvard released UniGeo with code, a report, a project page, and an HF Space. UniGeo injects geometry guidance into representation, architecture, and loss layers; it reports SOTA on DL3DV, RE10K, and Tanks against five methods. The key is video priors plus geometry-anchor attention, not just using a video model.

Why it matters: HKR-H and HKR-K pass: open code, HF Space, and three geometry-guidance layers make it testable. HKR-R is weak because it is specialized vision-generation research, so this sits near the featured floor.

Xinzhiyuan · WeChat

Claude Managed Agents Add Dreaming, With Reported Task Completion Up to 6x

Anthropic added Dreaming, Outcomes, and multi-agent orchestration to Claude managed agents; Harvey reports about 6x higher task completion. Dreaming reads up to 100 sessions; one demo distilled 5.3M tokens into 98 rules, while Outcomes raised success by up to 10 points. Opus 4.7 and Sonnet 4.6 require access, with $0.08 per session-hour runtime fees.

Why it matters: HKR-H/K/R all pass: Anthropic adds Dreaming, Outcomes, and multi-agent orchestration with 100-session memory, $0.08/session-hour runtime, and Harvey’s ~6x completion claim. This is a same-day Claude agent update.

Synced · WeChat

Claude, GPT and Gemini score 0% completion on ProgramBench

ProgramBench tested Claude Opus 4.7, GPT-5.4 and Gemini 3.1 Pro, with 0% full completion on rebuilding software projects. It gives only executables and usage docs, removes source/tests, and grades behavioral equivalence via agent-driven fuzzing. The key signal is system-level engineering, not function-level code generation.

Why it matters: HKR-H/K/R all pass: the 0% result is clickable, the setup is concrete, and the coding-agent gap matters to practitioners. Still, it is a single benchmark report, below a major model or product release.

r/LocalLLaMA

Exaggerated PCI-E Bandwidth Concerns?

Reddit user ziphnor tested 2x RTX 5060 Ti 16GB with vLLM TP=2 and 32k-context prefill. PCIe peaked at 3–4 GB/s, about 40–50% of a PCIe 4.0 x4 link. Prefill reached ~840–850, 1500, and 1600–1700 t/s; the post does not disclose decode bandwidth.

Why it matters: HKR-H/K/R all pass: a myth-busting PCIe bandwidth test with concrete vLLM conditions and numbers. Single Reddit source limits authority, but the named first-person experiment lifts it to the featured threshold.