Skip to content

Benchmarks

Who is actually stronger: benchmark results, methodology disputes and leaderboard changes.

Latest picks

201–220 of 453

Jun 2Tuesday

AI HOT (Curated Pool)

NVIDIA Cosmos 3 Tops Open-Weight Image and Video Generation Rankings

NVIDIA Cosmos 3 ranked first in Artificial Analysis’s open-weight text-to-image and image-to-video categories, with 16B Nano and 64B Super variants, and the release includes weights, code, curated datasets, and fine-tuning recipes under the OpenMDW 1.1 license.

Why it matters: HKR-H/K/R all pass: Cosmos 3 leads both Artificial Analysis open-weight image and video charts, with 16B/64B variants and OpenMDW 1.1 artifacts disclosed. Single-source benchmark news keeps it in the 78–84 featured band.

Jun 1Monday

AI HOT (Curated Pool)

MiniMax Releases Open-Source M3 with Coding, Long-Context, and Multimodal Capabilities

MiniMax released the open-source M3 model with coding, a 1M-token context window, and native multimodal support; M3 scores 59.0% on SWE-Bench Pro, 83.5% on BrowseComp, and costs about one-twelfth per token versus GPT-5.5.

Why it matters: HKR-H/K/R all pass: M3 has open source, 1M context, multimodal support, and 59.0% on SWE-Bench Pro. A single X post without official docs or third-party tests keeps it in the 78–84 band.

Import AI (Jack Clark)

Import AI 459: AI oversight is difficult; scaling laws for protein folding models; and pricing the extinction risk of AI systems

Import AI 459 summarizes papers on AI-economy measurement and AI oversight: one estimates U.S. nominal AI GDP at about $250 billion in 2025, with quality-adjusted real growth near 2,600% per year.

Why it matters: HKR-H/K/R all pass: the extinction-risk pricing hook is unusual, the summary gives $250B and 2600% as concrete figures, and oversight risk has practitioner resonance. It is still a secondary roundup, not a same-day must-write release.

r/LocalLLaMA

Deepseek V4 Flash performance on DGX Spark

A Reddit user ran DeepSeek-V4-Flash with vLLM on two ASUS GX10 DGX Spark nodes and reported 1,680 prefill tokens/s plus 39.8 decode tokens/s at a 256K context with MTP=2; the setup uses TP=2 over RoCE, fp8 KV cache, and fits about 1M tokens safely in KV cache.

Why it matters: This is not broad industry news, but it is a first-person benchmark with reproducible details: TP=2, RoCE, fp8 KV cache, 256K context, and ~1M KV. HKR-H/K/R all pass, so it lands at low featured.

AI HOT (Curated Pool)

MWC26 Shanghai to Host First Humanoid Robot Penalty Shootout With Unitree and 7 Other Teams

MWC26 Shanghai will host a humanoid robot penalty shootout in June 2026, with eight Chinese embodied intelligence teams competing under rules that require autonomous play without human control or preset scripts.

Why it matters: HKR-H/K/R all pass: the robot penalty shootout is clickable, with rules banning teleoperation and scripts. It stays in 72–77 because this is an event preview, not a model release or reproducible result.

May 31Sunday

r/LocalLLaMA

13 abliterated Gemma 4 E2B variants, 44 GPU hours, benchmark and comparison

Abliterlitics tested 13 abliterated Gemma 4 E2B variants using 44 RTX 5090 GPU hours, and HarmBench ASR rose from the base model’s 32.2% to 82%–100%, while coder3101 scored 84.8% on GSM8K versus the base model’s 83.5%.

Why it matters: HKR-H/K/R all pass, with a named first-person benchmark and concrete numbers. Scope stays narrow around abliterated Gemma 4 E2B variants, so it lands at the featured threshold rather than a must-write item.

r/LocalLLaMA

PolyRange: Contamination-resistant offensive-AI benchmark for web targets

PolyRange v1.0 ships 84 WSTG-derived classes across 12 OWASP testing-guide categories. It generates fresh targets per deploy with a chosen LLM, adds two defense tiers, uses an agent-submits-flag oracle, and runs via a single-command CLI on Fly.io or Docker.

Why it matters: HKR-H/K/R all pass: PolyRange turns web-security targets into a dynamic agent benchmark with 84 WSTG classes and two defense levels. Single-source Reddit origin and security niche keep it at 78.

Synced · WeChat

Rubrics Survey: How to Define a Good Answer in the Agent Era

Renmin University Gaoling School of Artificial Intelligence released a 40-page survey on rubrics for LLMs, organizing the topic into five parts: definitions, construction methods, training uses, evaluation scenarios, and open challenges.

Why it matters: HKR-H/K/R all pass, but this is a survey rather than a model or product launch. The 40-page rubric framework is useful for agent evaluation, placing it at the featured threshold.

Synced · WeChat

Microsoft open-sources SkillOpt for training Agent skill documents, reaching 3.3k stars in a week

Microsoft open-sourced SkillOpt, a text-space optimization framework that trains Agent skill documents without changing model weights; the paper reports best or tied-best results across 52 combinations covering 7 target models, 6 benchmarks, and 3 execution environments.

Why it matters: Microsoft’s open-source SkillOpt is a strong Agent tooling and research release. HKR-H has the 3.3k-star/trainable-skill hook, HKR-K has the text-parameter mechanism and 52 eval setups, and HKR-R hits agent engineering pain, so it lands in featured at 82.

May 30Saturday

Xinzhiyuan · WeChat

Claude AI fluency scorecard surfaces, with strong users scoring 7.5

Anthropic is testing a Claude AI Fluency scorecard that analyzes Chat, Cowork, and Claude Code history against 11 observable behaviors, with an 11-point maximum score. The underlying study used 9,830 anonymized multi-turn conversations, and iteration appeared in 85.7% of high-quality conversations.

Why it matters: HKR-H/K/R all land: the angle is clickable, the scorecard has concrete numbers, and Claude users will debate being graded. This is not a model launch or major capability release, so it stays in the 78–84 featured band.

QbitAI · WeChat

RUC and Zhizhi Institute Open-Source Claw Agent Data, Training, and Evaluation Pipeline

Renmin University of China and Zhizhi Institute open-sourced ClawGym, a Claw Agent framework with 13.5K synthetic executable tasks, 200 benchmark tasks, model checkpoints, training data, and training code; ClawGym-30B-A3B scores 56.82 on ClawGym-Bench and exceeds Qwen3-235B-A23B in the reported evaluation.

Why it matters: HKR-H/K/R all pass: ClawGym bundles data, code, checkpoints, and eval tasks rather than just a leaderboard. Its impact is developer-facing, below a major lab model release or market-moving event.

Synced · WeChat

CUHK Pion optimizer updates LLMs on iso-spectral manifolds to address AdamW and Muon instability

CUHK and collaborators introduced Pion, an optimizer that preserves weight singular values through orthogonal equivalence transformations, and reported that it kept a 60M normalization-free LLaMA-like model stable for 9.6B training tokens while AdamW and Muon collapsed with NaNs.

Why it matters: HKR-H/K/R pass: the hook is AdamW/Muon NaN instability, with a concrete isospectral update and 9.6B-token run. Niche optimizer math keeps it in 78–84, not same-day product news.

r/LocalLLaMA

Testing MTP on vLLM and llama.cpp for Gemma 4 and Qwen 3.6

The author tested MTP on an RTX PRO 6000 Blackwell setup, where Gemma 4 31B on vLLM reached 132.52 tok/s versus a 39.69 tok/s baseline, a 3.34x speedup; the post reports 10 runs of 1,500 tokens each but does not provide a full quality or VRAM evaluation.

Why it matters: HKR-H/K/R all pass via a first-person speed test with hardware, model, and tok/s numbers. Source authority is limited, and missing quality/VRAM evaluation keeps it at the low featured band.

May 29Friday

New York Times Chinese

Anthropic Tops OpenAI Valuation to Become the Most Valuable AI Startup

Anthropic raised $65 billion at a $900 billion pre-money valuation, above OpenAI’s last $730 billion valuation. Claude Opus 4.8 also scored 10% higher than Anthropic’s previous model on Vals AI’s vibe-coding benchmark.

Why it matters: Anthropic topping OpenAI with $65B financing and a $900B pre-money valuation is a foundation-model market-structure event. HKR-H/K/R all pass, with NYT source authority supporting p1.

Xinzhiyuan · WeChat

Three DeepSeek Models Enter OpenRouter Monthly Top 10 With Over 17 Trillion Tokens

DeepSeek placed three models in OpenRouter’s monthly top 10 with more than 17 trillion tokens combined, including V4 Flash at 9.13T tokens; the article says Ascend’s MegaMoE operator raised Prefill throughput by 20% to 30% on DeepSeek V3.1 and Qwen3-235B tests.

Why it matters: HKR-H/K/R all pass: the story has a 17T-token hook plus concrete OpenRouter and MegaMoE Prefill numbers. It stays at 82 because the compute-sovereignty framing is strong, while reproducible test conditions are not disclosed.

AI HOT (Curated Pool)

Adam's Law: Prompts Written with High-Frequency Words Work Better

FaceMind tested 100 languages and four core tasks, finding that, with semantics unchanged, prompts or fine-tuning text using higher-frequency expressions from pretraining data improves large language model performance.

Why it matters: HKR-H/K/R all pass: the claim is counterintuitive and backed by 100 languages and four task types. Missing models, datasets, and effect sizes keep it in the low featured band.

AI HOT (Curated Pool)

Tesla FSD Safety Claims Face Scrutiny

Tesla claimed FSD can be up to 10 times safer than humans, but Reuters found flaws in the comparison, with 11 traffic safety researchers saying Tesla used inappropriate baselines against broader federal crash data.

Why it matters: HKR-H/K/R all pass: the Reuters-backed challenge to Tesla’s 10x FSD safety claim has conflict, numbers, and safety resonance. The article does not disclose full samples or formulas, so it stays in the 72–77 band.

May 28Thursday

QbitAI · WeChat

A New Paradigm for GUI Agent Trajectories: FSMs Generate Trajectories at $0.04 Each

AutoWebWorld synthesized 29 web environments, 875 pages, and 11,663 verified trajectories at about $0.04 per trajectory, using FSM-defined states, preconditions, and transitions to verify GUI agent tasks instead of human labeling or an LLM judge.

Why it matters: HKR-H/K/R all pass: $0.04 per trace, 11,663 verified traces, and FSM state checks give concrete hooks for GUI-agent data and eval cost. The source is not a top lab release, so it stays in the 78–84 research-tool band.

r/LocalLLaMA

Qwen3.6-35B-A3B-APEX Runs 128K Context on RTX 3060 12GB

old-mike ran mudler/Qwen3.6-35B-A3B-APEX-MTP-I-Compact.gguf through spiritbuun’s llama.cpp fork on one RTX 3060 12GB, offloading a 17.3GB model and reaching 37.17 t/s generation at 72K filled context, 28.08 t/s at 129K, and PPL 3.2529 on an enwik8 64K-context perplexity test.

Why it matters: HKR-H/K/R all pass via a concrete consumer-GPU inference result with speed and PPL. Source is a single Reddit post and the impact stays within local inference, so it lands in featured, not P1.

Latent Space

Cognition Raises $1B in $26B Series D

Cognition raised a $1B Series D at a $26B valuation and projects more than $1B ARR by year-end; the post says its valuation rose 2.5× from the $10B Series C eight months earlier, while the rest of the issue summarizes agent, inference, benchmark, and multimodal AI updates from May 26–27, 2026.

Why it matters: HKR-H/K/R all pass: Cognition’s $1B Series D at a $26B valuation is large, and projected year-end ARR above $1B gives a concrete business signal. This is not a model launch, but it is must-write funding news for AI coding agents.