Skip to content

Benchmarks

Who is actually stronger: benchmark results, methodology disputes and leaderboard changes.

Latest picks

241–260 of 453

May 22Friday

AI HOT (Curated Pool)

Text Degeneration: A Production Failure Mode Most Benchmarks Do Not Track

Dharma-AI says in a Hugging Face post that large language models can produce repeated, incoherent, or logically confused text in production, and most mainstream benchmarks do not track this failure mode.

Why it matters: HKR-H/K/R all pass, but the post only discloses the failure pattern and benchmark blind spot, with no sample size, metric, or reproduction setup. This fits the lower featured threshold.

AI HOT (Curated Pool)

BitCPM-CANN Released as First 1.58-bit Open Model Fully Trained on Huawei Ascend 910B NPU

ModelBest, Tsinghua University, and OpenBMB released BitCPM-CANN, a 0.5B-8B open model family trained natively on Huawei Ascend 910B NPUs with 1.58-bit ternary weights, cutting memory use by about 6x versus BF16 while retaining 95-97% of full-precision benchmark performance.

Why it matters: HKR-H/K/R all pass: the Ascend 910B plus 1.58-bit open model angle is novel and metric-rich. It stays below P1 because the post offers release facts, not independent replication or adoption signal.

Bloomberg Technology

Pentagon Tests Rival AI Models in Race to Replace Anthropic

The Pentagon is testing rival AI models with 25 departmental “power users” as it seeks alternatives to Anthropic’s Claude, according to a senior defense official; the RSS snippet does not disclose the candidate model list, evaluation criteria, or deployment timeline.

Why it matters: Bloomberg sourcing plus Pentagon testing rivals to Anthropic clears HKR-H/K/R. Candidate models, contract size, and timeline are not disclosed, so it sits just above the featured threshold.

May 21Thursday

r/LocalLLaMA

Agent Execution Tax: New Procurement Metric for Browser Agent Benchmarks?

Fireworks ran 720 browser-agent tasks on WebVoyager and reported a 22.9% Agent Execution Tax, defined as wasted over productive inference; MiniMax M2.5 cost 2.3x less per successful task than Gemini, while GLM-5 reached 57.1% accuracy and Kimi K2.5 had 0% parse retries across 852 calls.

Why it matters: HKR-H/K/R all pass: the post adds a named procurement metric plus concrete benchmark numbers. Source scope is Reddit/Fireworks, so it stays in the 72–77 featured band rather than 78+.

r/LocalLLaMA

LLM planner: pick a rig by use case, model, or budget, or pick models for your rig

totosse17 published the LLMRequirements hardware planner with 60+ build configs, 50+ models, 130 cited tokens-per-second sources, 150+ reviewer videos, multi-region prices, idle and active watts, and a public GitHub data repo.

Why it matters: HKR-H/K/R all pass, but this is a Reddit community tool for local LLM rigs, not a broad platform release. The concrete dataset earns a featured-threshold score, not the 78+ band.

r/LocalLLaMA

Tencent Hy-MT2 30B/7B/1.8B

Tencent released Hy-MT2 translation models in 1.8B, 7B, and 30B-A3B sizes, supporting translation across 33 languages; AngelSlim 1.25-bit quantization reduces the 1.8B model’s storage requirement to 440 MB and raises inference speed by 1.5x.

Why it matters: HKR-H/K/R pass via the 440MB quantized 1.8B model, 33-language support, and local inference cost angle. Sparse Reddit sourcing keeps it at the featured threshold, not the 78+ band.

Xinzhiyuan · WeChat

USTC Papers Study Lifelong Learning for LMMs via Multimodal Knowledge Injection

USTC researchers released MMEVOKE and KORE: MMEVOKE contains 9,422 samples across 159 subcategories, while KORE uses knowledge-tree augmentation and null-space constrained fine-tuning to reduce catastrophic forgetting during multimodal knowledge injection.

Why it matters: HKR-K and HKR-R are solid: the post gives dataset size plus a concrete fine-tuning mechanism. It stays in the low featured band because this is paper-level knowledge injection without production evidence or full reproducibility details.

Latent Space

OpenAI GPT-next Disproves 80-Year-Old Erdős Planar Unit Distance Problem for Under $1000

OpenAI said an internal general-purpose reasoning model disproved the 1946 Erdős planar unit distance problem by finding a new family of constructions; the reasoning summary reportedly spans about 125 pages, while outside observers speculate the run used under 32 hours or under $1,000.

Why it matters: HKR-H/K/R all pass: an OpenAI internal reasoning model allegedly refuting the 1946 Erdős problem with ~125 pages is a major capability signal. Cost and runtime are still external estimates, keeping it below 95.

Synced · WeChat

Xie Saining’s Team Releases Second-Generation Representation Autoencoder RAEv2

Xie Saining’s team, Adobe Research, and the Australian National University released RAEv2, which reaches gFID 1.06 after 80 epochs on ImageNet-256 and reduces EPFID@2 from 177 epochs to 35 epochs while keeping compute at 189 GFLOPs.

Why it matters: HKR-K and HKR-R pass with concrete benchmark and training-efficiency claims. HKR-H is weak because the angle is a normal research release, so it lands at the featured threshold rather than a must-write item.

r/LocalLLaMA

HalBench: Custom sycophancy and hallucination benchmark tests 4 frontier models

HalBench tested 4 frontier models on 3,200 false-premise prompts, with Sonnet 4.6 ranking first at a 0.565 mean score and Gemini 3.1 Pro last at 0.339; higher scores mean the model more often named the false premise and pushed back instead of complying.

Why it matters: HKR-H/K/R all pass: HalBench has a clear custom-eval hook, 3,200 prompts with scores, and a live trust/safety angle. Single Reddit sourcing and an unvalidated benchmark keep it at the low featured band.

May 20Wednesday

AI HOT (Curated Pool)

Gemini 3.5 Flash launches with stronger performance and speed

Google opened Gemini 3.5 Flash after Google I/O across its products and API; the post says it outperforms Gemini 3.1 Pro on most benchmarks and generates tokens 4x faster than other frontier models.

Why it matters: HKR-H/K/R all pass: Sundar Pichai announced Gemini 3.5 Flash with product/API access and a 4x token-speed claim. This is same-day model-release signal, though price, context window, and full evals are not disclosed.

AI HOT (Curated Pool)

Google releases Gemini 3.5 Flash with 55 intelligence score

Google released Gemini 3.5 Flash, raising its intelligence score by 9 points to 55, exceeding 280 output tokens per second, and increasing operating cost by 5.5 times versus the previous generation.

Why it matters: A Google Gemini 3.5 Flash release is a top-lab model update, backed by Artificial Analysis numbers for speed, intelligence, and cost. HKR-H/K/R all pass, with the 5.5x cost jump making it more than a routine launch.

AI HOT (Curated Pool)

Google releases Gemini 3.5 Flash for complex agent workflows

Google introduced Gemini 3.5 Flash at Google I/O for long-running agent workflows; it outscored 3.1 Pro on Terminal-Bench and MCP Atlas, runs up to 4x faster than other frontier models, and reaches up to 12x speed gains in Google Antigravity.

Why it matters: HKR-H/K/R all pass: Google launched Gemini 3.5 Flash for long-horizon agents with benchmark and speed claims. This is a same-day major model update, below industry-shaking tier.

AI HOT (Curated Pool)

Google releases Gemini 3.5 Flash with output speed about 4x GPT-5.5

Google introduced Gemini 3.5 Flash at I/O 2026, with output speed reaching 289 tokens per second, about 4x faster than Claude Opus 4.7 and GPT-5.5 xhigh under the cited comparison.

Why it matters: HKR-H/K/R all pass: Google ships Gemini 3.5 Flash with a 289 tokens/sec claim and 4x speed comparison against GPT-5.5 xhigh. Details on price, context window, and capability limits are not disclosed, so it stays in the low 85-94 band.

r/LocalLLaMA

KV cache quantization benchmarks: TurboQuant is overrated, q5 deserves attention, q8 may waste VRAM

Anbeeld benchmarked KV cache quantization for Qwen 3.6 27B on one RTX 3090 at 64k and 128k context, reporting q4_0 tail KLD 32% worse than q5_0 and turbo4 running 17% slower than q4_0 with little memory saving.

Why it matters: HKR-H/K/R all pass, with a first-person benchmark and concrete deltas. Scope is narrow: one RTX 3090, one model, and a Reddit source, so it stays near the featured threshold.

May 19Tuesday

r/LocalLLaMA

Sapient Intelligence releases HRM-Text 1B: 40B tokens, ~$1k pretrain

Sapient Intelligence released HRM-Text 1B, a 1B-parameter model trained from scratch on 16 GPUs for 1.9 days with 40B tokens and a reported ~$1,000 budget; its self-reported chart shows MATH 56.2 and DROP 82.2, while independent evaluation remains pending.

Why it matters: HKR-H/K/R all pass: low-cost pretraining plus a smaller model beating a larger one is clickable, with concrete training and benchmark numbers. Independent eval is unfinished, so this stays at 78, not 85.

AI Chat-Group Daily (群聊日报)

May 18, 2026 Chat Group Daily

The chat group daily says AI21 Labs cut 60% of staff and stopped selling model access, and cites a University of Waterloo paper where GPT-5.4 accuracy dropped from 100% to 23% after false peer-consensus injection; the snippet also mentions Meta layoff talk at 10%, but does not disclose source details or confirmation conditions.

Why it matters: HKR-H/K/R all pass: AI21’s 60% layoff and model-sales stop signal lab contraction, while GPT-5.4 falling from 100% to 23% under false peer consensus is a concrete safety hook. The chat-digest source keeps it at 78.

QbitAI · WeChat

JD and CAS IIE Publish Three Papers Defining Self-Taught RLVR

JD and CAS IIE released three Self-Taught RLVR papers covering RLSD, NPO, and CoPD; RLSD reports that 200 training steps on Qwen3-VL-8B-Instruct exceed GRPO at 400 steps across 8 benchmarks.

Why it matters: HKR-H/K/R pass: self-taught RLVR is a clear hook; RLSD reports 8 benchmarks and a 200-vs-400-step GRPO comparison; it hits reasoning fine-tuning cost. Not a top-lab model launch and replication heat is undisclosed, so it stays low featured.

AI HOT (Curated Pool)

Qwen3.7 Preview lands on Arena; Alibaba rises to fifth in vision ranking

Alibaba says Qwen3.7-Plus-Preview has landed on Arena and that Alibaba now ranks fifth in vision; the post does not disclose benchmark scores, the number of competing models, or a release timeline for the Qwen3.7 series.

Why it matters: HKR-H/K/R pass: Qwen3.7-Plus-Preview appears on Arena with a #5 vision rank. Score stays in the low featured band because the vendor post omits scores, model count, access, and timeline.

r/LocalLLaMA

21 GPUs benchmarked running a small TTS model, with 5GB peak VRAM

A Reddit user rented 21 GPUs on vast.ai to benchmark OmniVoice, a small TTS model with about 5GB peak VRAM, using xRT as the audio generation speed metric and averaging 3 voice-cloning runs with reference audio.

Why it matters: HKR-H/K/R pass: a 21-GPU TTS benchmark with 5GB peak VRAM and 3-run xRT averaging is useful to local-inference builders. Scope is niche, so it sits at the low featured band.