Skip to content

Benchmarks

Who is actually stronger: benchmark results, methodology disputes and leaderboard changes.

Latest picks

261–280 of 453

May 19Tuesday

r/LocalLLaMA

llama.cpp MTP support landed: Qwen3.6 27B reaches 2.44× on Strix Halo

llama.cpp merged MTP speculative decoding in PR #22673; Qwen3.6 27B Q8_0 rose from 7.4 to 18.1 tok/s on Strix Halo, while a dual RTX 3090 Q8_0 setup rose from 25.7 to 55.9 tok/s.

Why it matters: HKR-H/K/R all pass: llama.cpp adds MTP speculative decoding with Qwen3.6 27B speedups on Strix Halo and RTX 3090. The scope is local inference, not a broad model release, so 78 fits featured.

May 18Monday

AI HOT (Curated Pool)

The Open Agent Leaderboard

IBM Research published the Open Agent Leaderboard on Hugging Face to evaluate agents across language understanding, tool use, and multi-step reasoning tasks; the post does not disclose dataset size, model scores, or the evaluation date.

Why it matters: HKR-H and HKR-R pass because an open agent leaderboard speaks to agent-eval pain. HKR-K fails: the article lacks scores, dataset size, and evaluation date, so it sits at the featured threshold.

r/LocalLLaMA

I Tested 42 LLMs on Their Willingness to Build the Apocalypse

DystopiaBench tested 42 open and closed models across 36 escalating scenarios and 6 dystopia types, using 3 LLM-as-judge scorers and an average over 3 runs; the post says many models catch obvious dangerous requests but fail when risk is hidden behind dual-use framing and normalization.

Why it matters: HKR-H/K/R all pass: the hook is sharp, the test setup has concrete numbers, and the topic hits safety trust. Reddit single-post sourcing and limited disclosed results keep it in featured, not P1.

r/LocalLLaMA

Qwen 3.6 27B on 24GB VRAM: backend comparisons, quant choice, and settings

The author tested Qwen 3.6 27B on an RTX 3090 24GB and kept ik_llama.cpp with Qwen3.6-27B-MTP-IQ4_KS.gguf; at 156k context with q8_0 KV and MTP, a ~5.9k-token prompt plus 1024-token output reached about 1261 tok/s prefill and 72.9 tok/s decode, while vLLM lacked a clean single-card long-context run.

Why it matters: HKR-H/K/R all pass: this is a first-person local inference benchmark with concrete VRAM, quant, context, and speed numbers. Its reach is narrower than a model release, so it sits at the featured threshold.

r/LocalLLaMA

I trained TIME: short context-triggered thinking on Qwen instead of overthinking

An independent author trained TIME with QLoRA on Qwen3 4B/8B/14B/32B to trigger short mid-response reasoning when context changes; the post says datasets, notebooks, scripts, curriculum, and TIMEBench are public, with 24GB VRAM enough for training up to 14B.

Why it matters: HKR-H/K/R all pass: the post has a clear tuning hook, concrete reproducible details, and strong local-LLM resonance. Reddit single-post sourcing keeps it in the 72-77 featured band, below lab-level releases.

AI HOT (Curated Pool)

Open-source tool exposes security risks and detection gaps in AI API relays

api-relay-audit audits AI API relay risks with verifiable three-state decisions and transparent logs, covering AC-1 tool-call rewriting, AC-2 error-response leakage, and context truncation, while the author has published the methodology, comparison results, quick-reference table, and the open-source tool.

Why it matters: HKR-H/K/R all pass because the tool targets real AI API relay risks with concrete checks. Source is a single X post, and adoption or incident data is not disclosed, so it stays in the low featured band.

r/LocalLLaMA

Benchmarking vLLM vs SGLang vs llama.cpp on a mixed Blackwell/Ada cluster

The author benchmarked long-context prefill on a 7-GPU mixed Blackwell/Ada cluster; on Qwen3.5-397B-A17B with 75k tokens, vLLM reached 9.8s TTFT and 7,683 t/s, while llama.cpp took 57.2s and 1,319 t/s.

Why it matters: Single-source Reddit benchmark, so source authority keeps it near the threshold. HKR-H/K/R pass on the mixed 7-GPU setup, 397B at 75k tokens, and concrete TTFT/throughput numbers.

r/LocalLLaMA

LLMs on Android: Snapdragon 8 Elite MoE Experience

A Reddit user tested MoE LLMs on an Honor Magic 7 Pro with Snapdragon 8 Elite and 24GB RAM; under Q4 quantization, LFM2-24b-a2b reached about 24 tokens/s while Gemma reached about 11 tokens/s, and CPU inference was still faster than NPU or GPU in the reported setup.

Why it matters: HKR-H/K/R all pass: a named Reddit test gives hardware, quantization, and token/s figures. Single-device anecdote and weak source authority keep it at the low featured band.

May 17Sunday

r/LocalLLaMA

85 GPU-hours comparing 5 abliteration methods on Qwen3.6-27B

Abliterlitics compared five Qwen3.6-27B abliteration variants against the base model using 85 GPU-hours of benchmarks, HarmBench, KL divergence, and weight forensics; Huihui had the smallest benchmark deltas, Heretic had the lowest KL divergence, and all five variants reached near-complete safety removal.

Why it matters: HKR-H/K/R all pass: the post gives an 85-GPU-hour comparison across five abliteration methods on Qwen3.6-27B. Niche open-model safety work, not a lab release, so it stays at the featured threshold.

r/LocalLLaMA

DeepSeek V4's 1M Context Window: The Breaking Point

A Reddit user tested DeepSeek V4 on 45k, 180k, and 520k-token codebases and found 150k-250k tokens best for coding work. Past 300k tokens, line-number precision degraded; at 520k, outputs shifted toward architecture summaries and skipped implementation details.

Why it matters: A single Reddit post limits authority, but HKR-H/K/R all pass: it is a numbered first-person test with a concrete long-context failure pattern. The right band is featured, not 78+, because replication and model details are thin.

Synced · WeChat

AI agents may spend 1,000x more tokens without better results: the hidden bill

Researchers used OpenHands to analyze traces from 8 frontier models on 500 swe-bench-verified tasks, finding that agentic coding reached a 154:1 input-output token ratio and that human difficulty labels correlated weakly with token use at Kendall tau 0.32.

Why it matters: All HKR axes pass: strong cost-performance hook, concrete benchmark setup and correlation numbers, and direct resonance with coding-agent economics. It is not a model or platform launch, so it fits the 78–84 quality-recommendation band.

r/LocalLLaMA

Same Models Tested Across Strix Halo, RTX 3090, and RTX 5070

C_Coffie published 55 local inference benchmark runs across Strix Halo, RTX 3090, RTX 5070, five backends, and 0.35B to 35B-A3B models; RTX 5070 beats RTX 3090 on models fitting 12GiB, while RTX 3090 leads in the 14–31B band that exceeds 12GiB but fits 24GiB.

Why it matters: Hits HKR-H/K/R with a named first-person benchmark: 55 runs and concrete GPU crossover points. Source is a single Reddit post, so it stays in the low featured band.

Dwarkesh Patel podcast

Notes on Pretraining Parallelisms and Failed Training Runs

Dwarkesh documents pretraining failure modes and parallelism tradeoffs: expert choice and token dropping can break causality in MoE routing, FP16 collectives can bias repeated additions after values exceed 1024, pretraining FLOPs are given as 6ND, B300 HBM is listed as 288GB, and FSDP communication can reach params × 3 with reduce-scatter.

Why it matters: HKR-H/K/R all pass: Dwarkesh’s notes expose concrete pretraining failure modes and numbers. The systems-training focus is specialized, so it sits in the high-quality band rather than same-day must-write.

AI HOT (Curated Pool)

Latest Open Artifacts #21: Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1, and More

Open AI model teams released Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1, and other versions this month, and the post says they were tested under CAISI’s V4 evaluation framework, but the RSS snippet does not disclose scores.

Why it matters: HKR-H/K/R all pass: a dense open-model roster, a named CAISI V4 evaluation frame, and clear practitioner relevance for model choice. Missing scores and reproducible detail keep it in the 78–84 band.

May 16Saturday

Synced · WeChat

Why Robots Need World Models: Top Institutions Release Joint Survey

NTU MARS Lab and collaborators released a 43-page survey on robot world models, covering definitions, architectures, applications, benchmarks, and challenges around action-conditioned consistency, inference efficiency, and physical grounding.

Why it matters: HKR-H and HKR-K pass: the hook is robot world models, and the post cites a 43-page survey with benchmarks and action-consistency framing. HKR-R is weak, so this stays at the featured threshold.

r/LocalLLaMA

Qwen3.6-35B-A3B and 9B land on the public Terminal-Bench 2.0 leaderboard

little-coder × Qwen3.6-35B-A3B scored 24.6% ±3.2 on Terminal-Bench 2.0, above Gemini 2.5 Pro on Gemini CLI at 19.6% and Qwen3-Coder-480B on Terminus 2 at 23.9%.

Why it matters: HKR-H/K/R all pass, but this is a Reddit post with leaderboard numbers only; test setup and reproducibility details are not disclosed. Strong code-agent benchmark signal, not a 78+ release story.

Google DeepMind

How WeatherNext helped the US National Hurricane Center forecast Hurricane Melissa's Jamaica landfall

Google DeepMind's AI weather model WeatherNext helped the US National Hurricane Center forecast five days ahead that Hurricane Melissa would hit Jamaica at Category 5 strength, with 80% confidence. Three days out, that rose to near 100%.

Why it matters: The Hurricane Melissa case shows how an AI weather model called a rapid intensification five days ahead, a concrete look at AI in extreme-weather warnings.

r/LocalLLaMA

Orthrus-Qwen3-8B: Up to 7.8× tokens/forward on Qwen3-8B with frozen backbone

Orthrus-Qwen3-8B adds a trainable diffusion attention head to a frozen Qwen3-8B backbone, reaches up to 7.8× tokens per forward and about 6× wall-clock speedup on MATH-500, while training 16% of parameters with under 1B tokens over 24 hours on 8×H200 GPUs.

Why it matters: HKR-H/K/R all pass: 7.8× speed is a strong hook, the post gives parameter, hardware, and benchmark conditions, and inference cost is a live practitioner pain. Single-source Reddit research keeps it in the lower good-quality band.

The Verge · AI

AI radio hosts demonstrate why AI can’t be trusted alone

Andon Labs had Claude, ChatGPT, Gemini, and Grok run separate radio stations with $20 in seed money each; the RSS snippet says all failed, but the post does not disclose the full experimental results.

Why it matters: HKR-H/R are strong because the agent-failure setup is memorable and relevant. HKR-K is present but thin: it gives four models and $20 budgets, while full experimental results are not disclosed.

May 15Friday

r/LocalLLaMA

Evaluated a RAG Chatbot: The Most Expensive Model Was the Worst Performer

The author evaluated a customer-support RAG bot and raised the quality score from 6.62 to 7.88 while cutting per-session cost from $0.002420 to $0.000509, using retrieval logging, LLM-as-judge scoring, chunk deduplication, stricter grounding, and a five-model sweep.

Why it matters: HKR-H/K/R all pass: counterintuitive model ranking, concrete quality and cost deltas, and direct RAG production relevance. Reddit source authority keeps it near the featured floor despite the first-person experiment signal.