Skip to content

All news

8 today

May 23Saturday

r/LocalLLaMA

club-rdna16: Practical 16GB AMD/Radeon local LLM testing repo

club-rdna16 publishes a practical 16GB Radeon local LLM testing repo, with an RX 6900 XT running llama.cpp on ROCm/HIP and Qwen3.6 35B-A3B reaching a stable 131k context using q8 KV cache.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post and the body only discloses test conditions, not speed, VRAM curves, or reproducible logs. It clears the featured floor as a practical local-LLM repo.

r/LocalLLaMA

Experts first llama.cpp

comanderxv published a llama.cpp fork that caches MoE experts in 12GB VRAM; on an RTX 2060 with Qwen3.6-35B-A3B, throughput rose from 19/22 tk/s to 26 tk/s at about a 62% expert-cache hit rate.

Why it matters: HKR-H/K/R all pass: the hook is a 35B MoE speedup on a 12GB RTX 2060, with concrete caching and hit-rate data. Scope stays niche to local inference, so it lands at the featured threshold rather than must-write.

May 22Friday

AI HOT (Curated Pool)

NetEase Youdao Open-Sources Ziyue 4 Multimodal and Text-to-Speech Models

NetEase Youdao open-sourced its Ziyue 4.0 multimodal and text-to-speech models, with the 27B multimodal model reporting 81.4% accuracy on Chinese math reasoning tasks and the speech model supporting 14 languages.

Why it matters: HKR-H/K/R pass: the story has a concrete open-source hook, specific model numbers, and practitioner relevance. NetEase Youdao is not a frontier lab, so it stays below the 78+ good-quality band.

AI HOT (Curated Pool)

Zhipu releases GLM-5.1-highspeed, claiming a large-model API speed record

Zhipu released the GLM-5.1-highspeed API to selected enterprise customers on May 22, with a claimed output speed of 400 tokens/s, built by the GLM team and TileRT team through system-level optimization.

Why it matters: HKR-H/K/R all pass: Zhipu’s GLM-5.1 high-speed API has a concrete 400 tokens/s claim and domestic flagship-model relevance. Test setup, pricing, and availability are not disclosed, so it stays in the 78–84 band.

Bloomberg Technology

Pentagon Tests Rival AI Models in Race to Replace Anthropic

The Pentagon is testing rival AI models with 25 departmental “power users” as it seeks alternatives to Anthropic’s Claude, according to a senior defense official; the RSS snippet does not disclose the candidate model list, evaluation criteria, or deployment timeline.

Why it matters: Bloomberg sourcing plus Pentagon testing rivals to Anthropic clears HKR-H/K/R. Candidate models, contract size, and timeline are not disclosed, so it sits just above the featured threshold.

May 21Thursday

r/LocalLLaMA

Agent Execution Tax: New Procurement Metric for Browser Agent Benchmarks?

Fireworks ran 720 browser-agent tasks on WebVoyager and reported a 22.9% Agent Execution Tax, defined as wasted over productive inference; MiniMax M2.5 cost 2.3x less per successful task than Gemini, while GLM-5 reached 57.1% accuracy and Kimi K2.5 had 0% parse retries across 852 calls.

Why it matters: HKR-H/K/R all pass: the post adds a named procurement metric plus concrete benchmark numbers. Source scope is Reddit/Fireworks, so it stays in the 72–77 featured band rather than 78+.

r/LocalLLaMA

Tencent Hy-MT2 30B/7B/1.8B

Tencent released Hy-MT2 translation models in 1.8B, 7B, and 30B-A3B sizes, supporting translation across 33 languages; AngelSlim 1.25-bit quantization reduces the 1.8B model’s storage requirement to 440 MB and raises inference speed by 1.5x.

Why it matters: HKR-H/K/R pass via the 440MB quantized 1.8B model, 33-language support, and local inference cost angle. Sparse Reddit sourcing keeps it at the featured threshold, not the 78+ band.

AI HOT (Curated Pool)

Tencent open-sources Hy-MT2 multilingual translation model

Tencent open-sourced the Hy-MT2 multilingual translation model with support for translation across 33 languages; its 1.8B version uses AngelSlim 1.25-bit quantization, occupies 440 MB of storage, and runs locally on mainstream mobile chipsets.

Why it matters: HKR-H/K/R all pass: Tencent gives a specific edge-AI hook with 33 languages, 1.25-bit quantization, and a 440MB phone-local build. Benchmarks, latency, and license terms are not disclosed, so it stays below major flagship releases.

r/LocalLLaMA

HalBench: Custom sycophancy and hallucination benchmark tests 4 frontier models

HalBench tested 4 frontier models on 3,200 false-premise prompts, with Sonnet 4.6 ranking first at a 0.565 mean score and Gemini 3.1 Pro last at 0.339; higher scores mean the model more often named the false premise and pushed back instead of complying.

Why it matters: HKR-H/K/R all pass: HalBench has a clear custom-eval hook, 3,200 prompts with scores, and a live trust/safety angle. Single Reddit sourcing and an unvalidated benchmark keep it at the low featured band.

r/LocalLLaMA

What happened to Cohere’s Command-A series of models?

Cohere launched Command A+, describing it as its first MoE model under the Apache 2.0 license, with quantization work that lets it run well on 1 or 2 GPUs; the post says top-line performance still needs work.

Why it matters: HKR-H/K/R pass: Cohere open model news has clear local deployment facts. Reddit-level sourcing and missing parameter count, benchmarks, and context window keep it in the low featured band.

May 20Wednesday

AI HOT (Curated Pool)

Stability AI Launches Stability Audio 3.0 for Songs Up to 6 Minutes

Stability AI launched the Stability Audio 3.0 audio generation model family with four sizes ranging from 459 million to 2.7 billion parameters; the small model targets on-device use and generates audio under 2 minutes locally, while medium and large models support full music creation beyond 6 minutes and 20 seconds.

Why it matters: HKR-H/K pass because Stability AI gives concrete duration and model-size details. HKR-R is weak: no benchmarks, licensing, pricing, or access terms are disclosed, so this sits at the featured threshold.

AI HOT (Curated Pool)

Kling AI Launches the First Native 4K Video Generation Model

Kling AI launched a native 4K video generation model on April 23, supporting one-click true 4K video generation; the post says Hollywood teams and Wonder Studios have adopted it, but does not disclose pricing, inference cost, or access limits.

Why it matters: HKR-H and HKR-K pass: Kling AI’s native 4K video model has a concrete capability and named adoption. Source is product-side, with no benchmark, pricing, or clip-duration data, so it sits at the featured threshold.

AI HOT (Curated Pool)

Smarter Google AI Edge Gallery: MCP Integration, Notifications, and Session Continuity

Google AI Edge Gallery adds experimental MCP support on Android, letting Gemma 4 coordinate external data sources including Google Workspace and Google Maps; the update also adds scheduled notifications and persistent chat history for faster restoration of long-session context.

Why it matters: HKR-H/K/R all pass: Google’s developer update adds experimental MCP, notifications, and session continuity to AI Edge Gallery. It is a mid-weight product update, not a model release or major capability launch.

AI HOT (Curated Pool)

Gemini 3.5 Released: A New Model Family Combining Intelligence and Action

Google AI Developers announced the Gemini 3.5 model family, saying it combines intelligence with action capabilities; the post does not disclose parameters, benchmarks, pricing, availability, or context window details.

Why it matters: HKR-H and HKR-R pass: an official Gemini 3.5 family launch has flagship-model pull and competitive resonance. HKR-K fails because the post gives no params, benchmarks, pricing, or context window, so this stays below the 85+ band.

Hacker News front page

Gemini 3.5 Flash

The title names Gemini 3.5 Flash, while the RSS body only includes a documentation link; the Hacker News item has 196 points and 179 comments, and the post does not disclose parameters, pricing, or context-window details.

Why it matters: HKR-H/R pass on an official Google Gemini 3.5 Flash release with HN traction; HKR-K fails because price, benchmarks, parameters, and context window are absent. That keeps it in the lower 78–84 band.

r/LocalLLaMA

KV cache quantization benchmarks: TurboQuant is overrated, q5 deserves attention, q8 may waste VRAM

Anbeeld benchmarked KV cache quantization for Qwen 3.6 27B on one RTX 3090 at 64k and 128k context, reporting q4_0 tail KLD 32% worse than q5_0 and turbo4 running 17% slower than q4_0 with little memory saving.

Why it matters: HKR-H/K/R all pass, with a first-person benchmark and concrete deltas. Scope is narrow: one RTX 3090, one model, and a Reddit source, so it stays near the featured threshold.

r/LocalLLaMA

Public repository Codegraph claims 94% fewer Claude, Cursor, Codex, and OpenCode tool calls locally

Codegraph uses a pre-indexed knowledge graph for symbol relationships, call graphs, and code structure. In the VS Code test, it reduced tool calls from 52 to 3 and runtime from 1m37s to 17s.

Why it matters: All HKR axes pass, but evidence is a Reddit/public-repo self-test without independent replication. The 94% reduction and 52→3 call count clear featured, not p1.

May 19Tuesday

Hacker News front page

Show HN: Forge takes an 8B model from 53% to 99% on agentic tasks

Forge adds five guardrail layers to self-hosted LLM tool calling, raising Ministral 8B to 99.3% across 18 multi-step agentic scenarios, with the accepted ACM CAIS ’26 paper covering 97 model/backend configurations and 50 runs per scenario.

Why it matters: HKR-H/K/R all pass: the 53%→99.3% jump is clickable, the test setup has concrete numbers, and self-hosted agent reliability is a live practitioner pain. Single-source Show HN/GitHub evidence keeps it in the 78–84 open-source-tool band, not P1.

r/LocalLLaMA

ByteDance released an open-source model that attempts broad multimodal tasks with 3B parameters

ByteDance released Lance, an open-source unified multimodal model with 3B active parameters that supports image and video understanding, generation, and editing, and the post says it was trained from scratch with a staged multi-task recipe under a 128-A100-GPU budget.

Why it matters: HKR-H/K/R all pass: ByteDance’s open Lance has a compact multimodal hook, concrete 3B/128-A100 facts, and clear cost/deployment resonance. Reddit-sourced details lack benchmarks, license terms, and official context, so it stays featured, not P1.

AI HOT (Curated Pool)

Horizon Open-Sources 400M-Parameter Robot Control Model HoloMotion-1

Horizon Robotics Lab open-sourced HoloMotion-1, a 400M-parameter full-body humanoid control model that uses MoE sparse activation and KV-cache inference to reach about 300 FPS on-device, with code and a technical report released.

Why it matters: HKR-H/K/R all pass: HoloMotion-1 has an open-source robotics hook plus 400M params and about 300FPS edge inference. Its reach is narrower than a frontier model release, so it fits the 78 featured band.

AI HOT (Curated Pool)

Cursor releases Composer 2.5, calling it its strongest model yet

Cursor released Composer 2.5, claiming a 10x efficiency gain at comparable capability, with larger training scale, more complex reinforcement-learning environments, and a text-feedback mechanism.

Why it matters: Cursor Composer 2.5 is a substantive model update for a front-line AI coding tool, with HKR-H/K/R from the 10x efficiency and RL-training details. The single social-source summary lacks benchmarks, pricing, and reproducible tests, keeping it in the 78–84 band.

r/LocalLLaMA

21 GPUs benchmarked running a small TTS model, with 5GB peak VRAM

A Reddit user rented 21 GPUs on vast.ai to benchmark OmniVoice, a small TTS model with about 5GB peak VRAM, using xRT as the audio generation speed metric and averaging 3 voice-cloning runs with reference audio.

Why it matters: HKR-H/K/R pass: a 21-GPU TTS benchmark with 5GB peak VRAM and 3-run xRT averaging is useful to local-inference builders. Scope is niche, so it sits at the low featured band.

May 18Monday

AI HOT (Curated Pool)

The Open Agent Leaderboard

IBM Research published the Open Agent Leaderboard on Hugging Face to evaluate agents across language understanding, tool use, and multi-step reasoning tasks; the post does not disclose dataset size, model scores, or the evaluation date.

Why it matters: HKR-H and HKR-R pass because an open agent leaderboard speaks to agent-eval pain. HKR-K fails: the article lacks scores, dataset size, and evaluation date, so it sits at the featured threshold.

r/LocalLLaMA

I Tested 42 LLMs on Their Willingness to Build the Apocalypse

DystopiaBench tested 42 open and closed models across 36 escalating scenarios and 6 dystopia types, using 3 LLM-as-judge scorers and an average over 3 runs; the post says many models catch obvious dangerous requests but fail when risk is hidden behind dual-use framing and normalization.

Why it matters: HKR-H/K/R all pass: the hook is sharp, the test setup has concrete numbers, and the topic hits safety trust. Reddit single-post sourcing and limited disclosed results keep it in featured, not P1.

r/LocalLLaMA

Qwen 3.6 27B on 24GB VRAM: backend comparisons, quant choice, and settings

The author tested Qwen 3.6 27B on an RTX 3090 24GB and kept ik_llama.cpp with Qwen3.6-27B-MTP-IQ4_KS.gguf; at 156k context with q8_0 KV and MTP, a ~5.9k-token prompt plus 1024-token output reached about 1261 tok/s prefill and 72.9 tok/s decode, while vLLM lacked a clean single-card long-context run.

Why it matters: HKR-H/K/R all pass: this is a first-person local inference benchmark with concrete VRAM, quant, context, and speed numbers. Its reach is narrower than a model release, so it sits at the featured threshold.

r/LocalLLaMA

I built a coding agent that gets 87% on benchmarks with a 4B parameter model

SmallCode passes 87 of 100 benchmark tasks with Gemma 4 activating 4B parameters per token. The author attributes the result to compound tools, compile and lint feedback, task decomposition after two repeated failures, and optional escalation to Claude or OpenAI for one task.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post and the benchmark identity plus replication details are incomplete. It fits a concrete first-person experiment above the featured bar, not the 78+ band.

r/LocalLLaMA

Benchmarking vLLM vs SGLang vs llama.cpp on a mixed Blackwell/Ada cluster

The author benchmarked long-context prefill on a 7-GPU mixed Blackwell/Ada cluster; on Qwen3.5-397B-A17B with 75k tokens, vLLM reached 9.8s TTFT and 7,683 t/s, while llama.cpp took 57.2s and 1,319 t/s.

Why it matters: Single-source Reddit benchmark, so source authority keeps it near the threshold. HKR-H/K/R pass on the mixed 7-GPU setup, 397B at 75k tokens, and concrete TTFT/throughput numbers.

r/LocalLLaMA

LLMs on Android: Snapdragon 8 Elite MoE Experience

A Reddit user tested MoE LLMs on an Honor Magic 7 Pro with Snapdragon 8 Elite and 24GB RAM; under Q4 quantization, LFM2-24b-a2b reached about 24 tokens/s while Gemma reached about 11 tokens/s, and CPU inference was still faster than NPU or GPU in the reported setup.

Why it matters: HKR-H/K/R all pass: a named Reddit test gives hardware, quantization, and token/s figures. Single-device anecdote and weak source authority keep it at the low featured band.

Google DeepMind

Google DeepMind releases Gemini Omni Flash video model

Google DeepMind released Gemini Omni Flash, the first model in the Gemini Omni family. It combines image, audio, video and text inputs to generate high-quality video, and supports multi-turn editing in natural language.

Why it matters: Gemini Omni Flash folds video generation and conversational editing into one model, a shift in how multimodal creation gets accessed.

May 17Sunday

r/LocalLLaMA

MiroThinker-1.7 Open-Weight Deep Research Agent Based on Qwen3 MoE

MiroMindAI released the MiroThinker-1.7-deepresearch and mini APIs, with the mini version using 30B total parameters and 3B active parameters, weights on HuggingFace, and context management based on sliding window K=5 plus episode restarts.

Why it matters: HKR-H/K/R all pass, but the source is a Reddit thread and the lab is not top-tier. Open weights, MoE sizing, and context-management details clear featured, not same-day must-write.

r/LocalLLaMA

85 GPU-hours comparing 5 abliteration methods on Qwen3.6-27B

Abliterlitics compared five Qwen3.6-27B abliteration variants against the base model using 85 GPU-hours of benchmarks, HarmBench, KL divergence, and weight forensics; Huihui had the smallest benchmark deltas, Heretic had the lowest KL divergence, and all five variants reached near-complete safety removal.

Why it matters: HKR-H/K/R all pass: the post gives an 85-GPU-hour comparison across five abliteration methods on Qwen3.6-27B. Niche open-model safety work, not a lab release, so it stays at the featured threshold.

r/LocalLLaMA

DeepSeek V4's 1M Context Window: The Breaking Point

A Reddit user tested DeepSeek V4 on 45k, 180k, and 520k-token codebases and found 150k-250k tokens best for coding work. Past 300k tokens, line-number precision degraded; at 520k, outputs shifted toward architecture summaries and skipped implementation details.

Why it matters: A single Reddit post limits authority, but HKR-H/K/R all pass: it is a numbered first-person test with a concrete long-context failure pattern. The right band is featured, not 78+, because replication and model details are thin.

r/LocalLLaMA

Same Models Tested Across Strix Halo, RTX 3090, and RTX 5070

C_Coffie published 55 local inference benchmark runs across Strix Halo, RTX 3090, RTX 5070, five backends, and 0.35B to 35B-A3B models; RTX 5070 beats RTX 3090 on models fitting 12GiB, while RTX 3090 leads in the 14–31B band that exceeds 12GiB but fits 24GiB.

Why it matters: Hits HKR-H/K/R with a named first-person benchmark: 55 runs and concrete GPU crossover points. Source is a single Reddit post, so it stays in the low featured band.

AI HOT (Curated Pool)

Latest Open Artifacts #21: Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1, and More

Open AI model teams released Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1, and other versions this month, and the post says they were tested under CAISI’s V4 evaluation framework, but the RSS snippet does not disclose scores.

Why it matters: HKR-H/K/R all pass: a dense open-model roster, a named CAISI V4 evaluation frame, and clear practitioner relevance for model choice. Missing scores and reproducible detail keep it in the 78–84 band.

AI HOT (Curated Pool)

Ring-2.6-1T Open-Sourced and Listed on OpenRouter for Agent Workflows

AntLingAGI open-sourced Ring-2.6-1T and listed it on OpenRouter with a 75% discount through the end of May; the trillion-scale reasoning model targets agent workflows, including planning, tool use, context maintenance, and complex task execution, using Async RL and IcePop training methods.

Why it matters: HKR-H/K/R all pass: a 1T open agent model is clickable, with OpenRouter access, discount, and training methods disclosed. Score stays at 74 because benchmarks, license, and context window are not given.

May 16Saturday

Hacker News front page

SANA-WM, a 2.6B open-source world model for 1-minute 720p video

SANA-WM’s title says the project is a 2.6B open-source world model for 1-minute 720p video; the RSS body only lists the project URL, Hacker News comments URL, 9 points, and 8 comments, and the post does not disclose training data, license terms, inference cost, evaluation setup, or benchmark results.

Why it matters: HKR-H/K/R pass on the concrete open-source world-model hook, 2.6B size, and video-model competition angle. Sparse body details keep it at the lower good-quality band.

r/LocalLLaMA

Qwen3.6-35B-A3B and 9B land on the public Terminal-Bench 2.0 leaderboard

little-coder × Qwen3.6-35B-A3B scored 24.6% ±3.2 on Terminal-Bench 2.0, above Gemini 2.5 Pro on Gemini CLI at 19.6% and Qwen3-Coder-480B on Terminus 2 at 23.9%.

Why it matters: HKR-H/K/R all pass, but this is a Reddit post with leaderboard numbers only; test setup and reproducibility details are not disclosed. Strong code-agent benchmark signal, not a 78+ release story.

Google DeepMind

Google DeepMind releases Gemini 3.5 Flash

Google DeepMind released the Gemini 3.5 model family, with the first model, Gemini 3.5 Flash, available the same day in the Gemini app, Google Search AI Mode, Google Antigravity, the Gemini API and Gemini Enterprise.

Why it matters: Google published 3.5 Flash's coding and agent benchmark scores and where it is available, enough to judge its place in long-horizon workflows.

The Verge · AI

AI radio hosts demonstrate why AI can’t be trusted alone

Andon Labs had Claude, ChatGPT, Gemini, and Grok run separate radio stations with $20 in seed money each; the RSS snippet says all failed, but the post does not disclose the full experimental results.

Why it matters: HKR-H/R are strong because the agent-failure setup is memorable and relevant. HKR-K is present but thin: it gives four models and $20 budgets, while full experimental results are not disclosed.

May 15Friday

r/LocalLLaMA

Evaluated a RAG Chatbot: The Most Expensive Model Was the Worst Performer

The author evaluated a customer-support RAG bot and raised the quality score from 6.62 to 7.88 while cutting per-session cost from $0.002420 to $0.000509, using retrieval logging, LLM-as-judge scoring, chunk deduplication, stricter grounding, and a five-model sweep.

Why it matters: HKR-H/K/R all pass: counterintuitive model ranking, concrete quality and cost deltas, and direct RAG production relevance. Reddit source authority keeps it near the featured floor despite the first-person experiment signal.