Skip to content

All news

8 today

May 15Friday

r/LocalLLaMA

Used over a million tokens in three sessions to test Qwen 3.6 35B MTP

A Reddit user tested Qwen3.6-35B-A3B MTP across three million-token-scale sessions, using 300k context and KV Q8_0, and reported about 1.5x the tok/sec of earlier tests.

Why it matters: HKR-H/K/R all pass: the million-token test is clickable, 300k context and KV Q8_0 add testable detail, and local speed maps to cost. Source is one Reddit post, so it stays below the high-importance band.

QbitAI · WeChat

Understand LeCun’s JEPA World Model in 160 Lines of Code

A developer released the keon/jepa teaching repository with five JEPA variants implemented as standalone PyTorch files, ranging from 160 to 278 lines, depending only on PyTorch and torchvision; the post reports iJEPA runs on CIFAR-10 for 100 epochs and reaches 52.7% linear-probe accuracy, while V-JEPA, C-JEPA, and LeWorldModel use toy or synthetic datasets.

Why it matters: HKR-H/K/R pass via the 160-line JEPA hook, reproducible repo, and non-LLM world-model angle. It is a tutorial artifact, not a model or paper release, so it sits at the featured threshold.

AI HOT (Curated Pool)

Granite Embedding Multilingual R2: Open Multilingual Embedding Model with 32K Context

IBM Granite released Granite Embedding Multilingual R2 on Hugging Face under Apache 2.0, with fewer than 100 million parameters, a 32K-token context length, and top same-scale retrieval performance on MTEB according to the post.

Why it matters: HKR-H/K/R pass: the 32K-context, sub-100M multilingual embedding model gives RAG builders a concrete open-source option. Impact is narrower than a frontier-model release, so it sits at the featured threshold.

May 14Thursday

AI HOT (Curated Pool)

MiMo V2.5 Pro Places Third on DesignArena

MiMo V2.5 Pro placed third on the DesignArena overall leaderboard; its Thinking version rose 8 spots over MiMo-V2.5 and matched Claude Sonnet 4.6 performance on frontend coding tasks.

Why it matters: HKR-H/K/R all pass, but the facts come from one official X post with no methodology, access, or pricing. This fits a mid-weight benchmark/product update, not a same-day must-write.

Xinzhiyuan · WeChat

Anthropic Overtakes OpenAI in Enterprise AI Adoption After Three Years

Ramp says Anthropic reached 34.4% enterprise adoption, surpassing OpenAI at 32.3% for the first time; the index is based on credit-card and invoice spending from more than 50,000 companies.

Why it matters: HKR-H/K/R all pass: a reversal hook, concrete 34.4%/32.3% figures, and a strong enterprise-AI rivalry angle. Score stays at 80 because Ramp spending data is not global market share.

AI HOT (Curated Pool)

OpenSquilla Open-Source Project Uses Smart Routing and Local Retrieval to Cut LLM Costs

OpenSquilla combines local model routing, vector retrieval, incremental sending, and cache hits to reduce transmitted tokens by more than 90%, while routing simple tasks to cheaper models and complex tasks to stronger models without spending tokens on the routing decision.

Why it matters: HKR-H/K/R all pass, but the source appears to be a single X project post; repo traction, test setup, and limits are not disclosed. Score lands at the featured threshold for practical open-source cost tooling.

AI HOT (Curated Pool)

Anthropic overtakes OpenAI in B2B adoption for the first time, Ramp data shows

Ramp AI Index data shows Anthropic reached 34.4% adoption among U.S. enterprise customers, surpassing OpenAI’s 32.3% for the first time, while its business coverage grew fourfold over one year.

Why it matters: HKR-H/K/R all pass: Ramp reports Anthropic at 34.4% enterprise adoption versus OpenAI at 32.3%. This is a strong market signal, but it is one spending dataset rather than a model or product release, so it stays in 78–84.

r/LocalLLaMA

sensenova/SenseNova-U1-A3B-MoT · Hugging Face

SenseNova published SenseNova-U1-A3B-MoT on Hugging Face; the post lists A3B MoT, 8B MoT, and 0.4B LoRA weight links, and says the NEO-unify architecture unifies multimodal understanding, reasoning, and generation in one model family.

Why it matters: HKR-H/K/R all pass: an open multimodal model release with multiple weight sizes and a named NEO-unify mechanism. Source authority and missing benchmarks/license details keep it in the lower featured band.

May 13Wednesday

TechCrunch · AI

Anthropic now has more business customers than OpenAI, according to Ramp data

Ramp’s survey of client expense data shows 34.4% of participating businesses pay for Anthropic services, while 32.3% pay for OpenAI; the snippet does not disclose sample size, customer segments, or spend levels.

Why it matters: HKR-H/K/R all pass: Ramp reports 34.4% paid business usage for Anthropic versus 32.3% for OpenAI. The sample is one payments platform, not official revenue or full-market share, so it sits at the featured threshold.

Xinzhiyuan · WeChat

Tsinghua-affiliated team open-sources MiniCPM-V 4.6, a 1.3B model tunable on one RTX 4090

ModelBest, Tsinghua University, and OpenBMB open-sourced MiniCPM-V 4.6, a 1.3B multimodal model that supports full fine-tuning on one RTX 4090 and offers 4x/16x visual token compression for accuracy or speed trade-offs.

Why it matters: HKR-H/K/R all pass: the story gives a concrete open-source multimodal release with size, hardware condition, and token-compression details. It lowers local fine-tuning cost, but it is not a frontier-lab flagship release, so 78–84 fits.

r/LocalLLaMA

A real transformer language model running locally on a stock Game Boy Color

maddiedreese ran Andrej Karpathy’s TinyStories-260K on a stock Game Boy Color with INT8 weights, fixed-point math, an MBC5 ROM, bank-switched cartridge storage, and KV cache in cartridge SRAM; the demo uses no phone, PC, Wi‑Fi, link cable, or cloud inference, but output is extremely slow and gibberish.

Why it matters: HKR-H/K/R all pass: a named first-person experiment with concrete model and memory details. Impact stays low-featured because it is a Reddit hardware hack with slow, garbled output, not a usable product or model release.

Hacker News front page

Show HN: Needle Distills Gemini Tool Calling into a 26M Model

Cactus open-sourced Needle, a 26M-parameter tool-calling model that reaches 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, with MIT-licensed weights released on Hugging Face.

Why it matters: HKR-H/K/R all pass: the tiny Gemini-style tool-calling angle is clickable, with concrete speed and license claims. Source is still Show HN/GitHub self-reporting, not an independent benchmark or major lab release, so it stays below the 78–84 band.

r/LocalLLaMA

Needle: We Distilled Gemini Tool Calling Into a 26M Model

Cactus Compute open-sourced Needle, a 26M-parameter tool-calling model that reaches 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, using an attention-and-gating architecture with no MLPs.

Why it matters: HKR-H/K/R all pass: a 26M tool-calling model has a strong hook and concrete speed/design claims. Single Reddit source and a less-known team keep it in the lower 78–84 band.

May 12Tuesday

AI HOT (Curated Pool)

How Open Model Ecosystems Compound

China’s open AI model community forms a self-reinforcing loop, with domestic open model downloads rising by more than 200% quarter over quarter.

Why it matters: HKR-H/K/R all pass: the flywheel framing is clickable, the article gives a >200% QoQ download claim, and the topic hits China open-model competition. It is strong commentary, not a model launch, so 78 featured.

Synced · WeChat

ByteDance Open-Sources DreamLite for Offline Mobile Image Generation and Editing

ByteDance open-sourced DreamLite, a 0.39B-parameter unified diffusion model that generates or edits a 1024×1024 image on an iPhone 17 Pro in about 3 seconds, using 4-step DMD2 distillation and on-device offline inference without cloud dependency.

Why it matters: HKR-H/K/R all pass: 3-second on-device 1024×1024 generation is a strong hook, with 0.39B params and 4-step DMD2 as concrete claims. As a ByteDance open-source vision model, it sits below a general foundation-model release.

AI HOT (Curated Pool)

Thinking Machines Releases Native Multimodal Interaction Model for Real-Time Human-AI Collaboration

Thinking Machines released an interaction model that natively receives audio, video, and text input, processes foreground interaction at 200-millisecond intervals, and uses a background reasoning model for long-horizon planning and tool calls.

Why it matters: HKR-H/K/R all pass: this is more than a model notice, with a two-layer foreground/background interaction design. Pricing, access scope, and benchmarks are missing, so it sits at the lower end of 85-94.

r/LocalLLaMA

I catalogued every way local models break JSON output and built a repair library across 288 model calls

Reddit user kexxty ran 288 structured-output calls through OpenRouter models, including Llama 3, Mistral, Command R, DeepSeek, and Qwen, and found similar JSON failure categories across local and API-only models. The MIT-licensed Python library outputguard validates against JSON Schema, applies 15 ordered repair strategies, includes 2,001 tests, and has no LLM provider dependency.

Why it matters: HKR-H/K/R all pass: 288 tests, the outputguard library, and a 15-step repair chain give practitioners reusable detail. Source is a single Reddit post, so it stays in the 72–77 featured band, not 78+.

May 11Monday

AI HOT (Curated Pool)

Fields Medalist Tests ChatGPT 5.5 Pro: Paper-Level Result in 17 Minutes

Timothy Gowers tested ChatGPT 5.5 Pro and said it independently solved an open additive number theory problem in 17 minutes with only a simple prompt, producing PhD thesis-level work; he warned that this pace threatens mathematics research training, while Terence Tao said human value lies in digesting and deeply understanding proofs.

Why it matters: HKR-H/K/R all pass: a named mathematician, a 17-minute result, and a PhD-training warning. The exact problem, prompt, and verification path are not disclosed, keeping it below P1.

AI HOT (Curated Pool)

Pareto Code Reorders Model Selection Using Market Demand

OpenRouter says Pareto Code observes the Pareto frontier using real market demand; DeepSeek V4 Pro ranks first, followed by GPT 5.4 Mini and Gemini 3.1 Pro, while the post does not disclose the scoring formula or evaluation sample size.

Why it matters: HKR-H/K/R all pass, but the source is a single OpenRouter post with no sample size, time window, or pricing basis disclosed. It clears featured as a model-selection benchmark, not the 78+ band.

AI HOT (Curated Pool)

AntLingAGI Releases Trillion-Parameter Ring-2.6-1T Model

AntLingAGI released Ring-2.6-1T, a trillion-parameter thinking model available for free on OpenRouter until May 15, with adjustable thinking intensity, agent-oriented multi-step execution, tool calling, and tasks covering math logic and scientific research.

Why it matters: HKR-H/K/R all pass, but the post is thin: no benchmarks, pricing, architecture, or training details. Treat as a mid-weight model launch on OpenRouter, not a same-day must-write.

Xinzhiyuan · WeChat

The Second Half of Agent Evaluation: Why a Live Benchmark Is Needed

Claw-Eval-Live evaluates 13 frontier models on 105 tasks, and the top model stays below a 70% pass rate, while HR tasks average only 6.8% pass rate.

Why it matters: HKR-H/K/R all pass: the live benchmark hook is specific, and the post gives 105 tasks, 13 models, HR at 6.8%. Claw-Eval-Live still lacks proven field impact, so this sits in the lower featured band.

Xinzhiyuan · WeChat

Claude Mythos Hits 50% Success on 16-Hour Tasks in METR Time Horizons

Claude Mythos Preview reached a 50% success rate on METR Time Horizons tasks that take humans 16 hours, while only 5 of 228 tasks exceeded the 16-hour range, so the article says METR lacks enough samples to quantify longer-horizon performance.

Why it matters: HKR-H/K/R all pass: the 16-hour task result is a strong hook, and the METR sample caveat adds substance. Capped at 82 because only 5 tasks exceed 16 hours, so the 2027 extrapolation is not same-day P1 material.

AI HOT (Curated Pool)

Local models handle half of daily tasks and respond faster than cloud models

A five-week experiment tested about 1,400 daily work tasks, where local 35B models such as Qwen 3.6 35B handled about 50% and averaged 2.8-second responses, 2.1 times faster than Claude Opus 4.5, while the cloud model still led complex reasoning by about 20%.

Why it matters: HKR-H/K/R all pass: Tom Tunguz’s experiment reports ~1,400 tasks, ~50% success, 2.8s latency, and a speed comparison to Claude Opus 4.5. Strong practitioner signal, but not a model launch or platform-level update.

r/LocalLLaMA

MTP benchmark results: task type determines speculative inference speedups or slowdowns

A Reddit LocalLLaMA user ran 300+ tests on Qwen 3.6 27B MTP quants, finding coding draft acceptance at 79-89% and F16 coding speed up 171%, while Q4_K_M creative writing slowed down 9%.

Why it matters: HKR-H/K/R all pass: this is a single Reddit experiment, not a market event, but 300+ Qwen 3.6 27B MTP quantization tests give practical numbers for local inference tuning.

May 9Saturday

AI HOT (Curated Pool)

Redis founder uses a C inference engine to run a large model on a personal computer

Antirez open-sourced ds4, a native inference engine for DeepSeek V4 Flash that uses a few thousand lines of C to run a 1M-context model on a 128GB MacBook Pro at a reported 27 tok/s.

Why it matters: HKR-H/K/R all pass: Antirez open-sourced a native C inference engine with hardware, model, context, and speed numbers. Single-source X provenance keeps it below P1, but it is strong open-source inference signal.

r/LocalLLaMA

80 tok/sec and 128K context on 12GB VRAM with Qwen3.6 35B A3B and llama.cpp MTP

Reddit user janvitos ran Qwen3.6-35B-A3B-MTP-GGUF with a llama.cpp MTP PR on an RTX 4070 Super. The posted benchmark shows 69.2-81.9 tok/s, 0.694-0.947 draft acceptance, 131072 context, and a -fitt 1536 setting that reserves 1536 MB for the draft model and KV cache.

Why it matters: HKR-H/K/R all pass with concrete single-user benchmark data and reproducible settings. Source is one Reddit post, so verification is thin; this lands above featured threshold, not in must-write range.

AI HOT (Curated Pool)

Baidu releases ERNIE 5.1 with compressed parameters and training cost

Baidu released ERNIE 5.1 with total parameters reduced to about one third of the original scale, active parameters to about one half, and pretraining cost to about 6% of same-scale models; the model is available on the ERNIE platform and Baidu AI Studio.

Why it matters: HKR-H/K/R all pass: Baidu ERNIE 5.1 is a domestic flagship-model release with concrete compression and 6% pretraining-cost claims. That puts it in the must-write band.

AI HOT (Curated Pool)

ERNIE 5.1 Released With Pretraining Cost at 6% of Comparable Models

Baidu released ERNIE 5.1, saying it builds on ERNIE 5.0 pretraining and improves search, reasoning, knowledge QA, creative writing, and agent capabilities, with pretraining cost at about 6% of comparable models.

Why it matters: Baidu released ERNIE 5.1 with a concrete “6% of reference pretraining cost” claim. HKR-H/K/R all pass, with a domestic flagship-model bump, but sparse technical detail keeps it below the 90s.

Xinzhiyuan · WeChat

CUHK Open-Sources ArbiterOS Agent Governance Kernel With 92.95% High-Risk Interception

CUHK CURE Lab open-sourced ArbiterOS, an agent runtime governance kernel that intercepts, parses, governs, and observes actions before execution, raising high-risk step interception on OpenClaw tasks from 6.17% to 92.95%.

Why it matters: HKR-H/K/R all pass: the story has a sharp execution-control hook, a concrete 6.17%→92.95% result, and clear agent-safety resonance. It is a strong open-source research tool, not a top-lab model release, so it stays in the 78–84 band.

Synced · WeChat

StarVLA Open-Sources a Unified VLA Framework from HKUST and the Community

HKUST and the open-source community released StarVLA, a unified Vision-Language-Action framework that integrates backbones, action heads, training strategies, and evaluation interfaces; the repository has 2.2k GitHub stars and supports benchmarks including LIBERO, SimplerEnv, RoboTwin 2.0, RoboCasa-GR1, and BEHAVIOR-1K.

Why it matters: HKR-H/K/R all pass: StarVLA ships a concrete open-source VLA framework with unified interfaces, 2.2k stars, and named robotics benchmarks. The robotics scope keeps it in the 78–84 band, below model-release weight.

AI HOT (Curated Pool)

Claude Mythos Evaluation Shows 16-Hour Risk Horizon

METR evaluated an early Claude Mythos Preview build during a limited March 2026 window and estimated its 50% time horizon at at least 16 hours, with a 95% confidence interval of 8.5 to 55 hours.

Why it matters: HKR-H/K/R all pass: METR reports a concrete 16h risk-horizon estimate for Claude Mythos Preview. The single X-source and limited eval window keep it below P1, but it is strong featured safety signal.

May 8Friday

r/LocalLLaMA

Gemma 4 26B Hits 600 Tok/s on One RTX 5090

chain-77 benchmarked Gemma 4 26B with vLLM 0.19.2rc1, and DFlash raised output throughput on one RTX 5090 from 228 tok/s to 578 tok/s under 256 input tokens, 1024 output tokens, concurrency 1, and num_speculative_tokens=13.

Why it matters: HKR-H/K/R all pass: the single-GPU throughput hook is strong, and the post gives reproducible settings plus before/after speed. Reddit single-post evidence and one hardware setup keep it in the featured-threshold band.

r/LocalLLaMA

11.67% ARC-AGI-2 Local Eval on a Single 4090: The TOPAS Recursive Architecture

Doug_Bitterbot says TOPAS scored 11.67% on ARC-AGI-2 using one RTX 4090 after about 14 days of training. The 100M-parameter checkpoint hit 36% locally, but recursive TTT caused null outputs on nearly half of Kaggle puzzles. The key detail is time management: the author expects 20% after threshold tuning and 3-5 more weeks of training.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post with unstable Kaggle submissions. It clears featured, not the higher research-release band.

May 7Thursday

AI HOT (Curated Pool)

SenseNova-U1 Open-Sources 8-Step Distilled LoRA, Speeds Diffusion Inference by 11x

SenseNova-U1 open-sourced an 8-step distilled LoRA that cuts diffusion generation from 100 steps to 8. GPU inference time drops from 23 seconds to 2 seconds, with ComfyUI workflows for text-to-image, image editing, and interleaved generation. The key signal is distillation for latency, not parameter scale.

Why it matters: HKR-H/K/R all pass: the 11x speedup hooks attention, the post gives step and latency numbers, and open LoRA affects diffusion deployment cost. Scope stays within image generation, so this is featured, not P1.

OpenAI News

Advancing Voice Intelligence with New Models in the API

OpenAI introduced new realtime voice models in its API for voice intelligence. The RSS snippet says they reason, translate, and transcribe speech; the post does not disclose counts, pricing, or limits.

Why it matters: OpenAI’s official voice API update hits HKR-H/K/R, but the available body gives capability direction only. Model count, pricing, latency, and context limits are not disclosed, so it stays at the top of 78–84.

Synced · WeChat

Claude, GPT and Gemini score 0% completion on ProgramBench

ProgramBench tested Claude Opus 4.7, GPT-5.4 and Gemini 3.1 Pro, with 0% full completion on rebuilding software projects. It gives only executables and usage docs, removes source/tests, and grades behavioral equivalence via agent-driven fuzzing. The key signal is system-level engineering, not function-level code generation.

Why it matters: HKR-H/K/R all pass: the 0% result is clickable, the setup is concrete, and the coding-agent gap matters to practitioners. Still, it is a single benchmark report, below a major model or product release.

r/LocalLLaMA

Exaggerated PCI-E Bandwidth Concerns?

Reddit user ziphnor tested 2x RTX 5060 Ti 16GB with vLLM TP=2 and 32k-context prefill. PCIe peaked at 3–4 GB/s, about 40–50% of a PCIe 4.0 x4 link. Prefill reached ~840–850, 1500, and 1600–1700 t/s; the post does not disclose decode bandwidth.

Why it matters: HKR-H/K/R all pass: a myth-busting PCIe bandwidth test with concrete vLLM conditions and numbers. Single Reddit source limits authority, but the named first-person experiment lifts it to the featured threshold.

r/LocalLLaMA

Analysis of 922 Agentic Task Traces Finds DeepSeek v4’s Cost Edge in Caching

A Reddit user analyzed 922 agentic task traces and reported $0.01 per task for DeepSeek v4 Flash versus $1.52 for Opus 4.7. Both used about 960K tokens per task, but DeepSeek showed a 97% cache hit rate versus 87%, with a 0.02 cache read/write price ratio versus 0.08. The key issue is caching, not headline pricing.

Why it matters: HKR-H/K/R all pass: 922 agent traces tie a large cost gap to cache hit rate and cache read/write pricing. Reddit single-source data and incomplete method detail keep it in the 78–84 band.

May 6Wednesday

r/LocalLLaMA

Qwen3.6 27B NVFP4 + MTP on a Single RTX 5090: 200k Context in vLLM

A Reddit user ran Qwen3.6 27B NVFP4 on one RTX 5090 32GB and validated 200k context in vLLM. The setup used fp8_e4m3 KV cache, FlashInfer, and MTP with 3 speculative tokens; a 10-run 200k pass completed with 73.6 tok/s mean generation and 70.2s TTFT. The key constraint is 32GB VRAM: logs showed 8.3GiB KV cache and about 30478MiB total GPU use.

Why it matters: HKR-H/K/R all pass: the hook is single-GPU 200k context, with concrete vLLM settings and 10-run stability data. Reddit sourcing keeps it in the 78–84 band, not P1.

r/LocalLLaMA

An Open Benchmark for Testing RAG on Realistic Company-Internal Data

EnterpriseRAG-Bench released a 500k-document corpus for testing RAG on company-internal data. It simulates Redwood Inference across 9 sources and includes 500 questions over 10 retrieval failure modes. Baselines show BM25 beats vector search overall, while agentic/bash retrieval has the best completeness at higher cost and latency.

Why it matters: HKR-H/K/R all pass: the benchmark targets a real enterprise RAG pain point, with 500k docs and testable BM25-vs-vector results. Single Reddit-source benchmark release keeps it below same-day must-write.