Skip to content

All news

8 today

Jun 3Wednesday

AI HOT (Curated Pool)

Intelligence Cost-Performance

Microsoft added average token usage to its model release card; the model scored 71.6 on SWE-Bench Verified while using about one-third of Claude Haiku 4.5’s tokens.

Why it matters: HKR-H/K/R all pass: the score-per-token angle is clickable, with concrete 71.6 and one-third-token claims. The article is thin on full test setup and pricing, so it lands at 78.

Hacker News front page

Microsoft's MAI-Code-1-Flash Scores 51% SWE-Bench Pro with Just 5B Active Params

The title says Microsoft's MAI-Code-1-Flash scores 51% on SWE-Bench Pro with 5B active parameters; the post does not disclose the evaluation setup, training data, release date, or deployment conditions.

Why it matters: HKR-H/K/R pass on the 51% SWE-Bench Pro with 5B active params claim from Microsoft. Missing eval setup, training data, and release timing keep it in the 72–77 band.

AI HOT (Curated Pool)

Microsoft releases its first advanced reasoning AI model, MAI-Thinking-1

Microsoft released MAI-Thinking-1 at Build 2026, describing it as a medium-sized reasoning model that matches leading models on key software engineering benchmarks.

Why it matters: HKR-H/K/R all pass: Microsoft released its first advanced reasoning model with a mid-sized design and SWE benchmark claim. Exact scores, access, and pricing are not disclosed, so it stays below 85.

r/LocalLLaMA

Using Gemma 4 E4B with LiteRT: about 2.4× faster text generation than Q4 GGUF

The author tested Gemma 4 E4B on an RTX 4060 Ti 16GB, where LiteRT averaged 157.2 tok/s for text generation versus 66.3 tok/s for llama.cpp Q4 GGUF; image captioning on 111 full-resolution images improved only 1.1×, at about 72 seconds versus 80 seconds.

Why it matters: HKR-H/K/R all pass, with a first-person benchmark including hardware, throughput, and sample count. Source authority is limited to one Reddit test, so it sits at the featured threshold rather than the 78+ band.

r/LocalLLaMA

Benchmarks of 20 Small LLMs on a 6GB RTX 4050

The author benchmarked 20 small LLMs on a 6GB RTX 4050 using LM Studio’s OpenAI-compatible API, with N=5 speed runs at 1k, 8k, and 32k context; unsloth/lfm2.5-vl-1.6b led throughput at 207 tok/s on 1k context while using 3.0GB VRAM.

Why it matters: HKR-H/K/R all pass: the low-VRAM GPU hook is concrete, the post gives speed/context/VRAM numbers, and it speaks to local-inference cost pressure. Source authority is a Reddit post, so it stays in the lower featured band.

Jun 2Tuesday

AI HOT (Curated Pool)

StepFun releases Step 3.7 Flash as an open-weight model for agentic coding

StepFun released the open-weight Step 3.7 Flash model for fast agentic coding, with tool calling and multimodal understanding, and the model is already available in Kilo alongside MiniMax M3.

Why it matters: HKR-H/K/R pass on the open-weight agentic-coding angle and Kilo availability. Missing benchmarks, size, license, and pricing keep it at the lower featured threshold.

r/LocalLLaMA

Replaced Claude with local Qwen3.6-27B in my multi-agent orchestrator for 2 weeks

The author ran Qwen3.6-27B on one RTX 3090 across 47 multi-step coding workflows. Plan generation reached about 95% schema validity, but tool-call formatting errors were about 12%, and practical long-context use degraded past about 12k tokens.

Why it matters: HKR-H/K/R all pass: a named first-person local-vs-Claude experiment with concrete numbers. The single Reddit source and 47-workflow scope keep it below the 78–84 band.

AI HOT (Curated Pool)

NVIDIA Cosmos 3 Tops Open-Weight Image and Video Generation Rankings

NVIDIA Cosmos 3 ranked first in Artificial Analysis’s open-weight text-to-image and image-to-video categories, with 16B Nano and 64B Super variants, and the release includes weights, code, curated datasets, and fine-tuning recipes under the OpenMDW 1.1 license.

Why it matters: HKR-H/K/R all pass: Cosmos 3 leads both Artificial Analysis open-weight image and video charts, with 16B/64B variants and OpenMDW 1.1 artifacts disclosed. Single-source benchmark news keeps it in the 78–84 featured band.

Jun 1Monday

AI HOT (Curated Pool)

MiniMax Releases Open-Source M3 with Coding, Long-Context, and Multimodal Capabilities

MiniMax released the open-source M3 model with coding, a 1M-token context window, and native multimodal support; M3 scores 59.0% on SWE-Bench Pro, 83.5% on BrowseComp, and costs about one-twelfth per token versus GPT-5.5.

Why it matters: HKR-H/K/R all pass: M3 has open source, 1M context, multimodal support, and 59.0% on SWE-Bench Pro. A single X post without official docs or third-party tests keeps it in the 78–84 band.

AI HOT (Curated Pool)

NVIDIA Open-Sources Cosmos 3, Its First Generalist Model for Physical AI

NVIDIA open-sourced Cosmos 3 at GTC Taipei, releasing two variants, Super 32B and Nano 8B, with model weights, code, and datasets made available.

Why it matters: HKR-H/K/R all pass: the concrete hook is NVIDIA opening Cosmos 3 with 32B/8B variants and released artifacts. The post is sparse and single-source, with no benchmarks or license details, so it stays in the 78–84 band.

QbitAI · WeChat

How Cloud Models Reach the Physical World: CMG Lion Rock AI Lab Uses LiOS for Embodied AI

CMG Lion Rock AI Lab released the LiOS edge-cloud architecture for embodied robotics, reporting about 30 ms one-way latency from local camera to cloud GPU memory in cross-machine tests, and open-sourced the low-latency video transmission module plus the LeFold laundry-folding dataset.

Why it matters: HKR-H/K/R pass: LiOS offers a concrete latency claim and open artifacts for embodied AI. Impact stays mid-tier because the lab is not a top platform vendor and no cross-source cluster is shown.

r/LocalLLaMA

Deepseek V4 Flash performance on DGX Spark

A Reddit user ran DeepSeek-V4-Flash with vLLM on two ASUS GX10 DGX Spark nodes and reported 1,680 prefill tokens/s plus 39.8 decode tokens/s at a 256K context with MTP=2; the setup uses TP=2 over RoCE, fp8 KV cache, and fits about 1M tokens safely in KV cache.

Why it matters: This is not broad industry news, but it is a first-person benchmark with reproducible details: TP=2, RoCE, fp8 KV cache, 256K context, and ~1M KV. HKR-H/K/R all pass, so it lands at low featured.

AI HOT (Curated Pool)

Cosmos 3 Released: First Open Physical AI Generalist Model

NVIDIA released Cosmos 3 as an open physical AI generalist model with native visual reasoning, world generation, and action generation, offering two variants: Super at 32B parameters and Nano at 8B parameters.

Why it matters: HKR-H/K/R all pass: NVIDIA names two Cosmos 3 variants and concrete physical-AI capabilities. Source is a single launch post with no benchmark or license detail, so it stays in the 78–84 band.

AI HOT (Curated Pool)

MiniMax M3: Frontier coding, 1M-token context, and native multimodal model

MiniMax released M3 as an open-source unified model with coding, agent, and native multimodal capabilities, supporting a 1M-token context window and using MiniMax Sparse Attention to cut per-token compute at 1M context to 1/20 of its predecessor, with over 9x faster prefill and over 15x faster decoding.

Why it matters: HKR-H/K/R all pass: MiniMax M3 has a 1M-token context hook, MSA with a claimed 20x cost cut, and open-source China-model resonance. Single official-source release keeps it in the 78–84 band, not P1.

AI HOT (Curated Pool)

MWC26 Shanghai to Host First Humanoid Robot Penalty Shootout With Unitree and 7 Other Teams

MWC26 Shanghai will host a humanoid robot penalty shootout in June 2026, with eight Chinese embodied intelligence teams competing under rules that require autonomous play without human control or preset scripts.

Why it matters: HKR-H/K/R all pass: the robot penalty shootout is clickable, with rules banning teleoperation and scripts. It stays in 72–77 because this is an event preview, not a model release or reproducible result.

r/LocalLLaMA

I ported NVIDIA Parakeet speech-to-text to ggml: same output as NeMo, faster, GGUF-quantized, no Python

mudler_it ported NVIDIA Parakeet speech-to-text models to C++/ggml with no Python or PyTorch, reporting byte-for-byte NeMo parity on f32/f16, up to about 5x GPU speedups on larger TDT and hybrid models, and GGUF quantization across f16, q8_0, q6_k, q5_k, and q4_k.

Why it matters: HKR-H/K/R all pass: the port has a concrete local-inference hook, byte-parity and speed claims, and clear practitioner resonance. Source scope keeps it at the low featured band, not P1.

May 31Sunday

r/LocalLLaMA

13 abliterated Gemma 4 E2B variants, 44 GPU hours, benchmark and comparison

Abliterlitics tested 13 abliterated Gemma 4 E2B variants using 44 RTX 5090 GPU hours, and HarmBench ASR rose from the base model’s 32.2% to 82%–100%, while coder3101 scored 84.8% on GSM8K versus the base model’s 83.5%.

Why it matters: HKR-H/K/R all pass, with a named first-person benchmark and concrete numbers. Scope stays narrow around abliterated Gemma 4 E2B variants, so it lands at the featured threshold rather than a must-write item.

r/LocalLLaMA

PolyRange: Contamination-resistant offensive-AI benchmark for web targets

PolyRange v1.0 ships 84 WSTG-derived classes across 12 OWASP testing-guide categories. It generates fresh targets per deploy with a chosen LLM, adds two defense tiers, uses an agent-submits-flag oracle, and runs via a single-command CLI on Fly.io or Docker.

Why it matters: HKR-H/K/R all pass: PolyRange turns web-security targets into a dynamic agent benchmark with 84 WSTG classes and two defense levels. Single-source Reddit origin and security niche keep it at 78.

Xinzhiyuan · WeChat

Fudan-Linked Team Releases STI-WM Spatiotemporally Integrated World Model

MouShen Intelligence released STI-WM, a spatiotemporally integrated world-action model for robotics, claiming support for RGB, point-cloud, and proprioceptive inputs, hundred-second task planning, and disclosing five funding rounds in six months plus a RMB 300 million Pre-A round.

Why it matters: HKR-H/K/R pass: STI-WM combines RGB, point clouds, and proprioception for 100-second planning, plus 5 funding rounds and a RMB300m Pre-A. Company-claim framing lacks public benchmarks or reproducible access, so it stays near the featured threshold.

QbitAI · WeChat

Robot-Native World Action Model Debuts With Spatiotemporal Architecture From Fudan-Linked Team

Moushen Intelligence released STI-WM, a spatiotemporally integrated world action model for robotics, with RGB, depth point cloud, and proprioceptive inputs; the post says it supports hundred-second-scale long-horizon task rollout and closed-loop replanning, but does not disclose benchmark scores or deployment costs.

Why it matters: HKR-H/K/R all pass: the STI-WM angle is novel, with concrete input modalities and hundred-second rollouts. Kept near the featured floor because public weights, benchmark results, and reproducible tests are not disclosed.

May 30Saturday

QbitAI · WeChat

RUC and Zhizhi Institute Open-Source Claw Agent Data, Training, and Evaluation Pipeline

Renmin University of China and Zhizhi Institute open-sourced ClawGym, a Claw Agent framework with 13.5K synthetic executable tasks, 200 benchmark tasks, model checkpoints, training data, and training code; ClawGym-30B-A3B scores 56.82 on ClawGym-Bench and exceeds Qwen3-235B-A23B in the reported evaluation.

Why it matters: HKR-H/K/R all pass: ClawGym bundles data, code, checkpoints, and eval tasks rather than just a leaderboard. Its impact is developer-facing, below a major lab model release or market-moving event.

r/LocalLLaMA

Testing MTP on vLLM and llama.cpp for Gemma 4 and Qwen 3.6

The author tested MTP on an RTX PRO 6000 Blackwell setup, where Gemma 4 31B on vLLM reached 132.52 tok/s versus a 39.69 tok/s baseline, a 3.34x speedup; the post reports 10 runs of 1,500 tokens each but does not provide a full quality or VRAM evaluation.

Why it matters: HKR-H/K/R all pass via a first-person speed test with hardware, model, and tok/s numbers. Source authority is limited, and missing quality/VRAM evaluation keeps it at the low featured band.

AI HOT (Curated Pool)

OpenAI launches real-time translation model with 70+ input languages

OpenAI launched gpt-realtime-translate, a speech translation model that accepts 70+ input languages and outputs speech in 13 target languages; the post says the feature is running on smart glasses.

Why it matters: HKR-H/K/R all pass: OpenAI has a concrete realtime-translation model with numbers and a wearable demo. Missing latency, pricing, and API availability keep it below P1.

May 29Friday

AI HOT (Curated Pool)

Xiaomi Open-Sources Controllable Video Foley Model ControlFoley

Xiaomi’s large model application team open-sourced ControlFoley, a controllable video Foley model supporting three tasks: text-guided video dubbing, text-controlled video dubbing, and reference-audio-controlled video dubbing, with code, model weights, and an online demo released.

Why it matters: ControlFoley clears HKR-H/K/R with controllable video Foley plus code, weights, and demo. It is a useful multimodal-audio release from Xiaomi, but not a flagship foundation-model launch, so it sits near the featured threshold.

Xinzhiyuan · WeChat

Three DeepSeek Models Enter OpenRouter Monthly Top 10 With Over 17 Trillion Tokens

DeepSeek placed three models in OpenRouter’s monthly top 10 with more than 17 trillion tokens combined, including V4 Flash at 9.13T tokens; the article says Ascend’s MegaMoE operator raised Prefill throughput by 20% to 30% on DeepSeek V3.1 and Qwen3-235B tests.

Why it matters: HKR-H/K/R all pass: the story has a 17T-token hook plus concrete OpenRouter and MegaMoE Prefill numbers. It stays at 82 because the compute-sovereignty framing is strong, while reproducible test conditions are not disclosed.

AI HOT (Curated Pool)

Nano Banana Pro and Nano Banana 2 officially released

Google AI Developers released Nano Banana Pro and Nano Banana 2, two image models available for production use through the Gemini API; the post names gemini-3-pro-image and gemini-3.1-flash-image but does not disclose pricing, benchmarks, or rate limits.

Why it matters: HKR-H/K/R all pass: Google shipped two production image models via Gemini API. The post gives no benchmarks, pricing, or safety mechanism, so this stays in the 78–84 band rather than p1.

The Verge · AI

Claude’s New Model Is More ‘Honest’ When It Messes Up

Anthropic will release Claude Opus 4.8 on Thursday, emphasizing its claimed “honesty.” The company says early testers found it flags uncertainty more often. It also says internal evaluations show Opus 4.8 is around 4x less likely than its predecessor to make unsupported claims, while the RSS snippet does not disclose the full benchmark setup.

Why it matters: HKR-H/K/R all pass: an Anthropic Claude model update with a concrete “4x fewer unsupported claims” eval claim. Details are thin: benchmark set, pricing, and context window are not disclosed, so it sits in the low 85–94 band.

May 28Thursday

r/LocalLLaMA

Qwen3.6-35B-A3B-APEX Runs 128K Context on RTX 3060 12GB

old-mike ran mudler/Qwen3.6-35B-A3B-APEX-MTP-I-Compact.gguf through spiritbuun’s llama.cpp fork on one RTX 3060 12GB, offloading a 17.3GB model and reaching 37.17 t/s generation at 72K filled context, 28.08 t/s at 129K, and PPL 3.2529 on an enwik8 64K-context perplexity test.

Why it matters: HKR-H/K/R all pass via a concrete consumer-GPU inference result with speed and PPL. Source is a single Reddit post and the impact stays within local inference, so it lands in featured, not P1.

AI HOT (Curated Pool)

Mistral AI launches physics AI model for industrial engineering

Mistral AI integrated the Emmi AI team and launched a physics AI foundation model for industrial engineering, with the post saying it can learn from geometry, boundary conditions, or measurement data and predict full physical fields on a single GPU in seconds.

Why it matters: HKR-H/K/R pass: a major model lab entering physics simulation with a concrete single-GPU seconds claim. The score stays in the lower featured band because model name, benchmarks, pricing, and access are not disclosed.

Synced · WeChat

Chinese pretrained embodied model Wall-OSS-0.5 is open sourced

X Square Robot open sourced Wall-OSS-0.5, a VLA model whose 400k pretraining checkpoint scored above 80 on 4 of 17 real-robot zero-shot tasks, with weights, code, training recipe, ablations, and a DMuon optimizer implementation released.

Why it matters: Clear HKR-H/K/R: a 400k checkpoint and 17 real-robot zero-shot tasks add substance, while “post-training not required” is a sharp hook. X Square Robot is not a top foundation-model lab, so this stays at 79.

Computing Life · Share · Yage

After SWE-Bench Pro Saturation, Someone Built a New Benchmark

DeepSWE says SWE-Bench Pro lost discrimination because of data contamination and verifier flaws; the same model set showed a 62-point spread on the benchmark, while the snippet does not disclose the audited models or test protocol.

Why it matters: HKR-H/K/R all pass: the “new ruler” hook, contamination/verifier claims, and 62-point spread give this real signal for code-agent evaluation. Source reach and impact are below the 85+ same-day tier.

r/LocalLLaMA

Inferencing at 10.33 t/s on Qwen 3.5 35B on a $300 laptop

A Reddit user ran Qwen 3.5 35B Q4_K_S on a $300 Lenovo Ideapad Slim 3i and reported 10.33 t/s inference using ik_llama.cpp with two pinned CPU cores, MTP speculative decoding, 64 batch size, and Q8_0 KV cache.

Why it matters: HKR-H/K/R all pass, with a concrete first-person benchmark. Reddit single-post sourcing and limited reproducibility details keep it at the lower featured threshold.

Hugging Face Blog

ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks

Artificial Analysis and IBM published the ITBench-AA title, saying frontier models scored below 50% on an enterprise IT agent task benchmark; the post does not disclose tested models, sample size, or scoring method.

Why it matters: HKR-H/R pass: frontier models under 50% on enterprise IT agent tasks is clickable and deployment-relevant. HKR-K is weak because models, sample size, and scoring are not disclosed, so it stays near the featured floor.

May 27Wednesday

AI HOT (Curated Pool)

AI Builds AI: ModelBest Open-Sources ForgeTrain, a Training Framework Written by AI

ModelBest, Tsinghua University, and OpenBMB open-sourced ForgeTrain, described as the first production-grade LLM training framework written entirely by AI with zero human code, and ModelBest used it to pretrain MiniCPM5-1B on Huawei Ascend chips.

Why it matters: HKR-H/K/R all pass: an open-source training framework, AI-written code, and MiniCPM5-1B pretraining on Ascend give concrete hooks. This is a strong tooling story, not a top-model launch, so 80 fits featured rather than P1.

May 26Tuesday

r/LocalLLaMA

[OSS] dlmserve: First Serving Engine for Diffusion Language Models

dlmserve released an MIT-licensed serving engine for diffusion language models, with LLaDA-8B-Instruct support and 2.5x HF throughput at batch=4. It exposes an OpenAI-compatible /v1/chat/completions API, batches at the denoising-step level, runs in 12GB VRAM, and adds about 1.8x throughput with optional LocalLeap acceleration.

Why it matters: HKR-H/K/R all pass: an open-source DLM serving engine with concrete throughput and VRAM claims. Single Reddit source and an early ecosystem keep it in low featured, not 78+.

AI HOT (Curated Pool)

Qwen3.7-Max Becomes the World’s No. 2 AI Coding Model

Qwen3.7-Max scored 1541 on Code Arena and ranked behind Claude; the post says it can run 35-hour tasks and perform more than 1,000 tool calls.

Why it matters: HKR-H/K/R all pass, but the source is a single Alibaba Cloud post and the evidence is benchmark plus vendor claims. This fits a strong product/benchmark update, not P1 without independent validation.

AI HOT (Curated Pool)

ModelBest open-sources MiniCPM5-1B, topping sub-2B models on AA-Index

ModelBest open-sourced MiniCPM5-1B, a 1B-parameter edge language model that beats all sub-2B models on AA-Index, uses a 0.5GB weight file after INT4 quantization, and runs on phones and browsers.

Why it matters: HKR-H/K/R all pass: MiniCPM5-1B has concrete params, quantized size, and edge runtime claims. It is still a small-model release, below flagship-model impact.

May 25Monday

r/LocalLLaMA

NuExtract3 released: open-weight 4B VLM for Markdown, OCR and structured extraction

Numind released NuExtract3, a 4B open-weight VLM based on Qwen3.5-4B under Apache-2.0, supporting image and text to Markdown, OCR, and JSON-template extraction, with self-hosting from 4GB VRAM and weights in Safetensors, GGUF, and MLX formats.

Why it matters: HKR-H/K/R all pass: NuExtract3 packages OCR, Markdown, and structured extraction into a 4B open-weight VLM with a 4GB self-hosting condition. Source and lab reach keep it in the low featured band.

May 24Sunday

r/LocalLLaMA

Vision-capable LLMs vs. OCR for long-document QA with charts, images, and tables

The author tested Claude Sonnet 4.5 on 171 questions from 30 image-heavy MMLongBench-Doc PDFs, comparing native PDF vision use with OCR pipelines. Native PDF ranked fifth of six at 52.0% accuracy and cost $0.2552 per query, while LlamaCloud premium with full context reached 59.6% at $0.1885 per query.

Why it matters: HKR-H/K/R pass: the post gives 30 PDFs, 171 questions, accuracy, and per-question cost for long-document QA. Limited sample and Reddit sourcing keep it in the featured-threshold band.

r/LocalLLaMA

It's OK to Quantize the KV Cache; Model Quant Matters More in Qwen3.6 27B KLD Tests

Reddit user hopbel tested Qwen3.6 27B with approximate KLD on wikitext-2 at 16k context, using Q5_K_M as the proxy baseline; Q5_K_S weights with q4_0 KV cache scored 0.016304, while Q4_K_XL with f16 KV cache scored 0.026067, so weight quant tier dominated KV-cache quant in this setup.

Why it matters: HKR-H/K/R all pass, backed by first-person test numbers. Source is a single Reddit post, the metric is approximated KLD, and the claim is narrow, so it sits at the featured threshold.