Skip to content

Deployment & engineering

Running models in practice: inference optimization, memory and cost, serving architecture and infrastructure choices.

Latest picks

161–180 of 455

May 23Saturday

QbitAI · WeChat

DeepSeek V4 cuts prices as CATL, JD.com and NetEase discuss investment; Liang Wenfeng targets AGI

DeepSeek-V4-Pro API will keep its promotional pricing from June 1, with cached input at RMB 0.025 per million tokens, while Bloomberg says DeepSeek is pursuing a RMB 70 billion round at a USD 45 billion pre-money valuation.

Why it matters: HKR-H/K/R all pass: DeepSeek V4 API price cuts and Bloomberg’s RMB 70B raise at a $45B pre-money valuation are same-day material. The cost and capital angles directly affect China model competition.

r/LocalLLaMA

club-rdna16: Practical 16GB AMD/Radeon local LLM testing repo

club-rdna16 publishes a practical 16GB Radeon local LLM testing repo, with an RX 6900 XT running llama.cpp on ROCm/HIP and Qwen3.6 35B-A3B reaching a stable 131k context using q8 KV cache.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post and the body only discloses test conditions, not speed, VRAM curves, or reproducible logs. It clears the featured floor as a practical local-LLM repo.

AI HOT (Curated Pool)

Jensen Huang Says Annual AI Infrastructure Spending Will Reach $4 Trillion

Jensen Huang predicted hyperscale cloud providers’ annual AI infrastructure spending will rise from $1 trillion to $3 trillion–$4 trillion, while Nvidia reported $81.6 billion in fiscal 2027 Q1 revenue and $75.2 billion from data centers.

Why it matters: HKR-H/K/R all pass: Jensen Huang’s $3-4T annual AI infrastructure forecast is specific and tied to NVIDIA revenue. It is strong industry signal, but a CEO forecast rather than a model or product launch, so it stays in the 78-84 band.

AI HOT (Curated Pool)

Agent Workloads Quietly Reshape Inference Economics

SemiAnalysis analyzed 432,000 real coding-agent requests and found a median input length of 96,000 tokens, not 32,000 or 64,000. The post does not disclose the model mix, cost curve, sampling method, or time window.

Why it matters: HKR-H/K/R all pass: SemiAnalysis adds a 432k coding-agent request dataset and 96k-token median input. Missing models, cost curves, and sampling keep it in the strong-data-point band, not must-write.

r/LocalLLaMA

Experts first llama.cpp

comanderxv published a llama.cpp fork that caches MoE experts in 12GB VRAM; on an RTX 2060 with Qwen3.6-35B-A3B, throughput rose from 19/22 tk/s to 26 tk/s at about a 62% expert-cache hit rate.

Why it matters: HKR-H/K/R all pass: the hook is a 35B MoE speedup on a 12GB RTX 2060, with concrete caching and hit-rate data. Scope stays niche to local inference, so it lands at the featured threshold rather than must-write.

May 22Friday

Hacker News front page

DeepSeek Makes the V4 Pro Price Discount Permanent

DeepSeek will set deepseek-v4-pro API pricing to one quarter of the original price after the 75% promotion ends on 2026-05-31 at 15:59 UTC; the post does not disclose the exact per-token price.

Why it matters: HKR-H/K/R all pass: the hook is a permanent DeepSeek price cut, the new fact is 1/4 pricing after a stated UTC time, and the nerve is API cost. Missing unit pricing keeps it at the featured floor, not a must-write release.

Dwarkesh Patel podcast

Reiner Pope – Chip Design from the Bottom Up

Dwarkesh Patel interviews MatX CEO Reiner Pope on chip design, starting with a 4-bit multiply and 8-bit accumulate example that uses 16 AND gates, then covering systolic arrays, pipeline registers, FPGAs versus ASICs, cache versus scratchpad, and why GPU cores are smaller than CPU cores.

Why it matters: Dwarkesh’s MatX CEO interview clears HKR-H/K/R with a bottom-up hardware hook, concrete mechanisms, and compute-cost resonance. It is educational rather than breaking news, so it sits in the 72–77 band.

Mistral AI

Mistral launches Connectors in Studio with built-in and custom MCP

Mistral launched Connectors in Studio. All built-in connectors and custom MCP are now callable through the API/SDK by every model and agent. New features include direct tool calling, human-in-the-loop approval flows, and programmatic access to create, modify, list and delete connectors.

Why it matters: The original gives the API usage and code examples for Connectors, enough to judge how enterprise MCP integration gets built.

AI HOT (Curated Pool)

BitCPM-CANN Released as First 1.58-bit Open Model Fully Trained on Huawei Ascend 910B NPU

ModelBest, Tsinghua University, and OpenBMB released BitCPM-CANN, a 0.5B-8B open model family trained natively on Huawei Ascend 910B NPUs with 1.58-bit ternary weights, cutting memory use by about 6x versus BF16 while retaining 95-97% of full-precision benchmark performance.

Why it matters: HKR-H/K/R all pass: the Ascend 910B plus 1.58-bit open model angle is novel and metric-rich. It stays below P1 because the post offers release facts, not independent replication or adoption signal.

Latent Space

[AINews] New AI Infra Unicorns: Exa, Modal, TurboPuffer

Latent Space summarized AI News for May 20-21, 2026, confirming TurboPuffer reached $100 million ARR and profitability, Exa raised a $250 million Series C at a $2.2 billion valuation, and Modal raised a $355 million Series C at a $4.7 billion valuation.

Why it matters: HKR-H/K/R all pass because the roundup gives concrete AI-infra funding and ARR numbers. It stays below 78 because it is market aggregation, not a new model, product capability, or technical release.

AI HOT (Curated Pool)

Zhipu releases GLM-5.1-highspeed, claiming a large-model API speed record

Zhipu released the GLM-5.1-highspeed API to selected enterprise customers on May 22, with a claimed output speed of 400 tokens/s, built by the GLM team and TileRT team through system-level optimization.

Why it matters: HKR-H/K/R all pass: Zhipu’s GLM-5.1 high-speed API has a concrete 400 tokens/s claim and domestic flagship-model relevance. Test setup, pricing, and availability are not disclosed, so it stays in the 78–84 band.

New York Times Chinese

Trump Approved Nvidia Chip Sales to China. Why Is Beijing Reluctant?

Trump approved Nvidia H200 sales to China six months ago, but Beijing has not allowed any company to buy even one chip and is steering firms toward domestic alternatives from Huawei and Cambricon.

Why it matters: HKR-H/K/R all pass: six months of zero H200 purchases after approval, plus Beijing steering firms toward Huawei and Cambricon. This is strong chip-policy signal, but not a model launch or major product release, so it sits in 78–84.

Computing Life · Share · Yage

How to Run DeepSeek V4 Flash Locally on Mac: DS4 Engine Explained

DS4 provides a macOS local runtime path for DeepSeek V4 Flash; the post only discloses three mechanisms—multi-agent integration, KV cache disk persistence, and activation steering—and does not disclose performance numbers, hardware requirements, or pricing.

Why it matters: HKR-H/K/R all pass, but the body only names DS4 mechanisms and omits performance, model size, Mac support, and reproducible tests; this fits the featured threshold for a local-inference tutorial.

Computing Life · Share · Yage

The Technology Behind GLM-5.1 Reaching 400 Tokens/s

Zhipu GLM-5.1 high-speed API claims 400 tokens/s, and the post says TileRT reconstructs GPU inference at the execution-model level; the RSS snippet does not disclose benchmark conditions, hardware, pricing, or latency distribution.

Why it matters: HKR-H/K/R all pass: 400 tokens/s is a concrete hook, TileRT adds mechanism, and latency/cost resonates with builders. It stays at 78 because the speed is claimed, with no independent test or pricing condition disclosed.

r/LocalLLaMA

Interesting Paper Advocates Quantized Prefilling and Precise Decoding

arXiv 2605.20315 argues for W4A4 quantization during prefilling to target a theoretical 4x gain, while keeping decoding on the original high-precision path because activation errors can perturb sampled tokens and accumulate across autoregressive generation.

Why it matters: HKR-H/K/R all pass, but the item only gives the paper claim and theoretical gain; measured throughput, perplexity, and hardware setup are not disclosed, so it stays at the featured threshold.

NVIDIA Blog

NVIDIA GTC Taipei at COMPUTEX: Live Updates on What’s Next in AI

NVIDIA won four COMPUTEX 2026 Best Choice Awards for Vera Rubin NVL72, Jetson Thor, and Alpamayo; Vera Rubin NVL72 connects 36 Vera CPUs and 72 Rubin GPUs, and NVIDIA says it delivers up to 10x higher inference performance per watt and 10x lower cost per token.

Why it matters: HKR-H/K/R all pass: NVIDIA gives concrete Vera Rubin NVL72 specs and a 10x inference-efficiency claim, directly tied to AI compute costs. The source is NVIDIA’s event blog, so this stays below the 85 same-day must-write band.

May 21Thursday

r/LocalLLaMA

Agent Execution Tax: New Procurement Metric for Browser Agent Benchmarks?

Fireworks ran 720 browser-agent tasks on WebVoyager and reported a 22.9% Agent Execution Tax, defined as wasted over productive inference; MiniMax M2.5 cost 2.3x less per successful task than Gemini, while GLM-5 reached 57.1% accuracy and Kimi K2.5 had 0% parse retries across 852 calls.

Why it matters: HKR-H/K/R all pass: the post adds a named procurement metric plus concrete benchmark numbers. Source scope is Reddit/Fireworks, so it stays in the 72–77 featured band rather than 78+.

r/LocalLLaMA

LLM planner: pick a rig by use case, model, or budget, or pick models for your rig

totosse17 published the LLMRequirements hardware planner with 60+ build configs, 50+ models, 130 cited tokens-per-second sources, 150+ reviewer videos, multi-region prices, idle and active watts, and a public GitHub data repo.

Why it matters: HKR-H/K/R all pass, but this is a Reddit community tool for local LLM rigs, not a broad platform release. The concrete dataset earns a featured-threshold score, not the 78+ band.

The Verge · AI

Anthropic is paying $15 billion a year for access to Elon Musk’s data centers

SpaceX said in its S-1 filing that Anthropic agreed to pay $1.25 billion per month through May 2029 for access to Colossus I and II AI training centers in Memphis, totaling $15 billion annually.

Why it matters: HKR-H/K/R all pass: the Anthropic–Musk pairing is a strong hook, and the S-1 gives $1.25B/month through May 2029. Compute cost and data-center dependence make this a same-day must-write story.

r/LocalLLaMA

Tencent Hy-MT2 30B/7B/1.8B

Tencent released Hy-MT2 translation models in 1.8B, 7B, and 30B-A3B sizes, supporting translation across 33 languages; AngelSlim 1.25-bit quantization reduces the 1.8B model’s storage requirement to 440 MB and raises inference speed by 1.5x.

Why it matters: HKR-H/K/R pass via the 440MB quantized 1.8B model, 33-language support, and local inference cost angle. Sparse Reddit sourcing keeps it at the featured threshold, not the 78+ band.