Skip to content

Deployment & engineering

Running models in practice: inference optimization, memory and cost, serving architecture and infrastructure choices.

Latest picks

21–40 of 455

Jul 9Thursday

Hacker News front page

Colibrì: A pure-C engine that runs GLM 5.2 on a 32 GB laptop with no GPU

Colibrì is a single-file C engine that runs the 744B-parameter GLM 5.2 model on a 32 GB laptop with no GPU. The dense part stays in RAM at int4 (~9.9 GB), while 21,504 routed experts stream from disk on demand with an LRU cache. Cold-start speed is 0.1 tok/s. The author built and tested it on a 12-core, 25 GB machine. The post does not include quality benchmarks against the full-precision model.

Why it matters: A personal experiment with concrete numbers and a reproducible approach — not marketing fluff. Splits GLM 5.2's 744B parameters into a dense part resident in RAM and routed experts loaded from disk on demand, with int4 quantization bringing memory down to 9.9GB. Not scoring hi...

Jul 7Tuesday

AI HOT (Curated Pool)

SGLang Integrates DSpark: Confidence-Driven, Variable-Length Speculative Decoding

LMSYS integrated DSpark speculative decoding into SGLang. The draft model's confidence head decides how many tokens to verify per request, instead of always verifying a full block. On H200 with DeepSeek-V4-Flash, it beats both MTP and non-spec baselines in throughput and per-user decode speed. The key engineering move is a ragged, variable-length verify under full CUDA graphs: when the scheduler trims the verify budget, the engine replays a genuinely smaller graph, not a padded one. A cost-table profiler lets the scheduler size each request's budget online, and a cap-accept mode exposes the acceptance ceiling hidden by trimming.

Why it matters: LMSYS integrated DSpark into SGLang with real H200 benchmarks—not a paper announcement but an engine-level landing. H and K are solid, but the topic is inference optimization, which doesn't resonate broadly, so R is absent and the score stays at the featured threshold of 78.

Jul 5Sunday

Computing Life · Share · Yage

Scaling Law's three corrections in five years: from bigger models to smaller models with more data

Scaling law is an empirically fitted curve, not a physical law. OpenAI's 2020 Kaplan paper concluded 'prioritize parameters' due to experimental biases, shaping GPT-3. DeepMind's 2022 Chinchilla corrected the ratio to 20:1, showing smaller models with more data outperform. Two 2024 replication studies confirmed that fixing Kaplan's setup reproduces Chinchilla's result—no fraud, just calibration. Since 2023, Meta and others deliberately deviate from Chinchilla: Llama 3 8B was trained on 15T tokens because the optimization target shifted from training cost to total cost of training plus inference. Tsinghua's Densing Law shows the parameter count needed for equal capability halves roughly every 3.5 months, but there is a floor: each parameter stores only ~2 bits of knowledge. The viral 'collapse' article cited a blog comment posted the same day as if it were peer-reviewed research; the post does not provide a paper source for that claim.

Why it matters: A high-quality explainer and fact-check on scaling laws, debunking a recent viral post with specific numbers and paper citations while tracing three key revisions over five years. Hits all three HKR axes, but as commentary/education rather than a first-party product release, i...

Jun 30Tuesday

Hacker News front page

Moondream's Photon engine uses pipelined decoding to pop GPU bubbles, boosting decode throughput up to 35%

Moondream's Photon inference engine hits ~33ms VLM inference on NVIDIA B200. The bottleneck is GPU bubbles: the GPU idles while the CPU finishes bookkeeping between tokens. Photon pipelines the decode loop so the GPU starts the next forward before the CPU commits the current token, overlapping CPU housekeeping under GPU work. Three mechanisms make it safe: ping-pong slots prevent buffer collisions, forward-now-sample-later handles constrained decoding, and zombie cleanup deals with finished requests. The result is up to 35% higher decode throughput.

Why it matters: Moondream's Photon engine hits ~33ms VLM inference with 35% higher decode throughput — solid technical detail. But it's a single-company optimization, not an industry event, and the audience is narrow, so it lands right at the featured threshold.

Hacker News front page

vLLM Semantic Router beats frontier models by making multiple models collaborate inside one API call

vLLM's Micro-Agent makes multiple models collaborate behind a single API endpoint. Users call one model name as usual; the router decides whether to try a cheaper model first, run several in parallel and aggregate, or let models review each other. The core idea is turning one API call into a bounded micro-collaboration with budget, topology, and failure policy. The post describes five looper patterns—Confidence, Ratings, ReMoM, Fusion, and Workflows—and claims they beat individual frontier models on benchmarks, though it does not name the benchmarks or provide scores.

Why it matters: vLLM upgrades semantic routing into a model-collaboration engine with five concrete patterns and budget controls — highly relevant for inference engineers. Held back from 80+ because it's a blog post, not a shipped product, and performance numbers aren't fully disclosed in the...

Jun 19Friday

Hacker News front page

I restarted a 10-year-old Xeon 174 times to delete 12 flags and gain 4 tps

The author ablated all 25 flags from a previous Gemma 4 setup on a 2016 Xeon. Most did nothing. The real levers are flash attention, physical-core thread count, and workload-gated speculative decoding: on for code, off for long-doc summarization (54% faster without it). The auto-tune drafter setting is a crude per-request router that currently hurts long-context throughput.

Why it matters: A solid local-inference tuning experiment that ablates 25 launch flags one by one, delivering scene-dependent numbers on speculative decoding and the finding that most flags are placebo. Hits H and K, but narrow audience and zero industry conversation value mean R misses — lan...

Jun 14Sunday

r/LocalLLaMA

Xiaomi serves MiMo V2.5 at 1000–3000 tps with DFlash and Persistent Kernel

Xiaomi's MiMo V2.5 is live, claiming 1000–3000 tps via DFlash and Persistent Kernel. The DFlash model weights are out, and an open-source release is promised soon. The post body is blocked by Reddit security, so only the headline is available—no details on measured latency, concurrency, or hardware. I'd discount that tps figure: headline peaks usually assume optimal batching, and real single-user throughput is likely lower.

Why it matters: MiMo V2.5's claimed 1000-3000 tps and the two named acceleration mechanisms (DFlash, Persistent Kernel) carry real information density; weights are out and open-source code is promised, directly relevant to local model deployers. Score capped because the Reddit body was blocke...

Jun 13Saturday

Hacker News front page

Can I Buy Your KV Cache?

This paper proposes letting publishers precompute a document's KV cache so AI agents can buy and load it, skipping the most compute-heavy step: prefill. On Qwen3-4B, reuse is 9–50x cheaper than prefill with zero accuracy loss—token outputs match exactly. Shipping the KV cache fails because it's nearly incompressible and egress costs more than the prefill saved. The fix: host it provider-side, like production prompt caching. Serving one 3,774-token document to 80M agents costs ~$1.5M to re-prefill but only ~$30K via reuse, a 49.7x gap. The paper frames this as an agent-native prefill CDN and leaves lossless KV compression and cross-party payments as open problems.

Why it matters: Selling precomputed KV caches is a practical idea with a 9–50× cost gap and zero accuracy loss. Held back by single-model experiments (Qwen3-4B only) and no detail on cache security or pricing in the excerpt.

Jun 12Friday

r/LocalLLaMA

MiniMax open-sources MSA, a sparse attention method that cuts attention compute by 28.4× at 1M tokens on a 109B model

MiniMax published a paper introducing MSA, a blockwise sparse attention built on GQA. A lightweight index branch scores KV blocks and picks a top-k subset per GQA group, then the main branch runs exact attention only on those blocks. With a co-designed GPU kernel, a 109B-parameter multimodal model achieves 14.2× prefill and 7.6× decoding wall-clock speedups on H800 at 1M context, matching full GQA quality. Code and inference kernel are open-sourced, along with a model called MiniMax-M3. The Reddit poster is curious whether the 109B model can run on consumer GPUs; the post doesn't say if weights will be released.

Why it matters: The paper has concrete mechanisms and measured numbers, not just theory—real knowledge for inference-optimization folks. But the audience is narrow (R missed), and the low-level CUDA details raise the accessibility bar for generalist readers, so I docked 3 points, landing righ...

Jun 10Wednesday

AI HOT (Curated Pool)

OpenRouter Launches Advisor Tool for Low-Cost Models to Consult Stronger Models

OpenRouter released the Advisor server tool, letting GPT-4o Mini consult Claude Fable during generation, but the post does not disclose pricing, latency, or the routing policy.

Why it matters: HKR-H/K/R all pass: OpenRouter turns cheap-model plus strong-model advising into a callable server tool. Price, latency, and call policy are not disclosed, so this stays in the upper mid-weight product-update band.

r/LocalLLaMA

ICML paper on predictable hallucination gate and ntkMirror open-weight implementation

An ICML 2026 paper presents an ISR=1 answer-abstain gate for evidence-grounded QA, and ntkMirror implements it for local open-weight models with multiple evidence orderings, reporting 0.0–0.7% hallucination at about 24% abstention in the held-out audit.

Why it matters: HKR-H/K/R all pass: an ICML paper with an open implementation, a concrete ISR=1 gate, and measured abstention-vs-hallucination tradeoff. Scope stays within evidence QA/RAG reliability, so it sits below must-write level.

Jun 9Tuesday

AI HOT (Curated Pool)

Google Releases Gemini 3.5 Live Translate for Real-Time Speech Translation

Google released Gemini 3.5 Live Translate, a speech-to-speech translation model that supports more than 70 languages, starts translating before the speaker finishes, uses streaming updates, and runs through Gemini Live API, Google Meet preview, and Google Translate apps on iOS and Android.

Why it matters: HKR-H/K/R all pass: Google ties real-time speech translation to 70+ languages and streaming output before the speaker finishes. It stays at 82 because rollout scope, pricing, and benchmarks are not disclosed.

AI HOT (Curated Pool)

Google DeepMind Releases Gemma 4 12B, a Unified Encoder-Free Multimodal Model

Google DeepMind released Gemma 4 12B, a multimodal model with a unified encoder-free architecture, native audio input, Apache 2.0 licensing, and local laptop runtime with 16GB of VRAM or unified memory.

Why it matters: HKR-H/K/R all pass: the hook is local multimodal audio in 16GB VRAM, and the new architecture is concrete. It is a strong Google DeepMind open-model release, but not a frontier-model launch, so it stays below p1.

r/LocalLLaMA

Apple Announced New On-Device Inference Engine for Apple Silicon

Apple announced CoreAI at WWDC as a future CoreML replacement for Apple Silicon on-device inference; models require Python-script conversion, the supported list is mostly mid-2025 models, and the post does not disclose performance data.

Why it matters: HKR-H/K/R pass, but the post is thin: CoreAI, CoreML successor status, and Python conversion are disclosed; throughput, latency, and model coverage are not. Apple on-device inference merits featured, capped in the 72–77 band.

AI HOT (Curated Pool)

China Prepares $295 Billion Plan to Fund Nationwide AI Infrastructure Buildout

China plans to invest about 2 trillion yuan, or $295 billion, over five years to build nationwide data centers, with funding covering large-scale data center infrastructure for domestic AI development.

Why it matters: Bloomberg reports China is preparing a five-year RMB 2T AI data-center plan, clearing HKR-H/K/R. This is national compute supply and geopolitical competition news, not routine policy; the preparation status keeps it at 90.

AI HOT (Curated Pool)

Xiaomi MiMo and TileRT Release UltraSpeed Mode, 1T Model Exceeds 1,000 Tokens/s

Xiaomi MiMo and TileRT released MiMo-V2.5-Pro-UltraSpeed, a 1T-parameter model mode exceeding 1,000 tokens/s, with API access open from June 9 to June 23, 2026, at 3× the MiMo-V2.5-Pro price and about 10× the speed.

Why it matters: HKR-H/K/R all pass, with a domestic flagship-model bump for Xiaomi. Missing hardware, batch, concurrency, and test conditions keep it in the 78-84 band rather than p1.

AI HOT (Curated Pool)

Elon Musk Details SpaceX's AI1 Orbital AI Data Center Satellite Plan

Elon Musk detailed SpaceX’s AI1 orbital AI data center satellite plan, with 150 kW peak power per satellite, about 120 kW sustained compute power, and 6-8 ms round-trip latency at 600-800 km low Earth orbit.

Why it matters: HKR-H/K/R all pass: the angle is unusual and the post gives power, orbit, and latency figures. It stays near the featured floor because launch timing, cost, and workload tests are not disclosed.

r/LocalLLaMA

2X tk/s on 1× MI50: Qwen3.6-27B inference rises from 19.4 to 38.1 tk/s

bigattichouse raised Qwen3.6-27B throughput on a single MI50 from 19.4 to 38.1 tk/s by running same-model parallel computations for Q8-or-lower quantization, exploiting unused compute lanes instead of adding a smaller speculative decoding model.

Why it matters: HKR-H/K/R all pass, but this is a Reddit first-person experiment with numbers and a hypothesis, not a validated release. No code or broader replication is disclosed, so it stays at the featured threshold.

r/LocalLLaMA

New MLX LM Server From Apple

A Reddit post says Apple’s MLX LM Server uses continuous batching for concurrent sub-agent requests and supports distributed inference across multiple Macs via Thunderbolt RDMA.

Why it matters: HKR-H/K/R all pass, but the item is based on a Reddit summary and lacks throughput, latency, model-size, or release details. Treat it as a mid-weight Apple/MLX inference update, just above the featured threshold.

r/LocalLLaMA

Levi: Run AlphaEvolve on Your Local Qwen 30B

LEVI runs an AlphaEvolve-like search system with Qwen3-30B-A3B and reports tests on ADRS, IFBench, and HotpotQA, claiming up to 35x lower cost overall and up to 12x fewer evals under the same single-model, same-budget comparison.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post with model, benchmarks, and cost ratios only; code maturity and reproducibility details are not disclosed. Scores as a strong open-source agent/inference item, not a major release.