Skip to content

Deployment & engineering

Running models in practice: inference optimization, memory and cost, serving architecture and infrastructure choices.

Latest picks

1–20 of 455

Yesterday · Sep 29Tuesday

Hacker News front page

World Labs to join AMD; Fei-Fei Li becomes Chief Scientist

World Labs signed a definitive agreement to join AMD. Fei-Fei Li will become AMD's EVP and Chief Scientist, reporting to CEO Lisa Su. Justin Johnson and Ben Mildenhall will keep leading the team as it forms a new frontier research org inside AMD. The two companies started a technical partnership last year on model training and inference optimization on AMD GPUs. The goal is an end-to-end open AI ecosystem spanning hardware, software, and open models. The deal is expected to close by end of 2026, pending regulatory approvals.

Why it matters: Fei-Fei Li's startup joining AMD with her as Chief Scientist is one of the year's biggest personnel + strategy mergers. All three HKR axes hit: the personnel pairing creates suspense, the technical partnership has concrete timeline and goals, and the academia-to-industry arc r...

Sep 24Thursday

Google DeepMind

Google DeepMind adds secure server-side memory to Private AI Compute

Google DeepMind detailed a new capability for Private AI Compute: private, server-side persistent memory that lets an AI assistant keep context across devices. Data sits sealed in encrypted storage, and the unlock key stays only on the user's device. When the model needs access, an end-to-end encrypted channel carries it into a secure cloud enclave, where it is briefly decrypted in isolated memory and immediately re-encrypted.

Why it matters: The post explains how cloud persistent memory uses secure enclaves and device-held keys for privacy, a look at the privacy architecture behind cloud AI memory.

Sep 23Wednesday

AI HOT (Curated Pool)

Modal details how to serve trillion-parameter coding agents at trillion-token scale

Modal's engineering team published a deep-dive on serving Moonshot AI's Kimi K2.6 for coding agents. They boosted per-replica performance by 2.8x per user and 5.6x across users, turning a ruinously expensive service into a price-competitive one. One service processed hundreds of billions of tokens per day and trillions in aggregate. The post walks through workload analysis for hybrid-attention MoE models and the engineering optimizations applied. Exact GPU models and per-request latency numbers are not disclosed in the body.

Why it matters: Modal's engineering breakdown of inference acceleration for Moonshot AI's Kimi K2.6 delivers two hard numbers — 2.8x and 5.6x speedups — directly useful for inference engineers. Not scored higher because this is infra optimization, not a model capability or product update; aud...

Sep 21Monday

AI HOT (Curated Pool)

AI Comes for the If Statement: Specialized Deciders Cut Classification Cost ~100x

Tomasz Tunguz tested Jev and SemIf on 98 production emails: classification accuracy jumped from 47% to over 80%, cost dropped to $0.0004 per call—76x to 209x cheaper than frontier models. These deciders skip text generation, run attention once, and output choice probabilities in hundreds of milliseconds. Tunguz sees this as a bifurcation: frontier models for discovery, specialized models for production, with if-then as the first optimized programming primitive.

Why it matters: Tunguz ran a real production test with concrete accuracy and cost numbers—not just trend talk. The piece flags a meaningful fork in AI infra: general-purpose generators vs. specialized deciders. Score stays at 78 because both tools are brand-new with no large-scale validation ...

Sep 19Saturday

Hacker News front page

C2C lets LLMs talk via KV-cache, 2.5× faster and 3–5% more accurate than text

This ICLR'26 paper proposes Cache-to-Cache (C2C), where multiple LLMs communicate by directly exchanging KV-cache instead of generating text. A neural network projects and fuses the source model's KV-cache into the target model, with a learnable gate selecting which layers benefit. C2C beats single models by 6.4–14.2% in average accuracy, outperforms text-based communication by ~3.1–5.4%, and delivers an average 2.5× latency speedup. Code is open at thu-nics/C2C.

Why it matters: ICLR'26 paper with a clever idea: let collaborating models pass KV-cache directly instead of text. Has validation experiments and a concrete framework, so knowledge density is solid. Score capped because it's low-level optimization — not immediately actionable for most practit...

Sep 17Thursday

Hacker News front page

Ternary LLMs break the 1.58-bit floor: BITCOS hits 1.485 bits per weight by exploiting zero-weight density

Ternary models store weights as -1, 0, or +1, with a theoretical floor of ~1.585 bits and a practical 1.625 bits in five-trit packing. Intel authors measured 29 ternary LLMs and found up to 51.5% zeros. BITCOS replaces fixed packing with a presence bitmap plus a compacted sign vector, costing 2 minus zero-density bits per weight. It beats five-trit packing on 26 of 29 models and reaches 1.485 bits on the sparsest. Optimized unpacking on AVX-512, AVX2, and Xe2 GPUs yields up to 1.28× faster matrix-vector multiply; end-to-end decode throughput improves up to 1.18× on CPUs and 1.27× on GPUs. The paper does not name the models or disclose their parameter counts, nor whether they are publicly available.

Why it matters: Intel team measured 29 ternary models, found zero weights up to 51.5%, and proposed BITCOS encoding to break the 1.58-bit floor. Solid K with concrete numbers and a new mechanism; H works on title intrigue. But it's a narrow inference-opt topic with no R pull, so it lands at t...

Sep 15Tuesday

r/LocalLLaMA

DeepSeek V4.1 Flash Q4 hits 40 t/s on M3 Ultra with native DSpark multi-token prediction

A developer forked ds4 and tuned it for DeepSeek V4.1 Flash Q4 on a 512 GB M3 Ultra, lifting decode from 16.6 t/s to 31.3 t/s, and to 40.5 t/s with DSpark speculative decoding. The ~300 GB Q4 weights fit only the 512 GB M3 Ultra. The speedup comes from cutting Metal dispatch overhead: the 384-expert router went from 9 dispatches to 1, the shared expert gate+up+SwiGLU became a single kernel, and BF16 rounding moved inside producer kernels, removing ~770 re-round dispatches. At 300k context, compressed attention selection was the bottleneck; the fix scores only admitted blocks and uses a bounded radix select, keeping decode at 90% of the 8k rate. DSpark verifies 6 tokens per step with shared weight streams and overlapped Engram fetches, dropping verify latency from 177 ms to 112 ms. Output is byte-identical to upstream under greedy decode, with SHA-256 manifests provided. The branch is M3 Ultra only because the optimizations rely on measured behavior of this specific chip's dual-die memory, 80-core GPU scheduling, and Metal dispatch characteristics.

Why it matters: A solid local-inference optimization post: DeepSeek V4.1 Flash Q4 on M3 Ultra goes from 16tps to 40tps via Metal command-buffer merging and native DSpark speculative decoding. Concrete technical detail, directly useful for the local-LLM crowd. Score stays at 72 because the aud...

Sep 2Wednesday

Hacker News front page

Slotstream runs the 104GB Qwen3.8-Flash-Next on a 48GB Mac at ~12 tok/s

carloslfu open-sourced Slotstream, an MLX + Swift tool that runs the 125B-parameter MoE model Qwen3.8-Flash-Next (104GB at 4-bit) on a Mac with only 48GB RAM. It streams expert modules from SSD on demand instead of loading everything into memory. Speed is ~12 tok/s, and it exposes an Ollama-compatible API. The post doesn't disclose time-to-first-token or the SSD model used, so real-world feel is still an open question.

Why it matters: Streams MoE experts from SSD on demand via MLX + Swift, letting a 104GB Qwen 125B model hit ~12 tok/s on a 48GB Mac. Clean engineering with an Ollama-compatible API that lowers the trial barrier. Docked a few points because it's a solo project with no community validation or m...

Aug 26Wednesday

Computing Life · Share · Yage

The term 'local LLM' conflates two separate markets

Yage breaks down 'local LLM' into two markets: a cost market buying 5–20× price gaps, and a control market buying 25–33-year certainty. Using a four-quadrant framework (open/closed weights × time/token billing), the piece explains why surging open-weight model usage on OpenRouter doesn't mean local deployment is winning. Self-hosting payback depends entirely on which cloud billing mode you replace—decades for subscriptions, months for high-cache-hit agent API calls. In July–August 2026, Anthropic and others made four moves at the inference layer: silently remapping parameters, repeatedly extending usage boosts, adding watermarks, and redefining self-hosting as 'your harness plus my inference.' But simultaneous deep price cuts mean the misalignment is real but direction is unresolved.

Why it matters: Splits 'local LLM' into cost vs control markets with OpenRouter data and hardware payback math — directly useful for infra decision-makers. Not scored higher because it's commentary rather than a product launch or research breakthrough, but hits all three HKR axes and earns a ...

Aug 21Friday

Hacker News front page

Nari Labs pushes Qwen3-TTS to sub-50 ms time-to-first-audio at 10 RPS on a single H100

Nari Labs open-sourced a Qwen3-TTS 1.7B CustomVoice serving implementation that hits sub-50 ms p95 time-to-first-audio at 10 RPS on a single H100 SXM with zero underruns. They benchmarked against vLLM-Omni, SGLang-Omni, VoxServe, and M*—default p95 latencies ranged from 277 to 1,160 ms at 1 RPS. At full utilization the system costs roughly $2 per 1M characters, compared to $100 for ElevenLabs V3 and $49 for Cartesia Sonic 3.5. Key optimizations include dynamic leading-silence trimming (~80 ms saved) and tuned codec-frame accumulation. Code and benchmarks are public; the post does not disclose underrun details at higher concurrency or long-form performance.

Why it matters: Nari Labs open-sourced a deployment recipe for Qwen3-TTS 1.7B that hits sub-50 ms p95 time-to-first-audio at 10 concurrent requests on a single H100—an order-of-magnitude improvement over vLLM-Omni and others. The post includes concrete benchmarks and reproducible optimization...

Aug 20Thursday

Hacker News front page

DFlash 2 pushes parallel drafting further: over 20% more output per verification pass for ~1% added latency

Inco AI released DFlash 2, adding a lightweight path selector on top of parallel speculative decoding. Instead of keeping only the top-1 candidate per position, it picks a coherent path from the top 16, raising accepted tokens per verification from 4.27 to 6.79. On Qwen3.8-27B with SGLang, throughput reaches 2.7–3.4× autoregressive decoding at batch size 1, with roughly 1% added cycle latency. SGLang, vLLM, llama.cpp, and oMLX already support it; DFlash models have been downloaded over 3.5 million times on Hugging Face.

Why it matters: DFlash 2 is a clear technical improvement on an already-adopted inference method, with measured results. Score isn't higher because this is a single technical blog post, not a model launch or product release — its reach is limited to the inference-stack crowd.

Aug 11Tuesday

Mistral AI

Mistral launches regional inference endpoints and a Priority Tier, adding third-party open models like GLM-5.2

Mistral announced general availability of Mistral Regional Endpoints, letting customers choose whether inference runs in Europe or the US. Mistral Priority Tier also entered public preview, offering custom rate limits and an availability commitment backed by an SLA.

Why it matters: Mistral puts regional inference endpoints, an SLA service tier and third-party open models on one infrastructure stack, a read on how European sovereign AI is being delivered.

Aug 8Saturday

Dwarkesh Patel podcast

The Era of Continual Learning: AI That Learns From Every Session

Dwarkesh Patel argues that once models can update weights continuously from deployment, the whole AI landscape shifts. Instead of train-then-deploy, models will learn from every interaction like a human practicing saxophone—notes alone can't transfer the skill. This breaks the current regulatory assumption of pre-deployment checks; monthly or quarterly risk inspections make more sense. Alignment research must pivot from controlling frozen weights to preventing jailbreaks or backdoors during constant updates. Commercially, the leading lab's advantage compounds: more usage yields more feedback, making the model smarter and pushing labs to ship their best models earlier. Switching costs become massive—ditching a model that has learned your org's context for months is like firing a veteran employee for a clueless intern, creating durable high margins. Enterprises will face a trade-off: accept lock-in for a model that improves with use, or lose access to top-tier AI. Labs may subsidize users who allow training on their sessions. Continual learning also increases AI mind diversity, breaking today's monoculture of a few similar base models. On the inference side, per-company full weight updates create huge batching economies; for a sparse model like DeepSeek v3, optimal batch size exceeds 2,400 concurrent sequences.

Why it matters: Dwarkesh himself is a high-credibility source in the AI podcast space, and this is his own prediction essay rather than an interview recap, with high opinion density. If continual learning lands, it genuinely destabilizes current safety frameworks — both K and R are solid. The...

Aug 5Wednesday

AI HOT (Curated Pool)

SpecForge v0.3: LMSYS releases a disaggregated speculative decoding training stack and new open draft models

SpecForge v0.3 decouples target-model inference from draft-model training. Patched SGLang servers capture features, Mooncake transports tensors, and trainer workers consume them independently. On an 8×H20 testbed, 3 servers + 5 trainers deliver ~10% higher end-to-end training throughput than the previous colocated design. The runtime now supports six speculative decoding families—EAGLE3, DFlash, Domino, DSpark, and more—and ships community-contributed draft models trained entirely on open data.

Why it matters: LMSYS disaggregated speculative decoding training into three independent contracts — capture, delivery, lifecycle — and showed ~10% throughput gain on 8×H20 with 3 servers + 5 training nodes. The architecture is clean and the numbers are concrete, but the audience is narrow (i...

Aug 4Tuesday

Latent Space

The Inference Engineering Masterclass with Baseten's Philip Kiely and Ali Taha

Baseten just raised a $13B Series F. Philip Kiely and Ali Taha explain why inference engineering is now its own discipline, covering quantization, speculative decoding, KV-cache movement, and disaggregated prefill/decode. In one GLM-5.2 experiment, quantizing more layers preserved benchmark quality while boosting throughput 20% because errors across layers canceled out. They also detail grafting Kimi's vision encoder onto GLM-5.2 without touching the language model, and note that inference optimizations can still deliver 20% to 200% gains. The conversation touches on NVIDIA Dynamo, Rubin, video generation, and local inference, but the post doesn't expand on those.

Why it matters: Baseten's $13B raise gives this deep-dive on inference engineering extra timeliness. The GLM-5.2 quantization experiment and Kimi vision encoder graft are concrete, novel details. Score stays at 78 rather than higher because it's a podcast transcript — high signal density but ...

Aug 3Monday

Hacker News front page

AirLLM runs 70B model inference on a single 4GB GPU

AirLLM is an open-source library that runs large models on consumer GPUs. It splits a model like Llama 3 70B into layers and loads them one at a time into VRAM, so a single 4GB GPU can handle inference without multi-GPU setups. It supports Llama, Mistral, ChatGLM, and other common architectures, and works with HuggingFace models. The trade-off is slower speed, but it lowers the hardware bar for local LLM inference to laptop level.

Why it matters: Fitting a 70B model onto a single 4GB GPU via layer-by-layer loading isn't a new idea, but the out-of-the-box engineering is solid. Speed is the obvious tradeoff, and the post doesn't give concrete latency numbers, so the score stays at the featured threshold.

Jul 31Friday

Latent Space

GPT-5.6 price cut by 20%-80%: March's flagship intelligence now costs 1/13th the token price

OpenAI slashed GPT-5.6 Luna to $0.20/$1.20 per million tokens, an 80% drop. Terra fell 20%, and Sol got a 2.5x faster mode at 2x the price. Luna now matches GPT-5.4's March xhigh score of 51 on the AA benchmark, at roughly 1/13th the token cost. The cuts follow GPT-5.6 rewriting its own Triton and Gluon production kernels, saving 20% end-to-end, plus speculative decoding and KV cache improvements. The post notes an annualized ~2000x cost decline but warns public benchmarks like AA may be partially trained on, so discount the headline a bit.

Why it matters: A 13x cost reduction for equivalent intelligence in four months is a major industry signal. The AA benchmark score of 51 directly ties Luna to GPT-5.4's full reasoning performance, making the price cut concrete rather than marketing fluff. The post doesn't detail the recursive...

Jul 24Friday

Hacker News front page

Echo routes prompts across open-weight models, claiming Fable-level results at 1/3 the cost

Echo is an experimental system from TracerML that pools open-weight models like GLM-5.2 and Kimi K2.7, then decides per request which models to invoke and how to combine their outputs. The author first computed a theoretical upper bound—if you could always pick the best model combination after seeing results, performance far exceeds any single model. Echo tries to approach that bound without knowing the answers in advance. On the author's own eval mix, Echo matched Fable's aggregate score at roughly one-third the inference cost. The post does not disclose specific benchmark names or absolute scores; methodology is at echo.tracerml.ai/eval. Known issues: routing and combination decisions sometimes fail, and the author is testing whether the approach holds for coding and agentic tasks. Caveat: the eval set is self-built, so saturation and representativeness are unknown—don't rush to benchmark against Fable yet.

Why it matters: Show HN project that dynamically routes across open-weight models (GLM-5.2, Kimi K2.7) to approximate Fable-level results at 1/3 cost. Concrete mechanism and numbers, but the eval set is self-built and the post doesn't disclose comparison details or sample size against Fable —...

Jul 15Wednesday

Hacker News front page

Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

The author got Google's Gemma 4 26B MoE model running on a dual Xeon E5-2690 v2 server from 2013 with no GPU, costing under $300. The CPUs only support AVX1, but ik_llama.cpp's optimized kernels require AVX2, causing silent gibberish output. Claude diagnosed that the graph builder unconditionally emitted MOE_FUSED_UP_GATE ops while the dispatcher had no matching case, leaving ~240 tensors per forward pass reading uninitialized memory. After the fix, decode reaches ~5.2 tokens/sec and prompt eval ~16 tokens/sec. A PR is open but not yet merged. The post doesn't disclose quantized model memory usage or power draw.

Why it matters: A first-person experiment with real numbers, not a generic 'run LLMs locally' tutorial. Gemma 4 26B MoE on a 13-year-old Xeon, no GPU, sub-$300 total cost — every detail is concrete. HKR all hit, but it's a personal blog experiment, not a product launch or research breakthroug...

Jul 14Tuesday

Hacker News front page

MemStitch: Zero-copy KV cache stitching for vLLM cuts multi-agent TTFT by up to 25x

DaqulaLin open-sourced MemStitch, a gateway that sits in front of vLLM and stitches KV caches across requests at the memory level using PagedAttention. It skips redundant prefill in multi-agent workflows, cutting TTFT by up to 25x and saving over 40% VRAM. The post doesn't disclose the test model or GPU setup, so I'd discount the 25x claim until those details surface.

Why it matters: Multi-agent inference latency is a real pain point, and MemStitch's approach of KV cache stitching at the vLLM memory layer is more fundamental than prompt-level solutions. The 25x claim lacks disclosed model/GPU config, so the number gets a discount, but the mechanism itself ...