Skip to content

#部署/工程

2 today

Today · Sep 30Wednesday · 2 items

Hacker News front page

What if Jev spoke Arrow?

TypeSafe AI 推出新模型 Jev,可将自然语言和应用状态转化为带类型决策,返回选项、分数和概率并以 JSON 输出。其采用并行采样器与名为 Reinforcement Learning for Calibrated Decisions 的训练方法,在决策工作流中相比通用 LLM 有显著速度和成本优势。

Yesterday · Sep 29Tuesday

MIT Technology Review · AI

Making AI an asset, not an expense

HPE 提出,当 AI 从试验走向客服、IT、研究等常驻生产负载,按 token 消费的模式会让支出变成难以预测的月度变动项,企业需按工作负载判断是否转向自有算力。

Hacker News front page

World Labs to join AMD; Fei-Fei Li becomes Chief Scientist

World Labs signed a definitive agreement to join AMD. Fei-Fei Li will become AMD's EVP and Chief Scientist, reporting to CEO Lisa Su. Justin Johnson and Ben Mildenhall will keep leading the team as it forms a new frontier research org inside AMD. The two companies started a technical partnership last year on model training and inference optimization on AMD GPUs. The goal is an end-to-end open AI ecosystem spanning hardware, software, and open models. The deal is expected to close by end of 2026, pending regulatory approvals.

Why it matters: Fei-Fei Li's startup joining AMD with her as Chief Scientist is one of the year's biggest personnel + strategy mergers. All three HKR axes hit: the personnel pairing creates suspense, the technical partnership has concrete timeline and goals, and the academia-to-industry arc r...

Sep 28Monday

Mistral AI

Hallo, Deutschland!

Mistral 在慕尼黑开设德国中心,组建专注 Physics AI 与工业 AI 的研究团队,并计划到 2030 年建成 1 吉瓦欧洲算力。该中心将携手 BMW 开展碰撞仿真与工程 AI 合作、与 Siemens Energy 推进工业 AI 应用,并与慕尼黑工业大学(TUM)合作利用风洞设施开发汽车空气动力学数字孪生。

Sep 26Saturday

r/LocalLLaMA

llama.cpp merges tiled mul_mat for k-quants, CPU inference speedup expected

PR #27851 by jbooth adds tiled matrix multiplication for k-quants in llama.cpp's ggml-cpu backend. The post body is blocked by Reddit, so no speedup or memory numbers are disclosed. Tiled mul_mat improves CPU cache utilization and reduces memory bandwidth pressure—a real win for local LLM inference.

Sep 25Friday

GitHub Blog · AI & ML

When chat is the wrong UI

GitHub Copilot 应用推出 canvas,一种运行在应用内、无浏览器外壳的全栈小应用,可与 Copilot 智能体双向通信,并能在本地执行代码、调用第三方 API。作者认为聊天只是 AI 的通用兜底界面,用户明确任务时更该让智能体生成可复用工具,而非把智能体本身当工具、白白消耗 token。示例包括 Connect 4 游戏、Winget 包管理、SQLite 操作和开发工作流自动化。

Sep 24Thursday

Ars Technica · AI

XPRIZE Wildfire winners spotted fires within 10 min—but couldn’t stop them

XPRIZE Wildfire 公布 1100 万美元竞赛结果,两大赛道大奖均空缺。太空探测赛道要求 10 分钟内识别澳大利亚大范围火情,SIRIUS Wildfire Alliance 获 50 万美元一等奖;自主响应赛道三支决赛队均在 10 分钟内探测到阿拉斯加高风险火情,但无一完全扑灭。

GitHub Blog · AI & ML

Rendering huge pull requests in the GitHub Copilot app

GitHub Copilot 应用重建了 pull request 视图,以流畅渲染含 2,200 个文件、超百万行改动和 400 多条行内评论的超大 PR。其做法是把文档高度拆成确定性的代码几何与动态评论块两套几何:代码行高提前精确算好,评论高度按块懒测量并锚定到文件、行与侧,避免滚动跳动。

Google DeepMind

Google DeepMind adds secure server-side memory to Private AI Compute

Google DeepMind detailed a new capability for Private AI Compute: private, server-side persistent memory that lets an AI assistant keep context across devices. Data sits sealed in encrypted storage, and the unlock key stays only on the user's device. When the model needs access, an end-to-end encrypted channel carries it into a secure cloud enclave, where it is briefly decrypted in isolated memory and immediately re-encrypted.

Why it matters: The post explains how cloud persistent memory uses secure enclaves and device-held keys for privacy, a look at the privacy architecture behind cloud AI memory.

Sep 23Wednesday

AI HOT (Curated Pool)

Modal details how to serve trillion-parameter coding agents at trillion-token scale

Modal's engineering team published a deep-dive on serving Moonshot AI's Kimi K2.6 for coding agents. They boosted per-replica performance by 2.8x per user and 5.6x across users, turning a ruinously expensive service into a price-competitive one. One service processed hundreds of billions of tokens per day and trillions in aggregate. The post walks through workload analysis for hybrid-attention MoE models and the engineering optimizations applied. Exact GPU models and per-request latency numbers are not disclosed in the body.

Why it matters: Modal's engineering breakdown of inference acceleration for Moonshot AI's Kimi K2.6 delivers two hard numbers — 2.8x and 5.6x speedups — directly useful for inference engineers. Not scored higher because this is infra optimization, not a model capability or product update; aud...

Simon Willison

llm 0.36

Simon Willison 发布 llm 0.36。该版本信息由其本人在 9 月 22 日发布,具体更新内容原文未作说明。

Sep 21Monday

AI HOT (Curated Pool)

AI Comes for the If Statement: Specialized Deciders Cut Classification Cost ~100x

Tomasz Tunguz tested Jev and SemIf on 98 production emails: classification accuracy jumped from 47% to over 80%, cost dropped to $0.0004 per call—76x to 209x cheaper than frontier models. These deciders skip text generation, run attention once, and output choice probabilities in hundreds of milliseconds. Tunguz sees this as a bifurcation: frontier models for discovery, specialized models for production, with if-then as the first optimized programming primitive.

Why it matters: Tunguz ran a real production test with concrete accuracy and cost numbers—not just trend talk. The piece flags a meaningful fork in AI infra: general-purpose generators vs. specialized deciders. Score stays at 78 because both tools are brand-new with no large-scale validation ...

Simon Willison

llm-keys-ui 0.1

Simon Willison 发布 llm-keys-ui 0.1 插件,用于在不向 ChatGPT 应用或智能体会话粘贴 API key 的情况下,把密钥配置到远程机器上。

Sep 19Saturday

Hacker News front page

Apple M6 Pro tops Geekbench 7 single-core chart

Apple M6 Pro scored 4150 in Geekbench 7 single-core, the highest public result so far. The 18-core chip runs at 4.78 GHz base, split into 6+12 clusters, with 48 GB RAM. Multi-core hit 37565. But this is just a benchmark—real-world power, thermals, and shipping devices aren't covered here.

Hacker News front page

C2C lets LLMs talk via KV-cache, 2.5× faster and 3–5% more accurate than text

This ICLR'26 paper proposes Cache-to-Cache (C2C), where multiple LLMs communicate by directly exchanging KV-cache instead of generating text. A neural network projects and fuses the source model's KV-cache into the target model, with a learnable gate selecting which layers benefit. C2C beats single models by 6.4–14.2% in average accuracy, outperforms text-based communication by ~3.1–5.4%, and delivers an average 2.5× latency speedup. Code is open at thu-nics/C2C.

Why it matters: ICLR'26 paper with a clever idea: let collaborating models pass KV-cache directly instead of text. Has validation experiments and a concrete framework, so knowledge density is solid. Score capped because it's low-level optimization — not immediately actionable for most practit...

Sep 17Thursday

Computing Life · Share · Yage

AWS partners Qualcomm for custom chips, Claude used for attack drone code, 25 Fields medalists oppose math benchmarking, Arena finds cross-vendor model similarity higher

Qualcomm and AWS announced a multi-generational collaboration to develop custom AI inference chips and 1.6 Tbps optical interconnects for data centers, but no chip model names or delivery timelines were disclosed. Anthropic's threat report revealed a Russian freelance team used Claude Code to develop autonomous attack drone software, with hardware-in-the-loop testing and TRL 3-4, but no evidence of deployment. 25 Fields Medalists signed a statement criticizing commercial math benchmarking that prioritizes answer scores over methodological contributions; OpenAI subsequently withdrew sponsorship from the Caltech Mathathon. Arena analyzed 30,086 model response pairs and found Fable 5 shared 59.2% conceptual overlap with DeepSeek V4 Pro, higher than with its sibling Claude Opus 5 at 41.2%, showing cross-vendor similarity can exceed in-family similarity.

Hacker News front page

Ternary LLMs break the 1.58-bit floor: BITCOS hits 1.485 bits per weight by exploiting zero-weight density

Ternary models store weights as -1, 0, or +1, with a theoretical floor of ~1.585 bits and a practical 1.625 bits in five-trit packing. Intel authors measured 29 ternary LLMs and found up to 51.5% zeros. BITCOS replaces fixed packing with a presence bitmap plus a compacted sign vector, costing 2 minus zero-density bits per weight. It beats five-trit packing on 26 of 29 models and reaches 1.485 bits on the sparsest. Optimized unpacking on AVX-512, AVX2, and Xe2 GPUs yields up to 1.28× faster matrix-vector multiply; end-to-end decode throughput improves up to 1.18× on CPUs and 1.27× on GPUs. The paper does not name the models or disclose their parameter counts, nor whether they are publicly available.

Why it matters: Intel team measured 29 ternary models, found zero weights up to 51.5%, and proposed BITCOS encoding to break the 1.58-bit floor. Solid K with concrete numbers and a new mechanism; H works on title intrigue. But it's a narrow inference-opt topic with no R pull, so it lands at t...

Sep 16Wednesday

NVIDIA Blog

From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production

NVIDIA published a blog on converting power efficiency into token output for AI factories. The key idea: measure tokens per watt, not just GPU flops. It covers full-stack optimization from data center design and cooling to inference tuning, aiming to run AI factories like production lines. The post does not disclose specific efficiency gains or new hardware SKUs.

Sep 15Tuesday

Hacker News front page

The Inference Hardware Revolution of 2026

IEEE Spectrum reports that 2026 is seeing a revolution in inference hardware. Specialized chips now focus on optimizing inference rather than just training, making deployed AI models faster and cheaper. The article claims inference efficiency has improved over 10x in the past two years, driven by architectural innovation and memory bandwidth breakthroughs. The post does not name specific companies or chip specs.

r/LocalLLaMA

DeepSeek V4.1 Flash Q4 hits 40 t/s on M3 Ultra with native DSpark multi-token prediction

A developer forked ds4 and tuned it for DeepSeek V4.1 Flash Q4 on a 512 GB M3 Ultra, lifting decode from 16.6 t/s to 31.3 t/s, and to 40.5 t/s with DSpark speculative decoding. The ~300 GB Q4 weights fit only the 512 GB M3 Ultra. The speedup comes from cutting Metal dispatch overhead: the 384-expert router went from 9 dispatches to 1, the shared expert gate+up+SwiGLU became a single kernel, and BF16 rounding moved inside producer kernels, removing ~770 re-round dispatches. At 300k context, compressed attention selection was the bottleneck; the fix scores only admitted blocks and uses a bounded radix select, keeping decode at 90% of the 8k rate. DSpark verifies 6 tokens per step with shared weight streams and overlapped Engram fetches, dropping verify latency from 177 ms to 112 ms. Output is byte-identical to upstream under greedy decode, with SHA-256 manifests provided. The branch is M3 Ultra only because the optimizations rely on measured behavior of this specific chip's dual-die memory, 80-core GPU scheduling, and Metal dispatch characteristics.

Why it matters: A solid local-inference optimization post: DeepSeek V4.1 Flash Q4 on M3 Ultra goes from 16tps to 40tps via Metal command-buffer merging and native DSpark speculative decoding. Concrete technical detail, directly useful for the local-LLM crowd. Score stays at 72 because the aud...

Sep 14Monday

r/LocalLLaMA

Llama.cpp PR adds Maple 20B-A1B ternary MoE architecture for CPU inference

Reddit user AlexGabbia submitted PR #27000 to llama.cpp, adding the Maple 20B-A1B ternary MoE architecture. The model has 20B total parameters but activates only 1B per inference, designed for CPU execution. Ternary weights (-1, 0, 1) cut memory and compute significantly, enabling faster large-model inference on ordinary CPUs. The post is blocked by Reddit and does not disclose specific performance numbers or release timeline.

Sep 10Thursday

Mistral AI

Cloudera and Mistral Partner to Bring Specialized, Sovereign Intelligence to Enterprise Data

Mistral 与 Cloudera 宣布合作,将 Mistral 模型集成进 Cloudera 混合数据平台,企业可在私有云、公有云、本地及完全气隙环境中部署推理并保持完全控制。企业还能在受控环境中用专有数据训练定制模型,数据与模型所有权均归企业,模型基于开放权重。Cloudera 平台上客户管理的数据规模达 30 exabytes。

Sep 7Monday

Hacker News front page

vLLM explores speculative decoding on AMD GPUs with five draft methods

vLLM's blog post benchmarks speculative decoding on AMD MI300X GPUs. The technique uses a lightweight draft model to propose tokens, then the target model verifies them in one pass, committing multiple tokens at once. The post compares five draft methods—native MTP, Gemma 4 MTP, EAGLE-3, DFlash, and DSpark—which differ in how they receive target-model info and generate candidates. Throughput gains vary by draft method, proposal length, model family, workload, and acceptance rate. The post also includes tuning guidance and a training workflow for custom speculators.

Sep 4Friday

Sep 2Wednesday

Hacker News front page

Slotstream runs the 104GB Qwen3.8-Flash-Next on a 48GB Mac at ~12 tok/s

carloslfu open-sourced Slotstream, an MLX + Swift tool that runs the 125B-parameter MoE model Qwen3.8-Flash-Next (104GB at 4-bit) on a Mac with only 48GB RAM. It streams expert modules from SSD on demand instead of loading everything into memory. Speed is ~12 tok/s, and it exposes an Ollama-compatible API. The post doesn't disclose time-to-first-token or the SSD model used, so real-world feel is still an open question.

Why it matters: Streams MoE experts from SSD on demand via MLX + Swift, letting a 104GB Qwen 125B model hit ~12 tok/s on a 48GB Mac. Clean engineering with an Ollama-compatible API that lowers the trial barrier. Docked a few points because it's a solo project with no community validation or m...

Aug 26Wednesday

Computing Life · Share · Yage

The term 'local LLM' conflates two separate markets

Yage breaks down 'local LLM' into two markets: a cost market buying 5–20× price gaps, and a control market buying 25–33-year certainty. Using a four-quadrant framework (open/closed weights × time/token billing), the piece explains why surging open-weight model usage on OpenRouter doesn't mean local deployment is winning. Self-hosting payback depends entirely on which cloud billing mode you replace—decades for subscriptions, months for high-cache-hit agent API calls. In July–August 2026, Anthropic and others made four moves at the inference layer: silently remapping parameters, repeatedly extending usage boosts, adding watermarks, and redefining self-hosting as 'your harness plus my inference.' But simultaneous deep price cuts mean the misalignment is real but direction is unresolved.

Why it matters: Splits 'local LLM' into cost vs control markets with OpenRouter data and hardware payback math — directly useful for infra decision-makers. Not scored higher because it's commentary rather than a product launch or research breakthrough, but hits all three HKR axes and earns a ...

Aug 25Tuesday

Mistral AI

Mistral x HUMAIN

Mistral 与 HUMAIN 宣布战略合作,覆盖 AI 基础设施、先进模型开发与 AI 解决方案部署,初期聚焦网络安全和语音,并计划开发阿拉伯语表现强劲的前沿模型。合作规模达数亿欧元,Mistral 将探索使用 HUMAIN 数据中心基础设施,双方还将在沙特面向受监管行业制定联合市场策略。

Aug 21Friday

Hacker News front page

Nari Labs pushes Qwen3-TTS to sub-50 ms time-to-first-audio at 10 RPS on a single H100

Nari Labs open-sourced a Qwen3-TTS 1.7B CustomVoice serving implementation that hits sub-50 ms p95 time-to-first-audio at 10 RPS on a single H100 SXM with zero underruns. They benchmarked against vLLM-Omni, SGLang-Omni, VoxServe, and M*—default p95 latencies ranged from 277 to 1,160 ms at 1 RPS. At full utilization the system costs roughly $2 per 1M characters, compared to $100 for ElevenLabs V3 and $49 for Cartesia Sonic 3.5. Key optimizations include dynamic leading-silence trimming (~80 ms saved) and tuned codec-frame accumulation. Code and benchmarks are public; the post does not disclose underrun details at higher concurrency or long-form performance.

Why it matters: Nari Labs open-sourced a deployment recipe for Qwen3-TTS 1.7B that hits sub-50 ms p95 time-to-first-audio at 10 concurrent requests on a single H100—an order-of-magnitude improvement over vLLM-Omni and others. The post includes concrete benchmarks and reproducible optimization...

Aug 20Thursday

Hacker News front page

DFlash 2 pushes parallel drafting further: over 20% more output per verification pass for ~1% added latency

Inco AI released DFlash 2, adding a lightweight path selector on top of parallel speculative decoding. Instead of keeping only the top-1 candidate per position, it picks a coherent path from the top 16, raising accepted tokens per verification from 4.27 to 6.79. On Qwen3.8-27B with SGLang, throughput reaches 2.7–3.4× autoregressive decoding at batch size 1, with roughly 1% added cycle latency. SGLang, vLLM, llama.cpp, and oMLX already support it; DFlash models have been downloaded over 3.5 million times on Hugging Face.

Why it matters: DFlash 2 is a clear technical improvement on an already-adopted inference method, with measured results. Score isn't higher because this is a single technical blog post, not a model launch or product release — its reach is limited to the inference-stack crowd.

Aug 11Tuesday

Mistral AI

Mistral launches regional inference endpoints and a Priority Tier, adding third-party open models like GLM-5.2

Mistral announced general availability of Mistral Regional Endpoints, letting customers choose whether inference runs in Europe or the US. Mistral Priority Tier also entered public preview, offering custom rate limits and an availability commitment backed by an SLA.

Why it matters: Mistral puts regional inference endpoints, an SLA service tier and third-party open models on one infrastructure stack, a read on how European sovereign AI is being delivered.

Aug 8Saturday

Dwarkesh Patel podcast

The Era of Continual Learning: AI That Learns From Every Session

Dwarkesh Patel argues that once models can update weights continuously from deployment, the whole AI landscape shifts. Instead of train-then-deploy, models will learn from every interaction like a human practicing saxophone—notes alone can't transfer the skill. This breaks the current regulatory assumption of pre-deployment checks; monthly or quarterly risk inspections make more sense. Alignment research must pivot from controlling frozen weights to preventing jailbreaks or backdoors during constant updates. Commercially, the leading lab's advantage compounds: more usage yields more feedback, making the model smarter and pushing labs to ship their best models earlier. Switching costs become massive—ditching a model that has learned your org's context for months is like firing a veteran employee for a clueless intern, creating durable high margins. Enterprises will face a trade-off: accept lock-in for a model that improves with use, or lose access to top-tier AI. Labs may subsidize users who allow training on their sessions. Continual learning also increases AI mind diversity, breaking today's monoculture of a few similar base models. On the inference side, per-company full weight updates create huge batching economies; for a sparse model like DeepSeek v3, optimal batch size exceeds 2,400 concurrent sequences.

Why it matters: Dwarkesh himself is a high-credibility source in the AI podcast space, and this is his own prediction essay rather than an interview recap, with high opinion density. If continual learning lands, it genuinely destabilizes current safety frameworks — both K and R are solid. The...

Aug 5Wednesday

AI HOT (Curated Pool)

SpecForge v0.3: LMSYS releases a disaggregated speculative decoding training stack and new open draft models

SpecForge v0.3 decouples target-model inference from draft-model training. Patched SGLang servers capture features, Mooncake transports tensors, and trainer workers consume them independently. On an 8×H20 testbed, 3 servers + 5 trainers deliver ~10% higher end-to-end training throughput than the previous colocated design. The runtime now supports six speculative decoding families—EAGLE3, DFlash, Domino, DSpark, and more—and ships community-contributed draft models trained entirely on open data.

Why it matters: LMSYS disaggregated speculative decoding training into three independent contracts — capture, delivery, lifecycle — and showed ~10% throughput gain on 8×H20 with 3 servers + 5 training nodes. The architecture is clean and the numbers are concrete, but the audience is narrow (i...

Aug 4Tuesday

Latent Space

The Inference Engineering Masterclass with Baseten's Philip Kiely and Ali Taha

Baseten just raised a $13B Series F. Philip Kiely and Ali Taha explain why inference engineering is now its own discipline, covering quantization, speculative decoding, KV-cache movement, and disaggregated prefill/decode. In one GLM-5.2 experiment, quantizing more layers preserved benchmark quality while boosting throughput 20% because errors across layers canceled out. They also detail grafting Kimi's vision encoder onto GLM-5.2 without touching the language model, and note that inference optimizations can still deliver 20% to 200% gains. The conversation touches on NVIDIA Dynamo, Rubin, video generation, and local inference, but the post doesn't expand on those.

Why it matters: Baseten's $13B raise gives this deep-dive on inference engineering extra timeliness. The GLM-5.2 quantization experiment and Kimi vision encoder graft are concrete, novel details. Score stays at 78 rather than higher because it's a podcast transcript — high signal density but ...

Aug 3Monday

Hacker News front page

AirLLM runs 70B model inference on a single 4GB GPU

AirLLM is an open-source library that runs large models on consumer GPUs. It splits a model like Llama 3 70B into layers and loads them one at a time into VRAM, so a single 4GB GPU can handle inference without multi-GPU setups. It supports Llama, Mistral, ChatGLM, and other common architectures, and works with HuggingFace models. The trade-off is slower speed, but it lowers the hardware bar for local LLM inference to laptop level.

Why it matters: Fitting a 70B model onto a single 4GB GPU via layer-by-layer loading isn't a new idea, but the out-of-the-box engineering is solid. Speed is the obvious tradeoff, and the post doesn't give concrete latency numbers, so the score stays at the featured threshold.

Jul 31Friday

Latent Space

GPT-5.6 price cut by 20%-80%: March's flagship intelligence now costs 1/13th the token price

OpenAI slashed GPT-5.6 Luna to $0.20/$1.20 per million tokens, an 80% drop. Terra fell 20%, and Sol got a 2.5x faster mode at 2x the price. Luna now matches GPT-5.4's March xhigh score of 51 on the AA benchmark, at roughly 1/13th the token cost. The cuts follow GPT-5.6 rewriting its own Triton and Gluon production kernels, saving 20% end-to-end, plus speculative decoding and KV cache improvements. The post notes an annualized ~2000x cost decline but warns public benchmarks like AA may be partially trained on, so discount the headline a bit.

Why it matters: A 13x cost reduction for equivalent intelligence in four months is a major industry signal. The AA benchmark score of 51 directly ties Luna to GPT-5.4's full reasoning performance, making the price cut concrete rather than marketing fluff. The post doesn't detail the recursive...

Jul 24Friday

Hacker News front page

Echo routes prompts across open-weight models, claiming Fable-level results at 1/3 the cost

Echo is an experimental system from TracerML that pools open-weight models like GLM-5.2 and Kimi K2.7, then decides per request which models to invoke and how to combine their outputs. The author first computed a theoretical upper bound—if you could always pick the best model combination after seeing results, performance far exceeds any single model. Echo tries to approach that bound without knowing the answers in advance. On the author's own eval mix, Echo matched Fable's aggregate score at roughly one-third the inference cost. The post does not disclose specific benchmark names or absolute scores; methodology is at echo.tracerml.ai/eval. Known issues: routing and combination decisions sometimes fail, and the author is testing whether the approach holds for coding and agentic tasks. Caveat: the eval set is self-built, so saturation and representativeness are unknown—don't rush to benchmark against Fable yet.

Why it matters: Show HN project that dynamically routes across open-weight models (GLM-5.2, Kimi K2.7) to approximate Fable-level results at 1/3 cost. Concrete mechanism and numbers, but the eval set is self-built and the post doesn't disclose comparison details or sample size against Fable —...

Jul 15Wednesday

Hacker News front page

Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

The author got Google's Gemma 4 26B MoE model running on a dual Xeon E5-2690 v2 server from 2013 with no GPU, costing under $300. The CPUs only support AVX1, but ik_llama.cpp's optimized kernels require AVX2, causing silent gibberish output. Claude diagnosed that the graph builder unconditionally emitted MOE_FUSED_UP_GATE ops while the dispatcher had no matching case, leaving ~240 tensors per forward pass reading uninitialized memory. After the fix, decode reaches ~5.2 tokens/sec and prompt eval ~16 tokens/sec. A PR is open but not yet merged. The post doesn't disclose quantized model memory usage or power draw.

Why it matters: A first-person experiment with real numbers, not a generic 'run LLMs locally' tutorial. Gemma 4 26B MoE on a 13-year-old Xeon, no GPU, sub-$300 total cost — every detail is concrete. HKR all hit, but it's a personal blog experiment, not a product launch or research breakthroug...