Skip to content

Deployment & engineering

Running models in practice: inference optimization, memory and cost, serving architecture and infrastructure choices.

Latest picks

281–300 of 455

May 10Sunday

r/LocalLLaMA

NVIDIA AI Releases Star Elastic: One Checkpoint Contains 30B, 23B, and 12B Reasoning Models

NVIDIA AI released Star Elastic, a single checkpoint that can zero-shot slice 30B, 23B, and 12B reasoning models in BF16, FP8, and NVFP4; when the 23B submodel handles thinking and the 30B model handles final answers, reported accuracy rises 16% and latency drops 1.9× on AIME-2025, GPQA, LiveCodeBench v5, and MMLU-Pro.

Why it matters: HKR-H/K/R all pass: Star Elastic has a concrete mechanism and testable numbers for inference deployment. Its reach is still narrower than a frontier-model release, so it sits in the high-quality featured band.

r/LocalLLaMA

BeeLlama.cpp: DFlash and TurboQuant with reasoning and vision support

Anbeeld released BeeLlama.cpp, a llama.cpp fork that runs Qwen 3.6 27B Q5 with 200k context and vision on a single RTX 3090 or 4090; the title claims 2–3x faster than baseline and a 135 tps peak.

Why it matters: HKR-H/K/R all pass, but the claims come from a Reddit title and summary without independent reproduction. Treat as a mid-weight open-source inference update, so it lands in the low featured band.

May 9Saturday

AI HOT (Curated Pool)

Redis founder uses a C inference engine to run a large model on a personal computer

Antirez open-sourced ds4, a native inference engine for DeepSeek V4 Flash that uses a few thousand lines of C to run a 1M-context model on a 128GB MacBook Pro at a reported 27 tok/s.

Why it matters: HKR-H/K/R all pass: Antirez open-sourced a native C inference engine with hardware, model, context, and speed numbers. Single-source X provenance keeps it below P1, but it is strong open-source inference signal.

r/LocalLLaMA

80 tok/sec and 128K context on 12GB VRAM with Qwen3.6 35B A3B and llama.cpp MTP

Reddit user janvitos ran Qwen3.6-35B-A3B-MTP-GGUF with a llama.cpp MTP PR on an RTX 4070 Super. The posted benchmark shows 69.2-81.9 tok/s, 0.694-0.947 draft acceptance, 131072 context, and a -fitt 1536 setting that reserves 1536 MB for the draft model and KV cache.

Why it matters: HKR-H/K/R all pass with concrete single-user benchmark data and reproducible settings. Source is one Reddit post, so verification is thin; this lands above featured threshold, not in must-write range.

AI HOT (Curated Pool)

Baidu releases ERNIE 5.1 with compressed parameters and training cost

Baidu released ERNIE 5.1 with total parameters reduced to about one third of the original scale, active parameters to about one half, and pretraining cost to about 6% of same-scale models; the model is available on the ERNIE platform and Baidu AI Studio.

Why it matters: HKR-H/K/R all pass: Baidu ERNIE 5.1 is a domestic flagship-model release with concrete compression and 6% pretraining-cost claims. That puts it in the must-write band.

Xinzhiyuan · WeChat

NVIDIA, AMD, and Intel Back RadixArk’s $100M Seed Round

RadixArk announced a $100 million seed round on May 5 at a $400 million post-money valuation, led by Accel and co-led by Spark Capital, with participation from NVentures, AMD, MediaTek, Databricks, and other investors tied to AI infrastructure.

Why it matters: HKR-H/K/R all pass: a $100M seed round, $400M post-money valuation, and chip/data investors create an AI-infra rivalry angle. It remains a single-company funding item with no product benchmarks or customer data, so it sits near the featured floor.

AI HOT (Curated Pool)

DeepSeek Raises $7 Billion at Record Scale, Founder Personally Invests $3 Billion

DeepSeek is raising up to $7 billion at a $50 billion valuation, with founder Liang Wenfeng personally contributing $3 billion, or 40% of the round, while the company says the funding will target large-scale compute, V4.1 model releases, enterprise products, and a path toward positive revenue.

Why it matters: HKR-H/K/R all pass: the $3B founder check is a strong hook, with concrete funding numbers and clear competitive resonance. Single X-source sourcing leaves lead investor, terms, and confirmation undisclosed, so it stays below P1.

r/LocalLLaMA

MTP + TurboQuant Running: Qwen3.6-27B Hits 80+ t/s on a Single RTX 4090

indrasmirror ran Qwen3.6-27B-Heretic-v2 on a single RTX 4090 with 262K context, TBQ4_0 KV cache, and MTP draft 3, improving throughput from about 43 t/s to 80-87 t/s with roughly 73% MTP draft acceptance.

Why it matters: HKR-H/K/R all pass, backed by a numbered first-person experiment. The Reddit-only source and niche local-inference focus keep it below the 78–84 band for broader industry releases.

The Verge · AI

All the Latest Updates on AI Data Centers

The Verge tracks AI data center disputes with specific updates: 43% of Americans blame data centers for rising power bills, a 40,000-acre Utah project won approval despite local opposition, and Anthropic says it will invest $50 billion in US AI data centers.

Why it matters: HKR-H/K/R all pass, but this is a Verge running roundup rather than a single breakout event. The concrete power-grid and capex numbers place it at the upper end of industry reporting.

Bloomberg Technology

Anthropic Inks $1.8 Billion Computing Deal With Akamai

Anthropic signed a $1.8 billion computing deal with Akamai to meet rising demand for its AI software; the post does not disclose capacity, contract duration, or deployment regions.

Why it matters: HKR-H/K/R all pass: the $1.8B number is concrete, the Anthropic-Akamai pairing is fresh, and the story maps to Claude compute pressure. Missing scale, term, and regions keep it just above the featured threshold.

AI HOT (Curated Pool)

EMO: Expert Mixture Models for Emergent Modular Pretraining

AllenAI introduced EMO, a mixture-of-experts model with 14B total parameters and 1B active parameters, trained on 1 trillion tokens and able to use only 12.5% of its experts for specific tasks while retaining near-full-model performance.

Why it matters: HKR-H/K/R all pass, but this is an AllenAI/Hugging Face research release rather than a frontier model launch. The 14B/1B and 12.5% expert-activation claims justify the low featured band.

May 8Friday

r/LocalLLaMA

Gemma 4 26B Hits 600 Tok/s on One RTX 5090

chain-77 benchmarked Gemma 4 26B with vLLM 0.19.2rc1, and DFlash raised output throughput on one RTX 5090 from 228 tok/s to 578 tok/s under 256 input tokens, 1024 output tokens, concurrency 1, and num_speculative_tokens=13.

Why it matters: HKR-H/K/R all pass: the single-GPU throughput hook is strong, and the post gives reproducible settings plus before/after speed. Reddit single-post evidence and one hardware setup keep it in the featured-threshold band.

Synced · WeChat

SGLang Team Launches RadixArk With $100M Seed Round

RadixArk announced a $100 million seed round on May 5 at a $400 million post-money valuation, while its SGLang inference project has 27K+ GitHub stars and deployments across 400K+ GPUs.

Why it matters: HKR-H/K/R all pass: the round size, valuation, and deployment numbers are concrete, and SGLang is a known inference stack. It is still a startup funding and infra-roadmap story, not a major model release, so it stays in the 78–84 featured band.

Xinzhiyuan · WeChat

Token-Level Length Control: 3B Model Beats GPT 5.4 and Claude

UC Santa Barbara and Apple researchers introduced LenVM, which models remaining generation length as a token-level value function; Qwen2.5-3B with a 1.5B LenVM scored 62.6 on LIFEBench length control, above GPT-5.4 at 37.4 and Claude-Opus-4-6 at 35.5.

Why it matters: HKR-H/K/R all pass: the headline has a sharp small-model-vs-frontier hook, and the post gives LenVM's mechanism plus 62.6/37.4 benchmark numbers. The topic is narrow research, not a model or major product release, so it fits the 78-84 band.

QbitAI · WeChat

HIT and Huawei propose Dynamic-dLLM, a training-free acceleration framework with 4.48x speedup

HIT Shenzhen, Huawei, and Shenzhen Hetao College proposed Dynamic-dLLM, a training-free dLLM acceleration framework that raises LLaDA-8B-Instruct throughput on GSM8k from 8.32 TPS to 37.29 TPS with almost no accuracy loss.

Why it matters: HKR-H/K/R all pass: the 4.48x speedup is clickable, and GSM8k TPS figures add concrete substance. It is inference-optimization research, not a mainstream model launch, so it fits the 78–84 band.

Financial Times · Technology

Big Tech’s $725bn AI Spending Spree Sends Free Cash Flow to a Decade Low

Big Tech is spending $725 billion on AI infrastructure, and the title says free cash flow has fallen to a decade low; the RSS snippet says Silicon Valley giants shifted from asset-light cash generators to infrastructure investors, but it does not disclose the company list, time period, or accounting basis.

Why it matters: HKR-H/K/R all pass: the FT angle ties $725bn in AI infrastructure spending to decade-low free cash flow, hitting cost and ROI anxiety. Missing company list, period, and accounting scope keeps it in the 78–84 band.

r/LocalLLaMA

11.67% ARC-AGI-2 Local Eval on a Single 4090: The TOPAS Recursive Architecture

Doug_Bitterbot says TOPAS scored 11.67% on ARC-AGI-2 using one RTX 4090 after about 14 days of training. The 100M-parameter checkpoint hit 36% locally, but recursive TTT caused null outputs on nearly half of Kaggle puzzles. The key detail is time management: the author expects 20% after threshold tuning and 3-5 more weeks of training.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post with unstable Kaggle submissions. It clears featured, not the higher research-release band.

Bloomberg Technology

Nvidia to Invest Up to $2.1 Billion in Data Center Firm IREN

Nvidia will invest up to $2.1 billion in IREN under an AI infrastructure partnership. The post discloses the cap and goal, but not equity terms, payment timing, or data center capacity.

Why it matters: Bloomberg source plus Nvidia’s up-to-$2.1B investment clears HKR-H/K/R. Details stop at amount and partnership direction, with no stake, payment schedule, or capacity, so it stays at the featured threshold.

The Verge · AI

SpaceX Has a $55 Billion Plan to Build AI Chips in Texas

SpaceX plans to invest at least $55 billion in its Terafab chip plant in Austin, Texas. A hearing notice says later phases could lift total investment to $119 billion. Musk said in March the target was chips for 200GW of compute per year; the post does not disclose process nodes.

Why it matters: HKR-H/K/R all pass on the SpaceX chip-plant hook, hard capex numbers, and compute-supply resonance. Not P1 because process node, timeline, and committed customers are not disclosed.

AI HOT (Curated Pool)

Readable behavioral signals remain in frozen LLM hidden states, Cygnus boosts accuracy

Proprioceptive AI says Cygnus adds adapters to frozen LLMs and raises Qwen-32B on ARC-Challenge from 82.2% to 94.97%. It projects hidden states into a gl(4,R) Lie-algebra space to isolate “dark modes.” Watch replication; the post does not disclose full eval sets or controls.

Why it matters: HKR-H/K/R pass: the claim is novel, quantified, and practitioner-relevant. Kept at low featured because the source is an X post and full eval set, training details, and controls are not disclosed.