Skip to content

Deployment & engineering

Running models in practice: inference optimization, memory and cost, serving architecture and infrastructure choices.

Latest picks

221–240 of 455

May 18Monday

Synced · WeChat

ICML 2026: Huawei GTS proposes EDCO for dynamic curriculum fine-tuning

Huawei GTS proposed EDCO, a dynamic curriculum method that selects fine-tuning samples by inference entropy; prefix entropy estimation cuts per-sample scoring time from 2.24 seconds to 0.37 seconds.

Why it matters: HKR-H/K/R pass: the story has a lab-race hook, a concrete entropy-based mechanism, and a 2.24s→0.37s efficiency claim. It stays below 78 because it is still a training-method paper, not a major model or product release.

r/LocalLLaMA

Benchmarking vLLM vs SGLang vs llama.cpp on a mixed Blackwell/Ada cluster

The author benchmarked long-context prefill on a 7-GPU mixed Blackwell/Ada cluster; on Qwen3.5-397B-A17B with 75k tokens, vLLM reached 9.8s TTFT and 7,683 t/s, while llama.cpp took 57.2s and 1,319 t/s.

Why it matters: Single-source Reddit benchmark, so source authority keeps it near the threshold. HKR-H/K/R pass on the mixed 7-GPU setup, 397B at 75k tokens, and concrete TTFT/throughput numbers.

r/LocalLLaMA

LLMs on Android: Snapdragon 8 Elite MoE Experience

A Reddit user tested MoE LLMs on an Honor Magic 7 Pro with Snapdragon 8 Elite and 24GB RAM; under Q4 quantization, LFM2-24b-a2b reached about 24 tokens/s while Gemma reached about 11 tokens/s, and CPU inference was still faster than NPU or GPU in the reported setup.

Why it matters: HKR-H/K/R all pass: a named Reddit test gives hardware, quantization, and token/s figures. Single-device anecdote and weak source authority keep it at the low featured band.

May 17Sunday

QbitAI · WeChat

A Robot Dog Challenges Nvidia's Compute Lead

Weilan Technology unveiled BabyAlpha A3, a consumer quadruped robot using a six-chip heterogeneous cluster that runs a 7B-parameter model on-device at 280 TPS; the article says it has 66MP vision, 2.232 million point-cloud samples per second, and a planned Q3 launch.

Why it matters: HKR-H/K/R pass: the robot-dog-versus-Nvidia angle is clickable, and 280 TPS on a local 7B model is concrete. Single-source summary lacks price, power draw, and benchmark setup, so it stays near the featured floor.

r/LocalLLaMA

Same Models Tested Across Strix Halo, RTX 3090, and RTX 5070

C_Coffie published 55 local inference benchmark runs across Strix Halo, RTX 3090, RTX 5070, five backends, and 0.35B to 35B-A3B models; RTX 5070 beats RTX 3090 on models fitting 12GiB, while RTX 3090 leads in the 14–31B band that exceeds 12GiB but fits 24GiB.

Why it matters: Hits HKR-H/K/R with a named first-person benchmark: 55 runs and concrete GPU crossover points. Source is a single Reddit post, so it stays in the low featured band.

Dwarkesh Patel podcast

Notes on Pretraining Parallelisms and Failed Training Runs

Dwarkesh documents pretraining failure modes and parallelism tradeoffs: expert choice and token dropping can break causality in MoE routing, FP16 collectives can bias repeated additions after values exceed 1024, pretraining FLOPs are given as 6ND, B300 HBM is listed as 288GB, and FSDP communication can reach params × 3 with reduce-scatter.

Why it matters: HKR-H/K/R all pass: Dwarkesh’s notes expose concrete pretraining failure modes and numbers. The systems-training focus is specialized, so it sits in the high-quality band rather than same-day must-write.

May 16Saturday

Latent Space

Cerebras' $60B IPO: Slowly, then All at Once

Cerebras closed its IPO at a $60 billion market cap. CFO Bob Komin said it serves trillion-parameter models, including internal OpenAI 5.4 and 5.5 workloads, but the post does not disclose traffic share, latency tier, or cost per token.

Why it matters: HKR-H/K/R all pass: a $60B Cerebras IPO is same-day material, and Bob Komin’s trillion-parameter/OpenAI 5.4–5.5 claim gives the story concrete compute-roadmap stakes.

r/LocalLLaMA

Orthrus-Qwen3-8B: Up to 7.8× tokens/forward on Qwen3-8B with frozen backbone

Orthrus-Qwen3-8B adds a trainable diffusion attention head to a frozen Qwen3-8B backbone, reaches up to 7.8× tokens per forward and about 6× wall-clock speedup on MATH-500, while training 16% of parameters with under 1B tokens over 24 hours on 8×H200 GPUs.

Why it matters: HKR-H/K/R all pass: 7.8× speed is a strong hook, the post gives parameter, hardware, and benchmark conditions, and inference cost is a live practitioner pain. Single-source Reddit research keeps it in the lower good-quality band.

Bloomberg Technology

Cerebras CEO Is Worth $3.2 Billion After Year’s Largest IPO

Cerebras Systems rose about 68% on Nasdaq after the year’s largest IPO, giving the company a market value of roughly $67 billion; the headline says its CEO is worth $3.2 billion after the listing.

Why it matters: HKR-H/K/R all pass: Cerebras' IPO gives AI compute a public-market price via a 68% jump and about $67B market cap. The Bloomberg item is wealth-video framed, so it lands in must-write range, not industry-shaking.

May 15Friday

r/LocalLLaMA

Fully Offline Suitcase Robot Built Around Jetson Orin NX SUPER 16GB

CreativelyBankrupt built Sparky as a fully offline suitcase robot on Jetson Orin NX SUPER 16GB, running Gemma 4 E4B Q4_K_M via llama.cpp with q8_0 KV cache, about 200 ms cached TTFT, 14-15 tok/s sustained output, 12K context, 30+ sensors, and no WiFi, Bluetooth, or cellular interface.

Why it matters: HKR-H/K/R all pass, with a named hands-on build and concrete latency/sensor numbers. It stays in low featured because this is a Reddit project post, not a product launch or research release.

AI HOT (Curated Pool)

X open-sources the “For You” feed recommendation algorithm

X open-sourced the For You recommendation pipeline on GitHub, using a Grok-based Phoenix Transformer to score candidate posts and predict engagement probabilities such as likes, replies, and reposts.

Why it matters: HKR-H/K/R all pass, but the item only gives the open-source claim and Phoenix Transformer ranking mechanism; repo details, license, and reproducible tests are not disclosed, so it stays low-featured.

r/LocalLLaMA

Evaluated a RAG Chatbot: The Most Expensive Model Was the Worst Performer

The author evaluated a customer-support RAG bot and raised the quality score from 6.62 to 7.88 while cutting per-session cost from $0.002420 to $0.000509, using retrieval logging, LLM-as-judge scoring, chunk deduplication, stricter grounding, and a five-model sweep.

Why it matters: HKR-H/K/R all pass: counterintuitive model ranking, concrete quality and cost deltas, and direct RAG production relevance. Reddit source authority keeps it near the featured floor despite the first-person experiment signal.

r/LocalLLaMA

Used over a million tokens in three sessions to test Qwen 3.6 35B MTP

A Reddit user tested Qwen3.6-35B-A3B MTP across three million-token-scale sessions, using 300k context and KV Q8_0, and reported about 1.5x the tok/sec of earlier tests.

Why it matters: HKR-H/K/R all pass: the million-token test is clickable, 300k context and KV Q8_0 add testable detail, and local speed maps to cost. Source is one Reddit post, so it stays below the high-importance band.

Bloomberg Technology

OpenAI May Raise More Money as Compute Crunch Deepens, CFO Says

OpenAI CFO Sarah Friar said the company may raise more capital after completing what she described as the largest private fundraising round ever, as OpenAI seeks computing power to meet rising AI demand; the RSS snippet does not disclose the round size, target amount, or timeline.

Why it matters: HKR-H/K/R pass: a named OpenAI CFO links more fundraising to the compute crunch. The score stays in the lower featured band because this is not a closed round and amount, investors, and timing are not disclosed.

AI HOT (Curated Pool)

Microsoft Has Invested Over $100 Billion in OpenAI, Nadella Says No One Wanted to Bet Then

Microsoft has invested more than $100 billion in OpenAI, including a $13 billion original investment and Azure infrastructure costs, while the partnership has generated about $30 billion in revenue; the renewed non-exclusive agreement caps OpenAI’s revenue share at $38 billion cumulatively through 2030.

Why it matters: HKR-H/K/R all pass: this is not a product launch, but the Microsoft-OpenAI economics include three concrete figures on spend, revenue, and revenue-share caps, placing it in the 78–84 quality band.

AI HOT (Curated Pool)

The First Derivative of Inference: Growth Logic in the AI Wave

Tom Tunguz says the AI inference market will reach $250 billion within seven years; Datadog’s LLM observability data volume nearly doubled in the latest quarter, and about 20% of its AI customers contribute roughly 80% of ARR.

Why it matters: HKR-H/K/R all pass: Tom Tunguz ties inference growth to Datadog volume and ARR concentration data. It stays in the 72–77 band because this is commentary, not a model, product, or protocol release.

AI HOT (Curated Pool)

API prompt precaching speeds up first-token generation

Claude API prewarms prompt cache with the system prompt, skips output, then hits cache on the real request.

Why it matters: HKR-H/K/R all pass: this is a concrete Claude API latency mechanism, not a vague product tease. It clears featured, but it is a mid-weight inference update rather than a major model or capability release.

Bloomberg Technology

Cerebras Goes Public in Year's Biggest IPO | Bloomberg Tech 5/14/2026

Bloomberg Tech says Cerebras is going public in the year’s biggest IPO, while the RSS snippet does not disclose the fundraising amount, valuation, offer price, or exchange.

Why it matters: HKR-H/K/R all pass, but the item only gives a title-level Bloomberg claim with no proceeds, valuation, price, or exchange. Cerebras’ IPO matters for AI infrastructure markets, so it lands in the 78–84 band.

Bloomberg Technology

AI Chipmaker Cerebras Climbs 68% After Year’s Biggest IPO

Cerebras Systems shares closed up 68% in their trading debut after raising $5.5 billion in the year’s largest IPO, ending Thursday at $311.07 in New York versus the $185 IPO price.

Why it matters: HKR-H/K/R all pass: Cerebras’ IPO gives AI infrastructure a public-market pricing signal, with a 68% debut gain and $5.5B raised. The chip-supply and NVIDIA-competition angle pushes it into p1.

Bloomberg Technology

Cerebras CEO Is Worth $3.2 Billion After Year’s Largest IPO

Cerebras’ IPO valued CEO Andrew Feldman’s fortune at $3.2 billion, and the title identifies the listing as the year’s largest IPO; the RSS snippet only says Feldman previously sold three companies and took another public, and does not disclose the IPO price, proceeds, valuation, or share count.

Why it matters: HKR-H/K/R all pass: Cerebras’ “year’s largest IPO” is a real AI-infra capital-market signal. Score stays in 78–84 because the feed gives $3.2B net worth but not offer price, proceeds, or valuation.