Skip to content

Deployment & engineering

Running models in practice: inference optimization, memory and cost, serving architecture and infrastructure choices.

Latest picks

241–260 of 455

May 15Friday

r/LocalLLaMA

I tracked EU GPU prices across 15 stores for 50+ days: RTX 5090 is the only card not dropping

Reddit user egudegi tracked EU GPU prices across 15 stores for more than 50 days with a 6-hour scrape cadence and about 126,000 readings; RTX 5090 average pricing rose from €3,392 to €3,487, a 3.0% increase.

Why it matters: HKR-H/K/R all pass, backed by a quantified first-person price scrape. Source authority is a single Reddit post, so it sits at the featured threshold rather than a higher band.

r/LocalLLaMA

The RTX 5000 PRO 48GB arrived and is better than expected

A Reddit user built a $5,600 RTX 5000 PRO 48GB PC and ran Qwen3.6-27B-FP8 with full-precision cache; they report up to 80 tok/s in TG, about 50–60 tok/s on very large prompts, 4,400 tok/s in prompt processing, and 200k tokens fitting in BF16 KV cache.

Why it matters: HKR-H/K/R all pass: a first-person local-inference test gives price and speed numbers, not vendor copy. Single Reddit source limits reach, so it lands in the featured-threshold band.

TechCrunch · AI

Cerebras raises $5.5B, then stock pops 108%, in the first huge tech IPO of 2026

The title says Cerebras raised $5.5 billion and its stock rose 108% after the first major tech IPO of 2026; the post does not disclose the offer price, valuation, share count, or use of proceeds.

Why it matters: Cerebras opens the 2026 tech IPO window with a $5.5B raise and 108% stock pop, clearing HKR-H/K/R. It is not a model launch, but an AI-chip IPO directly hits compute capital markets, so 90 and p1.

AI HOT (Curated Pool)

Accelerating On-Device AI: Arm and Google AI Edge Optimization Practices

Arm SME2 and Google AI Edge integrate with LiteRT, XNNPACK, and KleidiAI to optimize Stability AI’s stable-audio-open-small, delivering over 2x faster audio generation and 4x lower memory use on Arm-based mobile devices and laptops.

Why it matters: HKR-H/K/R pass via concrete 2x speed and 4x memory gains, plus an edge-deployment cost hook. Scope stays narrow to one audio model on Arm devices, so it lands at the featured threshold.

May 14Thursday

Bloomberg Technology

AI Chipmaker Cerebras Climbs 68% After Year’s Biggest IPO

Cerebras Systems rose 68% in its trading debut after raising $5.5 billion in the year’s largest IPO; the post does not disclose the IPO price or valuation.

Why it matters: Cerebras pairs a $5.5B IPO with a 68% first-day jump, giving AI infrastructure a fresh public-market price signal. HKR-H/K/R all pass; no hard-exclusion rule applies.

QbitAI · WeChat

Chinese GPU Vendor Hosts Open Source Meetup With SGLang Core Developers

Moore Threads said at the SGLang × MUSA Meetup that the MUSA backend has been merged into SGLang mainline, with 47 PRs submitted and 41 merged as of May 12.

Why it matters: HKR-H/K/R all pass, but this is an inference-backend ecosystem update rather than a model launch or platform shift. The 47 PRs and 41 merges make it concrete enough for featured, not P1.

AI HOT (Curated Pool)

OpenSquilla Open-Source Project Uses Smart Routing and Local Retrieval to Cut LLM Costs

OpenSquilla combines local model routing, vector retrieval, incremental sending, and cache hits to reduce transmitted tokens by more than 90%, while routing simple tasks to cheaper models and complex tasks to stronger models without spending tokens on the routing decision.

Why it matters: HKR-H/K/R all pass, but the source appears to be a single X project post; repo traction, test setup, and limits are not disclosed. Score lands at the featured threshold for practical open-source cost tooling.

AI HOT (Curated Pool)

UnslothAI Releases Qwen3.6 MTP GGUF Models With Over 1.4x Faster Inference

Daniel Han released experimental Qwen3.6 MTP GGUF models, with the 27B model reaching 140 tokens/s on one GPU and the 35B-A3B version reaching 220 tokens/s, using two draft tokens for speculative decoding.

Why it matters: HKR-H/K/R pass via concrete single-GPU speed claims and local-inference relevance. Score stays in low featured because the post is a single X source and does not disclose GPU, quantization settings, or repro steps.

AI HOT (Curated Pool)

Moonshot AI founder Yang Zhilin releases a 40-minute video

Yang Zhilin explains Kimi K2 training in a 40-minute video, saying the model cost $4.6 million and beat GPT-5.5 and other competitors on coding tasks.

Why it matters: HKR-H/K/R all pass: the founder-led Kimi K2 training breakdown adds a $4.6M cost figure and GPT-5.5 coding comparison. Single-source X relay and missing benchmark names keep it in 78-84, not P1.

AI HOT (Curated Pool)

Cost Analysis of AI Email

Top AI models process email at about $22 to $130 per month, with a $26 median; smaller models cut costs by 10 to 20 times, while local GPU execution can bring marginal cost close to zero.

Why it matters: HKR-H/K/R pass via a concrete cost spread and deployment-cost nerve. It is a useful opinion analysis, not a major product or model release, so it sits at 73.

AI HOT (Curated Pool)

Unlocking Asynchrony in Continuous Batching

Hugging Face says an 8B model generating 8K tokens leaves the GPU idle for 24% of the time, and asynchronous batching uses CUDA streams to overlap CPU preparation for batch N+1 with GPU computation for batch N.

Why it matters: HKR-H/K/R all pass, but this is inference-systems engineering rather than a major model release. The Hugging Face post provides a concrete 24% idle-rate number and CUDA-stream overlap mechanism, placing it in low featured.

Financial Times · Technology

Cerebras boosts IPO price to raise $5.5bn

Cerebras raised its IPO price to seek $5.5bn in proceeds; the post discloses a $40bn valuation, but does not disclose the price range, share count, or listing date.

Why it matters: HKR-H/K/R all pass: FT reports Cerebras lifting its IPO price with a $5.5bn raise target and $40bn valuation. A large AI-chip listing is same-day material for compute supply, NVIDIA competition, and AI capital markets.

r/LocalLLaMA

2x RTX 3090 setup for local Qwen 3.6 27B inference

A Reddit user ran Qwen 3.6 27B on a dual RTX 3090 Ubuntu setup, reporting 48GB VRAM, a 262k context window, no NVLink, about 4000 pp/s prompt processing, and 113 tk/s generation.

Why it matters: All HKR axes pass, and this is a first-person local-inference run with concrete numbers. Source is a single Reddit post with limited reproducibility detail, so it sits at the low featured threshold.

r/LocalLLaMA

24+ tok/s from ~30B MoE models on an old GTX 1080

User mdda ran Qwen 3.6 35B-A3B on an i7-6700, GTX 1080, and 32GB RAM machine at about 24 tok/s with 128k context; the setup uses llama.cpp MoE offloading plus TurboQuant/RotorQuant KV cache quantization, with PCIe 3.0 x16 saturated and GPU utilization at about 40–50%.

Why it matters: Single Reddit source limits authority, but the GTX 1080 + Qwen 3.6 35B-A3B + 128k + 24 tok/s setup gives a concrete local-inference result. HKR-H/K/R all pass; this is a practical featured item, not a major model or product launch.

Bloomberg Technology

AI Chipmaker Cerebras Raises $5.55 Billion in Year’s Biggest IPO

Cerebras Systems raised $5.55 billion in a US IPO, and the title calls it the year’s biggest IPO; the RSS snippet does not disclose the offering price or valuation.

Why it matters: Cerebras’ $5.55B U.S. IPO clears HKR-H/K/R and is same-day material for AI infrastructure. Missing offer price and valuation keeps it below the industry-shaking band for a foundation-model-company IPO.

AI HOT (Curated Pool)

Meta AI chief announces Incognito Chat for WhatsApp and Meta AI

Meta’s AI chief announced Incognito Chat for WhatsApp and Meta AI, with conversation inference running inside the phone’s hardware secure enclave, no server logs generated, and session data permanently deleted after the chat ends.

Why it matters: HKR-H/K/R all pass: the hook is Incognito Chat in WhatsApp, with secure-enclave inference and no server logs. Single-source brevity limits verification, so it sits below model releases and major capability launches.

May 13Wednesday

QbitAI · WeChat

ByteDance Proposes Generative Refinement Networks as a Third Route for Visual Generation

ByteDance’s commercial technology team proposed GRN, a visual generation architecture using HBQ, global refinement, and complexity-aware sampling to address quantization loss, error accumulation, and fixed-step inference; on a 130M model, adaptive sampling reduced inference from 50 steps to an average of 24, while gFID changed from 3.56 to 3.79.

Why it matters: HKR-H/K/R all pass: ByteDance’s GRN has a concrete hook plus 130M, 24-step inference and gFID 3.79. It is a strong research release, not a flagship model launch, so it stays in the 78–84 band.

r/LocalLLaMA

The Trillion-Parameter Dilemma: MiMo-V2.5-Pro Open-Sourced at 1.02T Parameters

Xiaomi open-sourced MiMo-V2.5-Pro with 1.02T parameters, 42B active parameters, a 1M context window, and an MIT license; the author ran 125 Claude Code sessions through the API, spending $70.12 for 387,380,436 tokens with a 96.3% cache hit rate.

Why it matters: HKR-H/K/R all pass: a Xiaomi 1.02T open model plus a concrete Claude Code API cost experiment. Reddit sourcing keeps it at the low end of the 85+ band, but the domestic flagship-model signal clears p1.

New York Times Chinese

Jensen Huang Gets Last-Minute Invitation to Join Trump’s China Trip

Trump called Jensen Huang on Tuesday morning to invite him to join the China trip; the White House’s Monday list of 16 CEOs did not include him, while Nvidia is still seeking approval to sell AI chips to China.

Why it matters: HKR-H/K/R all pass: NYT reports a last-minute Jensen Huang invite tied to Nvidia’s China AI-chip license push. No disclosed policy change or license outcome, so this stays near the featured threshold.

New York Times Chinese

China Seeks AI Technology Self-Reliance, Weakening Washington’s Leverage Over Beijing

DeepSeek optimized its latest model for inference on Huawei chips for the first time, while two semiconductor sources said training still relies on Nvidia chips; Huawei says it plans to release a training chip this year, but matching current Nvidia performance will take another year.

Why it matters: HKR-H/K/R all pass: NYT ties DeepSeek-Huawei chip optimization and Huawei's training-chip timeline to US export-control leverage. It is not a model launch and lacks benchmark results, so it stays in the 78–84 band.