Skip to content

Deployment & engineering

Running models in practice: inference optimization, memory and cost, serving architecture and infrastructure choices.

Latest picks

41–60 of 455

Jun 8Monday

r/LocalLLaMA

Xiaomi claims 1,000+ tps on a 1T model using a standard 8-GPU server

Xiaomi MiMo claims MiMo-V2.5-Pro UltraSpeed runs a 1T-parameter MoE model above 1,000 output tokens per second on one standard 8-GPU node; the post does not disclose the GPU model, batch settings, or reproducible configuration.

Why it matters: HKR-H/K/R all pass: the 1T MoE and 1,000+ tps claim is a strong inference-cost hook. Kept below P1 because the post lacks GPU model, batch size, quantization, and reproducible setup.

Hacker News front page

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

The title says Xiaomi MiMo-v2.5-Pro-UltraSpeed is a 1T model running at 1,000 tokens per second; the RSS body only provides the URL, Hacker News comments link, 66 points, and 14 comments, and the post does not disclose hardware, precision, context window, benchmark setup, or availability.

Why it matters: HKR-H/K/R all pass: Xiaomi’s MiMo update has a sharp 1T/1,000 tokens/s claim and clear cost-speed resonance. Missing hardware, precision, context window, and test setup keep it in the 78–84 band, not p1.

r/LocalLLaMA

Luce Spark: a 35B MoE on a 16 GB GPU, without the offload tax

Luce Spark runs Qwen3.6 35B-A3B at 13.3 GiB peak VRAM on an RTX 3090 by keeping hot experts on GPU, swapping cold experts through a bounded async cache, and using one fused graph for decode at about 100 tok/s.

Why it matters: HKR-H/K/R all pass: the hook is a 35B MoE on a 16 GB GPU, with 13.3 GiB peak use and ~100 tok/s. Reddit-source and no third-party replication keep it at 78.

AI HOT (Curated Pool)

Xiaomi MiMo-V2.5-Pro-UltraSpeed Exceeds 1,000 Tokens/s

Xiaomi MiMo and TileRT_AI released MiMo-V2.5-Pro-UltraSpeed, running a 1T MoE model above 1,000 tokens/s on a single standard 8-GPGPU node, with UltraSpeed API priced at 3x and applications open from June 8 to 23 PDT.

Why it matters: HKR-H/K/R all pass: Xiaomi MiMo gives a concrete claim of a 1T MoE exceeding 1,000 tokens/s on one 8-GPGPU node. The score stays at 80 because this is single-source and lacks task mix, precision, latency, and cost details.

r/LocalLLaMA

DFlash Speculative Decoding and KV Cache Compression on RTX 5090 Show 3.26x Speedup

The author tested Qwen3.6-27B on an RTX 5090 with DFlash plus KV cache compression, reaching up to 3.26x speedup; q4_0/turbo4 delivered 3.18x speedup with only +0.02% PPL on WikiText-2.

Why it matters: HKR-H/K/R all pass: RTX 5090 testing, DFlash speculative decoding, KV cache compression, 3.26x speedup, and PPL delta are concrete. Single Reddit source keeps it near the featured floor.

r/LocalLLaMA

Weird to get near-linear scaling by adding another GPU?

A Reddit user benchmarked qwen3.6-27b-autoround-int4 on 1x3090 versus 2x3090. Narrative decode rose from 53 TPS to 94 TPS, and code decode rose from 62 TPS to 120 TPS, under no NVLink, 8x/8x PCIe, P2P automatically enabled, tensor parallelism set to 2, and different KV-cache settings.

Why it matters: HKR-H/K/R all pass: the result is counterintuitive, includes concrete TPS and TP conditions, and speaks to local-inference cost. Single Reddit test lacks multi-model replication and full setup details, so it stays near the featured threshold.

Synced · WeChat

openJiuwen proposes MANGO for multi-agent flow networks

openJiuwen proposed MANGO, a multi-agent flow-network framework that combines reinforcement learning, textual gradients, and a Skip-k mechanism; using GPT-4o-mini, it reports a 12.8% accuracy gain over MaAS on MATH500 and a 5.1% F1 gain over AFlow on DROP.

Why it matters: HKR-K is strong: the post gives mechanisms and a MATH500 delta. HKR-H/R pass for the multi-agent flow-network angle, but this remains a research-framework story, not a major model or platform release.

Synced · WeChat

Alibaba RTPurboV2 uses hundreds of training steps for 10x sparse attention

Alibaba’s RTP team released RTPurboV2, replacing 85% of attention heads with SWA and compressing the remaining 15% retrieval heads using low-rank projection, clustering, and dynamic top-p; the adaptation uses about 600 training steps and roughly 1M label tokens, with reported Prefill speedup up to 9.36x.

Why it matters: HKR-H/K/R all pass: the hook is 100-step training and 10x sparse attention, with concrete mechanisms and 9.36x Prefill speedup. Strong engineering signal from Alibaba, but not a flagship model release or major product launch.

AI HOT (Curated Pool)

Apple Releases Third-Generation Apple Foundation Models (AFM)

Apple released its third-generation AFM family with five models. The RSS snippet says they span on-device use and Private Cloud Compute servers, with Google involved in customization for Apple Intelligence, Siri, and system-level tools.

Why it matters: Official Apple model-family release with 5 models, on-device/PCC deployment, and Google customization clears HKR-H/K/R. Missing benchmark and pricing details keep it at the low end of the 85+ band.

AI HOT (Curated Pool)

Nvidia and SK Hynix Sign Multi-Year Pact to Develop Next-Generation AI Memory Chips

Nvidia and SK Hynix signed a multi-year pact to co-design future generations of memory chips for AI applications; the RSS snippet does not disclose product specifications, production timelines, or financial terms.

Why it matters: HKR-H and HKR-R pass: Bloomberg plus Nvidia/SK Hynix matters for AI memory supply. HKR-K is weak because specs, production timing, and financial terms are missing, so this sits at the low featured band.

Jun 7Sunday

r/LocalLLaMA

Qwen3.6 35B-A3B on a Laptop: My Zero-to-One Moment

A Reddit user ran Qwen3.6 35B-A3B on an ASUS Zenbook Pro 14 with RTX 4060 8GB VRAM and 64GB RAM, reaching about 27 TPS at 32k context and 18 TPS at 256k context. The setup uses llama.cpp, unsloth’s IQ3_XXS GGUF quantization, and a 262144-token context flag.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit experiment, not an official release or paper. Concrete hardware, quantization, context, and TPS clear the featured bar, but keep it in the 72–77 band.

QbitAI · WeChat

Chinese open-source framework targets stable 5-minute AI long-video generation

JD open-sourced JoyAI-Echo, a long audio-video generation framework for 5-minute consistent videos, using cross-modal memory, DMD post-training for about 7.5x faster inference, and real-time upscaling from 720P to 1K or 2K output.

Why it matters: HKR-H/K/R all pass: the story has a clear 5-minute video hook, concrete speed and SR claims, and open-source competition resonance. Missing third-party evaluation keeps it in the lower 78–84 band.

AI HOT (Curated Pool)

AI Substitution Wave: Three Forces Reshape Cost Structures

Coinbase, Lindy, Harvey, and Cursor shifted workloads to cheaper models; Harvey reported Kimi 2.6 reached a 15% all-pass rate on Legal Agent Benchmark, versus Opus at 14%, with 100 tasks costing $84 versus $954.

Why it matters: HKR-H/K/R all pass: the $84 vs $954 cost delta and named cases from Coinbase, Lindy, Harvey, and Cursor give it concrete signal. It is a strong cost-structure commentary, not a major model or product release, so it fits the 72-77 band.

Jun 6Saturday

AI HOT (Curated Pool)

OpenCV 5 Released with New DNN Engine and Native LLM Support

OpenCV 5 introduces a graph-based DNN engine, raising ONNX operator coverage from under 23% in 4.x to over 80%, with native support for Transformer, VLM, and LLM workloads.

Why it matters: HKR-H/K/R all pass for a substantive OpenCV major release: graph DNN engine, ONNX coverage jump, and native Transformer/VLM/LLM support. Strong featured item, but below must-write model-lab release territory.

r/LocalLLaMA

Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding

Domino reports up to 5.8x throughput speedup on Qwen3 by decoupling causal modeling from autoregressive drafting in speculative decoding. The Reddit snippet links the arXiv paper, GitHub code, and Hugging Face models, but does not disclose hardware, baseline settings, dataset, or acceptance-rate details.

Why it matters: HKR-H/K/R all pass: 5.8x throughput is a concrete hook with open artifacts. Missing hardware, baseline config, and task set keep it in the good featured band, not same-day must-write.

Computing Life · Share · Yage

Google pays SpaceX $920M a month for GPUs, but compute is not the main story

Google pays SpaceX $920 million per month for GPU rentals, and the post says the contract includes 11% GPU utilization, a 90-day cancellation clause, and methane gas turbines used to bypass environmental approval.

Why it matters: HKR-H/K/R all pass: the deal size, utilization term, and energy workaround are concrete. I keep it below P1 because the provided item is a single-source summary with no contract file or cross-source confirmation.

AI HOT (Curated Pool)

Building a Multi-Agent Economy with Qwen2.5-3B: Engineering Report

A developer used Qwen2.5-3B to build a five-agent forest economy, and across 15 simulation rounds honey prices fell from 10 to 3, firewood rose from 4 to 7, and the Gini coefficient increased from 0.14 to 0.38.

Why it matters: HKR-H/K/R pass: the 3B multi-agent economy has a hook and concrete price/Gini results. It remains a single engineering experiment, not a product or framework launch, so it stays at the featured floor.

AI HOT (Curated Pool)

SpaceX and Google Reach New Cloud Computing Agreement

SpaceX disclosed a cloud services agreement with Google: Google will pay SpaceX $920 million per month for computing capacity tied to xAI data centers, while the post does not disclose contract duration, GPU scale, or delivery terms.

Why it matters: HKR-H/K/R all pass: the hook is a Google–SpaceX–xAI compute triangle, with $920M/month as the concrete fact. The single-post source and missing contract term, delivery scale, and filing details keep it at low P1.

r/LocalLLaMA

Running Qwen3.6-35B-A3B on a laptop RTX 4060 8GB

A Reddit user ran Qwen3.6-35B-A3B on an RTX 4060 8GB laptop and reported that --no-mmap raised generation from about 11 to 43 tok/s, while speculative decoding with a Qwen3.5-0.8B draft model improved throughput by 26%.

Why it matters: HKR-H/K/R all pass: the post has a clear laptop-35B hook, reproducible speed numbers, and local-LLM resonance. Reddit single-post sourcing keeps it below the 78+ good-quality band.

Hacker News front page

Google to Pay SpaceX $920M a Month for Compute Capacity at xAI Data Centers

The title says Google will pay SpaceX $920 million per month for compute capacity at xAI data centers; the RSS snippet does not disclose contract duration, GPU scale, or the capacity delivery mechanism.

Why it matters: HKR-H/K/R all pass: $920M/month is a hard compute-market number, and the Google-SpaceX-xAI structure is unusual. Missing duration and GPU details keep it below 90.