Skip to content

Deployment & engineering

Running models in practice: inference optimization, memory and cost, serving architecture and infrastructure choices.

Latest picks

121–140 of 455

May 28Thursday

QbitAI · WeChat

Behind DeepSeek V4's Chip-Model Co-Design, China's Compute Ecosystem Gains Speed

QbitAI says DeepSeek V4 validated Ascend chip-model co-design, with CANN open-sourcing 65 repositories and supporting day-zero adaptation for more than 70 mainstream models, while AIGCode reported 65% MFU in MoE pretraining on Ascend.

Why it matters: HKR-H/K/R all pass, but this is mainly a compute-ecosystem progress story, not a DeepSeek V4 capability release. Concrete repo, adaptation, and MFU numbers lift it into featured, below must-write.

r/LocalLLaMA

Zai replaced the network architecture for GLM-5.1 inference, lifting throughput 15%

Zai replaced the ROFT network topology with ZCube on a thousand-GPU GLM-5.1 coding inference cluster, keeping the same GPUs, software stack, and model; the Reddit post cites 33% lower switch and optical module costs, 15% higher GPU inference throughput, and a 40.6% drop in first-token P99 tail latency under prefill-decode disaggregated inference.

Why it matters: HKR-H/K/R all pass: the GLM-5.1 inference cluster has concrete cost, throughput, and P99 latency numbers. Reddit single-source sourcing and infra-niche scope keep it at 78.

r/LocalLLaMA

Qwen3.6-35B-A3B-APEX Runs 128K Context on RTX 3060 12GB

old-mike ran mudler/Qwen3.6-35B-A3B-APEX-MTP-I-Compact.gguf through spiritbuun’s llama.cpp fork on one RTX 3060 12GB, offloading a 17.3GB model and reaching 37.17 t/s generation at 72K filled context, 28.08 t/s at 129K, and PPL 3.2529 on an enwik8 64K-context perplexity test.

Why it matters: HKR-H/K/R all pass via a concrete consumer-GPU inference result with speed and PPL. Source is a single Reddit post and the impact stays within local inference, so it lands in featured, not P1.

AI HOT (Curated Pool)

AI Now Summit 2026

Mistral AI announced industrial AI work, a Vibe upgrade, and a 10 MW inference data center in Les Ulis at AI Now Summit 2026; it is working with Airbus, BMW Group, and ASML, and the data center is scheduled to start operating in Q3 2026.

Why it matters: HKR-H/K/R pass: Mistral gives a concrete 10 MW inference site, Q3 2026 timing, and major industrial partners. No new model capability or pricing is disclosed, so it stays just above the featured threshold.

AI HOT (Curated Pool)

Mistral AI launches physics AI model for industrial engineering

Mistral AI integrated the Emmi AI team and launched a physics AI foundation model for industrial engineering, with the post saying it can learn from geometry, boundary conditions, or measurement data and predict full physical fields on a single GPU in seconds.

Why it matters: HKR-H/K/R pass: a major model lab entering physics simulation with a concrete single-GPU seconds claim. The score stays in the lower featured band because model name, benchmarks, pricing, and access are not disclosed.

Latent Space

Cognition Raises $1B in $26B Series D

Cognition raised a $1B Series D at a $26B valuation and projects more than $1B ARR by year-end; the post says its valuation rose 2.5× from the $10B Series C eight months earlier, while the rest of the issue summarizes agent, inference, benchmark, and multimodal AI updates from May 26–27, 2026.

Why it matters: HKR-H/K/R all pass: Cognition’s $1B Series D at a $26B valuation is large, and projected year-end ARR above $1B gives a concrete business signal. This is not a model launch, but it is must-write funding news for AI coding agents.

Xinzhiyuan · WeChat

Tsinghua Team Open-Sources PilotDeck Agent System, Claims 70% Token Cost Reduction

Tsinghua THUNLP, ModelBest, OpenBMB, and AI9stars open-sourced PilotDeck; the article says its sub-agent routing reduced cost from $12.58 to $2.83 in a Xiaohongshu content-generation test, while preserving separate WorkSpaces, editable memory, and per-session routing logs.

Why it matters: HKR-H/K/R all pass: PilotDeck has a clear agent-cost hook, a concrete routing mechanism, and $12.58 to $2.83 data. It stays in the 78–84 band because this is a tool release, not a major model or platform launch.

Synced · WeChat

Mila and DeepMind Propose UNSL for Unified Multivariate Neural Scaling Laws

Mila and Google DeepMind proposed Unified Neural Scaling Law, a multivariate scaling-law form that models parameter count, token count, training steps, bottlenecks, overfitting, and adverse hyperparameter effects; UNSL achieved the best extrapolation on 60.87% of vision tasks and 88.89% of language tasks in the reported experiments.

Why it matters: HKR-H/K/R all pass: UNSL unifies parameters, tokens, steps, bottlenecks, overfitting, and hyperparameter feedback, with vision/language extrapolation numbers. Technical density keeps it in the 78–84 band.

TechCrunch · AI

Snowflake signs $6B deal with AWS for AI CPU chips

Snowflake signed a five-year, $6 billion AWS deal to secure chips for AI use; the post does not disclose chip models, delivery timing, or whether the capacity targets training, inference, or both.

Why it matters: HKR-H/K/R all pass: the 5-year, $6B AWS deal is a concrete AI-infra signal. Missing chip model, delivery cadence, and training/inference split keep it at the featured threshold.

r/LocalLLaMA

Inferencing at 10.33 t/s on Qwen 3.5 35B on a $300 laptop

A Reddit user ran Qwen 3.5 35B Q4_K_S on a $300 Lenovo Ideapad Slim 3i and reported 10.33 t/s inference using ik_llama.cpp with two pinned CPU cores, MTP speculative decoding, 64 batch size, and Q8_0 KV cache.

Why it matters: HKR-H/K/R all pass, with a concrete first-person benchmark. Reddit single-post sourcing and limited reproducibility details keep it at the lower featured threshold.

AI HOT (Curated Pool)

Open-source FastVideo Dreamverse real-time video generation tool

Hao AI Lab open-sourced FastVideo Dreamverse, a real-time video generation tool that generates a 30-second 1080p video in 7 seconds under the stated setup of one NVIDIA B200 GPU and LTX-2.

Why it matters: HKR-H/K/R all pass: the 7s-for-30s-1080p claim is concrete and practitioner-relevant. Single-source X sourcing and missing independent benchmarks keep it in the 78–84 band.

May 27Wednesday

AI HOT (Curated Pool)

Perplexity open-sources Unigram tokenizer to reduce CPU usage

Perplexity open-sourced a rebuilt Unigram tokenizer that reduces CPU usage by 5-6x, targeting tokenization latency when small rerankers and embedding models run on GPUs in single-digit milliseconds.

Why it matters: HKR-H/K/R all pass: the 5-6x CPU claim and tokenizer bottleneck are concrete for production RAG/search teams. It stays in the featured-threshold band because the post lacks independent benchmarks, repo details, and deployment scale.

QbitAI · WeChat

Language Models Need Sleep: Let AI Nap Before Continuing Inference

Carnegie Mellon University and the University of Maryland propose a “sleep” mechanism for language models: when the context window is nearly full, the model stops accepting new tokens, runs multiple offline recursive forward passes to compress accumulated context into fast weights, clears the KV cache, and then resumes inference; tests cover cellular automata, multi-hop graph retrieval, and GSM-Infinite reasoning tasks.

Why it matters: HKR-H/K/R all pass: the sleep metaphor is clickable, and the mechanism is concrete. Score stays below 78 because the provided body lacks benchmark gains, code, or deployment evidence.

Xinzhiyuan · WeChat

OpenRouter processes 100 trillion tokens monthly and raises $113M Series B

OpenRouter raised a $113 million Series B led by CapitalG, lifting its valuation to $1.3 billion; the platform processes 25 trillion tokens per week, about 100 trillion per month, and provides one API for more than 400 models.

Why it matters: HKR-H comes from the 100T-token/month hook; HKR-K has funding, valuation, usage, and model-count numbers; HKR-R maps to routing and API-cost competition. Still, this is infra funding news, not an 85+ must-write release.

Synced · WeChat

AMD paper: FP4 training instability is not caused by insufficient randomness

AMD and Penn State pretrained Llama 3.1-8B with MXFP4 on MI355X native FP4 hardware, achieving 9-10% end-to-end speedup over an FP8 baseline, while the paper identifies Wgrad quantization as the bottleneck that raises token overhead to 26-27% without deterministic Hadamard stabilization.

Why it matters: HKR-H/K/R all pass: a counterintuitive FP4 claim, concrete Llama 3.1-8B numbers, and a cost/hardware nerve. The topic is narrower training-infra research, so it stays in the 78-84 band.

Latent Space

[AINews] New AI Infra Decacorns: Fireworks, Baseten, with OpenRouter on the Way

Latent Space says Fireworks is in talks for a $15 billion valuation round, Baseten is raising at an $11 billion valuation, and OpenRouter closed a $113 million Series C after volume grew 5x in six months.

Why it matters: HKR-H/K/R all pass: the decacorn hook is clickable, the post gives valuation, round, and usage figures, and the topic speaks to inference economics. Fireworks and Baseten are still reported as in talks or raising, so this stays in the 78–84 band.

AI HOT (Curated Pool)

Qualcomm and ByteDance reportedly partner on AI ASIC chips with millions of units planned

The title says Qualcomm and ByteDance reached an AI ASIC chip partnership with procurement in the millions of units; the post does not disclose chip specifications, unit pricing, delivery timing, or production conditions.

Why it matters: HKR-H/K/R all pass: the rumored Qualcomm–ByteDance AI ASIC deal has a concrete million-unit volume hook. Thin sourcing and missing specs, pricing, delivery, and production terms keep it in the 72–77 band.

Bloomberg Technology

Fireworks AI in Talks for Funding at $15 Billion Valuation

Fireworks AI is in talks to raise a new funding round at a $15 billion valuation, according to people familiar with the matter; the post does not disclose the round size, investor names, or timeline.

Why it matters: HKR-H/K/R all pass on the $15B AI-inference valuation, backed by Bloomberg. The deal is still in talks, with round size, investors, and timeline undisclosed, so it stays in low featured.

AI HOT (Curated Pool)

Shipping a Trillion Parameters With a Hub Bucket: Delta Weight Sync in TRL

Hugging Face merged TRL PR 5417 for delta weight sync, sending only changed weights as sparse safetensors via a Hugging Face Bucket; on Qwen3-0.6B, the per-step payload falls from 1.2GB to 20–35MB.

Why it matters: HKR-H/K/R all pass: TRL gets delta weight sync with a concrete sparse-safetensors mechanism and a 1.2GB to 20–35MB example. Scope is training infra, so it stays below must-write.

AI HOT (Curated Pool)

MiMo 2.5 Pro Gets Major Price Cut, Matching DeepSeek V4 Pro

Xiaomi permanently cut MiMo-V2.5 API prices by up to 99%, matched DeepSeek V4 Pro pricing, increased same-price token allowances by 5–8x, reset existing user quotas in full, and set the new pricing to take effect on May 26.

Why it matters: HKR-H/K/R all pass: the 99% cut creates a price-war hook, the post gives 5-8x token economics, and API cost pressure resonates. It remains a pricing update, not a model or capability release, so it stays below the 78+ band.