Skip to content

#部署/工程

3 today

May 27Wednesday

AI HOT (Curated Pool)

Shipping a Trillion Parameters With a Hub Bucket: Delta Weight Sync in TRL

Hugging Face merged TRL PR 5417 for delta weight sync, sending only changed weights as sparse safetensors via a Hugging Face Bucket; on Qwen3-0.6B, the per-step payload falls from 1.2GB to 20–35MB.

Why it matters: HKR-H/K/R all pass: TRL gets delta weight sync with a concrete sparse-safetensors mechanism and a 1.2GB to 20–35MB example. Scope is training infra, so it stays below must-write.

AI HOT (Curated Pool)

MiMo 2.5 Pro Gets Major Price Cut, Matching DeepSeek V4 Pro

Xiaomi permanently cut MiMo-V2.5 API prices by up to 99%, matched DeepSeek V4 Pro pricing, increased same-price token allowances by 5–8x, reset existing user quotas in full, and set the new pricing to take effect on May 26.

Why it matters: HKR-H/K/R all pass: the 99% cut creates a price-war hook, the post gives 5-8x token economics, and API cost pressure resonates. It remains a pricing update, not a model or capability release, so it stays below the 78+ band.

r/LocalLLaMA

PrismML Released Binary and Ternary Bonsai Image 4B

PrismML released Binary and Ternary Bonsai Image 4B, 1-bit and ternary text-to-image diffusion transformers around 3GB, compared with FLUX.2 Klein 4B at about 16GB, with browser-local WebGPU demo links and an Apache-2.0 license disclosed in the Reddit snippet.

Why it matters: HKR-H/K/R all pass: low-bit image DiT plus local browser inference is a strong hook, backed by 4B, ~3GB, WebGPU, and license details. Reddit sourcing and limited lab weight keep it in the low featured band.

May 26Tuesday

Bloomberg Technology

Qualcomm to Supply Chips to TikTok Owner ByteDance

Qualcomm will supply chips to ByteDance for artificial intelligence data centers, according to people familiar with the matter; the post does not disclose chip models, order volume, pricing, or delivery timing.

Why it matters: Bloomberg sourcing ties Qualcomm, ByteDance, and AI data-center supply, so HKR-H/K/R pass. Missing chip model, volume, and delivery timing keep it in low featured, not P1.

r/LocalLLaMA

[OSS] dlmserve: First Serving Engine for Diffusion Language Models

dlmserve released an MIT-licensed serving engine for diffusion language models, with LLaDA-8B-Instruct support and 2.5x HF throughput at batch=4. It exposes an OpenAI-compatible /v1/chat/completions API, batches at the denoising-step level, runs in 12GB VRAM, and adds about 1.8x throughput with optional LocalLeap acceleration.

Why it matters: HKR-H/K/R all pass: an open-source DLM serving engine with concrete throughput and VRAM claims. Single Reddit source and an early ecosystem keep it in low featured, not 78+.

AI HOT (Curated Pool)

OpenRouter Raises $113M Series B

OpenRouter raised a $113 million Series B led by CapitalG; its weekly volume rose from 5 trillion to 25 trillion tokens over the past 6 months.

Why it matters: HKR-H/K/R all pass: OpenRouter is a common model-routing layer, and $113M plus 25T weekly tokens gives hard scale. Kept below 85 because this is funding plus growth data, not a new model or capability release.

Alibaba Technology · WeChat

Nearly 9x training speedup: residual streams in DiT are becoming a convergence bottleneck

Nanjing University LAMDA and Alibaba Intelligent Engine proposed DAR, a timestep-aware cross-layer routing method that replaces fixed residual accumulation in DiT; on ImageNet 256x256, it reduced SiT-XL/2 FID from 9.67 to 7.56 and reached baseline convergence quality with 8.75x fewer training iterations.

Why it matters: HKR-H/K/R all pass, but the topic is a narrow DiT training method rather than a broad model or product launch. Concrete ImageNet metrics and the Alibaba/LAMDA mechanism clear the featured bar, not the 78+ band.

QbitAI · WeChat

Chinese AI-Written Pretraining Framework ForgeTrain Trains MiniCPM5-1B

ModelBest released ForgeTrain and MiniCPM5-1B, saying ForgeTrain was written by AI and trains 10% faster than NVIDIA Megatron under the same hardware conditions. MiniCPM5-1B is a 1B-parameter edge model with about 2GB FP16 weights and about 0.5GB INT4/Q4 weights.

Why it matters: HKR-H/K/R all pass: an AI-written trainer, a 10% same-hardware Megatron speed claim, and a 0.5GB 1B edge model are concrete hooks. Score stays at 80 because the first-ever claim and benchmark lack third-party reproduction.

Synced · WeChat

AI-written training framework trains 1B edge model MiniCPM5-1B

ModelBest open-sourced MiniCPM5-1B and ForgeTrain; the 1B edge model scores 17.9 on AA-Index, while the AI-written ForgeTrain framework matches Megatron’s training results and runs 10% faster on Nvidia H100 under the article’s reported setup.

Why it matters: HKR-H/K/R all pass: the AI-written training framework hook is strong, with concrete AA-Index and H100 speed claims. It is not a flagship model release, so it stays in the 78–84 band.

AI HOT (Curated Pool)

ModelBest open-sources MiniCPM5-1B, topping sub-2B models on AA-Index

ModelBest open-sourced MiniCPM5-1B, a 1B-parameter edge language model that beats all sub-2B models on AA-Index, uses a 0.5GB weight file after INT4 quantization, and runs on phones and browsers.

Why it matters: HKR-H/K/R all pass: MiniCPM5-1B has concrete params, quantized size, and edge runtime claims. It is still a small-model release, below flagship-model impact.

Xinzhiyuan · WeChat

OpenAI Nearly Collapsed? President Says He Resigned the Day Altman Was Ousted

Greg Brockman recounted OpenAI’s 72-hour crisis: on November 17, 2023, the board removed Sam Altman as CEO and took Brockman off the board, after which Brockman resigned the same day and said he initially put the chance of taking the company back at 10%.

Why it matters: HKR-H/K/R all pass via an insider crisis hook, a 10% recovery-odds detail, and OpenAI governance resonance. It is still a retrospective on a heavily covered 2023 event, so it stays in the 72–77 band.

r/LocalLLaMA

Shard - Getting to 10× KV Cache Compression

Shard reduces Llama-3.1-8B KV memory by about 10× at 8K context and 11× at 32K, with no measured drop on NIAH or LongBench, using PCA plus int4 quantization for K and Hadamard rotation plus vector quantization for V.

Why it matters: HKR-H/K/R all pass: the 10× KV-cache claim has a strong hook and concrete model/context/benchmark details. Reddit-only sourcing and limited validation keep it in the 78–84 band.

AI HOT (Curated Pool)

OpenAI GPT-5.6 Reportedly Set for Next Month With 1.5M-Token Context

Developers found an unannounced OpenAI GPT-5.6 entry in Codex backend logs under the codename iris-alpha, with a 1.5 million-token context window, about 43% higher than GPT-5.5’s 1.05 million-token limit.

Why it matters: HKR-H/K/R all pass: the Codex-log leak, 1.5M-token window, and 43% increase are concrete and practitioner-relevant. It stays below 85 because this is not an official GPT-5.6 launch.

AI HOT (Curated Pool)

Apple reportedly uses a custom 1.2T-parameter Google model for next-generation Siri

Apple is reportedly using a custom 1.2T-parameter Google model to run parts of the next-generation Siri, while simpler queries are expected to run on-device; the post says response speed for everyday questions is the key constraint.

Why it matters: HKR-H/K/R all pass, but this is a single X-sourced reported claim; the post gives architecture details but not sourcing documents, rollout timing, or scope. Keep it at the featured threshold, below the 78+ band.

May 25Monday

r/LocalLLaMA

Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

RTPurbo converts full-attention LLMs to sparse inference with a few hundred adaptation steps. It keeps the full KV cache only for retrieval heads, uses a 16-dimensional token indexer, and reports up to 9.36x prefill speedup at 1M context plus about 2.01x decode speedup on long-context and reasoning benchmarks.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, the post gives 1M-context speedup numbers, and inference cost resonates. Reddit-only sourcing and missing model/code details keep it in the 78–84 band.

QbitAI · WeChat

Reasonix for DeepSeek V4 reaches 99.82% cache hit rate and cuts costs to 20%

Reasonix uses an append-only loop for DeepSeek V4 and reports a 99.82% cache hit rate in long coding sessions, cutting an example 400M-token bill from $61 to $12.

Why it matters: HKR-H/K/R all pass, but this is a third-party cost tool around DeepSeek V4, not a model launch or platform update. Concrete mechanism and billing numbers put it in the 72–77 featured band.

Hacker News front page

Memory has grown to nearly two-thirds of AI chip component costs

Epoch AI says memory has grown to nearly two-thirds of AI chip component costs; the RSS body only lists the article URL, 68 points, and 71 comments, and the post does not disclose the methodology or sample scope.

Why it matters: HKR-H/K/R all pass: the cost-share claim is clickable, specific, and relevant to infra economics. Sparse body details keep it near the featured floor: method, sample, and timeline are not disclosed.

May 24Sunday

r/LocalLLaMA

BitCPM-CANN: Native 1.58-Bit Large Language Model Training on Ascend NPU

OpenBMB released BitCPM-CANN, a 1.58-bit QAT training stack on Ascend NPU with 0.5B, 1B, 3B, and 8B models trained from scratch, where the 1B to 8B variants retain 95.7%–97.2% of full-precision MiniCPM4 performance across 11 benchmarks.

Why it matters: HKR-H/K/R pass: low-bit native training on Ascend is novel, and the summary gives sizes plus retention rates. Reddit-only sourcing and no throughput or reproduction details keep it at the featured floor.

r/LocalLLaMA

It's OK to Quantize the KV Cache; Model Quant Matters More in Qwen3.6 27B KLD Tests

Reddit user hopbel tested Qwen3.6 27B with approximate KLD on wikitext-2 at 16k context, using Q5_K_M as the proxy baseline; Q5_K_S weights with q4_0 KV cache scored 0.016304, while Q4_K_XL with f16 KV cache scored 0.026067, so weight quant tier dominated KV-cache quant in this setup.

Why it matters: HKR-H/K/R all pass, backed by first-person test numbers. Source is a single Reddit post, the metric is approximated KLD, and the claim is narrow, so it sits at the featured threshold.

May 23Saturday

Synced · WeChat

FlashAR speeds up pretrained autoregressive image models by 22.9x using 0.05% data

Zhejiang University and the University of Adelaide introduced FlashAR, using 0.05% of the original training data to reduce Emu3.5-Image-34B 512×512 generation latency from 130.10 seconds to 5.68 seconds, while GenEval changed from 80.48 to 80.29.

Why it matters: HKR-H/K/R all pass: FlashAR gives speedup, data ratio, latency, and GenEval deltas for AR image inference. It is a strong research item, but not a top-lab model release, so 80 featured rather than P1.

Synced · WeChat

Bengio Paper Raises Recursive Reasoning Limits as Parallel Trajectories Beat Serial Reasoning

Yoshua Bengio’s team introduced GRAM, a generative recursive reasoning model that samples multiple latent trajectories; on Sudoku-Extreme, GRAM reached 97.0% accuracy with 16 recursive steps and 20 parallel samples, exceeding TRM’s 90.5% result at 320 serial recursive steps.

Why it matters: HKR-H/K/R all pass: the hook is parallel recursion beating long serial recursion, with concrete GRAM numbers. Importance stays in 78–84 because the evidence is benchmark-centered, not a major model or product release.

Mistral AI

Mistral to acquire physics AI company Emmi AI

Mistral AI said it has reached a definitive agreement to acquire Emmi AI, a physics AI pioneer, to strengthen its position as an AI transformation partner for industrial companies. Austria-based Emmi AI works on physics AI and large engineering models that speed up engineering workflows, replace multi-day computations with real-time simulation and build digital twins. Emmi's co-founders and more than 30 researchers and engineers will join Mistral's Science and Applied AI teams in May.

Why it matters: Mistral is buying physics AI company Emmi to add industrial simulation, showing how it extends into engineering and manufacturing.

QbitAI · WeChat

DeepSeek V4 cuts prices as CATL, JD.com and NetEase discuss investment; Liang Wenfeng targets AGI

DeepSeek-V4-Pro API will keep its promotional pricing from June 1, with cached input at RMB 0.025 per million tokens, while Bloomberg says DeepSeek is pursuing a RMB 70 billion round at a USD 45 billion pre-money valuation.

Why it matters: HKR-H/K/R all pass: DeepSeek V4 API price cuts and Bloomberg’s RMB 70B raise at a $45B pre-money valuation are same-day material. The cost and capital angles directly affect China model competition.

r/LocalLLaMA

club-rdna16: Practical 16GB AMD/Radeon local LLM testing repo

club-rdna16 publishes a practical 16GB Radeon local LLM testing repo, with an RX 6900 XT running llama.cpp on ROCm/HIP and Qwen3.6 35B-A3B reaching a stable 131k context using q8 KV cache.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post and the body only discloses test conditions, not speed, VRAM curves, or reproducible logs. It clears the featured floor as a practical local-LLM repo.

AI HOT (Curated Pool)

Jensen Huang Says Annual AI Infrastructure Spending Will Reach $4 Trillion

Jensen Huang predicted hyperscale cloud providers’ annual AI infrastructure spending will rise from $1 trillion to $3 trillion–$4 trillion, while Nvidia reported $81.6 billion in fiscal 2027 Q1 revenue and $75.2 billion from data centers.

Why it matters: HKR-H/K/R all pass: Jensen Huang’s $3-4T annual AI infrastructure forecast is specific and tied to NVIDIA revenue. It is strong industry signal, but a CEO forecast rather than a model or product launch, so it stays in the 78-84 band.

AI HOT (Curated Pool)

Agent Workloads Quietly Reshape Inference Economics

SemiAnalysis analyzed 432,000 real coding-agent requests and found a median input length of 96,000 tokens, not 32,000 or 64,000. The post does not disclose the model mix, cost curve, sampling method, or time window.

Why it matters: HKR-H/K/R all pass: SemiAnalysis adds a 432k coding-agent request dataset and 96k-token median input. Missing models, cost curves, and sampling keep it in the strong-data-point band, not must-write.

r/LocalLLaMA

Experts first llama.cpp

comanderxv published a llama.cpp fork that caches MoE experts in 12GB VRAM; on an RTX 2060 with Qwen3.6-35B-A3B, throughput rose from 19/22 tk/s to 26 tk/s at about a 62% expert-cache hit rate.

Why it matters: HKR-H/K/R all pass: the hook is a 35B MoE speedup on a 12GB RTX 2060, with concrete caching and hit-rate data. Scope stays niche to local inference, so it lands at the featured threshold rather than must-write.

May 22Friday

Hacker News front page

DeepSeek Makes the V4 Pro Price Discount Permanent

DeepSeek will set deepseek-v4-pro API pricing to one quarter of the original price after the 75% promotion ends on 2026-05-31 at 15:59 UTC; the post does not disclose the exact per-token price.

Why it matters: HKR-H/K/R all pass: the hook is a permanent DeepSeek price cut, the new fact is 1/4 pricing after a stated UTC time, and the nerve is API cost. Missing unit pricing keeps it at the featured floor, not a must-write release.

Dwarkesh Patel podcast

Reiner Pope – Chip Design from the Bottom Up

Dwarkesh Patel interviews MatX CEO Reiner Pope on chip design, starting with a 4-bit multiply and 8-bit accumulate example that uses 16 AND gates, then covering systolic arrays, pipeline registers, FPGAs versus ASICs, cache versus scratchpad, and why GPU cores are smaller than CPU cores.

Why it matters: Dwarkesh’s MatX CEO interview clears HKR-H/K/R with a bottom-up hardware hook, concrete mechanisms, and compute-cost resonance. It is educational rather than breaking news, so it sits in the 72–77 band.

Mistral AI

Mistral launches Connectors in Studio with built-in and custom MCP

Mistral launched Connectors in Studio. All built-in connectors and custom MCP are now callable through the API/SDK by every model and agent. New features include direct tool calling, human-in-the-loop approval flows, and programmatic access to create, modify, list and delete connectors.

Why it matters: The original gives the API usage and code examples for Connectors, enough to judge how enterprise MCP integration gets built.

AI HOT (Curated Pool)

BitCPM-CANN Released as First 1.58-bit Open Model Fully Trained on Huawei Ascend 910B NPU

ModelBest, Tsinghua University, and OpenBMB released BitCPM-CANN, a 0.5B-8B open model family trained natively on Huawei Ascend 910B NPUs with 1.58-bit ternary weights, cutting memory use by about 6x versus BF16 while retaining 95-97% of full-precision benchmark performance.

Why it matters: HKR-H/K/R all pass: the Ascend 910B plus 1.58-bit open model angle is novel and metric-rich. It stays below P1 because the post offers release facts, not independent replication or adoption signal.

Latent Space

[AINews] New AI Infra Unicorns: Exa, Modal, TurboPuffer

Latent Space summarized AI News for May 20-21, 2026, confirming TurboPuffer reached $100 million ARR and profitability, Exa raised a $250 million Series C at a $2.2 billion valuation, and Modal raised a $355 million Series C at a $4.7 billion valuation.

Why it matters: HKR-H/K/R all pass because the roundup gives concrete AI-infra funding and ARR numbers. It stays below 78 because it is market aggregation, not a new model, product capability, or technical release.

AI HOT (Curated Pool)

Zhipu releases GLM-5.1-highspeed, claiming a large-model API speed record

Zhipu released the GLM-5.1-highspeed API to selected enterprise customers on May 22, with a claimed output speed of 400 tokens/s, built by the GLM team and TileRT team through system-level optimization.

Why it matters: HKR-H/K/R all pass: Zhipu’s GLM-5.1 high-speed API has a concrete 400 tokens/s claim and domestic flagship-model relevance. Test setup, pricing, and availability are not disclosed, so it stays in the 78–84 band.

New York Times Chinese

Trump Approved Nvidia Chip Sales to China. Why Is Beijing Reluctant?

Trump approved Nvidia H200 sales to China six months ago, but Beijing has not allowed any company to buy even one chip and is steering firms toward domestic alternatives from Huawei and Cambricon.

Why it matters: HKR-H/K/R all pass: six months of zero H200 purchases after approval, plus Beijing steering firms toward Huawei and Cambricon. This is strong chip-policy signal, but not a model launch or major product release, so it sits in 78–84.

Computing Life · Share · Yage

How to Run DeepSeek V4 Flash Locally on Mac: DS4 Engine Explained

DS4 provides a macOS local runtime path for DeepSeek V4 Flash; the post only discloses three mechanisms—multi-agent integration, KV cache disk persistence, and activation steering—and does not disclose performance numbers, hardware requirements, or pricing.

Why it matters: HKR-H/K/R all pass, but the body only names DS4 mechanisms and omits performance, model size, Mac support, and reproducible tests; this fits the featured threshold for a local-inference tutorial.

Computing Life · Share · Yage

The Technology Behind GLM-5.1 Reaching 400 Tokens/s

Zhipu GLM-5.1 high-speed API claims 400 tokens/s, and the post says TileRT reconstructs GPU inference at the execution-model level; the RSS snippet does not disclose benchmark conditions, hardware, pricing, or latency distribution.

Why it matters: HKR-H/K/R all pass: 400 tokens/s is a concrete hook, TileRT adds mechanism, and latency/cost resonates with builders. It stays at 78 because the speed is claimed, with no independent test or pricing condition disclosed.

r/LocalLLaMA

Interesting Paper Advocates Quantized Prefilling and Precise Decoding

arXiv 2605.20315 argues for W4A4 quantization during prefilling to target a theoretical 4x gain, while keeping decoding on the original high-precision path because activation errors can perturb sampled tokens and accumulate across autoregressive generation.

Why it matters: HKR-H/K/R all pass, but the item only gives the paper claim and theoretical gain; measured throughput, perplexity, and hardware setup are not disclosed, so it stays at the featured threshold.

NVIDIA Blog

NVIDIA GTC Taipei at COMPUTEX: Live Updates on What’s Next in AI

NVIDIA won four COMPUTEX 2026 Best Choice Awards for Vera Rubin NVL72, Jetson Thor, and Alpamayo; Vera Rubin NVL72 connects 36 Vera CPUs and 72 Rubin GPUs, and NVIDIA says it delivers up to 10x higher inference performance per watt and 10x lower cost per token.

Why it matters: HKR-H/K/R all pass: NVIDIA gives concrete Vera Rubin NVL72 specs and a 10x inference-efficiency claim, directly tied to AI compute costs. The source is NVIDIA’s event blog, so this stays below the 85 same-day must-write band.

May 21Thursday

r/LocalLLaMA

Agent Execution Tax: New Procurement Metric for Browser Agent Benchmarks?

Fireworks ran 720 browser-agent tasks on WebVoyager and reported a 22.9% Agent Execution Tax, defined as wasted over productive inference; MiniMax M2.5 cost 2.3x less per successful task than Gemini, while GLM-5 reached 57.1% accuracy and Kimi K2.5 had 0% parse retries across 852 calls.

Why it matters: HKR-H/K/R all pass: the post adds a named procurement metric plus concrete benchmark numbers. Source scope is Reddit/Fireworks, so it stays in the 72–77 featured band rather than 78+.

r/LocalLLaMA

LLM planner: pick a rig by use case, model, or budget, or pick models for your rig

totosse17 published the LLMRequirements hardware planner with 60+ build configs, 50+ models, 130 cited tokens-per-second sources, 150+ reviewer videos, multi-region prices, idle and active watts, and a public GitHub data repo.

Why it matters: HKR-H/K/R all pass, but this is a Reddit community tool for local LLM rigs, not a broad platform release. The concrete dataset earns a featured-threshold score, not the 78+ band.