Skip to content

#部署/工程

3 today

May 13Wednesday

r/LocalLLaMA

The Trillion-Parameter Dilemma: MiMo-V2.5-Pro Open-Sourced at 1.02T Parameters

Xiaomi open-sourced MiMo-V2.5-Pro with 1.02T parameters, 42B active parameters, a 1M context window, and an MIT license; the author ran 125 Claude Code sessions through the API, spending $70.12 for 387,380,436 tokens with a 96.3% cache hit rate.

Why it matters: HKR-H/K/R all pass: a Xiaomi 1.02T open model plus a concrete Claude Code API cost experiment. Reddit sourcing keeps it at the low end of the 85+ band, but the domestic flagship-model signal clears p1.

New York Times Chinese

Jensen Huang Gets Last-Minute Invitation to Join Trump’s China Trip

Trump called Jensen Huang on Tuesday morning to invite him to join the China trip; the White House’s Monday list of 16 CEOs did not include him, while Nvidia is still seeking approval to sell AI chips to China.

Why it matters: HKR-H/K/R all pass: NYT reports a last-minute Jensen Huang invite tied to Nvidia’s China AI-chip license push. No disclosed policy change or license outcome, so this stays near the featured threshold.

New York Times Chinese

China Seeks AI Technology Self-Reliance, Weakening Washington’s Leverage Over Beijing

DeepSeek optimized its latest model for inference on Huawei chips for the first time, while two semiconductor sources said training still relies on Nvidia chips; Huawei says it plans to release a training chip this year, but matching current Nvidia performance will take another year.

Why it matters: HKR-H/K/R all pass: NYT ties DeepSeek-Huawei chip optimization and Huawei's training-chip timeline to US export-control leverage. It is not a model launch and lacks benchmark results, so it stays in the 78–84 band.

r/LocalLLaMA

A real transformer language model running locally on a stock Game Boy Color

maddiedreese ran Andrej Karpathy’s TinyStories-260K on a stock Game Boy Color with INT8 weights, fixed-point math, an MBC5 ROM, bank-switched cartridge storage, and KV cache in cartridge SRAM; the demo uses no phone, PC, Wi‑Fi, link cable, or cloud inference, but output is extremely slow and gibberish.

Why it matters: HKR-H/K/R all pass: a named first-person experiment with concrete model and memory details. Impact stays low-featured because it is a Reddit hardware hack with slow, garbled output, not a usable product or model release.

Bloomberg Technology

CME Plans Computing Power Futures Market

CME Group and Silicon Data plan to create a futures market for computing power; the RSS snippet says Bloomberg Tech discusses the rationale and mechanics, but the post does not disclose contract specifications, launch timing, or pricing methodology.

Why it matters: HKR-H and HKR-R pass: CME moving into compute futures is a fresh hook and targets AI compute-cost anxiety. HKR-K is weak because contract specs, timeline, and pricing method are not disclosed.

AI HOT (Curated Pool)

Claude Opus 4.7 Fast Mode Opens Research Preview

Claude Opus 4.7 Fast Mode is now available as a research preview in the API and Claude Code. The post does not disclose model parameters, pricing, rate limits, or a general availability date.

Why it matters: HKR-H/K/R pass because this is a Claude fast-mode preview in API and Claude Code, directly tied to developer latency and workflows. Thin disclosure on pricing, limits, parameters, and GA timing keeps it at the featured threshold, not 78+.

Hacker News front page

Show HN: Needle Distills Gemini Tool Calling into a 26M Model

Cactus open-sourced Needle, a 26M-parameter tool-calling model that reaches 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, with MIT-licensed weights released on Hugging Face.

Why it matters: HKR-H/K/R all pass: the tiny Gemini-style tool-calling angle is clickable, with concrete speed and license claims. Source is still Show HN/GitHub self-reporting, not an independent benchmark or major lab release, so it stays below the 78–84 band.

r/LocalLLaMA

Needle: We Distilled Gemini Tool Calling Into a 26M Model

Cactus Compute open-sourced Needle, a 26M-parameter tool-calling model that reaches 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, using an attention-and-gating architecture with no MLPs.

Why it matters: HKR-H/K/R all pass: a 26M tool-calling model has a strong hook and concrete speed/design claims. Single Reddit source and a less-known team keep it in the lower 78–84 band.

TechCrunch · AI

Report: Google and SpaceX in Talks to Put Data Centers Into Orbit

Google and SpaceX are discussing orbital data centers for AI compute, while the post says current space costs remain higher than ground infrastructure; the RSS snippet does not disclose cost gaps, deployment scale, or a timeline.

Why it matters: HKR-H and HKR-R are strong: orbital data centers are a sharp compute-infrastructure hook. HKR-K is weak because cost, scale, and timeline are missing, so this stays in the low featured band.

Financial Times · Technology

CME plans to launch futures market for AI computing power

CME plans to launch futures contracts tied to GPU rental prices, allowing traders and companies to bet on or hedge future costs; the RSS snippet does not disclose contract specifications, launch timing, or the reference index.

Why it matters: FT reports CME plans GPU rental-price futures, clearing HKR-H/K/R through novelty, mechanism, and compute-cost resonance. Missing contract specs, launch timing, and index details keep it at featured threshold, not P1.

May 12Tuesday

r/LocalLLaMA

Local LLM Autocomplete and Agentic Coding on a Single 16GB GPU + 64GB RAM

Reddit user grumd runs Qwen2.5-Coder-7B Q6 for autocomplete and Qwen3.6-35B-A3B Q8 for agentic coding on one RTX 5080 with RAM offloading; the post reports about 145k context, 56GB RAM used with other apps open, and Qwen3.6-35B-A3B speed of tg128 at 35.29 tokens/s.

Why it matters: HKR-H/K/R all pass: a named first-person local coding experiment with concrete model, quantization, context, and throughput data. Source is a single Reddit post without replication or comparisons, so it stays in the low featured band.

Synced · WeChat

ByteDance Open-Sources DreamLite for Offline Mobile Image Generation and Editing

ByteDance open-sourced DreamLite, a 0.39B-parameter unified diffusion model that generates or edits a 1024×1024 image on an iPhone 17 Pro in about 3 seconds, using 4-step DMD2 distillation and on-device offline inference without cloud dependency.

Why it matters: HKR-H/K/R all pass: 3-second on-device 1024×1024 generation is a strong hook, with 0.39B params and 4-step DMD2 as concrete claims. As a ByteDance open-source vision model, it sits below a general foundation-model release.

AI HOT (Curated Pool)

What Parameter Golf Taught Us About AI-Assisted Research

OpenAI’s Parameter Golf brought together over 1,000 participants and more than 2,000 submissions to test AI-assisted machine learning research, coding agents, model quantization, and model design under strict parameter constraints.

Why it matters: OpenAI’s Parameter Golf recap clears HKR-H/K/R with a concrete contest, 1,000+ participants, and 2,000+ submissions. It is research/benchmark signal, not a model or product launch, so 78 fits the lower featured band.

r/LocalLLaMA

Prompt caching for RL training: 7.5x speedup on long-prompt, short-response workloads

The author proposes prompt caching for RL training. On Qwen3.5-4B, it reports a 7.5x speedup with 16k-token prompts and 64-token outputs, and the G=8 example with 1000-token prompts and 100-token responses reduces 8800 processed tokens to 1800 unique tokens.

Why it matters: HKR-H/K/R all pass: the angle is novel, and the post gives 16k/64 plus G=8 token-dedup numbers. Kept at 78 because this is a single Reddit post without independent replication or a paper/code artifact disclosed.

r/LocalLLaMA

Computer Build Using Intel Optane Persistent Memory Runs a 1T-Parameter Model at Over 4 Tokens/s

Reddit user APFrisco ran the 1T-parameter Kimi K2.5 Q2_K_XL quant locally at about 4 tokens/s using 768GB Intel Optane PMem, 192GB DDR4 ECC DRAM, and a 12GB RTX 3060 with llama.cpp hybrid GPU/CPU inference.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, the post gives concrete hardware and speed numbers, and it hits local-inference cost concerns. Single Reddit anecdote and limited replication detail keep it at the featured floor.

May 11Monday

r/LocalLLaMA

ExLlamaV3 Major Updates

ExLlamaV3 added DFlash in v0.0.31, raising Coding throughput from 59.21 t/s to 177.67 t/s; v0.0.32 optimized five models, with Trinity-Nano gaining 72.4% on 6000 Pro², while v0.0.33 adds DFlash model quantization plus bug fixes and efficiency work.

Why it matters: HKR-H/K/R all pass, but the blast radius is mostly LocalLLaMA and ExLlama users. This fits a mid-weight open-source inference update, not a same-day industry-wide story.

Synced · WeChat

ICML 2026: PRISM Brings Efficient Test-Time Scaling to dLLMs

PRISM raises LLaDA-8B-Instruct on GSM8K from 67.58% to 85.30% by combining hierarchical trajectory search, partial remasking, and self-verified feedback, reducing dLLM test-time scaling cost from O(NT) toward O(N+KT) under a final candidate width K.

Why it matters: HKR-H/K/R all pass: the hook rejects brute-force scaling, the post gives GSM8K and complexity numbers, and it speaks to inference cost. Still an ICML framework paper, not a mainstream product release, so it sits in 78–84.

QbitAI · WeChat

OpenAI backs Cerebras as the Nvidia challenger targets a $35B IPO valuation

Cerebras raised its IPO price range to $150-$160 per share, targeting about a $35 billion valuation at the top end, after OpenAI signed a 750-megawatt AI compute purchase agreement with deliveries through 2028.

Why it matters: HKR-H/K/R all pass: this is not a routine IPO note, since OpenAI’s 750MW purchase agreement anchors Cerebras at a reported $35B valuation and feeds the NVIDIA-alternative compute story.

QbitAI · WeChat

SpaceXAI Takes Shape as Elon Musk Files Trademark Applications

SpaceX filed two SpaceXAI trademark applications covering satellite-based data centers, orbital computing, AI SaaS, cloud storage, telecom hardware, and social networking; the post says xAI became a SpaceX subsidiary through an all-stock deal and cites a $250 billion xAI valuation.

Why it matters: HKR-H/K/R all pass, but the hard fact is trademark filings; the claimed xAI-SpaceX merger lacks disclosed deal terms or an official announcement. Featured, not 85+, because this is signal rather than confirmed restructuring.

AI HOT (Curated Pool)

Cerebras IPO reportedly over 20 times oversubscribed, with pricing set to rise nearly 30%

Cerebras received more than 20 times oversubscription for its IPO and plans to raise the share count from 28 million to 30 million while increasing the price range to $150-$160.

Why it matters: HKR-H/K/R all pass: the Cerebras IPO repricing has rare demand numbers and clear AI-infrastructure resonance. It stays in the lower 85-94 band because this is pricing news, not the actual listing or a new chip launch.

AI HOT (Curated Pool)

Local models handle half of daily tasks and respond faster than cloud models

A five-week experiment tested about 1,400 daily work tasks, where local 35B models such as Qwen 3.6 35B handled about 50% and averaged 2.8-second responses, 2.1 times faster than Claude Opus 4.5, while the cloud model still led complex reasoning by about 20%.

Why it matters: HKR-H/K/R all pass: Tom Tunguz’s experiment reports ~1,400 tasks, ~50% success, 2.8s latency, and a speed comparison to Claude Opus 4.5. Strong practitioner signal, but not a model launch or platform-level update.

r/LocalLLaMA

MTP benchmark results: task type determines speculative inference speedups or slowdowns

A Reddit LocalLLaMA user ran 300+ tests on Qwen 3.6 27B MTP quants, finding coding draft acceptance at 79-89% and F16 coding speed up 171%, while Q4_K_M creative writing slowed down 9%.

Why it matters: HKR-H/K/R all pass: this is a single Reddit experiment, not a market event, but 300+ Qwen 3.6 27B MTP quantization tests give practical numbers for local inference tuning.

May 10Sunday

r/LocalLLaMA

I have DeepSeek V4 Pro at home

Reddit user fairydreaming ran DeepSeek V4 Pro Q4_K_M with a modified llama.cpp CUDA repo on one RTX PRO 6000 Blackwell Max-Q workstation GPU, using an 859GB model file; the shared log reports a 1M context window and 8.6 tokens per second generation speed.

Why it matters: HKR-H/K/R all pass: the hook is single-GPU local inference, with concrete file size, context, speed, and runtime path. Reddit single-source sourcing keeps it below must-write model-release territory.

r/LocalLLaMA

NVIDIA AI Releases Star Elastic: One Checkpoint Contains 30B, 23B, and 12B Reasoning Models

NVIDIA AI released Star Elastic, a single checkpoint that can zero-shot slice 30B, 23B, and 12B reasoning models in BF16, FP8, and NVFP4; when the 23B submodel handles thinking and the 30B model handles final answers, reported accuracy rises 16% and latency drops 1.9× on AIME-2025, GPQA, LiveCodeBench v5, and MMLU-Pro.

Why it matters: HKR-H/K/R all pass: Star Elastic has a concrete mechanism and testable numbers for inference deployment. Its reach is still narrower than a frontier-model release, so it sits in the high-quality featured band.

r/LocalLLaMA

BeeLlama.cpp: DFlash and TurboQuant with reasoning and vision support

Anbeeld released BeeLlama.cpp, a llama.cpp fork that runs Qwen 3.6 27B Q5 with 200k context and vision on a single RTX 3090 or 4090; the title claims 2–3x faster than baseline and a 135 tps peak.

Why it matters: HKR-H/K/R all pass, but the claims come from a Reddit title and summary without independent reproduction. Treat as a mid-weight open-source inference update, so it lands in the low featured band.

May 9Saturday

AI HOT (Curated Pool)

Redis founder uses a C inference engine to run a large model on a personal computer

Antirez open-sourced ds4, a native inference engine for DeepSeek V4 Flash that uses a few thousand lines of C to run a 1M-context model on a 128GB MacBook Pro at a reported 27 tok/s.

Why it matters: HKR-H/K/R all pass: Antirez open-sourced a native C inference engine with hardware, model, context, and speed numbers. Single-source X provenance keeps it below P1, but it is strong open-source inference signal.

r/LocalLLaMA

80 tok/sec and 128K context on 12GB VRAM with Qwen3.6 35B A3B and llama.cpp MTP

Reddit user janvitos ran Qwen3.6-35B-A3B-MTP-GGUF with a llama.cpp MTP PR on an RTX 4070 Super. The posted benchmark shows 69.2-81.9 tok/s, 0.694-0.947 draft acceptance, 131072 context, and a -fitt 1536 setting that reserves 1536 MB for the draft model and KV cache.

Why it matters: HKR-H/K/R all pass with concrete single-user benchmark data and reproducible settings. Source is one Reddit post, so verification is thin; this lands above featured threshold, not in must-write range.

AI HOT (Curated Pool)

Baidu releases ERNIE 5.1 with compressed parameters and training cost

Baidu released ERNIE 5.1 with total parameters reduced to about one third of the original scale, active parameters to about one half, and pretraining cost to about 6% of same-scale models; the model is available on the ERNIE platform and Baidu AI Studio.

Why it matters: HKR-H/K/R all pass: Baidu ERNIE 5.1 is a domestic flagship-model release with concrete compression and 6% pretraining-cost claims. That puts it in the must-write band.

Xinzhiyuan · WeChat

NVIDIA, AMD, and Intel Back RadixArk’s $100M Seed Round

RadixArk announced a $100 million seed round on May 5 at a $400 million post-money valuation, led by Accel and co-led by Spark Capital, with participation from NVentures, AMD, MediaTek, Databricks, and other investors tied to AI infrastructure.

Why it matters: HKR-H/K/R all pass: a $100M seed round, $400M post-money valuation, and chip/data investors create an AI-infra rivalry angle. It remains a single-company funding item with no product benchmarks or customer data, so it sits near the featured floor.

AI HOT (Curated Pool)

DeepSeek Raises $7 Billion at Record Scale, Founder Personally Invests $3 Billion

DeepSeek is raising up to $7 billion at a $50 billion valuation, with founder Liang Wenfeng personally contributing $3 billion, or 40% of the round, while the company says the funding will target large-scale compute, V4.1 model releases, enterprise products, and a path toward positive revenue.

Why it matters: HKR-H/K/R all pass: the $3B founder check is a strong hook, with concrete funding numbers and clear competitive resonance. Single X-source sourcing leaves lead investor, terms, and confirmation undisclosed, so it stays below P1.

r/LocalLLaMA

MTP + TurboQuant Running: Qwen3.6-27B Hits 80+ t/s on a Single RTX 4090

indrasmirror ran Qwen3.6-27B-Heretic-v2 on a single RTX 4090 with 262K context, TBQ4_0 KV cache, and MTP draft 3, improving throughput from about 43 t/s to 80-87 t/s with roughly 73% MTP draft acceptance.

Why it matters: HKR-H/K/R all pass, backed by a numbered first-person experiment. The Reddit-only source and niche local-inference focus keep it below the 78–84 band for broader industry releases.

The Verge · AI

All the Latest Updates on AI Data Centers

The Verge tracks AI data center disputes with specific updates: 43% of Americans blame data centers for rising power bills, a 40,000-acre Utah project won approval despite local opposition, and Anthropic says it will invest $50 billion in US AI data centers.

Why it matters: HKR-H/K/R all pass, but this is a Verge running roundup rather than a single breakout event. The concrete power-grid and capex numbers place it at the upper end of industry reporting.

Bloomberg Technology

Anthropic Inks $1.8 Billion Computing Deal With Akamai

Anthropic signed a $1.8 billion computing deal with Akamai to meet rising demand for its AI software; the post does not disclose capacity, contract duration, or deployment regions.

Why it matters: HKR-H/K/R all pass: the $1.8B number is concrete, the Anthropic-Akamai pairing is fresh, and the story maps to Claude compute pressure. Missing scale, term, and regions keep it just above the featured threshold.

AI HOT (Curated Pool)

EMO: Expert Mixture Models for Emergent Modular Pretraining

AllenAI introduced EMO, a mixture-of-experts model with 14B total parameters and 1B active parameters, trained on 1 trillion tokens and able to use only 12.5% of its experts for specific tasks while retaining near-full-model performance.

Why it matters: HKR-H/K/R all pass, but this is an AllenAI/Hugging Face research release rather than a frontier model launch. The 14B/1B and 12.5% expert-activation claims justify the low featured band.

May 8Friday

r/LocalLLaMA

Gemma 4 26B Hits 600 Tok/s on One RTX 5090

chain-77 benchmarked Gemma 4 26B with vLLM 0.19.2rc1, and DFlash raised output throughput on one RTX 5090 from 228 tok/s to 578 tok/s under 256 input tokens, 1024 output tokens, concurrency 1, and num_speculative_tokens=13.

Why it matters: HKR-H/K/R all pass: the single-GPU throughput hook is strong, and the post gives reproducible settings plus before/after speed. Reddit single-post evidence and one hardware setup keep it in the featured-threshold band.

Synced · WeChat

SGLang Team Launches RadixArk With $100M Seed Round

RadixArk announced a $100 million seed round on May 5 at a $400 million post-money valuation, while its SGLang inference project has 27K+ GitHub stars and deployments across 400K+ GPUs.

Why it matters: HKR-H/K/R all pass: the round size, valuation, and deployment numbers are concrete, and SGLang is a known inference stack. It is still a startup funding and infra-roadmap story, not a major model release, so it stays in the 78–84 featured band.

Xinzhiyuan · WeChat

Token-Level Length Control: 3B Model Beats GPT 5.4 and Claude

UC Santa Barbara and Apple researchers introduced LenVM, which models remaining generation length as a token-level value function; Qwen2.5-3B with a 1.5B LenVM scored 62.6 on LIFEBench length control, above GPT-5.4 at 37.4 and Claude-Opus-4-6 at 35.5.

Why it matters: HKR-H/K/R all pass: the headline has a sharp small-model-vs-frontier hook, and the post gives LenVM's mechanism plus 62.6/37.4 benchmark numbers. The topic is narrow research, not a model or major product release, so it fits the 78-84 band.

QbitAI · WeChat

HIT and Huawei propose Dynamic-dLLM, a training-free acceleration framework with 4.48x speedup

HIT Shenzhen, Huawei, and Shenzhen Hetao College proposed Dynamic-dLLM, a training-free dLLM acceleration framework that raises LLaDA-8B-Instruct throughput on GSM8k from 8.32 TPS to 37.29 TPS with almost no accuracy loss.

Why it matters: HKR-H/K/R all pass: the 4.48x speedup is clickable, and GSM8k TPS figures add concrete substance. It is inference-optimization research, not a mainstream model launch, so it fits the 78–84 band.

Financial Times · Technology

Big Tech’s $725bn AI Spending Spree Sends Free Cash Flow to a Decade Low

Big Tech is spending $725 billion on AI infrastructure, and the title says free cash flow has fallen to a decade low; the RSS snippet says Silicon Valley giants shifted from asset-light cash generators to infrastructure investors, but it does not disclose the company list, time period, or accounting basis.

Why it matters: HKR-H/K/R all pass: the FT angle ties $725bn in AI infrastructure spending to decade-low free cash flow, hitting cost and ROI anxiety. Missing company list, period, and accounting scope keeps it in the 78–84 band.

r/LocalLLaMA

11.67% ARC-AGI-2 Local Eval on a Single 4090: The TOPAS Recursive Architecture

Doug_Bitterbot says TOPAS scored 11.67% on ARC-AGI-2 using one RTX 4090 after about 14 days of training. The 100M-parameter checkpoint hit 36% locally, but recursive TTT caused null outputs on nearly half of Kaggle puzzles. The key detail is time management: the author expects 20% after threshold tuning and 3-5 more weeks of training.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post with unstable Kaggle submissions. It clears featured, not the higher research-release band.