Skip to content

#部署/工程

3 today

Jun 1Monday

AI HOT (Curated Pool)

OpenAI Starts Construction of Stargate 1GW Data Center in Michigan

OpenAI started the Stargate 1GW data center project in Michigan; the RSS snippet discloses the 1GW capacity but does not disclose investment size, construction timeline, or compute configuration.

Why it matters: HKR-H/K/R all pass: OpenAI disclosed a 1GW Stargate data-center build in Michigan. Missing investment, timeline, and GPU configuration keep it in the 78–84 band, not same-day P1.

r/LocalLLaMA

Deepseek V4 Flash performance on DGX Spark

A Reddit user ran DeepSeek-V4-Flash with vLLM on two ASUS GX10 DGX Spark nodes and reported 1,680 prefill tokens/s plus 39.8 decode tokens/s at a 256K context with MTP=2; the setup uses TP=2 over RoCE, fp8 KV cache, and fits about 1M tokens safely in KV cache.

Why it matters: This is not broad industry news, but it is a first-person benchmark with reproducible details: TP=2, RoCE, fp8 KV cache, 256K context, and ~1M KV. HKR-H/K/R all pass, so it lands at low featured.

AI HOT (Curated Pool)

NVIDIA Releases RTX Spark and Local AI Agent Security and Performance Updates

NVIDIA released RTX Spark, a Windows PC for local AI agents with 1 petaflops of AI compute and 128GB of unified memory. OpenShell uses new Windows security primitives with Microsoft, while llama.cpp optimizations raise Qwen 27B throughput by up to 2x.

Why it matters: HKR-H/K/R all pass: NVIDIA frames RTX Spark for local agents and gives hard specs: 1 petaflops, 128GB, and up to 2x llama.cpp throughput. Vendor-blog framing keeps it in the low 78–84 band.

AI HOT (Curated Pool)

Nvidia Enters Windows Laptop Market, Taking on Intel and AMD

Nvidia introduced one new PC-focused chip to enter the Windows laptop market and compete with Intel and AMD; the RSS snippet does not disclose specifications, pricing, launch timing, or AI compute metrics.

Why it matters: Bloomberg authority and Nvidia’s move into Windows laptops clear HKR-H/R and the featured floor. HKR-K fails because specs, pricing, launch timing, and AI performance are not disclosed.

r/LocalLLaMA

I ported NVIDIA Parakeet speech-to-text to ggml: same output as NeMo, faster, GGUF-quantized, no Python

mudler_it ported NVIDIA Parakeet speech-to-text models to C++/ggml with no Python or PyTorch, reporting byte-for-byte NeMo parity on f32/f16, up to about 5x GPU speedups on larger TDT and hybrid models, and GGUF quantization across f16, q8_0, q6_k, q5_k, and q4_k.

Why it matters: HKR-H/K/R all pass: the port has a concrete local-inference hook, byte-parity and speed claims, and clear practitioner resonance. Source scope keeps it at the low featured band, not P1.

May 31Sunday

AI HOT (Curated Pool)

Apple WWDC AI Upgrade: Gemini-Distilled Model Runs Locally, With Heavy External Dependencies

Apple will present Siri and on-device AI upgrades at next month’s WWDC, with iPhones running a smaller Gemini-distilled model locally while complex queries route to Google Cloud using Nvidia confidential computing.

Why it matters: HKR-H/K/R all pass: the Apple-Google-Nvidia stack is a strong WWDC AI hook with a concrete routing mechanism and clear industry tension. Capped at 82 because this is a single X-sourced claim with no model size, latency, pricing, or contract terms disclosed.

QbitAI · WeChat

NVIDIA’s MacBook Pro-like laptop reportedly uses an in-house CPU

NVIDIA, Microsoft, and Arm posted the same “new era of PC” teaser, and the article says the rumored N1X laptop may use a 20-core Arm CPU, a Blackwell GPU, 6,144 CUDA cores, and 128GB of LPDDR5X unified memory, while bandwidth and x86 translation remain the stated constraints.

Why it matters: HKR-H/K/R all pass, but the story rests on hints and rumored specs; launch date, price, and production plan are not confirmed. Treat it as a strong hardware rumor, not a same-day must-write release.

r/LocalLLaMA

Cost Analysis of My $6.4k Local LLM Server

The author runs Qwen3.6 27B on a $6,406.45 local server with 4 MI100 GPUs, processing 20.4M input tokens and 1.32M output tokens per day; using OpenRouter prices, the first-year local cost is $2,992.72 versus $3,701.10 for API use.

Why it matters: HKR-H/K/R all pass: a first-person local-LLM cost test gives hardware, token volume, and API comparison. Single Reddit post and workload-specific economics keep it in the lower featured band.

AI HOT (Curated Pool)

DynoSim: Simulation-Driven Inference Stack Optimization

NVIDIA released DynoSim for optimizing its Dynamo inference serving stack; the Rust-based tool models thousands of deployment configurations on a single virtual timeline and reached 1,500x real-time speed in tests.

Why it matters: HKR-H/K/R all pass: the hook is 1500x real-time simulation, with a concrete virtual-timeline mechanism and infra cost resonance. Single-source NVIDIA product update keeps it in the lower featured band.

May 30Saturday

r/LocalLLaMA

Project Blackwell: Making an RTX Pro 6000 Run in a Dell R730 at 650K Context

The author installed an RTX Pro 6000 Blackwell in a 2016 Dell PowerEdge R730 and claims a 650K-context local AI box; the post describes fan-shroud modification, dual-riser power, PCIe BAR allocation failures, ACPI/DSDT inspection, MMIO aperture work, and Linux PCIe boot-flag testing as required conditions.

Why it matters: HKR-H/K/R all pass: the 650K-context Blackwell-in-R730 build is novel, concrete, and cost-relevant. Still, it is a niche local-AI hardware experiment, not a broad product or model release.

AI HOT (Curated Pool)

xAI drops JAX GPU for an in-house training framework

SemiAnalysis says xAI dropped JAX GPU and moved to a C training framework written with Grok Build; the snippet claims xAI’s JAX stack had MFU below 10%, but the post does not disclose reproducible benchmark conditions.

Why it matters: HKR-H/K/R all pass: xAI changing its training stack is a strong hook, MFU <10% is a concrete claim, and infra cost will spark debate. Single-source tweet format and no reproducible setup keep it at 80, not P1.

Synced · WeChat

Apple Uses AI to Rework Image Compression: Same Visual Quality at One-Third the File Size

Apple’s team published PICO, a perceptual image codec that uses 57%-70% fewer bits than AV1, VVC, and JPEG AI at the same subjective visual quality, while encoding a 12MP photo in 230 ms and decoding it in 150 ms on an iPhone 17 Pro Max.

Why it matters: HKR-H/K/R all pass: Apple PICO has concrete 57%-70% bitrate savings and 230 ms on-device encoding data. It remains a research release, not a shipped platform feature, so it sits in the 78-84 band.

r/LocalLLaMA

Testing MTP on vLLM and llama.cpp for Gemma 4 and Qwen 3.6

The author tested MTP on an RTX PRO 6000 Blackwell setup, where Gemma 4 31B on vLLM reached 132.52 tok/s versus a 39.69 tok/s baseline, a 3.34x speedup; the post reports 10 runs of 1,500 tokens each but does not provide a full quality or VRAM evaluation.

Why it matters: HKR-H/K/R all pass via a first-person speed test with hardware, model, and tok/s numbers. Source authority is limited, and missing quality/VRAM evaluation keeps it at the low featured band.

AI HOT (Curated Pool)

OpenAI launches real-time translation model with 70+ input languages

OpenAI launched gpt-realtime-translate, a speech translation model that accepts 70+ input languages and outputs speech in 13 target languages; the post says the feature is running on smart glasses.

Why it matters: HKR-H/K/R all pass: OpenAI has a concrete realtime-translation model with numbers and a wearable demo. Missing latency, pricing, and API availability keep it below P1.

TechCrunch · AI

After Nvidia’s $20B Not-Acqui-Hire, AI Chip Startup Groq Reportedly Raising $650M

Axios says Groq is seeking $650 million in internal funding while shifting from hardware toward AI inference, after Nvidia’s reported $20 billion not-acqui-hire; the RSS snippet does not disclose Groq’s valuation, investor names, deal structure, or fundraising timeline.

Why it matters: HKR-H/K/R pass: the $650M Groq raise is a concrete AI-inference infrastructure signal. Missing valuation, investor names, and timing keep it at the featured threshold rather than a higher funding story.

May 29Friday

Xinzhiyuan · WeChat

Three DeepSeek Models Enter OpenRouter Monthly Top 10 With Over 17 Trillion Tokens

DeepSeek placed three models in OpenRouter’s monthly top 10 with more than 17 trillion tokens combined, including V4 Flash at 9.13T tokens; the article says Ascend’s MegaMoE operator raised Prefill throughput by 20% to 30% on DeepSeek V3.1 and Qwen3-235B tests.

Why it matters: HKR-H/K/R all pass: the story has a 17T-token hook plus concrete OpenRouter and MegaMoE Prefill numbers. It stays at 82 because the compute-sovereignty framing is strong, while reproducible test conditions are not disclosed.

r/LocalLLaMA

Liquid AI releases LFM2.5-8B-A1B

Liquid AI released LFM2.5-8B-A1B with a 128K context window, 38T pre-training tokens, large-scale reinforcement learning, doubled vocabulary for non-Latin tokenization, and availability on Hugging Face.

Why it matters: HKR-H/K/R pass: 8B/A1B, 128K context, and 38T tokens are concrete hooks for local inference. No benchmarks, license, or deployment limits are disclosed, so it stays in the mid featured band.

Synced · WeChat

A True 2-bit KV Quantization Algorithm for Long-context Reasoning Beyond TurboQuant

TogetherAI and collaborators released OSCAR, a 2.28 BPE INT2 KV Cache system integrated with SGLang, reporting up to 3× decode speedup at 100k context and up to 7× job-level throughput under a fixed memory budget.

Why it matters: HKR-H/K/R pass, but this is niche inference optimization rather than a broad model launch. The 100k-context and ~3×/~7× claims justify a featured score, not same-day must-write.

Synced · WeChat

The Ma Jiaqi Failure Exposed an LLM Issue He Spotted in the Shower a Year Earlier

FaceMind links low-frequency token degradation to two papers: SLoW appeared at EMNLP 2025, Adam's Law was accepted as an ACL 2026 Oral, and high-frequency rewriting raised DeepSeek-V3 math accuracy from 63.55% to 71.54%.

Why it matters: HKR-H/K/R all pass: the odd celebrity-token hook is clickable, and the post gives a mechanism plus a 63.55%→71.54% DeepSeek-V3 result. Practical research signal, but not a major model launch.

Bloomberg Technology

Samsung Takes Lead in Shipping Top-End AI Memory Chip Samples

Samsung Electronics has begun shipping samples of its most advanced memory to customers for AI accelerators from companies including Nvidia; the RSS snippet does not disclose the chip model, customer list, sample volume, pricing, or mass-production timeline.

Why it matters: HKR-H/K/R pass on the Samsung AI-memory supply-chain angle, but the facts stop at sample shipments; no model, customer list, or production window keeps it at the featured floor.

Bloomberg Technology

Apollo Shops $36 Billion Debt Deal to Buy Google Chips for Anthropic

Apollo and Blackstone are seeking additional investors for about $36 billion in debt financing for Anthropic’s AI infrastructure; the title says the deal would buy Google chips, while the post does not disclose chip models, purchase volume, or timeline.

Why it matters: Bloomberg supplies HKR-H/K/R: a $36B Anthropic compute-finance hook, concrete backers, and a Google-chip supply angle. It is not a model or product release, so the score stays at the low end of 85+.

AI HOT (Curated Pool)

Apple reportedly tries to fit Google's large Gemini model into iPhone for new Siri

Apple is trying to integrate a large Gemini model into the iPhone for new Siri features. The RSS snippet says full local processing is unlikely because of model size, and a cloud component is likely required; the post does not disclose parameters, latency targets, or a release timeline.

Why it matters: HKR-H/K/R all pass, but this is a reported Apple-Google Siri effort, not a shipped product. The post gives distillation and likely cloud dependency, but no timeline, size, or tests, so it stays at the featured threshold.

May 28Thursday

QbitAI · WeChat

Behind DeepSeek V4's Chip-Model Co-Design, China's Compute Ecosystem Gains Speed

QbitAI says DeepSeek V4 validated Ascend chip-model co-design, with CANN open-sourcing 65 repositories and supporting day-zero adaptation for more than 70 mainstream models, while AIGCode reported 65% MFU in MoE pretraining on Ascend.

Why it matters: HKR-H/K/R all pass, but this is mainly a compute-ecosystem progress story, not a DeepSeek V4 capability release. Concrete repo, adaptation, and MFU numbers lift it into featured, below must-write.

r/LocalLLaMA

Zai replaced the network architecture for GLM-5.1 inference, lifting throughput 15%

Zai replaced the ROFT network topology with ZCube on a thousand-GPU GLM-5.1 coding inference cluster, keeping the same GPUs, software stack, and model; the Reddit post cites 33% lower switch and optical module costs, 15% higher GPU inference throughput, and a 40.6% drop in first-token P99 tail latency under prefill-decode disaggregated inference.

Why it matters: HKR-H/K/R all pass: the GLM-5.1 inference cluster has concrete cost, throughput, and P99 latency numbers. Reddit single-source sourcing and infra-niche scope keep it at 78.

r/LocalLLaMA

Qwen3.6-35B-A3B-APEX Runs 128K Context on RTX 3060 12GB

old-mike ran mudler/Qwen3.6-35B-A3B-APEX-MTP-I-Compact.gguf through spiritbuun’s llama.cpp fork on one RTX 3060 12GB, offloading a 17.3GB model and reaching 37.17 t/s generation at 72K filled context, 28.08 t/s at 129K, and PPL 3.2529 on an enwik8 64K-context perplexity test.

Why it matters: HKR-H/K/R all pass via a concrete consumer-GPU inference result with speed and PPL. Source is a single Reddit post and the impact stays within local inference, so it lands in featured, not P1.

AI HOT (Curated Pool)

AI Now Summit 2026

Mistral AI announced industrial AI work, a Vibe upgrade, and a 10 MW inference data center in Les Ulis at AI Now Summit 2026; it is working with Airbus, BMW Group, and ASML, and the data center is scheduled to start operating in Q3 2026.

Why it matters: HKR-H/K/R pass: Mistral gives a concrete 10 MW inference site, Q3 2026 timing, and major industrial partners. No new model capability or pricing is disclosed, so it stays just above the featured threshold.

AI HOT (Curated Pool)

Mistral AI launches physics AI model for industrial engineering

Mistral AI integrated the Emmi AI team and launched a physics AI foundation model for industrial engineering, with the post saying it can learn from geometry, boundary conditions, or measurement data and predict full physical fields on a single GPU in seconds.

Why it matters: HKR-H/K/R pass: a major model lab entering physics simulation with a concrete single-GPU seconds claim. The score stays in the lower featured band because model name, benchmarks, pricing, and access are not disclosed.

Latent Space

Cognition Raises $1B in $26B Series D

Cognition raised a $1B Series D at a $26B valuation and projects more than $1B ARR by year-end; the post says its valuation rose 2.5× from the $10B Series C eight months earlier, while the rest of the issue summarizes agent, inference, benchmark, and multimodal AI updates from May 26–27, 2026.

Why it matters: HKR-H/K/R all pass: Cognition’s $1B Series D at a $26B valuation is large, and projected year-end ARR above $1B gives a concrete business signal. This is not a model launch, but it is must-write funding news for AI coding agents.

Xinzhiyuan · WeChat

Tsinghua Team Open-Sources PilotDeck Agent System, Claims 70% Token Cost Reduction

Tsinghua THUNLP, ModelBest, OpenBMB, and AI9stars open-sourced PilotDeck; the article says its sub-agent routing reduced cost from $12.58 to $2.83 in a Xiaohongshu content-generation test, while preserving separate WorkSpaces, editable memory, and per-session routing logs.

Why it matters: HKR-H/K/R all pass: PilotDeck has a clear agent-cost hook, a concrete routing mechanism, and $12.58 to $2.83 data. It stays in the 78–84 band because this is a tool release, not a major model or platform launch.

Synced · WeChat

Mila and DeepMind Propose UNSL for Unified Multivariate Neural Scaling Laws

Mila and Google DeepMind proposed Unified Neural Scaling Law, a multivariate scaling-law form that models parameter count, token count, training steps, bottlenecks, overfitting, and adverse hyperparameter effects; UNSL achieved the best extrapolation on 60.87% of vision tasks and 88.89% of language tasks in the reported experiments.

Why it matters: HKR-H/K/R all pass: UNSL unifies parameters, tokens, steps, bottlenecks, overfitting, and hyperparameter feedback, with vision/language extrapolation numbers. Technical density keeps it in the 78–84 band.

TechCrunch · AI

Snowflake signs $6B deal with AWS for AI CPU chips

Snowflake signed a five-year, $6 billion AWS deal to secure chips for AI use; the post does not disclose chip models, delivery timing, or whether the capacity targets training, inference, or both.

Why it matters: HKR-H/K/R all pass: the 5-year, $6B AWS deal is a concrete AI-infra signal. Missing chip model, delivery cadence, and training/inference split keep it at the featured threshold.

r/LocalLLaMA

Inferencing at 10.33 t/s on Qwen 3.5 35B on a $300 laptop

A Reddit user ran Qwen 3.5 35B Q4_K_S on a $300 Lenovo Ideapad Slim 3i and reported 10.33 t/s inference using ik_llama.cpp with two pinned CPU cores, MTP speculative decoding, 64 batch size, and Q8_0 KV cache.

Why it matters: HKR-H/K/R all pass, with a concrete first-person benchmark. Reddit single-post sourcing and limited reproducibility details keep it at the lower featured threshold.

AI HOT (Curated Pool)

Open-source FastVideo Dreamverse real-time video generation tool

Hao AI Lab open-sourced FastVideo Dreamverse, a real-time video generation tool that generates a 30-second 1080p video in 7 seconds under the stated setup of one NVIDIA B200 GPU and LTX-2.

Why it matters: HKR-H/K/R all pass: the 7s-for-30s-1080p claim is concrete and practitioner-relevant. Single-source X sourcing and missing independent benchmarks keep it in the 78–84 band.

May 27Wednesday

AI HOT (Curated Pool)

Perplexity open-sources Unigram tokenizer to reduce CPU usage

Perplexity open-sourced a rebuilt Unigram tokenizer that reduces CPU usage by 5-6x, targeting tokenization latency when small rerankers and embedding models run on GPUs in single-digit milliseconds.

Why it matters: HKR-H/K/R all pass: the 5-6x CPU claim and tokenizer bottleneck are concrete for production RAG/search teams. It stays in the featured-threshold band because the post lacks independent benchmarks, repo details, and deployment scale.

QbitAI · WeChat

Language Models Need Sleep: Let AI Nap Before Continuing Inference

Carnegie Mellon University and the University of Maryland propose a “sleep” mechanism for language models: when the context window is nearly full, the model stops accepting new tokens, runs multiple offline recursive forward passes to compress accumulated context into fast weights, clears the KV cache, and then resumes inference; tests cover cellular automata, multi-hop graph retrieval, and GSM-Infinite reasoning tasks.

Why it matters: HKR-H/K/R all pass: the sleep metaphor is clickable, and the mechanism is concrete. Score stays below 78 because the provided body lacks benchmark gains, code, or deployment evidence.

Xinzhiyuan · WeChat

OpenRouter processes 100 trillion tokens monthly and raises $113M Series B

OpenRouter raised a $113 million Series B led by CapitalG, lifting its valuation to $1.3 billion; the platform processes 25 trillion tokens per week, about 100 trillion per month, and provides one API for more than 400 models.

Why it matters: HKR-H comes from the 100T-token/month hook; HKR-K has funding, valuation, usage, and model-count numbers; HKR-R maps to routing and API-cost competition. Still, this is infra funding news, not an 85+ must-write release.

Synced · WeChat

AMD paper: FP4 training instability is not caused by insufficient randomness

AMD and Penn State pretrained Llama 3.1-8B with MXFP4 on MI355X native FP4 hardware, achieving 9-10% end-to-end speedup over an FP8 baseline, while the paper identifies Wgrad quantization as the bottleneck that raises token overhead to 26-27% without deterministic Hadamard stabilization.

Why it matters: HKR-H/K/R all pass: a counterintuitive FP4 claim, concrete Llama 3.1-8B numbers, and a cost/hardware nerve. The topic is narrower training-infra research, so it stays in the 78-84 band.

Latent Space

[AINews] New AI Infra Decacorns: Fireworks, Baseten, with OpenRouter on the Way

Latent Space says Fireworks is in talks for a $15 billion valuation round, Baseten is raising at an $11 billion valuation, and OpenRouter closed a $113 million Series C after volume grew 5x in six months.

Why it matters: HKR-H/K/R all pass: the decacorn hook is clickable, the post gives valuation, round, and usage figures, and the topic speaks to inference economics. Fireworks and Baseten are still reported as in talks or raising, so this stays in the 78–84 band.

AI HOT (Curated Pool)

Qualcomm and ByteDance reportedly partner on AI ASIC chips with millions of units planned

The title says Qualcomm and ByteDance reached an AI ASIC chip partnership with procurement in the millions of units; the post does not disclose chip specifications, unit pricing, delivery timing, or production conditions.

Why it matters: HKR-H/K/R all pass: the rumored Qualcomm–ByteDance AI ASIC deal has a concrete million-unit volume hook. Thin sourcing and missing specs, pricing, delivery, and production terms keep it in the 72–77 band.

Bloomberg Technology

Fireworks AI in Talks for Funding at $15 Billion Valuation

Fireworks AI is in talks to raise a new funding round at a $15 billion valuation, according to people familiar with the matter; the post does not disclose the round size, investor names, or timeline.

Why it matters: HKR-H/K/R all pass on the $15B AI-inference valuation, backed by Bloomberg. The deal is still in talks, with round size, investors, and timeline undisclosed, so it stays in low featured.