Skip to content

#部署/工程

3 today

May 18Monday

Import AI (Jack Clark)

Import AI 457: AI Stuxnet, Cursed Muon Optimizer, and Positive Alignment

Import AI 457 covers fast16, Aurora, and positive alignment: SentinelOne found fewer than 10 matching files for fast16 signatures, while Tilde Research reports Aurora reached 2.26 loss on 1.1B-parameter transformers versus Muon’s 2.31 under a ~100B-token setup.

Why it matters: HKR-H/K/R all pass: strong hooks plus concrete fast16 and Aurora numbers, with safety and optimizer stakes. It stays below 78 because this is a multi-topic newsletter roundup, not a single major release or industry event.

r/LocalLLaMA

Qwen 3.6 27B on 24GB VRAM: backend comparisons, quant choice, and settings

The author tested Qwen 3.6 27B on an RTX 3090 24GB and kept ik_llama.cpp with Qwen3.6-27B-MTP-IQ4_KS.gguf; at 156k context with q8_0 KV and MTP, a ~5.9k-token prompt plus 1024-token output reached about 1261 tok/s prefill and 72.9 tok/s decode, while vLLM lacked a clean single-card long-context run.

Why it matters: HKR-H/K/R all pass: this is a first-person local inference benchmark with concrete VRAM, quant, context, and speed numbers. Its reach is narrower than a model release, so it sits at the featured threshold.

Synced · WeChat

ICML 2026: Huawei GTS proposes EDCO for dynamic curriculum fine-tuning

Huawei GTS proposed EDCO, a dynamic curriculum method that selects fine-tuning samples by inference entropy; prefix entropy estimation cuts per-sample scoring time from 2.24 seconds to 0.37 seconds.

Why it matters: HKR-H/K/R pass: the story has a lab-race hook, a concrete entropy-based mechanism, and a 2.24s→0.37s efficiency claim. It stays below 78 because it is still a training-method paper, not a major model or product release.

r/LocalLLaMA

Benchmarking vLLM vs SGLang vs llama.cpp on a mixed Blackwell/Ada cluster

The author benchmarked long-context prefill on a 7-GPU mixed Blackwell/Ada cluster; on Qwen3.5-397B-A17B with 75k tokens, vLLM reached 9.8s TTFT and 7,683 t/s, while llama.cpp took 57.2s and 1,319 t/s.

Why it matters: Single-source Reddit benchmark, so source authority keeps it near the threshold. HKR-H/K/R pass on the mixed 7-GPU setup, 397B at 75k tokens, and concrete TTFT/throughput numbers.

r/LocalLLaMA

LLMs on Android: Snapdragon 8 Elite MoE Experience

A Reddit user tested MoE LLMs on an Honor Magic 7 Pro with Snapdragon 8 Elite and 24GB RAM; under Q4 quantization, LFM2-24b-a2b reached about 24 tokens/s while Gemma reached about 11 tokens/s, and CPU inference was still faster than NPU or GPU in the reported setup.

Why it matters: HKR-H/K/R all pass: a named Reddit test gives hardware, quantization, and token/s figures. Single-device anecdote and weak source authority keep it at the low featured band.

Google DeepMind

Introducing Google Antigravity 2.0

Google 发布智能体开发平台 Google Antigravity 2.0。该平台在 Google DeepMind 官网被列为面向开发者的 agentic development platform,与 Gemini 应用、Google AI Studio 并列。原文未披露版本功能、参数或可用性细节。

May 17Sunday

QbitAI · WeChat

A Robot Dog Challenges Nvidia's Compute Lead

Weilan Technology unveiled BabyAlpha A3, a consumer quadruped robot using a six-chip heterogeneous cluster that runs a 7B-parameter model on-device at 280 TPS; the article says it has 66MP vision, 2.232 million point-cloud samples per second, and a planned Q3 launch.

Why it matters: HKR-H/K/R pass: the robot-dog-versus-Nvidia angle is clickable, and 280 TPS on a local 7B model is concrete. Single-source summary lacks price, power draw, and benchmark setup, so it stays near the featured floor.

r/LocalLLaMA

Same Models Tested Across Strix Halo, RTX 3090, and RTX 5070

C_Coffie published 55 local inference benchmark runs across Strix Halo, RTX 3090, RTX 5070, five backends, and 0.35B to 35B-A3B models; RTX 5070 beats RTX 3090 on models fitting 12GiB, while RTX 3090 leads in the 14–31B band that exceeds 12GiB but fits 24GiB.

Why it matters: Hits HKR-H/K/R with a named first-person benchmark: 55 runs and concrete GPU crossover points. Source is a single Reddit post, so it stays in the low featured band.

Dwarkesh Patel podcast

Notes on Pretraining Parallelisms and Failed Training Runs

Dwarkesh documents pretraining failure modes and parallelism tradeoffs: expert choice and token dropping can break causality in MoE routing, FP16 collectives can bias repeated additions after values exceed 1024, pretraining FLOPs are given as 6ND, B300 HBM is listed as 288GB, and FSDP communication can reach params × 3 with reduce-scatter.

Why it matters: HKR-H/K/R all pass: Dwarkesh’s notes expose concrete pretraining failure modes and numbers. The systems-training focus is specialized, so it sits in the high-quality band rather than same-day must-write.

May 16Saturday

Latent Space

Cerebras' $60B IPO: Slowly, then All at Once

Cerebras closed its IPO at a $60 billion market cap. CFO Bob Komin said it serves trillion-parameter models, including internal OpenAI 5.4 and 5.5 workloads, but the post does not disclose traffic share, latency tier, or cost per token.

Why it matters: HKR-H/K/R all pass: a $60B Cerebras IPO is same-day material, and Bob Komin’s trillion-parameter/OpenAI 5.4–5.5 claim gives the story concrete compute-roadmap stakes.

r/LocalLLaMA

Orthrus-Qwen3-8B: Up to 7.8× tokens/forward on Qwen3-8B with frozen backbone

Orthrus-Qwen3-8B adds a trainable diffusion attention head to a frozen Qwen3-8B backbone, reaches up to 7.8× tokens per forward and about 6× wall-clock speedup on MATH-500, while training 16% of parameters with under 1B tokens over 24 hours on 8×H200 GPUs.

Why it matters: HKR-H/K/R all pass: 7.8× speed is a strong hook, the post gives parameter, hardware, and benchmark conditions, and inference cost is a live practitioner pain. Single-source Reddit research keeps it in the lower good-quality band.

Bloomberg Technology

Cerebras CEO Is Worth $3.2 Billion After Year’s Largest IPO

Cerebras Systems rose about 68% on Nasdaq after the year’s largest IPO, giving the company a market value of roughly $67 billion; the headline says its CEO is worth $3.2 billion after the listing.

Why it matters: HKR-H/K/R all pass: Cerebras' IPO gives AI compute a public-market price via a 68% jump and about $67B market cap. The Bloomberg item is wealth-video framed, so it lands in must-write range, not industry-shaking.

May 15Friday

r/LocalLLaMA

Fully Offline Suitcase Robot Built Around Jetson Orin NX SUPER 16GB

CreativelyBankrupt built Sparky as a fully offline suitcase robot on Jetson Orin NX SUPER 16GB, running Gemma 4 E4B Q4_K_M via llama.cpp with q8_0 KV cache, about 200 ms cached TTFT, 14-15 tok/s sustained output, 12K context, 30+ sensors, and no WiFi, Bluetooth, or cellular interface.

Why it matters: HKR-H/K/R all pass, with a named hands-on build and concrete latency/sensor numbers. It stays in low featured because this is a Reddit project post, not a product launch or research release.

AI HOT (Curated Pool)

X open-sources the “For You” feed recommendation algorithm

X open-sourced the For You recommendation pipeline on GitHub, using a Grok-based Phoenix Transformer to score candidate posts and predict engagement probabilities such as likes, replies, and reposts.

Why it matters: HKR-H/K/R all pass, but the item only gives the open-source claim and Phoenix Transformer ranking mechanism; repo details, license, and reproducible tests are not disclosed, so it stays low-featured.

r/LocalLLaMA

Evaluated a RAG Chatbot: The Most Expensive Model Was the Worst Performer

The author evaluated a customer-support RAG bot and raised the quality score from 6.62 to 7.88 while cutting per-session cost from $0.002420 to $0.000509, using retrieval logging, LLM-as-judge scoring, chunk deduplication, stricter grounding, and a five-model sweep.

Why it matters: HKR-H/K/R all pass: counterintuitive model ranking, concrete quality and cost deltas, and direct RAG production relevance. Reddit source authority keeps it near the featured floor despite the first-person experiment signal.

r/LocalLLaMA

Used over a million tokens in three sessions to test Qwen 3.6 35B MTP

A Reddit user tested Qwen3.6-35B-A3B MTP across three million-token-scale sessions, using 300k context and KV Q8_0, and reported about 1.5x the tok/sec of earlier tests.

Why it matters: HKR-H/K/R all pass: the million-token test is clickable, 300k context and KV Q8_0 add testable detail, and local speed maps to cost. Source is one Reddit post, so it stays below the high-importance band.

Bloomberg Technology

OpenAI May Raise More Money as Compute Crunch Deepens, CFO Says

OpenAI CFO Sarah Friar said the company may raise more capital after completing what she described as the largest private fundraising round ever, as OpenAI seeks computing power to meet rising AI demand; the RSS snippet does not disclose the round size, target amount, or timeline.

Why it matters: HKR-H/K/R pass: a named OpenAI CFO links more fundraising to the compute crunch. The score stays in the lower featured band because this is not a closed round and amount, investors, and timing are not disclosed.

AI HOT (Curated Pool)

Microsoft Has Invested Over $100 Billion in OpenAI, Nadella Says No One Wanted to Bet Then

Microsoft has invested more than $100 billion in OpenAI, including a $13 billion original investment and Azure infrastructure costs, while the partnership has generated about $30 billion in revenue; the renewed non-exclusive agreement caps OpenAI’s revenue share at $38 billion cumulatively through 2030.

Why it matters: HKR-H/K/R all pass: this is not a product launch, but the Microsoft-OpenAI economics include three concrete figures on spend, revenue, and revenue-share caps, placing it in the 78–84 quality band.

AI HOT (Curated Pool)

The First Derivative of Inference: Growth Logic in the AI Wave

Tom Tunguz says the AI inference market will reach $250 billion within seven years; Datadog’s LLM observability data volume nearly doubled in the latest quarter, and about 20% of its AI customers contribute roughly 80% of ARR.

Why it matters: HKR-H/K/R all pass: Tom Tunguz ties inference growth to Datadog volume and ARR concentration data. It stays in the 72–77 band because this is commentary, not a model, product, or protocol release.

AI HOT (Curated Pool)

API prompt precaching speeds up first-token generation

Claude API prewarms prompt cache with the system prompt, skips output, then hits cache on the real request.

Why it matters: HKR-H/K/R all pass: this is a concrete Claude API latency mechanism, not a vague product tease. It clears featured, but it is a mid-weight inference update rather than a major model or capability release.

Bloomberg Technology

Cerebras Goes Public in Year's Biggest IPO | Bloomberg Tech 5/14/2026

Bloomberg Tech says Cerebras is going public in the year’s biggest IPO, while the RSS snippet does not disclose the fundraising amount, valuation, offer price, or exchange.

Why it matters: HKR-H/K/R all pass, but the item only gives a title-level Bloomberg claim with no proceeds, valuation, price, or exchange. Cerebras’ IPO matters for AI infrastructure markets, so it lands in the 78–84 band.

Bloomberg Technology

AI Chipmaker Cerebras Climbs 68% After Year’s Biggest IPO

Cerebras Systems shares closed up 68% in their trading debut after raising $5.5 billion in the year’s largest IPO, ending Thursday at $311.07 in New York versus the $185 IPO price.

Why it matters: HKR-H/K/R all pass: Cerebras’ IPO gives AI infrastructure a public-market pricing signal, with a 68% debut gain and $5.5B raised. The chip-supply and NVIDIA-competition angle pushes it into p1.

Bloomberg Technology

Cerebras CEO Is Worth $3.2 Billion After Year’s Largest IPO

Cerebras’ IPO valued CEO Andrew Feldman’s fortune at $3.2 billion, and the title identifies the listing as the year’s largest IPO; the RSS snippet only says Feldman previously sold three companies and took another public, and does not disclose the IPO price, proceeds, valuation, or share count.

Why it matters: HKR-H/K/R all pass: Cerebras’ “year’s largest IPO” is a real AI-infra capital-market signal. Score stays in 78–84 because the feed gives $3.2B net worth but not offer price, proceeds, or valuation.

r/LocalLLaMA

I tracked EU GPU prices across 15 stores for 50+ days: RTX 5090 is the only card not dropping

Reddit user egudegi tracked EU GPU prices across 15 stores for more than 50 days with a 6-hour scrape cadence and about 126,000 readings; RTX 5090 average pricing rose from €3,392 to €3,487, a 3.0% increase.

Why it matters: HKR-H/K/R all pass, backed by a quantified first-person price scrape. Source authority is a single Reddit post, so it sits at the featured threshold rather than a higher band.

r/LocalLLaMA

The RTX 5000 PRO 48GB arrived and is better than expected

A Reddit user built a $5,600 RTX 5000 PRO 48GB PC and ran Qwen3.6-27B-FP8 with full-precision cache; they report up to 80 tok/s in TG, about 50–60 tok/s on very large prompts, 4,400 tok/s in prompt processing, and 200k tokens fitting in BF16 KV cache.

Why it matters: HKR-H/K/R all pass: a first-person local-inference test gives price and speed numbers, not vendor copy. Single Reddit source limits reach, so it lands in the featured-threshold band.

TechCrunch · AI

Cerebras raises $5.5B, then stock pops 108%, in the first huge tech IPO of 2026

The title says Cerebras raised $5.5 billion and its stock rose 108% after the first major tech IPO of 2026; the post does not disclose the offer price, valuation, share count, or use of proceeds.

Why it matters: Cerebras opens the 2026 tech IPO window with a $5.5B raise and 108% stock pop, clearing HKR-H/K/R. It is not a model launch, but an AI-chip IPO directly hits compute capital markets, so 90 and p1.

AI HOT (Curated Pool)

Accelerating On-Device AI: Arm and Google AI Edge Optimization Practices

Arm SME2 and Google AI Edge integrate with LiteRT, XNNPACK, and KleidiAI to optimize Stability AI’s stable-audio-open-small, delivering over 2x faster audio generation and 4x lower memory use on Arm-based mobile devices and laptops.

Why it matters: HKR-H/K/R pass via concrete 2x speed and 4x memory gains, plus an edge-deployment cost hook. Scope stays narrow to one audio model on Arm devices, so it lands at the featured threshold.

May 14Thursday

Bloomberg Technology

AI Chipmaker Cerebras Climbs 68% After Year’s Biggest IPO

Cerebras Systems rose 68% in its trading debut after raising $5.5 billion in the year’s largest IPO; the post does not disclose the IPO price or valuation.

Why it matters: Cerebras pairs a $5.5B IPO with a 68% first-day jump, giving AI infrastructure a fresh public-market price signal. HKR-H/K/R all pass; no hard-exclusion rule applies.

QbitAI · WeChat

Chinese GPU Vendor Hosts Open Source Meetup With SGLang Core Developers

Moore Threads said at the SGLang × MUSA Meetup that the MUSA backend has been merged into SGLang mainline, with 47 PRs submitted and 41 merged as of May 12.

Why it matters: HKR-H/K/R all pass, but this is an inference-backend ecosystem update rather than a model launch or platform shift. The 47 PRs and 41 merges make it concrete enough for featured, not P1.

AI HOT (Curated Pool)

OpenSquilla Open-Source Project Uses Smart Routing and Local Retrieval to Cut LLM Costs

OpenSquilla combines local model routing, vector retrieval, incremental sending, and cache hits to reduce transmitted tokens by more than 90%, while routing simple tasks to cheaper models and complex tasks to stronger models without spending tokens on the routing decision.

Why it matters: HKR-H/K/R all pass, but the source appears to be a single X project post; repo traction, test setup, and limits are not disclosed. Score lands at the featured threshold for practical open-source cost tooling.

AI HOT (Curated Pool)

UnslothAI Releases Qwen3.6 MTP GGUF Models With Over 1.4x Faster Inference

Daniel Han released experimental Qwen3.6 MTP GGUF models, with the 27B model reaching 140 tokens/s on one GPU and the 35B-A3B version reaching 220 tokens/s, using two draft tokens for speculative decoding.

Why it matters: HKR-H/K/R pass via concrete single-GPU speed claims and local-inference relevance. Score stays in low featured because the post is a single X source and does not disclose GPU, quantization settings, or repro steps.

AI HOT (Curated Pool)

Moonshot AI founder Yang Zhilin releases a 40-minute video

Yang Zhilin explains Kimi K2 training in a 40-minute video, saying the model cost $4.6 million and beat GPT-5.5 and other competitors on coding tasks.

Why it matters: HKR-H/K/R all pass: the founder-led Kimi K2 training breakdown adds a $4.6M cost figure and GPT-5.5 coding comparison. Single-source X relay and missing benchmark names keep it in 78-84, not P1.

AI HOT (Curated Pool)

Cost Analysis of AI Email

Top AI models process email at about $22 to $130 per month, with a $26 median; smaller models cut costs by 10 to 20 times, while local GPU execution can bring marginal cost close to zero.

Why it matters: HKR-H/K/R pass via a concrete cost spread and deployment-cost nerve. It is a useful opinion analysis, not a major product or model release, so it sits at 73.

AI HOT (Curated Pool)

Unlocking Asynchrony in Continuous Batching

Hugging Face says an 8B model generating 8K tokens leaves the GPU idle for 24% of the time, and asynchronous batching uses CUDA streams to overlap CPU preparation for batch N+1 with GPU computation for batch N.

Why it matters: HKR-H/K/R all pass, but this is inference-systems engineering rather than a major model release. The Hugging Face post provides a concrete 24% idle-rate number and CUDA-stream overlap mechanism, placing it in low featured.

Financial Times · Technology

Cerebras boosts IPO price to raise $5.5bn

Cerebras raised its IPO price to seek $5.5bn in proceeds; the post discloses a $40bn valuation, but does not disclose the price range, share count, or listing date.

Why it matters: HKR-H/K/R all pass: FT reports Cerebras lifting its IPO price with a $5.5bn raise target and $40bn valuation. A large AI-chip listing is same-day material for compute supply, NVIDIA competition, and AI capital markets.

r/LocalLLaMA

2x RTX 3090 setup for local Qwen 3.6 27B inference

A Reddit user ran Qwen 3.6 27B on a dual RTX 3090 Ubuntu setup, reporting 48GB VRAM, a 262k context window, no NVLink, about 4000 pp/s prompt processing, and 113 tk/s generation.

Why it matters: All HKR axes pass, and this is a first-person local-inference run with concrete numbers. Source is a single Reddit post with limited reproducibility detail, so it sits at the low featured threshold.

r/LocalLLaMA

24+ tok/s from ~30B MoE models on an old GTX 1080

User mdda ran Qwen 3.6 35B-A3B on an i7-6700, GTX 1080, and 32GB RAM machine at about 24 tok/s with 128k context; the setup uses llama.cpp MoE offloading plus TurboQuant/RotorQuant KV cache quantization, with PCIe 3.0 x16 saturated and GPU utilization at about 40–50%.

Why it matters: Single Reddit source limits authority, but the GTX 1080 + Qwen 3.6 35B-A3B + 128k + 24 tok/s setup gives a concrete local-inference result. HKR-H/K/R all pass; this is a practical featured item, not a major model or product launch.

Bloomberg Technology

AI Chipmaker Cerebras Raises $5.55 Billion in Year’s Biggest IPO

Cerebras Systems raised $5.55 billion in a US IPO, and the title calls it the year’s biggest IPO; the RSS snippet does not disclose the offering price or valuation.

Why it matters: Cerebras’ $5.55B U.S. IPO clears HKR-H/K/R and is same-day material for AI infrastructure. Missing offer price and valuation keeps it below the industry-shaking band for a foundation-model-company IPO.

AI HOT (Curated Pool)

Meta AI chief announces Incognito Chat for WhatsApp and Meta AI

Meta’s AI chief announced Incognito Chat for WhatsApp and Meta AI, with conversation inference running inside the phone’s hardware secure enclave, no server logs generated, and session data permanently deleted after the chat ends.

Why it matters: HKR-H/K/R all pass: the hook is Incognito Chat in WhatsApp, with secure-enclave inference and no server logs. Single-source brevity limits verification, so it sits below model releases and major capability launches.

May 13Wednesday

QbitAI · WeChat

ByteDance Proposes Generative Refinement Networks as a Third Route for Visual Generation

ByteDance’s commercial technology team proposed GRN, a visual generation architecture using HBQ, global refinement, and complexity-aware sampling to address quantization loss, error accumulation, and fixed-step inference; on a 130M model, adaptive sampling reduced inference from 50 steps to an average of 24, while gFID changed from 3.56 to 3.79.

Why it matters: HKR-H/K/R all pass: ByteDance’s GRN has a concrete hook plus 130M, 24-step inference and gFID 3.79. It is a strong research release, not a flagship model launch, so it stays in the 78–84 band.