Skip to content

#部署/工程

3 today

May 3Sunday

Xinzhiyuan · WeChat

Claude Code helps Anthropic double revenue pace in two months

Semi Analysis says Anthropic’s ARR reached $44B, adding $35B over 12 months. Claude Code hit $2.5B annualized revenue by Feb 2026, while inference gross margin rose from 38% to over 70%. The key test is keeping enterprise usage, coding-agent revenue, and inference margin together.

Why it matters: HKR-H/K/R all pass: SemiAnalysis gives hard ARR, Claude Code revenue, and inference-margin numbers. Not a model launch, but it materially shifts the view of Claude Code monetization.

QbitAI · WeChat

DeepSeek V4’s biggest omission

DeepSeek V4’s technical report omits Engram while listing mHC, CSA, HCA, Muon, and FP4. Engram was open-sourced by DeepSeek and Peking University in January, inserting lookup modules between Transformer layers 2 and 15; its 27B test raised MMLU by 3.4 and Multi-Query NIAH to 97.0%. The engineering signal is CXL pooling: 8 servers shared a 4TB memory pool with under 5% throughput loss.

Why it matters: HKR-H/K/R all pass: the omitted-Engram angle is clickable, with layer ranges, benchmark deltas, and CXL memory-pool numbers. It is analysis, not the V4 launch itself, so 78–84 fits.

r/LocalLLaMA

Implemented TurboQuant, but results do not fully match the paper

A Reddit user reimplemented TurboQuant and found the PROD variant reached about 95.8% correlation at 4-bit, below the paper’s 99%+ claim. They report degraded attention quality, with about 67% top-1 accuracy in a simple simulation. The key issue is correlation versus ranking preservation in KV cache quantization.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit reproduction, not a formal release. The 95.8% 4-bit correlation and ~67% top-1 result make it a low featured item.

r/LocalLLaMA

Built a C++17 transformer from scratch with 0.83M params and CPU training

Reddit user Suspicious_Gap1121 released Quadtrix.cpp, a C++17 GPT-style model with 0.83M parameters. It uses 4 layers, 4 heads, 200d width, and a 128-character context; one CPU core trained on 31.4M characters for 76.2 minutes to 1.6371 nats val loss. The key detail is handwritten backprop for LayerNorm, attention, Q/K/V, dropout, and AdamW without PyTorch, BLAS, or autograd.

Why it matters: HKR-H/K/R all pass: the no-framework C++17 build is clickable, the training setup is specific, and local-LLM builders care about dependency-free control. It stays in the 72–77 band because it is a small personal project.

May 2Saturday

r/LocalLLaMA

Qwen3.6-27B hits 72 tok/s on RTX 3090 with native vLLM on Windows

Reddit user One_Slip1455 released a native Windows vLLM launcher for Qwen3.6-27B, reaching 72 tok/s on an RTX 3090. It reports 64.5 tok/s at ~25k tokens, 53.4 tok/s at 127k ctx on one GPU, and 160k ctx with PP=2 on 2×3090. The key detail is no WSL or Docker, an OpenAI-compatible endpoint, and an INT4 quant path.

Why it matters: HKR-H/K/R all pass: native Windows on an RTX 3090 is the hook, the post gives tok/s and ctx figures, and it hits local-inference cost concerns. Reddit single-source limits it to the lower featured band.

QbitAI · WeChat

Tencent Hunyuan open-sources 440MB offline translation model, claims Google Translate quality lead

Tencent Hunyuan open-sourced Hy-MT1.5-1.8B-1.25bit, compressing a 1.8B translation model to 440MB. It supports 33 languages and 1,056 directions, with an Android demo running offline on Snapdragon 888 and 8GB RAM. The key detail is Sherry 1.25-bit quantization: 3 of every 4 weights use 1 bit and 1 is zeroed.

Why it matters: HKR-H/K/R all pass: the story has a strong offline-phone hook, concrete quantization details, and practitioner relevance around edge inference. It stays below P1 because this is a vertical translation model, not a major foundation-model release.

May 1Friday

r/LocalLLaMA

PFlash: 10x prefill speedup over llama.cpp at 128K on an RTX 3090

PFlash cuts Qwen3.6-27B Q4_K_M 128K TTFT to 24.8s on an RTX 3090, versus 248.4s cold for llama.cpp. It uses a Qwen3-0.6B drafter to score token importance, keeps 5% of spans, and runs C++/CUDA without Python, Triton, or PyTorch. The quality caveat is clear: only NIAH single-needle passes from 32K to 128K; RULER and multi-needle results are not disclosed.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit claim with quality evidence limited to single-needle NIAH 32K–128K. RULER and multi-needle results are not disclosed, so it stays at featured threshold.

r/LocalLLaMA

OpenAI's Privacy Filter vs GLiNER on 600 PII Samples

A Reddit user compared openai/privacy-filter and GLiNER large-v2.1 on 600 PII samples. On CPU, OpenAI's model ran 2.8 samples/s versus 1.1 for GLiNER; English boundary macro F1 was 0.498 versus 0.416. The key issue is tokenizer offset: strict matching drops openai/privacy-filter to 0.155.

Why it matters: HKR-H/K/R all pass: the Reddit test has a clear matchup, 600 PII samples, speed/F1 numbers, and a tokenizer-offset caveat. Source authority is limited, so it stays in the low featured band.

r/LocalLLaMA

16x Spark Cluster Build Update

Reddit user Kurcide finished a 16-node DGX Spark cluster, with all nodes hitting line rate on the fabric. Each node uses one QSFP56 link to an FS N8510, showing 100–111 Gbps per rail and about 200 Gbps aggregate. The key angle is unified memory: 8 nodes served 434GB GLM-5.1-NVFP4, with DeepSeek and Kimi tests next.

Why it matters: HKR-H/K/R all pass: the post gives first-person cluster numbers, networking conditions, and a live 434GB model test. Scope stays local-inference hardware, so it fits the 72–77 band rather than a broader product-release tier.

Financial Times · Technology

Huawei’s AI chip sales surge as Nvidia stalls in China

Huawei received large AI processor orders from Chinese tech companies as Nvidia stalls in China. The post does not disclose order value, chip models, or delivery timing. The key issue is China’s domestic compute substitution path, not one sales headline.

Why it matters: FT sourcing and the Huawei-vs-Nvidia China angle clear HKR-H and HKR-R. HKR-K is weak because value, chip model, and delivery timing are not disclosed, so this stays in the 78–84 band.

r/LocalLLaMA

32x AMD MI50 32GB runs Kimi K2.6 at 9.7 t/s TG and 264 t/s PP

Reddit user ai-infos ran Kimi K2.6 int4 on 32 AMD MI50 32GB GPUs, reaching 9.7 tok/s TG on 136 output tokens. PP hit 263 tok/s on 14,564 input tokens using vllm-gfx906-mobydick across two 16-GPU nodes over 10G Ethernet. Power was about 640W idle and 4,800W peak inference; PCIe bandwidth and the vLLM distributed stack are the real bottlenecks.

Why it matters: HKR-H/K/R all pass via an unusual 32x MI50 build with concrete throughput, power, and network conditions. It stays in the 72–77 band because it is a niche Reddit benchmark, not a broader product or model release.

r/LocalLLaMA

Follow-up: Qwen3.6-27B on 1× RTX 3090 reaches ~218K context and stable tool calls

A Reddit user ran Qwen3.6-27B on one RTX 3090, reporting ~218K context at 50/66 TPS. After fixing Genesis PN12 patch anchor drift, ~25K-token tool outputs stopped OOMing; 198K plus vision reached 51/68 TPS. Single-prompt single-GPU runs still hit a second memory cliff near 50–60K.

Why it matters: HKR-H/K/R all pass: the single-3090 context claim is catchy, the post gives measured TPS and OOM conditions, and local-inference cost pressure resonates. Reddit source keeps it in the low featured band.

r/LocalLLaMA

Long-context coding on RTX 5080 16GB: Qwen3.6-35B-A3B holds 30 t/s at 128K

A Reddit user tested a local coding-agent setup on RTX 5080 16GB; the title says Qwen3.6-35B-A3B reaches 30 t/s at 128K. The post lists Ryzen 9700X, 96GB DDR5, Windows 11, and CUDA 12.9.1 as required. Qwen3.6-27B dense hit only 3.2 t/s at 128K, so the key path is KV quantization plus MoE offload.

Why it matters: HKR-H/K/R all pass: 30 t/s at 128K on a 16GB RTX 5080 is a strong hook, with hardware/CUDA details and a dense baseline. Single Reddit run lacks multi-source reproduction, so featured not P1.

NVIDIA Blog

Nemotron Labs: What OpenClaw Agents Mean for Every Organization

NVIDIA says OpenClaw reached 250,000 GitHub stars by March 2026, passing React within 60 days. OpenClaw is Peter Steinberger’s self-hosted persistent agent; NVIDIA introduced NemoClaw with OpenShell sandboxing and Nemotron models. The key issue is governance: the post claims reasoning AI raised token use 100x, and autonomous agents add another 1,000x.

Why it matters: HKR-H/K/R all pass: OpenClaw’s GitHub growth is a hook, and NemoClaw names concrete sandbox and access-control mechanisms. NVIDIA’s own blog keeps it in the 78–84 band.

The Verge · AI

Elon Musk confirms xAI used OpenAI’s models to train Grok

Elon Musk testified Thursday in a California federal court that xAI used OpenAI models to improve Grok. The mechanism described is model distillation: a larger teacher model transfers knowledge to a smaller student model. The post does not disclose which OpenAI models, data volume, or training runs.

Why it matters: HKR-H/K/R all pass: Musk confirmed in court that xAI used OpenAI models to train Grok, with distillation as the mechanism. Missing model names, scale, and runs keeps it in 78–84, not P1.

Apr 30Thursday

Latent Space

[AINews] The Inference Inflection

Latent Space argues inference demand has hit an inflection point, citing its Apr 28-29, 2026 AINews roundup. Jensen Huang is quoted saying per-task compute rose about 10,000x in two years, with usage up about 100x. The key watchpoints are CPU sandboxes, agent harnesses, and split inference workloads.

Why it matters: HKR-H/K/R all pass, but this is a Latent Space AINews roundup and trend read, not a model launch or major product release. It fits the upper featured-threshold band for insightful commentary.

Bloomberg Technology

Samsung’s Chip Profit Soars 48-Fold Due to AI Spending Spree

Samsung Electronics’ chip unit posted a 48-fold profit jump in the March quarter, driven by AI data-center orders. The RSS snippet says profit hit a record and beat expectations, but the post does not disclose profit value, memory type, or customers.

Why it matters: HKR-H/K/R all pass: Bloomberg reports a 48x chip-profit jump tied to AI data-center demand. I keep it at 74 because the body lacks profit amount, memory category, and customer detail.

r/LocalLLaMA

Building a fully local PDF-to-audiobook workflow with Kokoro 82M, Qwen and llama.cpp

Reddit user purellmagents shared a local PDF-to-audiobook workflow using Kokoro 82M, Qwen 3.5 0.8B/2B, and llama.cpp. The Tauri 2.0 app runs on an M1 Mac, reads 15 initial sentences, then prepares the next 15. The hard parts are PDF-text alignment, code snippets, tables, and first-generation latency.

Why it matters: HKR-H/K/R all pass, but this is a Reddit personal workflow, not a model or platform release. Specific components and the 15-sentence pipeline keep it at the low featured band.

Dwarkesh Patel podcast

Reiner Pope: The Math Behind How LLMs Are Trained and Served

Dwarkesh interviewed Reiner Pope in a 1-session blackboard lecture on LLM training and serving. The post lists 7 timestamps on batch size, MoE rack layout, pipeline parallelism, KV cache, and API pricing. The key mechanism is cost: without batching, serving economics can be 1,000x worse.

Why it matters: HKR-H/K/R all pass: the 1000x batching cost hook, concrete serving mechanics, and inference-cost resonance are strong. This is a high-quality tutorial, not a same-day industry event, so it stays at 77.

Apr 29Wednesday

Xinzhiyuan · WeChat

Google Translate Turns 20 as Pichai Highlights Four AI Generations

Google Translate turned 20 on April 28, and Pichai said it now has 1B monthly users. The post traces four AI phases: SMT, GNMT, PaLM 2, and Gemini 2.5 Flash Native Audio, including 110 languages added in 2024. The key shift is native speech-to-speech translation that preserves intonation, pacing, and pitch.

Why it matters: HKR-H/K/R all pass, but the core event is a Google Translate anniversary and architecture recap, not a clear launch. The 1B MAU, 110-language expansion, and native speech-to-speech detail justify featured at the 72–77 band.

Financial Times · Technology

How OpenAI’s $500bn Data Centre Venture Stargate Has Shifted Shape

OpenAI’s Stargate data centre venture is valued at $500bn. The RSS snippet says Sam Altman’s flexible infrastructure approach unsettles partners but boosts compute lead; the post does not disclose structure, partners, or timeline.

Why it matters: HKR-H and HKR-R pass: FT’s OpenAI $500bn Stargate angle has a clear compute-arms-race hook. HKR-K is weak because structure, partners, and timing are not disclosed, so it stays in the 72–77 band.

r/LocalLLaMA

Qwen3.6 27B on Dual RTX 5060 Ti 16GB with vLLM: ~60 tok/s, 204k Context Working

A user ran Qwen3.6 27B with vLLM on dual RTX 5060 Ti 16GB cards, reaching ~62–66 tok/s at 8K. The setup used 32GB VRAM, TP=2, fp8 KV cache, MTP 3 tokens, and a 204800 context window. The tight part is memory: after a 168k prefill, each GPU used ~15.65GiB with max_num_seqs=1.

Why it matters: HKR-H/K/R all pass: the post gives a concrete local-inference benchmark with hardware, vLLM settings, speed, and context limits. Single Reddit sourcing caps it below the 78–84 band.

X · @dotey

Microsoft VibeVoice-ASR tested on Mac for a one-hour podcast

Simon Willison ran 4-bit VibeVoice-ASR on an M5 Max MacBook Pro and transcribed a one-hour podcast in 8m45s. The 9B MIT-licensed model supports 60-minute audio, 50+ languages, and structured speaker output. Memory is the constraint: prefill peaked at 61.5GB, making 32GB laptops impractical.

Why it matters: HKR-H/K/R all pass: Simon Willison’s local test gives speed, parameter size, and memory peak that practitioners can act on. It is a single benchmark, not a fresh model launch, so it stays at the featured threshold.

Computing Life · Share · Yage

DeepSeek V4 Explained: Engineering Decisions Around Agentic Workloads

DeepSeek V4 targets long-horizon agent tasks with a 1M context. The snippet cites hybrid attention, OPD, Muon, and mHC; the post does not disclose size, data, pricing, or release timing.

Why it matters: HKR-H/K/R all pass: DeepSeek V4, 1M context, and agentic workload engineering create a strong hook with concrete mechanisms. Missing params, data, price, and launch timing keep it at 78, not P1.

Sinocism (Bill Bishop)

April Politburo Meeting, Manus Mess, and Possible New US Semiconductor Restrictions

China’s April Politburo meeting called for full implementation of the “AI+” initiative and listed computing power networks among six infrastructure networks. The readout signals no new stimulus, but stresses AI governance, supply-chain control, and rectifying involution-style competition.

Why it matters: HKR-H/K/R all pass, but the body gives policy signals without budget, timeline, or agencies. China AI infrastructure priority merits 76, not same-day must-write.

r/LocalLLaMA

XiaomiMiMo MiMo-V2.5: Sparse MoE with 310B total and 15B activated parameters

XiaomiMiMo shared MiMo-V2.5 with 310B total parameters and 15B activated parameters. The post only links Hugging Face and says it runs on more “human” configs than its larger sibling. It does not disclose VRAM needs, quantization, or benchmarks.

Why it matters: HKR passes: the 310B/15B Sparse MoE hook is concrete and relevant to local deployment. Detail is thin: the post links Hugging Face but gives no VRAM, quantization, or benchmarks, so it stays near the featured threshold.

Apr 28Tuesday

Synced · WeChat

ACL 2026: Huawei Taylor Lab Proposes SHAPE, Adding a Reasoning Tax to LLM Inference

Huawei Taylor Lab, Peking University, and Shanghai University of Finance and Economics proposed SHAPE, accepted by ACL 2026, with about 3% average accuracy gain. It uses entropy segmentation, short rollouts for potential estimation, dynamic length discounts, and token-level credit assignment, cutting token use by about 30%. The key mechanism is a reasoning tax: long high-potential late-stage segments are penalized to reduce verbose confirmation loops.

Why it matters: HKR-H/K/R all pass: the paper gives testable gains of about +3% math accuracy and -30% tokens, with concrete mechanisms. It is a strong research item, not a same-day model-launch story.

Latent Space

Physical AI that Moves the World — Qasar Younis & Peter Ludwig, Applied Intuition

Applied Intuition’s founders reviewed a 10-year physical AI path, with the company valued at $15B. The post cites 30+ products, 18 of the top 20 non-Chinese automakers as customers, and L4 driverless trucks in Japan. The key constraint is onboard deployment: millisecond latency, low power, small models, and safety validation.

Why it matters: HKR-H/K/R all pass: the piece ties a major Physical AI company to real AV deployment with customer, valuation, and L4 details. No new model or major launch is disclosed, so it stays in the 78–84 band.

Apr 27Monday

Hacker News front page

Show HN: Utilyze — an open-source GPU monitoring tool claiming higher accuracy than nvtop

Systalyze open-sourced Utilyze to measure real GPU compute efficiency in production, with negligible overhead claimed. The post says nvidia-smi and nvtop only check whether any kernel runs during the sampling window; an H100 has 132 SMs and 17,424 cores. The key issue is real throughput headroom, not binary utilization dashboards.

Why it matters: HKR-H/K/R all pass: the hook is sharp, the post explains the sampling flaw, and GPU waste is a real practitioner nerve. Unknown vendor and single-tool scope keep it in the 72–77 band.

Hacker News front page

Running Local LLMs Offline on a Ten-Hour Flight

Dmitri Lerko ran Gemma 4 31B and Qwen 4.6 36B locally during a 10-hour flight with no Wi‑Fi. The MacBook Pro M5 Max had 128GB unified memory and a 40-core GPU; sustained load used about 1% battery per minute, and performance degraded past 100k tokens. The sharp finding is instrumentation: an iPhone cable delivered 60W, while a MacBook cable delivered 94W under the same load.

Why it matters: HKR-H/K/R all pass: this is a named first-person local-inference test with concrete hardware, model, battery, and power numbers. Scope stays practical rather than industry-shaking, so it lands in the 72–77 band.

Mistral AI

Mistral AI opens public preview of Workflows

Mistral AI has put Workflows, its enterprise AI orchestration layer, into public preview. It offers durable execution, observability and human-in-the-loop approvals. ASML, ABANCA and CMA-CGM are already using it to automate critical processes.

Why it matters: It lays out Workflows' orchestration features, deployment model and customer cases, showing the engineering bar for enterprise AI processes.

Hacker News front page

AI can cost more than human workers now

Axios says some firms now spend more on AI than salaries; Nvidia's Bryan Catanzaro says compute costs exceed employee costs. Gartner forecasts 2026 IT spending at $6.31T, up 13.5%, driven by AI infrastructure, software, and cloud. Watch token costs: Uber's CTO has already exhausted the 2026 AI budget.

Why it matters: HKR-H/K/R all pass: the piece turns AI cost anxiety into budget facts, including Nvidia compute costs and Uber’s token-budget issue. It stays in the 72–77 band because this is trend reporting, not a launch or hard news event.

QbitAI · WeChat

DeepSeek V4 Cuts Prices Permanently; Cached Inputs Get 90% Off, Coding Test Costs Drop 83%

DeepSeek V4 cut prices twice in two days: input/output pricing is 75% lower, with cached inputs getting another 90% off. QbitAI’s coding test fell from 31.73 yuan for 35M tokens to 5.34 yuan under new pricing, an 83% drop. The key case is high cache-hit workloads, with V4-Pro at about 95–96% cache hits.

Why it matters: HKR-H/K/R all pass: DeepSeek V4 pricing has a sharp cost hook, concrete test numbers, and strong cost resonance. It is still a pricing update, not a new model release, so it stays below the 85 P1 band.

Hacker News front page

TurboQuant: A First-Principles Walkthrough

TurboQuant walkthrough explains compressing AI vectors to 2–4 bits per coordinate. It uses random rotation to map high-dimensional coordinates to a fixed distribution, then reuses one codebook with no scale overhead, training, or calibration.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the mechanisms are new, and the cost angle is relevant. It stays below 78 because this is a technical walkthrough, not a model or product release.

Apr 26Sunday

Hacker News front page

DeepSeek-V4 on Day 0: From Fast Inference to Verified RL with SGLang and Miles

SGLang and Miles added day-0 inference and RL support for DeepSeek-V4, covering 1.6T Pro and 284B Flash. The post cites a 1M-token context, FP4 MoE expert weights, 128-token SWA, and 4:1 or 128:1 KV compression. The key systems detail is ShadowRadix coherence across three KV pools and two compression-state pools.

Why it matters: HKR-H/K/R all pass: a DeepSeek-V4 day-0 systems stack, concrete context/compression mechanisms, and clear deployment-cost stakes. The systems depth narrows reach, but no hard-exclusion rule is triggered.

Apr 25Saturday

Latent Space

DeepSeek V4 Pro and Flash released, runnable on Huawei Ascend chips

DeepSeek released V4 Pro and V4 Flash, with 1.6T/49B active and 284B/13B active parameters. Both support 1M-token context, Base/Instruct variants, and an MIT license; the report claims 27% FLOPs and 10% KV cache versus V3.2 at 1M tokens. The key point is Huawei CANN compatibility, not just benchmarks, because it reduces CUDA dependence.

Why it matters: HKR-H/K/R all pass: a major DeepSeek release adds concrete specs, 1M context, MIT licensing, and Huawei Ascend support. This sits in the 85–94 must-write band, with hardware independence pushing it upward.

Computing Life · Share · Yage

TPU vs. CUDA: A Post-Cloud Next 2026 Assessment

Google announced TPU 8t/8i, TorchTPU, and an Anthropic deal at Cloud Next 2026; TPU 8i is slated for H2 2027 volume production. 8i has 288GB HBM, 8.6TB/s bandwidth, and 384MB SRAM; TorchTPU runs PyTorch on TPU, but the post says independent benchmarks are missing. The key crack is vLLM inference, while the author says TPU will not replace NVIDIA within 18-24 months.

Why it matters: HKR-H/K/R all pass: clear TPU-vs-CUDA rivalry, concrete 8i specs and TorchTPU details, and strong NVIDIA cost/supply resonance. No independent benchmark and H2 2027 production keep it in 78–84, not P1.

Financial Times · Technology

Google to invest up to $40bn in Anthropic

Google plans to invest up to $40bn in Anthropic to add computing power for running its models. The RSS snippet confirms the funds are tied to compute expansion; the post does not disclose deal structure, timing, valuation, or compute source. The key signal is compute lock-in, not just capital.

Why it matters: FT reports Google plans to invest up to $40bn in Anthropic, and the feed says the money is for compute expansion rather than a routine financial round. HKR-H/K/R all clear; structure, valuation, and timing are still undisclosed, so it lands in must-write territory, not 95+.

Apr 24Friday

TechCrunch · AI

In another wild turn for AI chips, Meta signs deal for millions of Amazon AI CPUs

Meta signed a deal for millions of Amazon-built AI CPUs for agentic AI workloads. The snippet confirms CPUs, not GPUs, and a scale of “millions”; the post does not disclose chip model, price, delivery timeline, or deployment details. The signal to watch is agent workloads pulling demand beyond GPUs.

Why it matters: Meta buying millions of Amazon AI CPUs is an unusual infra move, so HKR-H and HKR-R are strong. HKR-K clears because the story gives scale, chip class, and agentic-workload use, but model, price, delivery, and deployment details are undisclosed, so it stays in the 78–84 band.

Synced · WeChat

Remember more, answer faster, use less: HERMES speeds real-time streaming video understanding by 10x

Fudan University, Shanghai Academy of AI for Science, and NUS proposed HERMES, a training-free framework that turns KV cache into hierarchical memory for streaming video understanding and cuts TTFT by up to 10x. The post lists three mechanisms: hierarchical cache management, cross-layer memory smoothing, and position re-indexing; it reports 68% fewer video tokens with comparable or better results, and Qwen2.5-VL-7B on StreamingBench rising from 73.31% to 79.44%. What matters for practitioners: it answers without external retrieval, with TTFT around 27/29/28 ms at 16/64/256 frames.

Why it matters: Strong HKR-H/K/R: the 10x speed claim is a real hook, and the article includes concrete mechanisms and numbers, including 68% fewer video tokens and 27-29 ms TTFT. It stays below major product-news bands because this is an academic research release, not a market-moving launch.