Skip to content

#部署/工程

3 today

Jul 15Wednesday

Hacker News front page

Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

The author got Google's Gemma 4 26B MoE model running on a dual Xeon E5-2690 v2 server from 2013 with no GPU, costing under $300. The CPUs only support AVX1, but ik_llama.cpp's optimized kernels require AVX2, causing silent gibberish output. Claude diagnosed that the graph builder unconditionally emitted MOE_FUSED_UP_GATE ops while the dispatcher had no matching case, leaving ~240 tensors per forward pass reading uninitialized memory. After the fix, decode reaches ~5.2 tokens/sec and prompt eval ~16 tokens/sec. A PR is open but not yet merged. The post doesn't disclose quantized model memory usage or power draw.

Why it matters: A first-person experiment with real numbers, not a generic 'run LLMs locally' tutorial. Gemma 4 26B MoE on a 13-year-old Xeon, no GPU, sub-$300 total cost — every detail is concrete. HKR all hit, but it's a personal blog experiment, not a product launch or research breakthroug...

Jul 14Tuesday

Hacker News front page

MemStitch: Zero-copy KV cache stitching for vLLM cuts multi-agent TTFT by up to 25x

DaqulaLin open-sourced MemStitch, a gateway that sits in front of vLLM and stitches KV caches across requests at the memory level using PagedAttention. It skips redundant prefill in multi-agent workflows, cutting TTFT by up to 25x and saving over 40% VRAM. The post doesn't disclose the test model or GPU setup, so I'd discount the 25x claim until those details surface.

Why it matters: Multi-agent inference latency is a real pain point, and MemStitch's approach of KV cache stitching at the vLLM memory layer is more fundamental than prompt-level solutions. The 25x claim lacks disclosed model/GPU config, so the number gets a discount, but the mechanism itself ...

Jul 9Thursday

Hacker News front page

Colibrì: A pure-C engine that runs GLM 5.2 on a 32 GB laptop with no GPU

Colibrì is a single-file C engine that runs the 744B-parameter GLM 5.2 model on a 32 GB laptop with no GPU. The dense part stays in RAM at int4 (~9.9 GB), while 21,504 routed experts stream from disk on demand with an LRU cache. Cold-start speed is 0.1 tok/s. The author built and tested it on a 12-core, 25 GB machine. The post does not include quality benchmarks against the full-precision model.

Why it matters: A personal experiment with concrete numbers and a reproducible approach — not marketing fluff. Splits GLM 5.2's 744B parameters into a dense part resident in RAM and routed experts loaded from disk on demand, with int4 quantization bringing memory down to 9.9GB. Not scoring hi...

Jul 7Tuesday

AI HOT (Curated Pool)

SGLang Integrates DSpark: Confidence-Driven, Variable-Length Speculative Decoding

LMSYS integrated DSpark speculative decoding into SGLang. The draft model's confidence head decides how many tokens to verify per request, instead of always verifying a full block. On H200 with DeepSeek-V4-Flash, it beats both MTP and non-spec baselines in throughput and per-user decode speed. The key engineering move is a ragged, variable-length verify under full CUDA graphs: when the scheduler trims the verify budget, the engine replays a genuinely smaller graph, not a padded one. A cost-table profiler lets the scheduler size each request's budget online, and a cap-accept mode exposes the acceptance ceiling hidden by trimming.

Why it matters: LMSYS integrated DSpark into SGLang with real H200 benchmarks—not a paper announcement but an engine-level landing. H and K are solid, but the topic is inference optimization, which doesn't resonate broadly, so R is absent and the score stays at the featured threshold of 78.

Jul 5Sunday

Computing Life · Share · Yage

Scaling Law's three corrections in five years: from bigger models to smaller models with more data

Scaling law is an empirically fitted curve, not a physical law. OpenAI's 2020 Kaplan paper concluded 'prioritize parameters' due to experimental biases, shaping GPT-3. DeepMind's 2022 Chinchilla corrected the ratio to 20:1, showing smaller models with more data outperform. Two 2024 replication studies confirmed that fixing Kaplan's setup reproduces Chinchilla's result—no fraud, just calibration. Since 2023, Meta and others deliberately deviate from Chinchilla: Llama 3 8B was trained on 15T tokens because the optimization target shifted from training cost to total cost of training plus inference. Tsinghua's Densing Law shows the parameter count needed for equal capability halves roughly every 3.5 months, but there is a floor: each parameter stores only ~2 bits of knowledge. The viral 'collapse' article cited a blog comment posted the same day as if it were peer-reviewed research; the post does not provide a paper source for that claim.

Why it matters: A high-quality explainer and fact-check on scaling laws, debunking a recent viral post with specific numbers and paper citations while tracing three key revisions over five years. Hits all three HKR axes, but as commentary/education rather than a first-party product release, i...

Jun 30Tuesday

Hacker News front page

Moondream's Photon engine uses pipelined decoding to pop GPU bubbles, boosting decode throughput up to 35%

Moondream's Photon inference engine hits ~33ms VLM inference on NVIDIA B200. The bottleneck is GPU bubbles: the GPU idles while the CPU finishes bookkeeping between tokens. Photon pipelines the decode loop so the GPU starts the next forward before the CPU commits the current token, overlapping CPU housekeeping under GPU work. Three mechanisms make it safe: ping-pong slots prevent buffer collisions, forward-now-sample-later handles constrained decoding, and zombie cleanup deals with finished requests. The result is up to 35% higher decode throughput.

Why it matters: Moondream's Photon engine hits ~33ms VLM inference with 35% higher decode throughput — solid technical detail. But it's a single-company optimization, not an industry event, and the audience is narrow, so it lands right at the featured threshold.

Hacker News front page

vLLM Semantic Router beats frontier models by making multiple models collaborate inside one API call

vLLM's Micro-Agent makes multiple models collaborate behind a single API endpoint. Users call one model name as usual; the router decides whether to try a cheaper model first, run several in parallel and aggregate, or let models review each other. The core idea is turning one API call into a bounded micro-collaboration with budget, topology, and failure policy. The post describes five looper patterns—Confidence, Ratings, ReMoM, Fusion, and Workflows—and claims they beat individual frontier models on benchmarks, though it does not name the benchmarks or provide scores.

Why it matters: vLLM upgrades semantic routing into a model-collaboration engine with five concrete patterns and budget controls — highly relevant for inference engineers. Held back from 80+ because it's a blog post, not a shipped product, and performance numbers aren't fully disclosed in the...

Jun 19Friday

Hacker News front page

I restarted a 10-year-old Xeon 174 times to delete 12 flags and gain 4 tps

The author ablated all 25 flags from a previous Gemma 4 setup on a 2016 Xeon. Most did nothing. The real levers are flash attention, physical-core thread count, and workload-gated speculative decoding: on for code, off for long-doc summarization (54% faster without it). The auto-tune drafter setting is a crude per-request router that currently hurts long-context throughput.

Why it matters: A solid local-inference tuning experiment that ablates 25 launch flags one by one, delivering scene-dependent numbers on speculative decoding and the finding that most flags are placebo. Hits H and K, but narrow audience and zero industry conversation value mean R misses — lan...

Jun 14Sunday

r/LocalLLaMA

Xiaomi serves MiMo V2.5 at 1000–3000 tps with DFlash and Persistent Kernel

Xiaomi's MiMo V2.5 is live, claiming 1000–3000 tps via DFlash and Persistent Kernel. The DFlash model weights are out, and an open-source release is promised soon. The post body is blocked by Reddit security, so only the headline is available—no details on measured latency, concurrency, or hardware. I'd discount that tps figure: headline peaks usually assume optimal batching, and real single-user throughput is likely lower.

Why it matters: MiMo V2.5's claimed 1000-3000 tps and the two named acceleration mechanisms (DFlash, Persistent Kernel) carry real information density; weights are out and open-source code is promised, directly relevant to local model deployers. Score capped because the Reddit body was blocke...

Jun 13Saturday

Hacker News front page

Can I Buy Your KV Cache?

This paper proposes letting publishers precompute a document's KV cache so AI agents can buy and load it, skipping the most compute-heavy step: prefill. On Qwen3-4B, reuse is 9–50x cheaper than prefill with zero accuracy loss—token outputs match exactly. Shipping the KV cache fails because it's nearly incompressible and egress costs more than the prefill saved. The fix: host it provider-side, like production prompt caching. Serving one 3,774-token document to 80M agents costs ~$1.5M to re-prefill but only ~$30K via reuse, a 49.7x gap. The paper frames this as an agent-native prefill CDN and leaves lossless KV compression and cross-party payments as open problems.

Why it matters: Selling precomputed KV caches is a practical idea with a 9–50× cost gap and zero accuracy loss. Held back by single-model experiments (Qwen3-4B only) and no detail on cache security or pricing in the excerpt.

Jun 12Friday

r/LocalLLaMA

MiniMax open-sources MSA, a sparse attention method that cuts attention compute by 28.4× at 1M tokens on a 109B model

MiniMax published a paper introducing MSA, a blockwise sparse attention built on GQA. A lightweight index branch scores KV blocks and picks a top-k subset per GQA group, then the main branch runs exact attention only on those blocks. With a co-designed GPU kernel, a 109B-parameter multimodal model achieves 14.2× prefill and 7.6× decoding wall-clock speedups on H800 at 1M context, matching full GQA quality. Code and inference kernel are open-sourced, along with a model called MiniMax-M3. The Reddit poster is curious whether the 109B model can run on consumer GPUs; the post doesn't say if weights will be released.

Why it matters: The paper has concrete mechanisms and measured numbers, not just theory—real knowledge for inference-optimization folks. But the audience is narrow (R missed), and the low-level CUDA details raise the accessibility bar for generalist readers, so I docked 3 points, landing righ...

Jun 10Wednesday

AI HOT (Curated Pool)

OpenRouter Launches Advisor Tool for Low-Cost Models to Consult Stronger Models

OpenRouter released the Advisor server tool, letting GPT-4o Mini consult Claude Fable during generation, but the post does not disclose pricing, latency, or the routing policy.

Why it matters: HKR-H/K/R all pass: OpenRouter turns cheap-model plus strong-model advising into a callable server tool. Price, latency, and call policy are not disclosed, so this stays in the upper mid-weight product-update band.

r/LocalLLaMA

ICML paper on predictable hallucination gate and ntkMirror open-weight implementation

An ICML 2026 paper presents an ISR=1 answer-abstain gate for evidence-grounded QA, and ntkMirror implements it for local open-weight models with multiple evidence orderings, reporting 0.0–0.7% hallucination at about 24% abstention in the held-out audit.

Why it matters: HKR-H/K/R all pass: an ICML paper with an open implementation, a concrete ISR=1 gate, and measured abstention-vs-hallucination tradeoff. Scope stays within evidence QA/RAG reliability, so it sits below must-write level.

Jun 9Tuesday

AI HOT (Curated Pool)

Google Releases Gemini 3.5 Live Translate for Real-Time Speech Translation

Google released Gemini 3.5 Live Translate, a speech-to-speech translation model that supports more than 70 languages, starts translating before the speaker finishes, uses streaming updates, and runs through Gemini Live API, Google Meet preview, and Google Translate apps on iOS and Android.

Why it matters: HKR-H/K/R all pass: Google ties real-time speech translation to 70+ languages and streaming output before the speaker finishes. It stays at 82 because rollout scope, pricing, and benchmarks are not disclosed.

AI HOT (Curated Pool)

Google DeepMind Releases Gemma 4 12B, a Unified Encoder-Free Multimodal Model

Google DeepMind released Gemma 4 12B, a multimodal model with a unified encoder-free architecture, native audio input, Apache 2.0 licensing, and local laptop runtime with 16GB of VRAM or unified memory.

Why it matters: HKR-H/K/R all pass: the hook is local multimodal audio in 16GB VRAM, and the new architecture is concrete. It is a strong Google DeepMind open-model release, but not a frontier-model launch, so it stays below p1.

r/LocalLLaMA

Apple Announced New On-Device Inference Engine for Apple Silicon

Apple announced CoreAI at WWDC as a future CoreML replacement for Apple Silicon on-device inference; models require Python-script conversion, the supported list is mostly mid-2025 models, and the post does not disclose performance data.

Why it matters: HKR-H/K/R pass, but the post is thin: CoreAI, CoreML successor status, and Python conversion are disclosed; throughput, latency, and model coverage are not. Apple on-device inference merits featured, capped in the 72–77 band.

AI HOT (Curated Pool)

China Prepares $295 Billion Plan to Fund Nationwide AI Infrastructure Buildout

China plans to invest about 2 trillion yuan, or $295 billion, over five years to build nationwide data centers, with funding covering large-scale data center infrastructure for domestic AI development.

Why it matters: Bloomberg reports China is preparing a five-year RMB 2T AI data-center plan, clearing HKR-H/K/R. This is national compute supply and geopolitical competition news, not routine policy; the preparation status keeps it at 90.

AI HOT (Curated Pool)

Xiaomi MiMo and TileRT Release UltraSpeed Mode, 1T Model Exceeds 1,000 Tokens/s

Xiaomi MiMo and TileRT released MiMo-V2.5-Pro-UltraSpeed, a 1T-parameter model mode exceeding 1,000 tokens/s, with API access open from June 9 to June 23, 2026, at 3× the MiMo-V2.5-Pro price and about 10× the speed.

Why it matters: HKR-H/K/R all pass, with a domestic flagship-model bump for Xiaomi. Missing hardware, batch, concurrency, and test conditions keep it in the 78-84 band rather than p1.

AI HOT (Curated Pool)

Elon Musk Details SpaceX's AI1 Orbital AI Data Center Satellite Plan

Elon Musk detailed SpaceX’s AI1 orbital AI data center satellite plan, with 150 kW peak power per satellite, about 120 kW sustained compute power, and 6-8 ms round-trip latency at 600-800 km low Earth orbit.

Why it matters: HKR-H/K/R all pass: the angle is unusual and the post gives power, orbit, and latency figures. It stays near the featured floor because launch timing, cost, and workload tests are not disclosed.

r/LocalLLaMA

2X tk/s on 1× MI50: Qwen3.6-27B inference rises from 19.4 to 38.1 tk/s

bigattichouse raised Qwen3.6-27B throughput on a single MI50 from 19.4 to 38.1 tk/s by running same-model parallel computations for Q8-or-lower quantization, exploiting unused compute lanes instead of adding a smaller speculative decoding model.

Why it matters: HKR-H/K/R all pass, but this is a Reddit first-person experiment with numbers and a hypothesis, not a validated release. No code or broader replication is disclosed, so it stays at the featured threshold.

r/LocalLLaMA

New MLX LM Server From Apple

A Reddit post says Apple’s MLX LM Server uses continuous batching for concurrent sub-agent requests and supports distributed inference across multiple Macs via Thunderbolt RDMA.

Why it matters: HKR-H/K/R all pass, but the item is based on a Reddit summary and lacks throughput, latency, model-size, or release details. Treat it as a mid-weight Apple/MLX inference update, just above the featured threshold.

r/LocalLLaMA

Levi: Run AlphaEvolve on Your Local Qwen 30B

LEVI runs an AlphaEvolve-like search system with Qwen3-30B-A3B and reports tests on ADRS, IFBench, and HotpotQA, claiming up to 35x lower cost overall and up to 12x fewer evals under the same single-model, same-budget comparison.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post with model, benchmarks, and cost ratios only; code maturity and reproducibility details are not disclosed. Scores as a strong open-source agent/inference item, not a major release.

Jun 8Monday

r/LocalLLaMA

Xiaomi claims 1,000+ tps on a 1T model using a standard 8-GPU server

Xiaomi MiMo claims MiMo-V2.5-Pro UltraSpeed runs a 1T-parameter MoE model above 1,000 output tokens per second on one standard 8-GPU node; the post does not disclose the GPU model, batch settings, or reproducible configuration.

Why it matters: HKR-H/K/R all pass: the 1T MoE and 1,000+ tps claim is a strong inference-cost hook. Kept below P1 because the post lacks GPU model, batch size, quantization, and reproducible setup.

Hacker News front page

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

The title says Xiaomi MiMo-v2.5-Pro-UltraSpeed is a 1T model running at 1,000 tokens per second; the RSS body only provides the URL, Hacker News comments link, 66 points, and 14 comments, and the post does not disclose hardware, precision, context window, benchmark setup, or availability.

Why it matters: HKR-H/K/R all pass: Xiaomi’s MiMo update has a sharp 1T/1,000 tokens/s claim and clear cost-speed resonance. Missing hardware, precision, context window, and test setup keep it in the 78–84 band, not p1.

r/LocalLLaMA

Luce Spark: a 35B MoE on a 16 GB GPU, without the offload tax

Luce Spark runs Qwen3.6 35B-A3B at 13.3 GiB peak VRAM on an RTX 3090 by keeping hot experts on GPU, swapping cold experts through a bounded async cache, and using one fused graph for decode at about 100 tok/s.

Why it matters: HKR-H/K/R all pass: the hook is a 35B MoE on a 16 GB GPU, with 13.3 GiB peak use and ~100 tok/s. Reddit-source and no third-party replication keep it at 78.

AI HOT (Curated Pool)

Xiaomi MiMo-V2.5-Pro-UltraSpeed Exceeds 1,000 Tokens/s

Xiaomi MiMo and TileRT_AI released MiMo-V2.5-Pro-UltraSpeed, running a 1T MoE model above 1,000 tokens/s on a single standard 8-GPGPU node, with UltraSpeed API priced at 3x and applications open from June 8 to 23 PDT.

Why it matters: HKR-H/K/R all pass: Xiaomi MiMo gives a concrete claim of a 1T MoE exceeding 1,000 tokens/s on one 8-GPGPU node. The score stays at 80 because this is single-source and lacks task mix, precision, latency, and cost details.

r/LocalLLaMA

DFlash Speculative Decoding and KV Cache Compression on RTX 5090 Show 3.26x Speedup

The author tested Qwen3.6-27B on an RTX 5090 with DFlash plus KV cache compression, reaching up to 3.26x speedup; q4_0/turbo4 delivered 3.18x speedup with only +0.02% PPL on WikiText-2.

Why it matters: HKR-H/K/R all pass: RTX 5090 testing, DFlash speculative decoding, KV cache compression, 3.26x speedup, and PPL delta are concrete. Single Reddit source keeps it near the featured floor.

r/LocalLLaMA

Weird to get near-linear scaling by adding another GPU?

A Reddit user benchmarked qwen3.6-27b-autoround-int4 on 1x3090 versus 2x3090. Narrative decode rose from 53 TPS to 94 TPS, and code decode rose from 62 TPS to 120 TPS, under no NVLink, 8x/8x PCIe, P2P automatically enabled, tensor parallelism set to 2, and different KV-cache settings.

Why it matters: HKR-H/K/R all pass: the result is counterintuitive, includes concrete TPS and TP conditions, and speaks to local-inference cost. Single Reddit test lacks multi-model replication and full setup details, so it stays near the featured threshold.

Synced · WeChat

openJiuwen proposes MANGO for multi-agent flow networks

openJiuwen proposed MANGO, a multi-agent flow-network framework that combines reinforcement learning, textual gradients, and a Skip-k mechanism; using GPT-4o-mini, it reports a 12.8% accuracy gain over MaAS on MATH500 and a 5.1% F1 gain over AFlow on DROP.

Why it matters: HKR-K is strong: the post gives mechanisms and a MATH500 delta. HKR-H/R pass for the multi-agent flow-network angle, but this remains a research-framework story, not a major model or platform release.

Synced · WeChat

Alibaba RTPurboV2 uses hundreds of training steps for 10x sparse attention

Alibaba’s RTP team released RTPurboV2, replacing 85% of attention heads with SWA and compressing the remaining 15% retrieval heads using low-rank projection, clustering, and dynamic top-p; the adaptation uses about 600 training steps and roughly 1M label tokens, with reported Prefill speedup up to 9.36x.

Why it matters: HKR-H/K/R all pass: the hook is 100-step training and 10x sparse attention, with concrete mechanisms and 9.36x Prefill speedup. Strong engineering signal from Alibaba, but not a flagship model release or major product launch.

AI HOT (Curated Pool)

Apple Releases Third-Generation Apple Foundation Models (AFM)

Apple released its third-generation AFM family with five models. The RSS snippet says they span on-device use and Private Cloud Compute servers, with Google involved in customization for Apple Intelligence, Siri, and system-level tools.

Why it matters: Official Apple model-family release with 5 models, on-device/PCC deployment, and Google customization clears HKR-H/K/R. Missing benchmark and pricing details keep it at the low end of the 85+ band.

AI HOT (Curated Pool)

Nvidia and SK Hynix Sign Multi-Year Pact to Develop Next-Generation AI Memory Chips

Nvidia and SK Hynix signed a multi-year pact to co-design future generations of memory chips for AI applications; the RSS snippet does not disclose product specifications, production timelines, or financial terms.

Why it matters: HKR-H and HKR-R pass: Bloomberg plus Nvidia/SK Hynix matters for AI memory supply. HKR-K is weak because specs, production timing, and financial terms are missing, so this sits at the low featured band.

Jun 7Sunday

r/LocalLLaMA

Qwen3.6 35B-A3B on a Laptop: My Zero-to-One Moment

A Reddit user ran Qwen3.6 35B-A3B on an ASUS Zenbook Pro 14 with RTX 4060 8GB VRAM and 64GB RAM, reaching about 27 TPS at 32k context and 18 TPS at 256k context. The setup uses llama.cpp, unsloth’s IQ3_XXS GGUF quantization, and a 262144-token context flag.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit experiment, not an official release or paper. Concrete hardware, quantization, context, and TPS clear the featured bar, but keep it in the 72–77 band.

QbitAI · WeChat

Chinese open-source framework targets stable 5-minute AI long-video generation

JD open-sourced JoyAI-Echo, a long audio-video generation framework for 5-minute consistent videos, using cross-modal memory, DMD post-training for about 7.5x faster inference, and real-time upscaling from 720P to 1K or 2K output.

Why it matters: HKR-H/K/R all pass: the story has a clear 5-minute video hook, concrete speed and SR claims, and open-source competition resonance. Missing third-party evaluation keeps it in the lower 78–84 band.

AI HOT (Curated Pool)

AI Substitution Wave: Three Forces Reshape Cost Structures

Coinbase, Lindy, Harvey, and Cursor shifted workloads to cheaper models; Harvey reported Kimi 2.6 reached a 15% all-pass rate on Legal Agent Benchmark, versus Opus at 14%, with 100 tasks costing $84 versus $954.

Why it matters: HKR-H/K/R all pass: the $84 vs $954 cost delta and named cases from Coinbase, Lindy, Harvey, and Cursor give it concrete signal. It is a strong cost-structure commentary, not a major model or product release, so it fits the 72-77 band.

Jun 6Saturday

AI HOT (Curated Pool)

OpenCV 5 Released with New DNN Engine and Native LLM Support

OpenCV 5 introduces a graph-based DNN engine, raising ONNX operator coverage from under 23% in 4.x to over 80%, with native support for Transformer, VLM, and LLM workloads.

Why it matters: HKR-H/K/R all pass for a substantive OpenCV major release: graph DNN engine, ONNX coverage jump, and native Transformer/VLM/LLM support. Strong featured item, but below must-write model-lab release territory.

r/LocalLLaMA

Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding

Domino reports up to 5.8x throughput speedup on Qwen3 by decoupling causal modeling from autoregressive drafting in speculative decoding. The Reddit snippet links the arXiv paper, GitHub code, and Hugging Face models, but does not disclose hardware, baseline settings, dataset, or acceptance-rate details.

Why it matters: HKR-H/K/R all pass: 5.8x throughput is a concrete hook with open artifacts. Missing hardware, baseline config, and task set keep it in the good featured band, not same-day must-write.

Computing Life · Share · Yage

Google pays SpaceX $920M a month for GPUs, but compute is not the main story

Google pays SpaceX $920 million per month for GPU rentals, and the post says the contract includes 11% GPU utilization, a 90-day cancellation clause, and methane gas turbines used to bypass environmental approval.

Why it matters: HKR-H/K/R all pass: the deal size, utilization term, and energy workaround are concrete. I keep it below P1 because the provided item is a single-source summary with no contract file or cross-source confirmation.

AI HOT (Curated Pool)

Building a Multi-Agent Economy with Qwen2.5-3B: Engineering Report

A developer used Qwen2.5-3B to build a five-agent forest economy, and across 15 simulation rounds honey prices fell from 10 to 3, firewood rose from 4 to 7, and the Gini coefficient increased from 0.14 to 0.38.

Why it matters: HKR-H/K/R pass: the 3B multi-agent economy has a hook and concrete price/Gini results. It remains a single engineering experiment, not a product or framework launch, so it stays at the featured floor.

AI HOT (Curated Pool)

SpaceX and Google Reach New Cloud Computing Agreement

SpaceX disclosed a cloud services agreement with Google: Google will pay SpaceX $920 million per month for computing capacity tied to xAI data centers, while the post does not disclose contract duration, GPU scale, or delivery terms.

Why it matters: HKR-H/K/R all pass: the hook is a Google–SpaceX–xAI compute triangle, with $920M/month as the concrete fact. The single-post source and missing contract term, delivery scale, and filing details keep it at low P1.