Skip to content

NVIDIA chips and ecosystem: new GPUs, CUDA, robotics platforms and the market for AI compute.

Latest picks

221–240 of 290

May 14Thursday

r/LocalLLaMA

2x RTX 3090 setup for local Qwen 3.6 27B inference

A Reddit user ran Qwen 3.6 27B on a dual RTX 3090 Ubuntu setup, reporting 48GB VRAM, a 262k context window, no NVLink, about 4000 pp/s prompt processing, and 113 tk/s generation.

Why it matters: All HKR axes pass, and this is a first-person local-inference run with concrete numbers. Source is a single Reddit post with limited reproducibility detail, so it sits at the low featured threshold.

Bloomberg Technology

Anthropic Eyeing Over $900 Billion Valuation

Bloomberg says Anthropic is seeking at least $30 billion in new financing at a valuation above $900 billion; the post does not disclose investors, deal terms, or a timeline.

Why it matters: Bloomberg discloses only the $30B raise target and $900B+ valuation, with no investors, terms, or timeline. HKR-H/K/R all pass; the unclosed status keeps it in the lower 85–94 band.

May 13Wednesday

New York Times Chinese

Jensen Huang Gets Last-Minute Invitation to Join Trump’s China Trip

Trump called Jensen Huang on Tuesday morning to invite him to join the China trip; the White House’s Monday list of 16 CEOs did not include him, while Nvidia is still seeking approval to sell AI chips to China.

Why it matters: HKR-H/K/R all pass: NYT reports a last-minute Jensen Huang invite tied to Nvidia’s China AI-chip license push. No disclosed policy change or license outcome, so this stays near the featured threshold.

New York Times Chinese

China Seeks AI Technology Self-Reliance, Weakening Washington’s Leverage Over Beijing

DeepSeek optimized its latest model for inference on Huawei chips for the first time, while two semiconductor sources said training still relies on Nvidia chips; Huawei says it plans to release a training chip this year, but matching current Nvidia performance will take another year.

Why it matters: HKR-H/K/R all pass: NYT ties DeepSeek-Huawei chip optimization and Huawei's training-chip timeline to US export-control leverage. It is not a model launch and lacks benchmark results, so it stays in the 78–84 band.

May 11Monday

QbitAI · WeChat

OpenAI backs Cerebras as the Nvidia challenger targets a $35B IPO valuation

Cerebras raised its IPO price range to $150-$160 per share, targeting about a $35 billion valuation at the top end, after OpenAI signed a 750-megawatt AI compute purchase agreement with deliveries through 2028.

Why it matters: HKR-H/K/R all pass: this is not a routine IPO note, since OpenAI’s 750MW purchase agreement anchors Cerebras at a reported $35B valuation and feeds the NVIDIA-alternative compute story.

May 10Sunday

r/LocalLLaMA

NVIDIA AI Releases Star Elastic: One Checkpoint Contains 30B, 23B, and 12B Reasoning Models

NVIDIA AI released Star Elastic, a single checkpoint that can zero-shot slice 30B, 23B, and 12B reasoning models in BF16, FP8, and NVFP4; when the 23B submodel handles thinking and the 30B model handles final answers, reported accuracy rises 16% and latency drops 1.9× on AIME-2025, GPQA, LiveCodeBench v5, and MMLU-Pro.

Why it matters: HKR-H/K/R all pass: Star Elastic has a concrete mechanism and testable numbers for inference deployment. Its reach is still narrower than a frontier-model release, so it sits in the high-quality featured band.

May 9Saturday

TechCrunch · AI

Nvidia has already committed $40B to equity AI deals this year

Nvidia has committed more than $40 billion to equity investments in AI companies in 2026, including a $30 billion investment in OpenAI, seven multi-billion-dollar public-company deals, and around two dozen private startup rounds, according to CNBC and FactSet data cited by TechCrunch.

Why it matters: HKR-H/K/R all pass: the $40B hook is strong, with $30B to OpenAI and about 24 deals disclosed. It is a capital-structure signal, not a model or product launch, so it sits in the 78–84 band.

Xinzhiyuan · WeChat

NVIDIA, AMD, and Intel Back RadixArk’s $100M Seed Round

RadixArk announced a $100 million seed round on May 5 at a $400 million post-money valuation, led by Accel and co-led by Spark Capital, with participation from NVentures, AMD, MediaTek, Databricks, and other investors tied to AI infrastructure.

Why it matters: HKR-H/K/R all pass: a $100M seed round, $400M post-money valuation, and chip/data investors create an AI-infra rivalry angle. It remains a single-company funding item with no product benchmarks or customer data, so it sits near the featured floor.

Synced · WeChat

StarVLA Open-Sources a Unified VLA Framework from HKUST and the Community

HKUST and the open-source community released StarVLA, a unified Vision-Language-Action framework that integrates backbones, action heads, training strategies, and evaluation interfaces; the repository has 2.2k GitHub stars and supports benchmarks including LIBERO, SimplerEnv, RoboTwin 2.0, RoboCasa-GR1, and BEHAVIOR-1K.

Why it matters: HKR-H/K/R all pass: StarVLA ships a concrete open-source VLA framework with unified interfaces, 2.2k stars, and named robotics benchmarks. The robotics scope keeps it in the 78–84 band, below model-release weight.

r/LocalLLaMA

MTP + TurboQuant Running: Qwen3.6-27B Hits 80+ t/s on a Single RTX 4090

indrasmirror ran Qwen3.6-27B-Heretic-v2 on a single RTX 4090 with 262K context, TBQ4_0 KV cache, and MTP draft 3, improving throughput from about 43 t/s to 80-87 t/s with roughly 73% MTP draft acceptance.

Why it matters: HKR-H/K/R all pass, backed by a numbered first-person experiment. The Reddit-only source and niche local-inference focus keep it below the 78–84 band for broader industry releases.

May 8Friday

Synced · WeChat

ICLR 2026: NVIDIA and Purdue Use an Agentic Loop for Text-to-3D Scene Generation

NVIDIA Cosmos Lab and Purdue University proposed Scenethesis, a language-and-vision agentic framework for text-to-3D scene generation that uses visual grounding, SDF-based physical constraints, and a judge module; experiments report about 72% first-pass success, 91% after self-checking, and collision rate reduction from 6.1% to 0.8%.

Why it matters: HKR-H/K/R all pass: NVIDIA/Purdue plus an agent loop is clickable, and the post gives SDF constraints, a judge module, and 72%→91% results. Strong research signal, but not a product release, so it stays in 78–84.

Bloomberg Technology

US Said to Suspect Nvidia Chips Smuggled to Alibaba Via Thailand

The US suspects a company tied to Thailand’s national AI effort helped smuggle billions of dollars of Super Micro servers with advanced Nvidia chips to China, with Alibaba named as one of multiple end customers, according to people familiar with the matter.

Why it matters: HKR-H/K/R all pass: Bloomberg reports a specific alleged route, dollar scale, and Alibaba link. The claim remains a US-suspicion report without enforcement outcome or cross-source confirmation, so it stays below P1.

Bloomberg Technology

Nvidia to Invest Up to $2.1 Billion in Data Center Firm IREN

Nvidia will invest up to $2.1 billion in IREN under an AI infrastructure partnership. The post discloses the cap and goal, but not equity terms, payment timing, or data center capacity.

Why it matters: Bloomberg source plus Nvidia’s up-to-$2.1B investment clears HKR-H/K/R. Details stop at amount and partnership direction, with no stake, payment schedule, or capacity, so it stays at the featured threshold.

NVIDIA Blog

Powering the Next American Century: Chris Wright and NVIDIA’s Ian Buck on Genesis Mission

The U.S. DOE and NVIDIA are building two AI supercomputers at Argonne; Equinox uses 10,000 Grace Blackwell GPUs. Solstice will use 100,000 Vera Rubin GPUs, which Buck said reach 5,000 exaflops. The key bottleneck is grid work: Wright said AI can cut interconnection studies from years to weeks or hours.

Why it matters: HKR-H/K/R all pass: the GPU counts, DOE-NVIDIA role, and grid bottleneck are concrete. NVIDIA-blog sourcing keeps it below must-write; this fits the 78–84 band.

May 7Thursday

r/LocalLLaMA

GB10 inference engine Atlas is open source, with Qwen3.6-35B-FP8 over 100 tok/s

Avarok open-sourced Atlas, an inference engine running Qwen3.5-35B at ~111 tok/s sustained on one DGX Spark. It uses Rust+CUDA, a ~2.5GB image, and sub-2-minute cold start; the author claims 3.0–3.3x vLLM in tests. The key details are Blackwell SM120/121 kernels, NVFP4/FP8, and MTP decoding.

Why it matters: HKR-H/K/R pass: open-source inference engine, 35B FP8 at 111 tok/s, and a direct vLLM comparison. Single Reddit sourcing and unreproduced benchmarks keep it at the lower featured band.

May 6Wednesday

r/LocalLLaMA

Qwen3.6 27B NVFP4 + MTP on a Single RTX 5090: 200k Context in vLLM

A Reddit user ran Qwen3.6 27B NVFP4 on one RTX 5090 32GB and validated 200k context in vLLM. The setup used fp8_e4m3 KV cache, FlashInfer, and MTP with 3 speculative tokens; a 10-run 200k pass completed with 73.6 tok/s mean generation and 70.2s TTFT. The key constraint is 32GB VRAM: logs showed 8.3GiB KV cache and about 30478MiB total GPU use.

Why it matters: HKR-H/K/R all pass: the hook is single-GPU 200k context, with concrete vLLM settings and 10-run stability data. Reddit sourcing keeps it in the 78–84 band, not P1.

NVIDIA Blog

NVIDIA Spectrum-X AI-Native Ethernet Fabric Adds MRC for Gigascale AI

NVIDIA added MRC support to Spectrum-X Ethernet, letting one RDMA connection spread traffic across multiple paths. MRC ran in Blackwell deployments, with microsecond failure bypass and hardware rerouting. The key detail is the OCP open specification and multiplane support for clusters up to hundreds of thousands of GPUs.

Why it matters: HKR-K/R are solid: MRC stripes one RDMA flow across paths, detects failures in microseconds, and is tied to Blackwell deployments. HKR-H is narrow and the source is vendor-owned, so this stays below major release level.

TechCrunch · AI

SAP Bets $1.16B on 18-Month-Old German AI Lab and Says Yes to NemoClaw

SAP plans to buy 18-month-old German AI startup Prior Labs in a $1.16B bet. The RSS snippet says SAP restricts customer agent use to a few options such as Nvidia NemoClaw; the post does not disclose deal structure, closing date, or technical details.

Why it matters: HKR-H/K/R all pass: $1.16B for an 18-month-old AI lab is a strong enterprise-AI hook. Kept at 76 because deal structure, closing timeline, and technical details are not disclosed.

NVIDIA Blog

NVIDIA and ServiceNow Partner on Autonomous AI Agents for Enterprises

NVIDIA and ServiceNow expanded their partnership with Project Arc, an enterprise desktop agent. It connects via Action Fabric and uses OpenShell for sandboxed, policy-governed execution. Blackwell delivers over 50x Hopper’s token output per watt and nearly 35x lower cost per million tokens.

Why it matters: HKR-K/R pass: the post gives mechanisms and Blackwell economics. HKR-H misses because the angle is a standard vendor partnership, so this sits in the 72–77 featured-threshold band.

May 5Tuesday

Synced · WeChat

Massive Idle Cluster: Musk’s 550,000 Nvidia GPUs Are Only 11% Utilized

The Information says xAI’s roughly 550,000 Nvidia GPUs have only 11% MFU, equal to about 60,000 effective GPUs. The post cites HBM I/O, inter-server communication, training idle time, and software-stack inconsistency; Meta and Google are listed at 43% and 46%.

Why it matters: HKR-H/K/R all pass: the 550k-GPU versus 11% MFU contrast is strong, with concrete efficiency numbers and bottlenecks. This is high-signal infra reporting, not a model or product release, so it fits 78–84.