Skip to content

NVIDIA chips and ecosystem: new GPUs, CUDA, robotics platforms and the market for AI compute.

Latest picks

201–220 of 290

May 20Wednesday

Xinzhiyuan · WeChat

Behind Jensen Huang’s Douzhi Moment, Chinese GPUs Are Filling CUDA’s Moat

Moore Threads presented progress on its MUSA GPU ecosystem, with SDK 5.1.0 targeting CUDA 12.8 and supporting 761 driver and runtime APIs. The post says MUSA has entered SGLang’s mainline, is listed for 2026 Q2 hardware support, and supports automated library migration via MUSACODE.

Why it matters: HKR-H/K/R all pass: the headline has a meme hook, the post gives 761 APIs plus SGLang mainline support, and CUDA-lock-in anxiety is real. It remains a single-vendor ecosystem update, so it sits in mid featured rather than P1.

Financial Times · Technology

Nvidia’s Huang Bankrolls AI Boom with $90bn Deal Spree

Nvidia’s Jensen Huang is funding an AI deal spree worth $90bn; the RSS snippet says the chipmaker’s spending rivals Big Tech’s largest venture operations and ties customers and start-ups to Nvidia’s technology.

Why it matters: HKR-H/K/R all pass: FT gives a $90bn figure and a customer-lock-in mechanism around Nvidia’s AI financing web. Strong industry analysis, but not a model launch or major product release, so it stays at the top of 78–84.

New York Times Chinese

China Wants to Expand AI, but Not at the Cost of Jobs

Chinese courts have publicly backed workers in AI-replacement dismissal cases for a third time; in the Hangzhou case, an employee’s monthly salary was cut from 25,000 yuan to 15,000 yuan before dismissal, and the court found the employer failed to provide a reasonable placement and had illegally terminated the worker.

Why it matters: HKR-H/K/R all pass: the NYT story links China’s AI push with labor-court precedents and gives concrete Hangzhou case numbers. It is not a model or product launch, so 78–84 fits better than the same-day must-write band.

r/LocalLLaMA

Running DeepSeek-V4 locally on 4 legacy RTX 2080 Ti GPUs with W8A8 at 255 prefill tok/s

A Reddit user ran DeepSeek-V4-Flash locally on 4 RTX 2080 Ti GPUs, reporting 284B total parameters, 13B active parameters, a sub-$2,500 build, custom Turing CUDA kernels, W8A8 quantization, 1TB DDR4 ECC RAM, and about 255 prefill tokens/s.

Why it matters: HKR-H/K/R all pass: this is a numeric first-person local-inference experiment. Single-source Reddit provenance and custom Turing kernels keep it in the lower featured band.

r/LocalLLaMA

Nemotron-Labs-Diffusion from NVIDIA

NVIDIA released the Nemotron-Labs-Diffusion 3B, 8B, and 14B dense model family with AR decoding, diffusion parallel decoding, and self-speculation; the 8B model reaches 850 tok/s on GB200 at concurrency 1, compared with 253 tok/s for AR and 360 tok/s for Eagle3.

Why it matters: HKR-H/K/R all pass: NVIDIA diffusion LLMs, concrete sizes/mechanisms, and an 850 tok/s GB200 claim. Single-source Reddit sourcing keeps it in the 78–84 band, not P1.

AI HOT (Curated Pool)

NVIDIA open-sources first 4-bit infrastructure for ultra-long video generation

NVIDIA researchers open-sourced LongLive 2.0, an end-to-end long-video generation infrastructure covering training and inference with 4-bit quantization, FP4 quantization, parallel acceleration, KV-cache optimization, and 45.7 FPS generation on a 5B model.

Why it matters: HKR-H/K/R all pass: NVIDIA researcher open-sources LongLive 2.0 with 4-bit long-video train/inference and 45.7 FPS on a 5B model. This is strong open-source infra, not a flagship model launch, so it fits the 78–84 band.

May 19Tuesday

Bloomberg Technology

Self-Improving AI Startup Recursive AI Valued at $4.65B

Recursive came out of stealth at a $4.65 billion valuation, building AI that runs experiments on safe self-improvement, with backers including Google Ventures, Greycroft, Nvidia, and AMD Ventures.

Why it matters: HKR-H/K/R all pass: Bloomberg gives a $4.65B valuation and named backers, with a self-improving AI safety angle. No model capability, experiment result, or product path is disclosed, so it stays below 85.

Bloomberg Technology

Nvidia’s CEO Sees China Opening Market to AI Chips From US

Nvidia CEO Jensen Huang said Chinese authorities will eventually allow imports of US AI chips; the RSS snippet only says he made the comment days after joining President Donald Trump’s summit in China and does not disclose a timetable.

Why it matters: HKR-H and HKR-R clear because a Nvidia CEO prediction on China reopening the AI-chip market is clickable and geopolitically loaded. HKR-K misses: no timeline, approval path, or chip model is disclosed, so it sits at the featured threshold.

r/LocalLLaMA

llama.cpp MTP support landed: Qwen3.6 27B reaches 2.44× on Strix Halo

llama.cpp merged MTP speculative decoding in PR #22673; Qwen3.6 27B Q8_0 rose from 7.4 to 18.1 tok/s on Strix Halo, while a dual RTX 3090 Q8_0 setup rose from 25.7 to 55.9 tok/s.

Why it matters: HKR-H/K/R all pass: llama.cpp adds MTP speculative decoding with Qwen3.6 27B speedups on Strix Halo and RTX 3090. The scope is local inference, not a broad model release, so 78 fits featured.

AI HOT (Curated Pool)

NVIDIA fine-tunes Cosmos Predict 2.5 with LoRA/DoRA for robot video generation

NVIDIA published a Hugging Face post on fine-tuning Cosmos Predict 2.5 with LoRA and DoRA to generate robot first-person videos from text prompts; the post does not disclose dataset size, training cost, or evaluation results.

Why it matters: HKR-H/K/R pass: the robot POV video angle is clickable, and LoRA/DoRA on Cosmos Predict 2.5 is a concrete mechanism. Missing dataset scale and metrics keep it in the low featured band.

May 17Sunday

QbitAI · WeChat

A Robot Dog Challenges Nvidia's Compute Lead

Weilan Technology unveiled BabyAlpha A3, a consumer quadruped robot using a six-chip heterogeneous cluster that runs a 7B-parameter model on-device at 280 TPS; the article says it has 66MP vision, 2.232 million point-cloud samples per second, and a planned Q3 launch.

Why it matters: HKR-H/K/R pass: the robot-dog-versus-Nvidia angle is clickable, and 280 TPS on a local 7B model is concrete. Single-source summary lacks price, power draw, and benchmark setup, so it stays near the featured floor.

Synced · WeChat

What Are World Models? Their History and the $10 Billion Bet

Jiqizhixin translated a MoE Capital blog tracing two world-model lineages. The article says more than $10 billion entered the category over 18 months, and cites DreamDojo as using 44,711 hours of first-person video pretraining to reach r=0.995 correlation with real-world robot policy outcomes.

Why it matters: HKR-H/K/R all pass: the hook is strong and the article gives concrete figures, but it is a compiled explainer rather than a new release. It fits the featured-threshold band for a strong commentary/tutorial.

r/LocalLLaMA

Same Models Tested Across Strix Halo, RTX 3090, and RTX 5070

C_Coffie published 55 local inference benchmark runs across Strix Halo, RTX 3090, RTX 5070, five backends, and 0.35B to 35B-A3B models; RTX 5070 beats RTX 3090 on models fitting 12GiB, while RTX 3090 leads in the 14–31B band that exceeds 12GiB but fits 24GiB.

Why it matters: Hits HKR-H/K/R with a named first-person benchmark: 55 runs and concrete GPU crossover points. Source is a single Reddit post, so it stays in the low featured band.

Dwarkesh Patel podcast

Notes on Pretraining Parallelisms and Failed Training Runs

Dwarkesh documents pretraining failure modes and parallelism tradeoffs: expert choice and token dropping can break causality in MoE routing, FP16 collectives can bias repeated additions after values exceed 1024, pretraining FLOPs are given as 6ND, B300 HBM is listed as 288GB, and FSDP communication can reach params × 3 with reduce-scatter.

Why it matters: HKR-H/K/R all pass: Dwarkesh’s notes expose concrete pretraining failure modes and numbers. The systems-training focus is specialized, so it sits in the high-quality band rather than same-day must-write.

May 16Saturday

Hacker News front page

SANA-WM, a 2.6B open-source world model for 1-minute 720p video

SANA-WM’s title says the project is a 2.6B open-source world model for 1-minute 720p video; the RSS body only lists the project URL, Hacker News comments URL, 9 points, and 8 comments, and the post does not disclose training data, license terms, inference cost, evaluation setup, or benchmark results.

Why it matters: HKR-H/K/R pass on the concrete open-source world-model hook, 2.6B size, and video-model competition angle. Sparse body details keep it at the lower good-quality band.

AI HOT (Curated Pool)

Nvidia CEO Says Skilled Trades Have Better Prospects Than CS Graduates

Jensen Huang told Carnegie Mellon’s 2026 CS graduates that skilled trades have better prospects; Randstad says trade demand is growing three times faster than white-collar roles, with robotics technician jobs up 107%.

Why it matters: HKR-H/K/R all pass: a sharp Jensen Huang career claim, two concrete labor-market numbers, and clear jobs anxiety for AI workers. It is still an X-sourced commentary item, not a model, product, or policy event, so it stays at low featured.

Bloomberg Technology

Cerebras CEO Is Worth $3.2 Billion After Year’s Largest IPO

Cerebras Systems rose about 68% on Nasdaq after the year’s largest IPO, giving the company a market value of roughly $67 billion; the headline says its CEO is worth $3.2 billion after the listing.

Why it matters: HKR-H/K/R all pass: Cerebras' IPO gives AI compute a public-market price via a 68% jump and about $67B market cap. The Bloomberg item is wealth-video framed, so it lands in must-write range, not industry-shaking.

May 15Friday

r/LocalLLaMA

Fully Offline Suitcase Robot Built Around Jetson Orin NX SUPER 16GB

CreativelyBankrupt built Sparky as a fully offline suitcase robot on Jetson Orin NX SUPER 16GB, running Gemma 4 E4B Q4_K_M via llama.cpp with q8_0 KV cache, about 200 ms cached TTFT, 14-15 tok/s sustained output, 12K context, 30+ sensors, and no WiFi, Bluetooth, or cellular interface.

Why it matters: HKR-H/K/R all pass, with a named hands-on build and concrete latency/sensor numbers. It stays in low featured because this is a Reddit project post, not a product launch or research release.

r/LocalLLaMA

I tracked EU GPU prices across 15 stores for 50+ days: RTX 5090 is the only card not dropping

Reddit user egudegi tracked EU GPU prices across 15 stores for more than 50 days with a 6-hour scrape cadence and about 126,000 readings; RTX 5090 average pricing rose from €3,392 to €3,487, a 3.0% increase.

Why it matters: HKR-H/K/R all pass, backed by a quantified first-person price scrape. Source authority is a single Reddit post, so it sits at the featured threshold rather than a higher band.

r/LocalLLaMA

The RTX 5000 PRO 48GB arrived and is better than expected

A Reddit user built a $5,600 RTX 5000 PRO 48GB PC and ran Qwen3.6-27B-FP8 with full-precision cache; they report up to 80 tok/s in TG, about 50–60 tok/s on very large prompts, 4,400 tok/s in prompt processing, and 200k tokens fitting in BF16 KV cache.

Why it matters: HKR-H/K/R all pass: a first-person local-inference test gives price and speed numbers, not vendor copy. Single Reddit source limits reach, so it lands in the featured-threshold band.