Skip to content

NVIDIA chips and ecosystem: new GPUs, CUDA, robotics platforms and the market for AI compute.

Latest picks

121–140 of 290

Jul 1Wednesday

TechCrunch · AI

Nvidia rival Etched hits $5B valuation with $1B in booked orders for its AI inference chip

After TSMC manufactured its chip, Etched disclosed $1B in contracted orders for full inference systems—not just chips. Total funding now stands at $800M, including an unannounced $500M round. The pitch: faster, cheaper, more power-efficient inference for frontier models, which is the biggest cost bottleneck for AI companies serving users at scale. The post doesn't name specific customers, disclose performance benchmarks, or give a delivery timeline, so I'd discount the $1B figure until we see actual shipments.

Why it matters: Etched has silicon fabbed and in customer hands, with $1B in contracts and a $5B valuation — one of the few Nvidia alternatives with both hardware and revenue signals. Score capped at 78 because it's a single-source report and the product hasn't been deployed at scale yet.

AI HOT (Curated Pool)

Anthropic launches Claude Science, an AI workbench that unifies scientific toolchains

Claude Science is an AI workbench for researchers, now in beta for Pro, Max, Team, and Enterprise users. It combines literature search, data analysis, figure generation, and manuscript editing in one session, with over 60 pre-configured skills and connectors for genomics, single-cell, proteomics, cheminformatics, and more. Every output includes auditable code, environment, and chat history for reproducibility. Compute runs locally on macOS or Linux, or on your lab's HPC cluster via SSH, with on-demand GPU scaling through Modal. The post does not disclose beta pricing or a general-availability timeline.

Why it matters: Anthropic ships Claude Science, a vertical workbench for researchers that unifies literature, data, visualization, and writing in one session with 60+ pre-built skills and auditable outputs. The product shape is differentiated and hits all three HKR axes. The score stays at 82...

Jun 30Tuesday

AI HOT (Curated Pool)

NVIDIA Rubin Ultra canceled, new version halved in size and performance

SemiAnalysis reports the original 4-die Rubin Ultra announced at GTC 2026 has been canceled after just 3 months due to manufacturing execution issues. The new Rubin Ultra is half the size with roughly half the performance. The post doesn't spell out which manufacturing steps failed or the new timeline.

Why it matters: SemiAnalysis reports NVIDIA's next-gen flagship Rubin Ultra has been canceled due to manufacturing issues, with the replacement halving the original specs. This is a major hardware roadmap shift that directly affects compute expectations for next year's training clusters. Scor...

Jun 29Monday

Import AI (Jack Clark)

NVIDIA builds a self-improving loop for robots; Tencent details its 10k-GPU debug tool

NVIDIA's ENPIRE lets physical robots self-improve through trial and error like coding agents, hitting 99% on tasks like GPU insertion and zip-tie cutting. The catch: auto-evaluation and auto-reset still break on harder tasks. Tencent open-sourced ARGUS, an always-on tracing system for 10k+ GPU training clusters, already battle-tested for six months. A separate law paper points out that top minds badly misjudged nuclear fission and the internet—today's AI hot takes will likely age just as poorly.

Why it matters: NVIDIA's ENPIRE ports the agent trial-and-error loop to physical robots, hitting 99% on GPU insertion but still failing on auto-eval and reset for harder tasks. HKR all hit, but this is a newsletter digest rather than the primary paper, so information density is diluted — capp...

Jun 26Friday

TechCrunch · AI

OpenAI's Jalapeño chip is Big Tech's spiciest move away from Nvidia

OpenAI revealed Jalapeño, a custom inference chip built with Broadcom, joining Google, Apple, and SpaceX in reducing single-supplier dependence on Nvidia. The move is a hedge, not a clean break—more control and workload-specific performance, similar to Apple's shift from Intel. The same podcast episode covers Groq's $650M raise after Nvidia poached its top talent, and AI agents entering loops that Claude Code creator Boris Cherny calls as big a step as the jump from source code to agents.

Why it matters: OpenAI's custom inference chip is a real signal, and the Jalapeño codename plus Broadcom partnership give it concrete hooks. But this is a podcast discussion, not an official announcement — the detail density is thin, so it lands at the featured threshold of 72.

AI HOT (Curated Pool)

SGLang adds Waterfill and LPLB to improve DeepEP MoE load balancing

LMSYS and NVIDIA introduced two dispatch-time load balancers in SGLang for DeepEP MoE inference. Waterfill handles the shared expert by sending it to the least-loaded GPU; it boosts throughput by 1.48–4.66% on DeepSeek V3/R1 and lifts DeepSeek V4 from 49,253 tok/s to 51,677 tok/s (+4.92%). LPLB uses linear programming to route tokens across redundant expert replicas, adding 0.84–7.34% throughput on the same benchmarks. Both methods preserve model accuracy. The post reports throughput gains but does not disclose latency impact.

Why it matters: Joint release from LMSYS and NVIDIA with concrete throughput gains and two named methods. Useful for anyone running DeepSeek-family models in production. Narrow audience and purely engineering-layer keeps it at the featured threshold.

Jun 25Thursday

Bloomberg Technology

Qualcomm projects $15B in data center chip sales by 2029

Qualcomm set an aggressive target at its investor day: $15B in annual data center chip revenue by 2029. That's larger than its entire automotive and IoT segments combined in fiscal 2025. The bet is on AI inference chips, pivoting from training to inference and competing with Nvidia and Broadcom. The post doesn't disclose current data center revenue, but the base is clearly small. Going from near-zero to $15B in five years is a tight timeline.

Why it matters: Qualcomm's first public data-center revenue target of $15B by 2029, starting from near-zero base and betting on AI inference against Nvidia and Broadcom. Bloomberg exclusive gives it sourcing authority. The catch: the article doesn't disclose current data-center revenue, so $1...

Jun 24Wednesday

The Verge · AI

OpenAI reveals its first AI processor: Jalapeño

OpenAI unveiled Jalapeño, its first custom chip built with Broadcom for ChatGPT inference. The chip is in mass production and already deploying on OpenAI's own servers, aiming to cut reliance on Nvidia and control costs. It only handles inference for now; a training chip is still in development. The post does not disclose performance, power, or cost figures.

Why it matters: OpenAI's first custom chip is in production and deploying — a real step toward reducing Nvidia reliance. No perf, power, or cost numbers disclosed, so capped below 85.

Financial Times · Technology

SK Hynix bets on AI demand with a $29bn US listing

SK Hynix plans a $29bn US IPO, the largest overseas listing by a Korean company. The capital will mainly expand HBM memory production for AI chip customers like Nvidia. The post doesn't disclose the timeline or pricing yet. The direction is clear: the AI arms race keeps pushing capital upstream.

Why it matters: SK Hynix's $29bn US IPO is hard evidence of AI demand flowing upstream. HBM expansion financing figure is concrete, but the post doesn't disclose timeline or pricing, so score stays below 85.

Jun 23Tuesday

TechCrunch · AI

Groq confirms $650M raise and re-staffs after Nvidia's $20B not-acqui-hire deal

AI chip startup Groq confirmed a $650M funding round and is rebuilding its executive team after Nvidia spent roughly $20B to hire away most of Groq's staff without acquiring the company. CEO Jonathan Ross says Groq is now leaning into a neocloud model—running inference services on its own chips instead of just selling hardware. New hires include a chief product officer and a VP of engineering, both from outside the group that left for Nvidia. The post does not disclose the round's valuation or investors.

Why it matters: Groq raises $650M after Nvidia's $20B talent grab and pivots to inference cloud — a story with reversal, concrete numbers, and strategic shift, hitting all three HKR axes. Not scoring higher because only the funding is confirmed; the new cloud model's actual performance and cu...

TechCrunch · AI

SpaceX inks $150M/month compute deal with open source AI lab Reflection AI

SpaceX's Colossus 2 data center near Memphis landed its third major AI compute customer. Open source lab Reflection AI will pay $150 million per month starting July 1, 2026 through 2029 for immediate access to Nvidia GB300 chips. The deal follows earlier contracts with Anthropic ($1.25B/month) and Google ($920M/month). The post doesn't disclose what models Reflection AI plans to train or its funding sources.

Why it matters: SpaceX lands a $150M/month compute deal with open-source lab Reflection AI through 2029. The story has novelty, hard numbers, and industry buzz, but Reflection AI's low profile and undisclosed funding keep it at the 78 featured threshold.

Jun 17Wednesday

New York Times Chinese

AI chip boom redraws global tech map, with China conspicuously absent

Nvidia's AI boom is making SK Hynix, Samsung, and Micron extremely rich. The most advanced memory chips come only from these three—none from China. Huang wrote "Please produce more :)" on a Hynix wafer at Computex because supply can't keep up. Memory prices more than doubled this year. Samsung and SK Hynix made South Korea the first country outside the US with two trillion-dollar companies. US tariffs and tech restrictions locked China out of this boom more effectively than subsidies ever did. The supply chain now sits almost entirely in Taiwan and South Korea, two geopolitical flashpoints.

Why it matters: NYT uses the memory supply chain as a lens—doubled wafer prices, two trillion-dollar Korean companies, and a Jensen Huang anecdote—to explain the AI hardware bottleneck. It's a trend piece rather than event-driven news, so it lands at the featured threshold.

Jun 16Tuesday

Hacker News front page

SubQ 1.1 Small: Sparse attention cuts long-context compute by 64.5x at 1M tokens

Subquadratic released the model card for SubQ 1.1 Small. It replaces quadratic dense attention with Subquadratic Sparse Attention (SSA) that routes based on content relevance, scaling linearly with context length. At 1M tokens, SSA uses 64.5x less compute than dense attention and runs 56x faster than FlashAttention-2. The model scores near-perfect on needle-in-a-haystack from 1M to 12M tokens and 99.12% on RULER at 128K. General reasoning holds: GPQA Diamond 85.4%, LiveCodeBench pass@4 89.7%, AutomationBench Finance 13%. Training started from an open-weight frontier model, replaced attention with SSA, then ran staged context extension up to 2M and ~1T tokens of continued pretraining on books, documents, and repo-scale code. The post does not name the base model. SubQ 1.1 Small is deploying with select design partners; a broader lineup from 2M to 12M tokens is planned later this year.

Why it matters: SubQ 1.1 Small ships a deployable sparse-attention model with a 64.5x compute reduction and near-perfect 12M-token retrieval. Held below 85 because it's still a model card + design-partner deployment — no open weights or public API yet, so the production story is incomplete.

Jun 10Wednesday

NVIDIA Blog

NVIDIA confidential computing will help Apple expand Private Cloud Compute

Apple is bringing NVIDIA's confidential computing into Private Cloud Compute, running AI inference inside encrypted GPU environments. The setup uses H100 GPUs and Hopper architecture with hardware-level trusted execution environments, so data stays encrypted during processing and even the cloud provider can't access it. Apple previously ran private cloud inference only on its own silicon; this deal signals a shift of some workloads to NVIDIA while keeping the same security isolation. The post doesn't give a launch date or scale numbers, but confirms deployment will start in Apple's own data centers.

Why it matters: Apple is moving Private Cloud Compute workloads to NVIDIA H100 for the first time, using hardware-level TEEs to keep inference data encrypted while switching the compute substrate. Score isn't higher because the post gives no launch timeline or deployment scale — it's a direct...

Jun 9Tuesday

AI HOT (Curated Pool)

Elon Musk Details SpaceX's AI1 Orbital AI Data Center Satellite Plan

Elon Musk detailed SpaceX’s AI1 orbital AI data center satellite plan, with 150 kW peak power per satellite, about 120 kW sustained compute power, and 6-8 ms round-trip latency at 600-800 km low Earth orbit.

Why it matters: HKR-H/K/R all pass: the angle is unusual and the post gives power, orbit, and latency figures. It stays near the featured floor because launch timing, cost, and workload tests are not disclosed.

Jun 8Monday

AI HOT (Curated Pool)

The Vanishing Crash in Five-Model Economies: Control and Emergence

The experiment used five models from OpenAI, NVIDIA, OpenBMB, and a self-fine-tuned 500M-parameter model to drive market agents; three interventions failed to reproduce the price crash, and the crash was created only by overriding prices during settlement.

Why it matters: HKR-H/K/R all pass: the angle is counterintuitive, the post gives 5 models, 3 interventions, and a settlement override mechanism, and it speaks to agent-eval reliability. Scope remains an experiment blog, not a major release.

r/LocalLLaMA

DFlash Speculative Decoding and KV Cache Compression on RTX 5090 Show 3.26x Speedup

The author tested Qwen3.6-27B on an RTX 5090 with DFlash plus KV cache compression, reaching up to 3.26x speedup; q4_0/turbo4 delivered 3.18x speedup with only +0.02% PPL on WikiText-2.

Why it matters: HKR-H/K/R all pass: RTX 5090 testing, DFlash speculative decoding, KV cache compression, 3.26x speedup, and PPL delta are concrete. Single Reddit source keeps it near the featured floor.

r/LocalLLaMA

Weird to get near-linear scaling by adding another GPU?

A Reddit user benchmarked qwen3.6-27b-autoround-int4 on 1x3090 versus 2x3090. Narrative decode rose from 53 TPS to 94 TPS, and code decode rose from 62 TPS to 120 TPS, under no NVLink, 8x/8x PCIe, P2P automatically enabled, tensor parallelism set to 2, and different KV-cache settings.

Why it matters: HKR-H/K/R all pass: the result is counterintuitive, includes concrete TPS and TP conditions, and speaks to local-inference cost. Single Reddit test lacks multi-model replication and full setup details, so it stays near the featured threshold.

AI HOT (Curated Pool)

Nvidia and SK Hynix Sign Multi-Year Pact to Develop Next-Generation AI Memory Chips

Nvidia and SK Hynix signed a multi-year pact to co-design future generations of memory chips for AI applications; the RSS snippet does not disclose product specifications, production timelines, or financial terms.

Why it matters: HKR-H and HKR-R pass: Bloomberg plus Nvidia/SK Hynix matters for AI memory supply. HKR-K is weak because specs, production timing, and financial terms are missing, so this sits at the low featured band.

Jun 7Sunday

AI HOT (Curated Pool)

Five Labs, Five Minds: Building a Multi-Model Financial Drama Game with Small Models

Thousand Token Wood v2 uses four small models from different labs to drive agents in a financial simulation game, with vLLM 0.22.1’s CUDA toolkit dependency identified as the main serving friction, while a fine-tuned 0.5B Qwen reached 0% self-trading and 100% valid quotes.

Why it matters: HKR-H/K/R all pass: the small-model finance game is a real hook, with vLLM and 0.5B Qwen metrics, plus agent-engineering resonance. Scope remains an experiment, so it sits in low featured.