Skip to content

All news

3 today

May 23Saturday

AI HOT (Curated Pool)

Agent Workloads Quietly Reshape Inference Economics

SemiAnalysis analyzed 432,000 real coding-agent requests and found a median input length of 96,000 tokens, not 32,000 or 64,000. The post does not disclose the model mix, cost curve, sampling method, or time window.

Why it matters: HKR-H/K/R all pass: SemiAnalysis adds a 432k coding-agent request dataset and 96k-token median input. Missing models, cost curves, and sampling keep it in the strong-data-point band, not must-write.

May 22Friday

Dwarkesh Patel podcast

Reiner Pope – Chip Design from the Bottom Up

Dwarkesh Patel interviews MatX CEO Reiner Pope on chip design, starting with a 4-bit multiply and 8-bit accumulate example that uses 16 AND gates, then covering systolic arrays, pipeline registers, FPGAs versus ASICs, cache versus scratchpad, and why GPU cores are smaller than CPU cores.

Why it matters: Dwarkesh’s MatX CEO interview clears HKR-H/K/R with a bottom-up hardware hook, concrete mechanisms, and compute-cost resonance. It is educational rather than breaking news, so it sits in the 72–77 band.

AI HOT (Curated Pool)

Text Degeneration: A Production Failure Mode Most Benchmarks Do Not Track

Dharma-AI says in a Hugging Face post that large language models can produce repeated, incoherent, or logically confused text in production, and most mainstream benchmarks do not track this failure mode.

Why it matters: HKR-H/K/R all pass, but the post only discloses the failure pattern and benchmark blind spot, with no sample size, metric, or reproduction setup. This fits the lower featured threshold.

最佳拍档 (BestPartners)

Nvidia reports Q1 2026 results: revenue 81.6B, shares down 2%

The title says Nvidia reported Q1 2026 revenue of 81.6 billion, profit of 58.3 billion, 92% data-center growth, and a 2% share-price drop; the post does not disclose the currency or profit metric.

Why it matters: HKR-H/K/R all pass, but the post only gives title-level earnings figures and omits currency, profit basis, and guidance. NVIDIA data-center +92% is strong enough for featured, kept in the lower featured band.

AI HOT (Curated Pool)

Plastic Interfaces: The Future Shape of AI-Driven Software

Salesforce has adopted a headless architecture that lets salespeople update data through AI; the post says MCPs, HTML, audio, and web interfaces can be generated dynamically by context, but it does not disclose implementation metrics or adoption numbers.

Why it matters: HKR-H/K/R all pass, but this is a software-form thesis without user metrics, launch timing, or a reproducible test. It fits the insightful-commentary band, not a must-write release.

May 21Thursday

AI HOT (Curated Pool)

Lessons from Building Cloud Agents

Cursor summarizes lessons from building cloud agents: after migrating to Temporal, reliability rose above 99.9%, and the platform processes more than 50 million operations per day.

Why it matters: HKR-H/K/R all pass: Cursor is central to coding agents, and the post gives Temporal, 99.9%+ reliability, and 50M daily operations. Not a launch, so it stays at low-end featured.

Alibaba Technology · WeChat

Building an Agent from 0 to 1: Principles and Personal Assistant Practice

Zhan Xupeng published a roughly 50-minute article on Agent theory and a personal assistant implementation, covering memory, ReAct planning, progressive skill loading, subagents, and harness-level fault recovery.

Why it matters: HKR-K/R pass via concrete agent mechanisms and practitioner reliability pain; HKR-H is weak because the headline is a standard tutorial frame. This fits the quality-tutorial threshold, not the 78+ news band.

r/LocalLLaMA

Moved from prompt-based output validation to schema-enforced execution, with significant reliability gains

A Reddit user tested Claude structured outputs and reported 90–95%+ first-pass parse rates with tool_use, typed schemas, enum constraints, and stepwise validation, versus 65–70% for prompt instructions followed by regex or JSON parsing and retries.

Why it matters: HKR-H/K/R all pass: the post has a clear reliability contrast and concrete 90–95%+ vs 65–70% numbers. Source authority is limited to one Reddit experiment, with sample and task details not disclosed, so it stays at low featured.

Financial Times · Technology

Nvidia lifts dividend as investors fret about growth prospects

Nvidia raised its dividend and reported revenue and forecasts above expectations, but its shares still fell; the RSS snippet does not disclose the dividend increase, revenue figures, forecast range, or trading move percentage.

Why it matters: HKR-H and HKR-R pass: FT frames NVIDIA beats against a share drop and AI compute-cycle anxiety. HKR-K fails because dividend, revenue, and guidance figures are not disclosed.

Financial Times · Technology

Anthropic on Track for First Profitable Quarter

Anthropic is on track to record its first profitable quarter ahead of OpenAI and xAI; the RSS snippet does not disclose the quarter, revenue, profit figure, or accounting basis.

Why it matters: HKR-H/K/R all pass: the FT claim reframes Anthropic’s business race against OpenAI and xAI. Missing quarter, revenue, and profit figures keeps it below P1.

May 20Wednesday

AI HOT (Curated Pool)

Unsustainable Subsidies

Google, OpenAI, and Anthropic diverged on model pricing: Gemini 3.1 Pro is priced at $2 input and $12 output, GPT-5.5 at $5 and $30 after a short subsidy, and Claude Opus 4.7 stayed at $5 and $25.

Why it matters: HKR-H/K/R all pass, but this is Tom Tunguz commentary on pricing rather than a primary model release. The concrete price spread makes it featured, not must-write.

r/LocalLLaMA

Cursor and Claude Code Are Not Getting Dumber; Agent Loops Are Suffocating Context

A Reddit user says an API-log audit showed Cursor and Claude Code recursively grep about 40 files in 10k-plus-line repositories, sometimes load 2k-line files for 5-line edits, and spend roughly 30k tokens on tool definitions and logs before generating code.

Why it matters: HKR-H/K/R all pass: the hook is contrarian, the API-log numbers are concrete, and coding-agent context waste is a live practitioner pain. Reddit single-post sourcing and no shared logs keep it at the featured threshold.

May 19Tuesday

AI HOT (Curated Pool)

Former executive says Microsoft’s AI strategy faltered, with Copilot paid usage below 3%

Former Microsoft executive Matt Veloso said Microsoft generated about $30 billion from its AI partnership between 2023 and 2025, while related costs reached $100 billion; he also said actual usage among paid Copilot users is below 3%.

Why it matters: HKR-H/K/R all pass: a former executive gives concrete Microsoft AI cost, revenue, and Copilot usage numbers. Kept at 80 because this is a single former-exec claim, not an official Microsoft disclosure.

AI HOT (Curated Pool)

I really want to praise HTML!

The author used Claude Code to generate a single-file HTML project plan page in 2 minutes, with a dark theme, timeline, and collapsible tables; the comparable Notion template previously took 30-40 minutes.

Why it matters: HKR-H/K/R all pass: the post has a concrete Claude Code workflow hook, a 2-minute vs 30-40-minute comparison, and clear practitioner resonance. Scope is small, so it sits at the featured threshold.

Latent Space

[AINews] How to Land a Job at a Frontier Lab (on Pretraining)

Latent Space says Vlad Feinberg’s pretraining job-prep notes reduce frontier-lab readiness to kernel-level performance work: derive Chinchilla laws, compare dense and MoE architectures, code the solution in JAX, then write a Pallas kernel that beats jax.lax.ragged_dot for F > D by fusing up/down projections.

Why it matters: HKR-H/K/R all pass: the career hook is strong and the prep list is concrete. It is not a model release or major product update, and the kernel-heavy angle keeps it at the lower featured band.

AI Chat-Group Daily (群聊日报)

May 18, 2026 Chat Group Daily

The chat group daily says AI21 Labs cut 60% of staff and stopped selling model access, and cites a University of Waterloo paper where GPT-5.4 accuracy dropped from 100% to 23% after false peer-consensus injection; the snippet also mentions Meta layoff talk at 10%, but does not disclose source details or confirmation conditions.

Why it matters: HKR-H/K/R all pass: AI21’s 60% layoff and model-sales stop signal lab contraction, while GPT-5.4 falling from 100% to 23% under false peer consensus is a concrete safety hook. The chat-digest source keeps it at 78.

r/LocalLLaMA

Tried Every Hermes Agent Alternative So You Don't Have To: 2026 Roundup

A Reddit user compared 11 Hermes Agent alternatives across open-source and managed options; OpenClaw is listed with 347k GitHub stars, 24+ integrations, and 9 CVEs in four days, while TrustClaw uses OAuth-only sandboxed execution and Perplexity Computer requires a $200/month Max tier.

Why it matters: HKR-H/K/R all pass: this is a practical agent-tool comparison with 11 items and concrete integration/security figures. Reddit single-post sourcing limits confidence, so it stays near the featured threshold.

May 18Monday

Latent Space

The Autonomous Drone Tech Stack and Economics of Drones — Yaroslav Azhnyuk

Latent Space interviewed The Fourth Law founder Yaroslav Azhnyuk for a two-hour episode covering FPV drones, five levels of autonomy, eight dimensions of the autonomous battlefield, and China’s manufacturing advantage; the transcript claims Ukraine produced 4 million FPV drones last year and discusses a hypothetical Chinese capacity of 4 billion.

Why it matters: HKR-H/K/R all pass: the Latent Space interview offers concrete autonomy and battlefield frameworks. It is still commentary, not a model release, product update, or research artifact, so it stays just above the featured threshold.

May 17Sunday

AI HOT (Curated Pool)

Microsoft AI CEO predicts AI will automate all white-collar jobs within 18 months

Mustafa Suleyman predicts AI will reach human-level performance within 18 months and automate most professional tasks, including accounting, law, marketing, and project management.

Why it matters: HKR-H and HKR-R are strong, and HKR-K passes on the testable 18-month timeline. The score stays in the low 78–84 band because this is a CEO forecast, not evidence, benchmarks, or a shipped capability.

Financial Times · Technology

Chinese AI Groups Pull Ahead of US Rivals in Video Generation Race

FT says Chinese AI groups have moved ahead of US rivals in video generation; the RSS snippet names ByteDance and Kuaishou and says they outshine western competitors in advertising and entertainment quality, but the post does not disclose benchmark metrics or model details.

Why it matters: FT authority plus a China-vs-US video-generation lead claim clears HKR-H and HKR-R. HKR-K fails because the body lacks metrics, samples, and eval method, so it sits at the low featured threshold.

Synced · WeChat

What Are World Models? Their History and the $10 Billion Bet

Jiqizhixin translated a MoE Capital blog tracing two world-model lineages. The article says more than $10 billion entered the category over 18 months, and cites DreamDojo as using 44,711 hours of first-person video pretraining to reach r=0.995 correlation with real-world robot policy outcomes.

Why it matters: HKR-H/K/R all pass: the hook is strong and the article gives concrete figures, but it is a compiled explainer rather than a new release. It fits the featured-threshold band for a strong commentary/tutorial.

Synced · WeChat

Peter Steinberger Says His Monthly Token Bill Hit $1.3M, Covered by OpenAI

Peter Steinberger used 603 billion tokens across 7.6 million requests in 30 days, with the bill exceeding $1.3 million; he said disabling fast mode cut the price by 70%, and OpenAI does not charge him for the tokens.

Why it matters: HKR-H/K/R all pass: the story has a sharp cost hook, concrete usage numbers, and strong practitioner resonance. It is a first-person bill disclosure, not an OpenAI pricing or product launch, so it sits just above the featured threshold.

AI HOT (Curated Pool)

Anthropic CEO discusses AI’s dual impact: high growth and high unemployment

Dario Amodei said AI may drive 5%-10% GDP growth while increasing unemployment and inequality, and near-free software costs would challenge the assumptions behind traditional software business models.

Why it matters: HKR-H/K/R all pass: Dario Amodei’s 5%-10% GDP and near-free software claims are concrete and highly discussable. The source is an X summary, not a full primary transcript, so it stays at 78.

AI HOT (Curated Pool)

Anthropic CEO predicts near-free software and major job shifts

Dario Amodei said in a Wall Street Journal YouTube interview that software costs will fall sharply toward near-free, and the traditional assumption that software needs millions of users to spread costs will no longer hold.

Why it matters: HKR-H/K/R all pass: Dario Amodei’s software-cost and labor-structure claim is highly discussable. The source is a secondhand X summary, with no full argument, timeline, or data disclosed, so it stays in the low featured band.

TechCrunch · AI

The Haves and Have-Nots of the AI Gold Rush

Deedy Das estimated that about 10,000 founders and employees at companies including OpenAI, Anthropic, and Nvidia have accumulated more than $20 million in wealth, while many software engineers face layoffs, sub-$500,000 career ceilings, and anxiety that their core skills are losing labor-market value.

Why it matters: HKR-H/K/R all pass: the wealth-gap angle is clickable, the $20M/10,000-person estimate is concrete, and the labor-market anxiety is strong. It is commentary, not a model, product, or funding event, so it stays at the featured threshold.

Dwarkesh Patel podcast

The mistake of conflating intelligence and power

Dwarkesh Patel argues that intelligence and power are being conflated: current AI systems improve through economically valuable tasks such as coding, while real-world power depends more on authority, trust, and large-scale cooperation than isolated strategic reasoning.

Why it matters: HKR-H/K/R all pass: Dwarkesh targets the capability-to-power link at the center of AI-safety debate. The summary gives no new data or empirical case, so this stays in the quality commentary band, not 85+.

Dwarkesh Patel podcast

Notes on Pretraining Parallelisms and Failed Training Runs

Dwarkesh documents pretraining failure modes and parallelism tradeoffs: expert choice and token dropping can break causality in MoE routing, FP16 collectives can bias repeated additions after values exceed 1024, pretraining FLOPs are given as 6ND, B300 HBM is listed as 288GB, and FSDP communication can reach params × 3 with reduce-scatter.

Why it matters: HKR-H/K/R all pass: Dwarkesh’s notes expose concrete pretraining failure modes and numbers. The systems-training focus is specialized, so it sits in the high-quality band rather than same-day must-write.

AI HOT (Curated Pool)

RLVR May Perform Disproportionately Poorly in Science

Dwarkesh argues that RLVR has a short-feedback weakness in scientific theory validation; the post says validation loops can span decades or centuries, and does not disclose experimental results or benchmark numbers.

Why it matters: HKR-H/K/R all pass: a sharp counter-narrative, a concrete feedback-loop mechanism, and strong resonance for RLVR/AI-for-science debates. It stays in 78–84 because this is commentary, not a release or empirical result.

AI HOT (Curated Pool)

Eric Jang shares lessons from building AlphaGo from scratch

Eric Jang spent several months implementing AlphaGo from scratch and says that in 2026, training a strong Go AI requires only a few thousand dollars in rented compute rather than DeepMind-scale resources.

Why it matters: All three HKR axes pass: the hook is a from-scratch AlphaGo rebuild, and K has concrete claims on months of work and few-thousand-dollar compute. It stays in 78-84 because this is a social post, not a model release or full paper.

May 16Saturday

AI HOT (Curated Pool)

Anthropic Founder’s Playbook warns AI can raise startup failure rates

Anthropic published Founder’s Playbook, arguing that AI tools such as Claude Code reduce prototyping cost but increase startup failure risk across the Idea, MVP, Launch, and Scale stages through false validation, confirmation bias, agentic technical debt, and founder decision bottlenecks.

Why it matters: HKR-H/K/R pass: the Anthropic founder playbook has a sharp counterintuitive angle, a four-stage mechanism, and clear founder resonance. It stays near the featured floor because no dataset or reproducible test is disclosed.

AI HOT (Curated Pool)

Nvidia CEO Says Skilled Trades Have Better Prospects Than CS Graduates

Jensen Huang told Carnegie Mellon’s 2026 CS graduates that skilled trades have better prospects; Randstad says trade demand is growing three times faster than white-collar roles, with robotics technician jobs up 107%.

Why it matters: HKR-H/K/R all pass: a sharp Jensen Huang career claim, two concrete labor-market numbers, and clear jobs anxiety for AI workers. It is still an X-sourced commentary item, not a model, product, or policy event, so it stays at low featured.

Bloomberg Technology

US Is Starting to See Heavy Job Losses in Roles Exposed to AI

Several US occupations expected to be exposed to AI recorded heavy job losses for a second year in 2025, led by customer service representatives and some secretary and salesperson roles; the RSS snippet does not disclose job-loss counts or the attribution method.

Why it matters: Strong HKR: Bloomberg frames AI-exposed roles as seeing job losses for a second straight year and names affected occupations. Exact loss counts and methodology are not disclosed in the summary, so this stays above featured threshold, not P1.

AI HOT (Curated Pool)

Yann LeCun interview: LLM limits, AI's future, and a new startup path

Yann LeCun discussed LLM limitations on the Unsupervised Learning podcast, covering his 2027 forecast, AMI’s bet on world models, his reasons for leaving Meta, and major disagreements with Geoffrey Hinton and Yoshua Bengio over Turing Award-era views.

Why it matters: HKR-H/K/R all pass: LeCun combines LLM limits, 2027 forecasts, world models, and Meta departure in one interview, matching the 85–94 band for major AGI-timeline commentary.

AI HOT (Curated Pool)

Eric Jang: Building AlphaGo from Scratch

Eric Jang uses AlphaGo to break down an intelligence system; the post only discloses three mechanisms: search, learning from experience, and self-play.

Why it matters: HKR-H/K/R pass, but this is a mechanism teardown/commentary rather than a model or product release. Dwarkesh + Eric Jang add authority, placing it at the featured threshold for a quality tutorial-style piece.

May 15Friday

The Verge · AI

AI research papers are getting better, and it’s a big problem for scientists

The Verge describes Peter Degen investigating unusual citations to a 2017 paper: it rose from a few dozen citations over several years to being cited every few days, while the RSS snippet does not disclose the full sample size or review findings.

Why it matters: HKR-H/K/R all pass: the paradoxical angle, named investigation, and citation spike give it signal. The post lacks full sample size, so it stays in the lower featured band rather than becoming must-write.

Bloomberg Technology

Enterprise 40% of Revenue Streams, Says OpenAI CRO

OpenAI CRO Denise Dresser said enterprise business makes up 40% of total revenue and is expected to reach 50% by year-end; the Bloomberg snippet does not disclose OpenAI’s total revenue size.

Why it matters: HKR-H/K/R all pass, but this is a short Bloomberg interview clip: it has OpenAI CRO revenue-mix numbers, not total revenue, margins, or customer scale. Featured threshold, not 78+.

AI HOT (Curated Pool)

The First Derivative of Inference: Growth Logic in the AI Wave

Tom Tunguz says the AI inference market will reach $250 billion within seven years; Datadog’s LLM observability data volume nearly doubled in the latest quarter, and about 20% of its AI customers contribute roughly 80% of ARR.

Why it matters: HKR-H/K/R all pass: Tom Tunguz ties inference growth to Datadog volume and ARR concentration data. It stays in the 72–77 band because this is commentary, not a model, product, or protocol release.

AI HOT (Curated Pool)

API prompt precaching speeds up first-token generation

Claude API prewarms prompt cache with the system prompt, skips output, then hits cache on the real request.

Why it matters: HKR-H/K/R all pass: this is a concrete Claude API latency mechanism, not a vague product tease. It clears featured, but it is a mid-weight inference update rather than a major model or capability release.

r/LocalLLaMA

I tracked EU GPU prices across 15 stores for 50+ days: RTX 5090 is the only card not dropping

Reddit user egudegi tracked EU GPU prices across 15 stores for more than 50 days with a 6-hour scrape cadence and about 126,000 readings; RTX 5090 average pricing rose from €3,392 to €3,487, a 3.0% increase.

Why it matters: HKR-H/K/R all pass, backed by a quantified first-person price scrape. Source authority is a single Reddit post, so it sits at the featured threshold rather than a higher band.

Bloomberg Technology

AI Buildout Drives 76% Power Bill Jump on Largest US Grid

Power prices on the largest US electric grid rose 76% in the first quarter, and the RSS snippet attributes the increase to data-center demand; the post does not disclose the grid operator’s name or a specific capacity shortfall.

Why it matters: HKR-H/K/R all pass: the 76% bill jump is a hard number, data-center demand gives a mechanism, and Bloomberg adds source weight. Missing grid-operator and capacity-gap details keep it in the 72–77 band.