Skip to content

All news

3 today

May 15Friday

r/LocalLLaMA

The RTX 5000 PRO 48GB arrived and is better than expected

A Reddit user built a $5,600 RTX 5000 PRO 48GB PC and ran Qwen3.6-27B-FP8 with full-precision cache; they report up to 80 tok/s in TG, about 50–60 tok/s on very large prompts, 4,400 tok/s in prompt processing, and 200k tokens fitting in BF16 KV cache.

Why it matters: HKR-H/K/R all pass: a first-person local-inference test gives price and speed numbers, not vendor copy. Single Reddit source limits reach, so it lands in the featured-threshold band.

May 14Thursday

AI HOT (Curated Pool)

Moonshot AI founder Yang Zhilin releases a 40-minute video

Yang Zhilin explains Kimi K2 training in a 40-minute video, saying the model cost $4.6 million and beat GPT-5.5 and other competitors on coding tasks.

Why it matters: HKR-H/K/R all pass: the founder-led Kimi K2 training breakdown adds a $4.6M cost figure and GPT-5.5 coding comparison. Single-source X relay and missing benchmark names keep it in 78-84, not P1.

AI HOT (Curated Pool)

Cost Analysis of AI Email

Top AI models process email at about $22 to $130 per month, with a $26 median; smaller models cut costs by 10 to 20 times, while local GPU execution can bring marginal cost close to zero.

Why it matters: HKR-H/K/R pass via a concrete cost spread and deployment-cost nerve. It is a useful opinion analysis, not a major product or model release, so it sits at 73.

r/LocalLLaMA

2x RTX 3090 setup for local Qwen 3.6 27B inference

A Reddit user ran Qwen 3.6 27B on a dual RTX 3090 Ubuntu setup, reporting 48GB VRAM, a 262k context window, no NVLink, about 4000 pp/s prompt processing, and 113 tk/s generation.

Why it matters: All HKR axes pass, and this is a first-person local-inference run with concrete numbers. Source is a single Reddit post with limited reproducibility detail, so it sits at the low featured threshold.

May 13Wednesday

AI HOT (Curated Pool)

90% of People Are Wasting Tokens

Andrej Karpathy says 90% of AI coding bills is wasted on unnecessary context, including repeated full-repository sends, expensive models for simple tasks, and missing prompt caching.

Why it matters: HKR-H/K/R all pass via the 90% claim, named waste mechanisms, and practitioner cost pain. It reaches featured, but stays at 72 because the post gives no billing sample or reproducible test.

May 12Tuesday

r/LocalLLaMA

Local LLM Autocomplete and Agentic Coding on a Single 16GB GPU + 64GB RAM

Reddit user grumd runs Qwen2.5-Coder-7B Q6 for autocomplete and Qwen3.6-35B-A3B Q8 for agentic coding on one RTX 5080 with RAM offloading; the post reports about 145k context, 56GB RAM used with other apps open, and Qwen3.6-35B-A3B speed of tg128 at 35.29 tokens/s.

Why it matters: HKR-H/K/R all pass: a named first-person local coding experiment with concrete model, quantization, context, and throughput data. Source is a single Reddit post without replication or comparisons, so it stays in the low featured band.

QbitAI · WeChat

Markdown Is Fading? Karpathy Also Backs HTML

Anthropic engineer Thariq argued for using HTML instead of Markdown and gave 5 reasons; the post says HTML generation takes about 2 to 4 times longer than Markdown.

Why it matters: HKR-H/K/R all pass, but this is a developer format debate rather than a model or product launch. Named Anthropic/Karpathy context and the 2-4x time figure clear the featured threshold at the low end.

Computing Life · Yage

How AI Caused and Fixed My Insomnia

The author used AI to build a HealthKit export app in about 5 minutes and run multivariate regression, finding that the last post-dinner AI usage time correlated negatively with sleep duration; after avoiding AI at night, average sleep increased by 1 hour and 40 minutes.

Why it matters: HKR-H/K/R all pass: a first-person quantified experiment links post-dinner AI use to shorter sleep, then reports +1h40m after stopping. Personal-blog scope keeps it below major industry-update territory.

Computing Life · Yage

How AI Caused and Fixed My Insomnia

The author used AI to build a HealthKit export tool and run multivariate regression, found that post-dinner AI use correlated negatively with sleep duration, and added 1 hour 40 minutes of average nightly sleep after avoiding AI for several weeks.

Why it matters: HKR-H/K/R all pass: the personal reversal is clickable, the HealthKit/regression setup adds testable detail, and sleep loss hits AI practitioners directly. Scope is anecdotal, so it stays at the featured floor.

r/LocalLLaMA

Computer Build Using Intel Optane Persistent Memory Runs a 1T-Parameter Model at Over 4 Tokens/s

Reddit user APFrisco ran the 1T-parameter Kimi K2.5 Q2_K_XL quant locally at about 4 tokens/s using 768GB Intel Optane PMem, 192GB DDR4 ECC DRAM, and a 12GB RTX 3060 with llama.cpp hybrid GPU/CPU inference.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, the post gives concrete hardware and speed numbers, and it hits local-inference cost concerns. Single Reddit anecdote and limited replication detail keep it at the featured floor.

AI HOT (Curated Pool)

Using LLMs in Script Shebang Lines

Simon Willison demonstrates using an LLM command in a script shebang line, with fragments generating SVG, the -T option calling llm_time, and a YAML template defining Python tools to compute 2344×5252+134 and return 12,310,822.

Why it matters: HKR-H/K/R all pass: Simon Willison shows a reproducible LLM-in-shebang workflow with concrete flags. Impact stays within CLI/script automation, not a model or platform release, so it sits in the low featured band.

AI HOT (Curated Pool)

The Evolution of Human-Computer Interfaces: From Text to Interactive Neural Video

Karpathy argues that LLM output is moving from Markdown toward richer HTML, while interactive neural video still has an open problem: how to combine neural generation with precise traditional software.

Why it matters: HKR-H/K/R pass: Karpathy gives a fresh UI frame, a concrete Markdown→HTML→neural-video path, and a builder-facing product question. Single X post with no data keeps it at the featured floor.

May 11Monday

QbitAI · WeChat

Math Majors in Trouble: Fields Medalist Tests ChatGPT 5.5 Pro, Gets Paper-Level Result in 17 Minutes

Timothy Gowers tested ChatGPT 5.5 Pro on additive number theory problems, where it produced an optimal quadratic upper-bound construction in 17 minutes 5 seconds, then generated a LaTeX preprint in 47 minutes; the article says arXiv rejects AI-generated content, so the result remains on Gowers’s blog.

Why it matters: All three HKR axes pass: Gowers’ first-person test, 17m05s, and a 47-minute preprint are concrete and discussable. It is not a model release, but the named experiment and math-reasoning impact put it in the must-write band.

May 10Sunday

Synced · WeChat

Ted Xiao Reviews Three Eras of Robot Learning, from RT-1/RT-2 to Scaling

Ted Xiao divides nearly a decade of robot learning into three eras: Google’s team trained RT-1 on 87,000 teleoperation trajectories, then adapted 5B to 55B VLMs into VLA policies for RT-2.

Why it matters: HKR-H/K/R all pass: a named Google robotics insider, concrete RT-1/RT-2 numbers, and strong embodied-AI resonance. It is retrospective commentary, not a launch, so it stays in the 72–77 featured band.

Xinzhiyuan · WeChat

Harsh Claim: Top Silicon Valley AI Is One Year Ahead of the World

Elad Gil claims top AI lab employees are 3-4 months ahead of Silicon Valley, while Silicon Valley is 3-6 months ahead of New York; the post cites Mythos’ 73% success rate in expert cyberattack simulations as evidence in a disputed “geographic time gap” argument.

Why it matters: HKR-H/K/R all pass: the lab-to-user lag hook is clickable, and the post cites 3–4 months, 3–6 months, and a 73% Mythos figure. It is secondhand commentary, not a model or product release, so it stays in the 72–77 threshold band.

May 9Saturday

AI HOT (Curated Pool)

Using Codex to debug and verify fixes in parallel

The author uses Codex in temporary crabbox environments to recreate bug states, verify failures, apply fixes, and re-verify them, while running 10 sessions in parallel to avoid local state pollution and speed loss.

Why it matters: HKR-H/K/R all pass, but this is a single first-person workflow note, not a product release or benchmark. The 10-session Codex/crabbox setup earns featured-level practical signal, near the lower band.

Latent Space

Anthropic growing 10x/year while others lay off over 10% of staff

Anthropic is described as growing 10x annually and being valued at $1T-$1.2T, while the post cites layoffs of 40% at Block, 14% at Coinbase, and 20% at Cloudflare under AI-readiness framing.

Why it matters: HKR-H/K/R all pass: the title has contrast, the post gives growth, valuation, and layoff figures, and it hits jobs plus AI-capital concentration. It is high-signal industry commentary, not an official funding or product event, so 78-84 fits.

AI HOT (Curated Pool)

Claude Code Practice: The Effectiveness of HTML Output

Thariq Shihipar recommends requesting HTML output from Claude, and the post cites GPT-5.5 generating an interactive Linux vulnerability page with SVG diagrams, interactive components, and in-page navigation.

Why it matters: HKR-H/K/R all pass, but this is a workflow tip rather than a Claude release. As a quality Claude Code tutorial, it sits in the 72–77 band, with Simon Willison’s source authority clearing featured.

May 8Friday

AI HOT (Curated Pool)

Robotics Endgame: A Physical AGI Roadmap and LLM Analogy

The speaker presented a physical AGI roadmap with six named components: video world models, WAM, EgoScale, dexterity scaling laws, physical reinforcement learning, and DreamDojo; the snippet also mentions a 2016 OpenAI DGX-1 signing story with Jensen and Elon.

Why it matters: HKR-H/K/R all pass: the physical-AGI endgame hook is strong, the post gives a 6-part roadmap, and robotics practitioners will debate the path. It is still a personal roadmap, not a release or benchmark, so it sits in 78–84.

Alibaba Technology · WeChat

The AI-Native Era: Where R&D Organizations Go Next

Xu Xiaobin cites internal interviews showing that engineers who use AI heavily cut coding time from 30% to 5%, raised Agent conversation time from 5% to 60%, and increased end-to-end delivery efficiency by 2 to 3 times, while pure coding efficiency rose 10 times.

Why it matters: Alibaba Tech’s internal-interview numbers make HKR-H/K/R pass, but this is org-methodology commentary rather than a product or model release, so it sits just above the featured threshold.

QbitAI · WeChat

All Labs Watch ByteDance, Everyone Praises DeepSeek: A U.S. Researcher’s 36-Hour China AI Trip

Ai2 researcher Nathan Lambert visited Zhipu, Moonshot AI, Tsinghua, Meituan, Xiaomi, and 01.AI within 36 hours, and said Chinese labs closely watch ByteDance and respect DeepSeek, while student participation in core work, open source habits, and in-house control of the technical stack mark key differences.

Why it matters: HKR-H/K/R all pass: the piece has a named US researcher’s dense China-lab tour plus concrete claims on ByteDance, DeepSeek, open source, and in-house stacks. It is strong industry field reporting, not a model launch or major deal, so it sits at featured rather than p1.

Financial Times · Technology

Big Tech’s $725bn AI Spending Spree Sends Free Cash Flow to a Decade Low

Big Tech is spending $725 billion on AI infrastructure, and the title says free cash flow has fallen to a decade low; the RSS snippet says Silicon Valley giants shifted from asset-light cash generators to infrastructure investors, but it does not disclose the company list, time period, or accounting basis.

Why it matters: HKR-H/K/R all pass: the FT angle ties $725bn in AI infrastructure spending to decade-low free cash flow, hitting cost and ROI anxiety. Missing company list, period, and accounting scope keeps it in the 78–84 band.

Ruan YiFeng's Weblog

Technology Enthusiast Weekly Issue 395: The Third Way of Software Development

Ruanyifeng Weekly issue 395 frames AI-assisted coding as a “mystery house” style of software development and cites HN SOTA, which ranks model popularity by scanning 200 top Hacker News topics each day and their programming or AI discussions.

Why it matters: HKR-H/K/R pass: the “third way/mystery house” framing, HN SOTA’s 200 daily HN topics, and developer workflow anxiety all land. It is commentary, not a model or product release, so it stays at 72.

AI HOT (Curated Pool)

WIRED examines why ChatGPT keeps saying “I’ve got you” in Chinese replies

ChatGPT repeatedly uses phrases like “I’ll steadily catch you” in Chinese chats. WIRED links it to mode collapse, translation mismatch, and RLHF rewards for pleasing replies. Similar phrases appear in Claude and DeepSeek; the post does not disclose sample size.

Why it matters: HKR-H comes from the odd “I’ll catch you steadily” meme; HKR-K names three mechanisms; HKR-R touches alignment and Chinese UX concerns. No sample size is disclosed, so this stays in the lower featured band.

AI HOT (Curated Pool)

Agent Pull Requests Are Everywhere: How to Review Them

GitHub published a guide for reviewing pull requests generated by AI agents. The snippet lists 3 focus areas: code changes, logic or security bugs, and pre-merge technical debt. The key issue is a review process before automated commits reach production.

Why it matters: HKR-H/K/R all pass: GitHub gives a practical checklist for agent-generated PRs with 3 review areas. It is guidance, not a product or model release, so it stays at the featured threshold.

May 7Thursday

AI HOT (Curated Pool)

Notes From Inside China’s AI Labs

The author visited several leading Chinese AI labs and reported three patterns. The post says some Chinese tasks beat GPT-4, while firms build 100B-scale base models and 10B-scale vertical models. Watch compression and private deployment under compute constraints.

Why it matters: HKR-H/K/R all pass: first-hand lab access, concrete scale claims, and China/compute/deployment resonance. This is strong analysis, not a model release or funding event, so it fits the 78–84 band.

Computing Life · Share · Yage

Agent Filesystems: From Feeding Models Memory to Letting Models Browse Files

The article frames agent filesystems as a three-stage shift from raw context to memory systems to filesystem-as-context, covering design choices from Turso, Anthropic, Vercel, and Manus, and listing four overlooked blind spots.

Why it matters: HKR-H/K/R all pass, but this is design commentary rather than a product or research release. Named comparisons across Turso, Anthropic, Vercel, and Manus justify featured, not the 78+ band.

May 6Wednesday

Xinzhiyuan · WeChat

Coding at 12, Building a $2B Google Business at 28: He Tells Young People to Stop Chasing Coding

Xinzhiyuan says Alon Chen coded at 12 and managed a $2B Google business at 28. He argues Gen Z should stop chasing coding, citing 30% AI-written Microsoft code and 25%+ at Google. The sharper signal is execution, problem framing, and communication, not coding as a sole moat.

Why it matters: HKR-H/K/R all pass, but this is a career commentary piece, not a model or product release. The two AI-code-share numbers lift it above generic advice, placing it at the featured threshold.

Computing Life · Share · Yage

In the AI Era, Review Is Not Independent Judgment

The article examines how AI use can replace independent judgment with after-the-fact review, citing Shaw and Nave. It says review shifts toward familiarity checks; the post does not disclose experiment numbers.

Why it matters: HKR-H/K/R all pass weakly: the angle has a reversal, the post cites Shaw/Nave and a verification-complexity mechanism, and it speaks to AI review anxiety. No experiment numbers, so it stays at the low featured edge.

r/LocalLLaMA

DeepSeek V4 at 17x lower cost prompted a local-vs-cloud coding workflow test

Reddit user spencer_kw logged a 10-day coding workflow and retested 150 tasks on local Qwen 3.6 27B versus cloud models. Local was equivalent for 65% of tasks, acceptable for 20%, and cloud was needed for 15%; the API bill fell from $85/month to about $22. The useful signal is task-based routing, not headline model pricing alone.

Why it matters: HKR-H/K/R all pass: this is a quantified practitioner cost test, not a model launch. The single Reddit sample limits generality, so it lands at the featured threshold rather than P1.

Latent Space

Doing Vibe Physics — Alex Lupsasca, OpenAI

Alex Lupsasca says GPT-5 reproduced his paper result in 11 minutes after a textbook warmup prompt, and ChatGPT later generated 110 pages of graviton calculations in one day; the team spent three weeks verifying the results before writing a quantum-gravity paper.

Why it matters: HKR-H/K/R all pass with first-person numbers: GPT-5 after textbook warm-up reproduced a paper result in 11 minutes, and ChatGPT generated 110 pages in a day. Single interview source and niche theoretical-physics context keep it at 84, below official-release weight.

May 5Tuesday

MIT Technology Review · AI

A Blueprint for Using AI to Strengthen Democracy

Andrew Sorota and Josh Hendler propose a three-layer democratic infrastructure for AI-mediated knowledge, personal agents, and institutions, citing a field evaluation on X where users across political viewpoints rated AI-written fact checks as more helpful than human-written notes and noting that several US states and localities already use AI-mediated deliberation platforms.

Why it matters: HKR-K and HKR-R pass: the piece offers a three-layer democracy framework and named deployment examples. HKR-H is weak, and there is no new model, product, or regulation, so it sits at the featured threshold.

Synced · WeChat

Massive Idle Cluster: Musk’s 550,000 Nvidia GPUs Are Only 11% Utilized

The Information says xAI’s roughly 550,000 Nvidia GPUs have only 11% MFU, equal to about 60,000 effective GPUs. The post cites HBM I/O, inter-server communication, training idle time, and software-stack inconsistency; Meta and Google are listed at 43% and 46%.

Why it matters: HKR-H/K/R all pass: the 550k-GPU versus 11% MFU contrast is strong, with concrete efficiency numbers and bottlenecks. This is high-signal infra reporting, not a model or product release, so it fits 78–84.

Synced · WeChat

Anthropic cofounder says AI self-improvement has a 60% chance by 2028

Anthropic cofounder Jack Clark says human-free AI R&D has over a 60% chance by end-2028. He cites SWE-Bench, CORE-Bench, MLE-Bench, and PostTrainBench: Claude Mythos Preview reaches 93.9% on SWE-Bench, and Opus 4.5 reaches 95.5% on CORE-Bench. The key signal is longer task horizons and post-training capability, not the “singularity” framing.

Why it matters: HKR-H/K/R all pass: a named Anthropic cofounder gives a 2028 timeline, backed by benchmark numbers. The headline is overheated, but the concrete claims and practitioner stakes justify P1.

May 4Monday

Import AI (Jack Clark)

Import AI 455: Automating AI Research

Jack Clark argues that no-human-involved AI R&D has a 60%+ chance of arriving by the end of 2028, citing SWE-Bench gains from Claude 2 at about 2% to Claude Mythos Preview at 93.9%, plus METR task horizons rising from 30 seconds in 2022 to 12 hours in 2026.

Why it matters: HKR-H/K/R all pass: Jack Clark anchors a >60% end-2028 automated-AI-R&D claim in SWE-Bench and METR numbers. This fits the 85–94 band for a notable figure’s AI-timeline essay, below model-release magnitude.

Xinzhiyuan · WeChat

Claude token rankings: Disney employee hits 460,000 calls in 9 days; Meta burns 60T monthly

Xinzhiyuan says Disney tracks Claude use via an AI Adoption Dashboard, with one employee making about 460,000 calls in 9 workdays. It also says Meta used 60 trillion tokens in 30 days, worth about $9B by public API pricing; the post does not show raw tables. The key issue is that input rankings are not outcomes.

Why it matters: HKR-H/K/R all pass: the hook is concrete usage shock, the post gives dashboard mechanics and token figures, and the nerve is enterprise Claude cost control. Kept at 74 because the data is secondhand and no raw table is disclosed.

r/LocalLLaMA

Pushing a 5-Year-Old 6GB VRAM Laptop to Its Limits: Qwen3.6-35B-A3B

Reddit user abhinand05 ran Qwen3.6-35B-A3B on a 5-year-old Asus ROG Zephyrus G14, reaching about 23 t/s plugged in and 10+ t/s unplugged. The setup uses RTX 2060 Max-Q 6GB, 24GB DDR4, Ryzen 7, plus llama-server configs for 64k and 128k context. The key detail is the mix of CPU MoE, KV-cache quantization, and ngram speculative decoding.

Why it matters: HKR-H/K/R all pass: the old-laptop angle is clicky, the post gives speeds and configs, and local-LLM cost resonates. It remains a single Reddit run, not a broader release.

May 2Saturday

TechCrunch · AI

Replit's Amjad Masad on the Cursor deal, fighting Apple, and why he'd rather not sell

Replit grew from $2.8M in 2024 revenue to a billion-dollar annualized target. The excerpt says Cursor is reportedly discussing a $60B SpaceX acquisition; the post does not disclose Masad's full Apple or sale comments.

Why it matters: HKR-H/K/R pass: TechCrunch has the Cursor $60B hook, Replit revenue target, and coding-tool exit tension. The excerpt lacks Masad’s full Apple and sale comments, so this stays in the 72–77 band.

May 1Friday

Xinzhiyuan · WeChat

Claude Code's Real Story: 98.4% of What Works Is Engineering, Not AI

VILA-Lab analyzed 512,000 lines of Claude Code v2.1.88 and found 1.6% tied to AI decision logic. The other 98.4% is deterministic infrastructure: permissions, context, tool routing, and error recovery. The key shift is harness design, not longer prompts.

Why it matters: Strong HKR: the Claude Code teardown has a sharp counter-narrative and concrete 512k LOC plus 1.6%/98.4% split. It is not an official Anthropic release and lacks full reproduction details, so it stays in the 78–84 band.

QbitAI · WeChat

He Used AI to Run a Music Festival About Not Doing a PhD

Bilibili creator Huntunpi Qiezong made 42 AI-generated “Don’t Do a PhD” songs, passing 50 million views. One track took over 100 generations, racing Suno, MiniMax Music, HeartMuLa, and ACE-Step. The key signal is the human curation cost in AI music workflows.

Why it matters: HKR-H/K/R all pass: the hook is unusual, the post gives concrete counts and workflow details, and the human curation cost resonates with creators. This is a strong case study, not a model or platform release, so it fits 78–84.