Skip to content

#大佬观点

1 today

May 10Sunday

Xinzhiyuan · WeChat

Harsh Claim: Top Silicon Valley AI Is One Year Ahead of the World

Elad Gil claims top AI lab employees are 3-4 months ahead of Silicon Valley, while Silicon Valley is 3-6 months ahead of New York; the post cites Mythos’ 73% success rate in expert cyberattack simulations as evidence in a disputed “geographic time gap” argument.

Why it matters: HKR-H/K/R all pass: the lab-to-user lag hook is clickable, and the post cites 3–4 months, 3–6 months, and a 73% Mythos figure. It is secondhand commentary, not a model or product release, so it stays in the 72–77 threshold band.

May 9Saturday

AI HOT (Curated Pool)

Using Codex to debug and verify fixes in parallel

The author uses Codex in temporary crabbox environments to recreate bug states, verify failures, apply fixes, and re-verify them, while running 10 sessions in parallel to avoid local state pollution and speed loss.

Why it matters: HKR-H/K/R all pass, but this is a single first-person workflow note, not a product release or benchmark. The 10-session Codex/crabbox setup earns featured-level practical signal, near the lower band.

Latent Space

Anthropic growing 10x/year while others lay off over 10% of staff

Anthropic is described as growing 10x annually and being valued at $1T-$1.2T, while the post cites layoffs of 40% at Block, 14% at Coinbase, and 20% at Cloudflare under AI-readiness framing.

Why it matters: HKR-H/K/R all pass: the title has contrast, the post gives growth, valuation, and layoff figures, and it hits jobs plus AI-capital concentration. It is high-signal industry commentary, not an official funding or product event, so 78-84 fits.

AI HOT (Curated Pool)

Claude Code Practice: The Effectiveness of HTML Output

Thariq Shihipar recommends requesting HTML output from Claude, and the post cites GPT-5.5 generating an interactive Linux vulnerability page with SVG diagrams, interactive components, and in-page navigation.

Why it matters: HKR-H/K/R all pass, but this is a workflow tip rather than a Claude release. As a quality Claude Code tutorial, it sits in the 72–77 band, with Simon Willison’s source authority clearing featured.

May 8Friday

AI HOT (Curated Pool)

Robotics Endgame: A Physical AGI Roadmap and LLM Analogy

The speaker presented a physical AGI roadmap with six named components: video world models, WAM, EgoScale, dexterity scaling laws, physical reinforcement learning, and DreamDojo; the snippet also mentions a 2016 OpenAI DGX-1 signing story with Jensen and Elon.

Why it matters: HKR-H/K/R all pass: the physical-AGI endgame hook is strong, the post gives a 6-part roadmap, and robotics practitioners will debate the path. It is still a personal roadmap, not a release or benchmark, so it sits in 78–84.

Alibaba Technology · WeChat

The AI-Native Era: Where R&D Organizations Go Next

Xu Xiaobin cites internal interviews showing that engineers who use AI heavily cut coding time from 30% to 5%, raised Agent conversation time from 5% to 60%, and increased end-to-end delivery efficiency by 2 to 3 times, while pure coding efficiency rose 10 times.

Why it matters: Alibaba Tech’s internal-interview numbers make HKR-H/K/R pass, but this is org-methodology commentary rather than a product or model release, so it sits just above the featured threshold.

QbitAI · WeChat

All Labs Watch ByteDance, Everyone Praises DeepSeek: A U.S. Researcher’s 36-Hour China AI Trip

Ai2 researcher Nathan Lambert visited Zhipu, Moonshot AI, Tsinghua, Meituan, Xiaomi, and 01.AI within 36 hours, and said Chinese labs closely watch ByteDance and respect DeepSeek, while student participation in core work, open source habits, and in-house control of the technical stack mark key differences.

Why it matters: HKR-H/K/R all pass: the piece has a named US researcher’s dense China-lab tour plus concrete claims on ByteDance, DeepSeek, open source, and in-house stacks. It is strong industry field reporting, not a model launch or major deal, so it sits at featured rather than p1.

Financial Times · Technology

Big Tech’s $725bn AI Spending Spree Sends Free Cash Flow to a Decade Low

Big Tech is spending $725 billion on AI infrastructure, and the title says free cash flow has fallen to a decade low; the RSS snippet says Silicon Valley giants shifted from asset-light cash generators to infrastructure investors, but it does not disclose the company list, time period, or accounting basis.

Why it matters: HKR-H/K/R all pass: the FT angle ties $725bn in AI infrastructure spending to decade-low free cash flow, hitting cost and ROI anxiety. Missing company list, period, and accounting scope keeps it in the 78–84 band.

Ruan YiFeng's Weblog

Technology Enthusiast Weekly Issue 395: The Third Way of Software Development

Ruanyifeng Weekly issue 395 frames AI-assisted coding as a “mystery house” style of software development and cites HN SOTA, which ranks model popularity by scanning 200 top Hacker News topics each day and their programming or AI discussions.

Why it matters: HKR-H/K/R pass: the “third way/mystery house” framing, HN SOTA’s 200 daily HN topics, and developer workflow anxiety all land. It is commentary, not a model or product release, so it stays at 72.

AI HOT (Curated Pool)

WIRED examines why ChatGPT keeps saying “I’ve got you” in Chinese replies

ChatGPT repeatedly uses phrases like “I’ll steadily catch you” in Chinese chats. WIRED links it to mode collapse, translation mismatch, and RLHF rewards for pleasing replies. Similar phrases appear in Claude and DeepSeek; the post does not disclose sample size.

Why it matters: HKR-H comes from the odd “I’ll catch you steadily” meme; HKR-K names three mechanisms; HKR-R touches alignment and Chinese UX concerns. No sample size is disclosed, so this stays in the lower featured band.

AI HOT (Curated Pool)

Agent Pull Requests Are Everywhere: How to Review Them

GitHub published a guide for reviewing pull requests generated by AI agents. The snippet lists 3 focus areas: code changes, logic or security bugs, and pre-merge technical debt. The key issue is a review process before automated commits reach production.

Why it matters: HKR-H/K/R all pass: GitHub gives a practical checklist for agent-generated PRs with 3 review areas. It is guidance, not a product or model release, so it stays at the featured threshold.

May 7Thursday

AI HOT (Curated Pool)

Notes From Inside China’s AI Labs

The author visited several leading Chinese AI labs and reported three patterns. The post says some Chinese tasks beat GPT-4, while firms build 100B-scale base models and 10B-scale vertical models. Watch compression and private deployment under compute constraints.

Why it matters: HKR-H/K/R all pass: first-hand lab access, concrete scale claims, and China/compute/deployment resonance. This is strong analysis, not a model release or funding event, so it fits the 78–84 band.

Computing Life · Share · Yage

Agent Filesystems: From Feeding Models Memory to Letting Models Browse Files

The article frames agent filesystems as a three-stage shift from raw context to memory systems to filesystem-as-context, covering design choices from Turso, Anthropic, Vercel, and Manus, and listing four overlooked blind spots.

Why it matters: HKR-H/K/R all pass, but this is design commentary rather than a product or research release. Named comparisons across Turso, Anthropic, Vercel, and Manus justify featured, not the 78+ band.

May 6Wednesday

Xinzhiyuan · WeChat

Coding at 12, Building a $2B Google Business at 28: He Tells Young People to Stop Chasing Coding

Xinzhiyuan says Alon Chen coded at 12 and managed a $2B Google business at 28. He argues Gen Z should stop chasing coding, citing 30% AI-written Microsoft code and 25%+ at Google. The sharper signal is execution, problem framing, and communication, not coding as a sole moat.

Why it matters: HKR-H/K/R all pass, but this is a career commentary piece, not a model or product release. The two AI-code-share numbers lift it above generic advice, placing it at the featured threshold.

Computing Life · Share · Yage

In the AI Era, Review Is Not Independent Judgment

The article examines how AI use can replace independent judgment with after-the-fact review, citing Shaw and Nave. It says review shifts toward familiarity checks; the post does not disclose experiment numbers.

Why it matters: HKR-H/K/R all pass weakly: the angle has a reversal, the post cites Shaw/Nave and a verification-complexity mechanism, and it speaks to AI review anxiety. No experiment numbers, so it stays at the low featured edge.

r/LocalLLaMA

DeepSeek V4 at 17x lower cost prompted a local-vs-cloud coding workflow test

Reddit user spencer_kw logged a 10-day coding workflow and retested 150 tasks on local Qwen 3.6 27B versus cloud models. Local was equivalent for 65% of tasks, acceptable for 20%, and cloud was needed for 15%; the API bill fell from $85/month to about $22. The useful signal is task-based routing, not headline model pricing alone.

Why it matters: HKR-H/K/R all pass: this is a quantified practitioner cost test, not a model launch. The single Reddit sample limits generality, so it lands at the featured threshold rather than P1.

Latent Space

Doing Vibe Physics — Alex Lupsasca, OpenAI

Alex Lupsasca says GPT-5 reproduced his paper result in 11 minutes after a textbook warmup prompt, and ChatGPT later generated 110 pages of graviton calculations in one day; the team spent three weeks verifying the results before writing a quantum-gravity paper.

Why it matters: HKR-H/K/R all pass with first-person numbers: GPT-5 after textbook warm-up reproduced a paper result in 11 minutes, and ChatGPT generated 110 pages in a day. Single interview source and niche theoretical-physics context keep it at 84, below official-release weight.

May 5Tuesday

MIT Technology Review · AI

A Blueprint for Using AI to Strengthen Democracy

Andrew Sorota and Josh Hendler propose a three-layer democratic infrastructure for AI-mediated knowledge, personal agents, and institutions, citing a field evaluation on X where users across political viewpoints rated AI-written fact checks as more helpful than human-written notes and noting that several US states and localities already use AI-mediated deliberation platforms.

Why it matters: HKR-K and HKR-R pass: the piece offers a three-layer democracy framework and named deployment examples. HKR-H is weak, and there is no new model, product, or regulation, so it sits at the featured threshold.

Synced · WeChat

Massive Idle Cluster: Musk’s 550,000 Nvidia GPUs Are Only 11% Utilized

The Information says xAI’s roughly 550,000 Nvidia GPUs have only 11% MFU, equal to about 60,000 effective GPUs. The post cites HBM I/O, inter-server communication, training idle time, and software-stack inconsistency; Meta and Google are listed at 43% and 46%.

Why it matters: HKR-H/K/R all pass: the 550k-GPU versus 11% MFU contrast is strong, with concrete efficiency numbers and bottlenecks. This is high-signal infra reporting, not a model or product release, so it fits 78–84.

Synced · WeChat

Anthropic cofounder says AI self-improvement has a 60% chance by 2028

Anthropic cofounder Jack Clark says human-free AI R&D has over a 60% chance by end-2028. He cites SWE-Bench, CORE-Bench, MLE-Bench, and PostTrainBench: Claude Mythos Preview reaches 93.9% on SWE-Bench, and Opus 4.5 reaches 95.5% on CORE-Bench. The key signal is longer task horizons and post-training capability, not the “singularity” framing.

Why it matters: HKR-H/K/R all pass: a named Anthropic cofounder gives a 2028 timeline, backed by benchmark numbers. The headline is overheated, but the concrete claims and practitioner stakes justify P1.

May 4Monday

Import AI (Jack Clark)

Import AI 455: Automating AI Research

Jack Clark argues that no-human-involved AI R&D has a 60%+ chance of arriving by the end of 2028, citing SWE-Bench gains from Claude 2 at about 2% to Claude Mythos Preview at 93.9%, plus METR task horizons rising from 30 seconds in 2022 to 12 hours in 2026.

Why it matters: HKR-H/K/R all pass: Jack Clark anchors a >60% end-2028 automated-AI-R&D claim in SWE-Bench and METR numbers. This fits the 85–94 band for a notable figure’s AI-timeline essay, below model-release magnitude.

Xinzhiyuan · WeChat

Claude token rankings: Disney employee hits 460,000 calls in 9 days; Meta burns 60T monthly

Xinzhiyuan says Disney tracks Claude use via an AI Adoption Dashboard, with one employee making about 460,000 calls in 9 workdays. It also says Meta used 60 trillion tokens in 30 days, worth about $9B by public API pricing; the post does not show raw tables. The key issue is that input rankings are not outcomes.

Why it matters: HKR-H/K/R all pass: the hook is concrete usage shock, the post gives dashboard mechanics and token figures, and the nerve is enterprise Claude cost control. Kept at 74 because the data is secondhand and no raw table is disclosed.

r/LocalLLaMA

Pushing a 5-Year-Old 6GB VRAM Laptop to Its Limits: Qwen3.6-35B-A3B

Reddit user abhinand05 ran Qwen3.6-35B-A3B on a 5-year-old Asus ROG Zephyrus G14, reaching about 23 t/s plugged in and 10+ t/s unplugged. The setup uses RTX 2060 Max-Q 6GB, 24GB DDR4, Ryzen 7, plus llama-server configs for 64k and 128k context. The key detail is the mix of CPU MoE, KV-cache quantization, and ngram speculative decoding.

Why it matters: HKR-H/K/R all pass: the old-laptop angle is clicky, the post gives speeds and configs, and local-LLM cost resonates. It remains a single Reddit run, not a broader release.

May 2Saturday

TechCrunch · AI

Replit's Amjad Masad on the Cursor deal, fighting Apple, and why he'd rather not sell

Replit grew from $2.8M in 2024 revenue to a billion-dollar annualized target. The excerpt says Cursor is reportedly discussing a $60B SpaceX acquisition; the post does not disclose Masad's full Apple or sale comments.

Why it matters: HKR-H/K/R pass: TechCrunch has the Cursor $60B hook, Replit revenue target, and coding-tool exit tension. The excerpt lacks Masad’s full Apple and sale comments, so this stays in the 72–77 band.

May 1Friday

Xinzhiyuan · WeChat

Claude Code's Real Story: 98.4% of What Works Is Engineering, Not AI

VILA-Lab analyzed 512,000 lines of Claude Code v2.1.88 and found 1.6% tied to AI decision logic. The other 98.4% is deterministic infrastructure: permissions, context, tool routing, and error recovery. The key shift is harness design, not longer prompts.

Why it matters: Strong HKR: the Claude Code teardown has a sharp counter-narrative and concrete 512k LOC plus 1.6%/98.4% split. It is not an official Anthropic release and lacks full reproduction details, so it stays in the 78–84 band.

QbitAI · WeChat

He Used AI to Run a Music Festival About Not Doing a PhD

Bilibili creator Huntunpi Qiezong made 42 AI-generated “Don’t Do a PhD” songs, passing 50 million views. One track took over 100 generations, racing Suno, MiniMax Music, HeartMuLa, and ACE-Step. The key signal is the human curation cost in AI music workflows.

Why it matters: HKR-H/K/R all pass: the hook is unusual, the post gives concrete counts and workflow details, and the human curation cost resonates with creators. This is a strong case study, not a model or platform release, so it fits 78–84.

Bloomberg Technology

The Audio Industry Is Grappling with the Rise of ‘Podslop’

Podcast Index says 39% of new podcasts over nine days were likely AI-generated. The title cites industry concern over “podslop”; the post does not disclose detection methods, sample size, or platform split.

Why it matters: HKR-H/K/R pass, but the post lacks detection method, sample size, and platform split. Bloomberg plus Podcast Index’s 39% claim clears featured, not a major industry event.

Bloomberg Technology

Meta Needs to Stop Spending as If It's a Cloud Giant: Lee

Dave Lee criticized Meta for spending on AI like a cloud giant, with capex reaching up to $145 billion. The post says Meta lacks Amazon- or Google-style cloud sales growth from AI. The key issue is capex without matching visible revenue.

Why it matters: HKR-H/K/R all pass: a sharp Meta capex mismatch, a $145B figure, and infra-spend anxiety. It is commentary rather than a major release, so it sits at the 72 threshold.

Apr 30Thursday

r/LocalLLaMA

Notes on what actually breaks when you run a coding agent on small local models

A Reddit user tested small local and free-tier cloud models for weeks on multi-file coding tasks. Sub-7B structured output was unreliable; failures included markdown fences, wrong-file edits, and read/write misclassification, with post-processing and validation as fixes.

Why it matters: HKR-H/K/R pass: the post names real local coding-agent failure points, a sub-7B threshold, four failure classes, and mitigations. Reddit single-post scope keeps it below release-tier news, so 75.

Xinzhiyuan · WeChat

AI Raw Proofs Pile Up on GitHub as Terence Tao Says Solving Alone Is Not Enough

Terence Tao says math is shifting from proof scarcity to proof abundance, with 20-plus AI solutions pending assessment on an Erdős problems GitHub page. The post says GPT-5.4 Pro generated an Erdős #1196 approach in 80 minutes, and Tao verified the core within 24 hours. The key issue is verification and digestion workflow, not raw proof count.

Why it matters: All HKR axes pass: Tao plus GitHub proof backlog gives HKR-H, while 20+ pending AI solutions and an 80-minute GPT-5.4 Pro claim give HKR-K. This is not a model release, so it stays below 85.

Financial Times · Technology

Google outpaces Big Tech rivals as AI spending plans rise to $725bn

Google outpaced Big Tech rivals as AI spending plans rose to $725bn. The snippet says Meta fell on higher capex, while Alphabet cloud grew faster than Amazon and Microsoft. The post does not disclose the spending split or timeframe.

Why it matters: HKR-H/K/R all pass: the FT gives a $725bn AI capex race and Alphabet cloud lead. Missing company split, time frame, and model-level spend keep it in the lower 78–84 band.

Latent Space

[AINews] The Inference Inflection

Latent Space argues inference demand has hit an inflection point, citing its Apr 28-29, 2026 AINews roundup. Jensen Huang is quoted saying per-task compute rose about 10,000x in two years, with usage up about 100x. The key watchpoints are CPU sandboxes, agent harnesses, and split inference workloads.

Why it matters: HKR-H/K/R all pass, but this is a Latent Space AINews roundup and trend read, not a model launch or major product release. It fits the upper featured-threshold band for insightful commentary.

OpenAI News

Where the goblins came from

OpenAI posted about goblin outputs in GPT-5; only an RSS snippet is available. The snippet names timeline, root cause, and fixes, but does not disclose mechanisms or conditions. The key issue is how personality-driven quirks enter model behavior.

Why it matters: HKR-H and HKR-R pass: OpenAI is addressing odd GPT-5 behavior with clear talk value. HKR-K fails because the RSS text lacks reproduction conditions, timeline, and fix details, so it stays in the low featured band.

Dwarkesh Patel podcast

Reiner Pope: The Math Behind How LLMs Are Trained and Served

Dwarkesh interviewed Reiner Pope in a 1-session blackboard lecture on LLM training and serving. The post lists 7 timestamps on batch size, MoE rack layout, pipeline parallelism, KV cache, and API pricing. The key mechanism is cost: without batching, serving economics can be 1,000x worse.

Why it matters: HKR-H/K/R all pass: the 1000x batching cost hook, concrete serving mechanics, and inference-cost resonance are strong. This is a high-quality tutorial, not a same-day industry event, so it stays at 77.

Apr 28Tuesday

The Verge · AI

Attack of the Killer Script Kiddies

The Verge discusses Claude Mythos and AI bug finding, citing DARPA AIxCC scans over 54 million code lines. Teams found most seeded flaws plus over a dozen unseeded bugs; the RSS snippet does not disclose Mythos benchmarks, pricing, or access terms.

Why it matters: HKR-H/K/R all pass: the hook is strong, DARPA AIxCC supplies concrete numbers, and the security angle resonates. No Claude Mythos benchmark, pricing, or access terms are disclosed, so it stays in the featured-threshold band.

Bloomberg Technology

OpenAI Misses Its Own User and Sales Goals, WSJ Reports

WSJ says OpenAI missed its own new-user and sales goals. The RSS snippet cites internal concern over AI infrastructure spending. The post does not disclose targets, gaps, timing, or spend size.

Why it matters: HKR-H/R are strong because OpenAI growth missed its plan and infra spend is the nerve. HKR-K is thin: WSJ reports the miss, but target size, gap, period, and spend are undisclosed.

Latent Space

Physical AI that Moves the World — Qasar Younis & Peter Ludwig, Applied Intuition

Applied Intuition’s founders reviewed a 10-year physical AI path, with the company valued at $15B. The post cites 30+ products, 18 of the top 20 non-Chinese automakers as customers, and L4 driverless trucks in Japan. The key constraint is onboard deployment: millisecond latency, low power, small models, and safety validation.

Why it matters: HKR-H/K/R all pass: the piece ties a major Physical AI company to real AV deployment with customer, valuation, and L4 details. No new model or major launch is disclosed, so it stays in the 78–84 band.

X · @dotey

The West forgot how to build things, and may forget how to write code

Denis Stetskov compares Western defense production gaps with AI coding, citing Stinger orders placed in 2022 for 2026 delivery. He says Europe’s 1M-shell target was 9 months late, and METR found senior developers 19% slower with AI. The key risk is the junior-engineer pipeline, not code generation speed.

Why it matters: HKR-H/K/R all pass: the analogy is clickable, the post gives concrete defense and METR numbers, and the junior-engineer pipeline resonates. X translation/commentary limits authority, so it sits just above the featured threshold.

Apr 27Monday

Dwarkesh Patel podcast

What I've been Thinking About This Weekend: Open Questions, Intelligence vs Power, Verification in Science

Dwarkesh lists open AI questions, including that five hyperscalers own over 70% of global AI compute. He asks about coding agents, KV cache costs, merging training with inference, and online learning; the post gives questions, not experimental answers.

Why it matters: HKR-H/K/R all pass: Dwarkesh adds a concrete compute-concentration claim and practitioner-relevant questions. No experiment, release, or policy change, so it stays in the 72–77 commentary band.

Hacker News front page

Running Local LLMs Offline on a Ten-Hour Flight

Dmitri Lerko ran Gemma 4 31B and Qwen 4.6 36B locally during a 10-hour flight with no Wi‑Fi. The MacBook Pro M5 Max had 128GB unified memory and a 40-core GPU; sustained load used about 1% battery per minute, and performance degraded past 100k tokens. The sharp finding is instrumentation: an iPhone cable delivered 60W, while a MacBook cable delivered 94W under the same load.

Why it matters: HKR-H/K/R all pass: this is a named first-person local-inference test with concrete hardware, model, battery, and power numbers. Scope stays practical rather than industry-shaking, so it lands in the 72–77 band.