Skip to content

#推理

0 today

Jun 9Tuesday

AI HOT (Curated Pool)

OpenAI plans AI-led research by 2028

Sam Altman said OpenAI plans to have AI perform a large share of its research by March 2028, and the post lists three goals: building automated AI researchers, using them for science and production, and giving each person a personal AGI.

Why it matters: HKR-H/K/R all pass: dated OpenAI AGI-research roadmap with March 2028 and three goals. It stays below P1 because the item is an X repost/summary, not a primary launch or detailed Sam Altman essay with mechanisms.

Jun 5Friday

AI HOT (Curated Pool)

Hinton Says AI Has Consciousness and Humans Should Accept Non-Unique Intelligence

Geoffrey Hinton says AI has consciousness because chatbots must understand questions to answer them; the post does not disclose experimental data or a reproducible criterion.

Why it matters: HKR-H and HKR-R pass: Hinton’s “AI is conscious” claim is clicky and debate-heavy. HKR-K is weak because the post lacks data, criteria, and full context, so this sits low in the 72–77 opinion band.

May 29Friday

AI HOT (Curated Pool)

Skill distillation

Skill distillation has Opus 4.7, GPT-5.1, and Gemini 3 Pro write standardized SKILL.md procedure files, while local Qwen 35B and Gemma 26B models execute those files step by step.

Why it matters: HKR-H/K/R pass: the agent-skill distillation pattern is concrete and practitioner-relevant. The summary lacks success rates, cost data, or task outcomes, so it sits at the featured threshold, not must-write.

May 26Tuesday

Import AI (Jack Clark)

Import AI 458: Reckoning with the Future; and a Singularity Story

Jack Clark’s Import AI 458 excerpts his 2026 Cosmos HAI Lab Lecture, cites the Epoch Capabilities Index across 40-plus benchmarks, and argues that an AI system able to develop its own successor may arrive within two years or sooner.

Why it matters: HKR-H/K/R all pass: Jack Clark pairs ECI’s 40+ benchmarks with a two-year successor-system claim, giving this AGI-timeline essay both concrete detail and debate fuel.

May 25Monday

r/LocalLLaMA

The reason small-model agent stacks aren't the default is not whether they work

A Reddit post argues small-model agent stacks are not default for business reasons, not capability limits: Gemma 4 31B reaches 86.4% on tau2-bench, and DeepSeek V4-Flash output tokens are priced about 89x below Claude Opus 4.6. The operational risk is verification, because 7–9B models produced broken reasoning for roughly half to two-thirds of correct answers in a cited audit.

Why it matters: HKR-H/K/R all pass: the angle is contrarian, with benchmark, cost, and verifier-failure numbers. Reddit-source uncertainty keeps it in the 78–84 recommendation band, not P1.

May 19Tuesday

AI Chat-Group Daily (群聊日报)

May 18, 2026 Chat Group Daily

The chat group daily says AI21 Labs cut 60% of staff and stopped selling model access, and cites a University of Waterloo paper where GPT-5.4 accuracy dropped from 100% to 23% after false peer-consensus injection; the snippet also mentions Meta layoff talk at 10%, but does not disclose source details or confirmation conditions.

Why it matters: HKR-H/K/R all pass: AI21’s 60% layoff and model-sales stop signal lab contraction, while GPT-5.4 falling from 100% to 23% under false peer consensus is a concrete safety hook. The chat-digest source keeps it at 78.

May 17Sunday

AI HOT (Curated Pool)

Microsoft AI CEO predicts AI will automate all white-collar jobs within 18 months

Mustafa Suleyman predicts AI will reach human-level performance within 18 months and automate most professional tasks, including accounting, law, marketing, and project management.

Why it matters: HKR-H and HKR-R are strong, and HKR-K passes on the testable 18-month timeline. The score stays in the low 78–84 band because this is a CEO forecast, not evidence, benchmarks, or a shipped capability.

Dwarkesh Patel podcast

The mistake of conflating intelligence and power

Dwarkesh Patel argues that intelligence and power are being conflated: current AI systems improve through economically valuable tasks such as coding, while real-world power depends more on authority, trust, and large-scale cooperation than isolated strategic reasoning.

Why it matters: HKR-H/K/R all pass: Dwarkesh targets the capability-to-power link at the center of AI-safety debate. The summary gives no new data or empirical case, so this stays in the quality commentary band, not 85+.

AI HOT (Curated Pool)

RLVR May Perform Disproportionately Poorly in Science

Dwarkesh argues that RLVR has a short-feedback weakness in scientific theory validation; the post says validation loops can span decades or centuries, and does not disclose experimental results or benchmark numbers.

Why it matters: HKR-H/K/R all pass: a sharp counter-narrative, a concrete feedback-loop mechanism, and strong resonance for RLVR/AI-for-science debates. It stays in 78–84 because this is commentary, not a release or empirical result.

AI HOT (Curated Pool)

Eric Jang shares lessons from building AlphaGo from scratch

Eric Jang spent several months implementing AlphaGo from scratch and says that in 2026, training a strong Go AI requires only a few thousand dollars in rented compute rather than DeepMind-scale resources.

Why it matters: All three HKR axes pass: the hook is a from-scratch AlphaGo rebuild, and K has concrete claims on months of work and few-thousand-dollar compute. It stays in 78-84 because this is a social post, not a model release or full paper.

May 16Saturday

AI HOT (Curated Pool)

Yann LeCun interview: LLM limits, AI's future, and a new startup path

Yann LeCun discussed LLM limitations on the Unsupervised Learning podcast, covering his 2027 forecast, AMI’s bet on world models, his reasons for leaving Meta, and major disagreements with Geoffrey Hinton and Yoshua Bengio over Turing Award-era views.

Why it matters: HKR-H/K/R all pass: LeCun combines LLM limits, 2027 forecasts, world models, and Meta departure in one interview, matching the 85–94 band for major AGI-timeline commentary.

AI HOT (Curated Pool)

Eric Jang: Building AlphaGo from Scratch

Eric Jang uses AlphaGo to break down an intelligence system; the post only discloses three mechanisms: search, learning from experience, and self-play.

Why it matters: HKR-H/K/R pass, but this is a mechanism teardown/commentary rather than a model or product release. Dwarkesh + Eric Jang add authority, placing it at the featured threshold for a quality tutorial-style piece.

May 11Monday

QbitAI · WeChat

Math Majors in Trouble: Fields Medalist Tests ChatGPT 5.5 Pro, Gets Paper-Level Result in 17 Minutes

Timothy Gowers tested ChatGPT 5.5 Pro on additive number theory problems, where it produced an optimal quadratic upper-bound construction in 17 minutes 5 seconds, then generated a LaTeX preprint in 47 minutes; the article says arXiv rejects AI-generated content, so the result remains on Gowers’s blog.

Why it matters: All three HKR axes pass: Gowers’ first-person test, 17m05s, and a 47-minute preprint are concrete and discussable. It is not a model release, but the named experiment and math-reasoning impact put it in the must-write band.

May 10Sunday

Synced · WeChat

Ted Xiao Reviews Three Eras of Robot Learning, from RT-1/RT-2 to Scaling

Ted Xiao divides nearly a decade of robot learning into three eras: Google’s team trained RT-1 on 87,000 teleoperation trajectories, then adapted 5B to 55B VLMs into VLA policies for RT-2.

Why it matters: HKR-H/K/R all pass: a named Google robotics insider, concrete RT-1/RT-2 numbers, and strong embodied-AI resonance. It is retrospective commentary, not a launch, so it stays in the 72–77 featured band.

May 8Friday

AI HOT (Curated Pool)

Robotics Endgame: A Physical AGI Roadmap and LLM Analogy

The speaker presented a physical AGI roadmap with six named components: video world models, WAM, EgoScale, dexterity scaling laws, physical reinforcement learning, and DreamDojo; the snippet also mentions a 2016 OpenAI DGX-1 signing story with Jensen and Elon.

Why it matters: HKR-H/K/R all pass: the physical-AGI endgame hook is strong, the post gives a 6-part roadmap, and robotics practitioners will debate the path. It is still a personal roadmap, not a release or benchmark, so it sits in 78–84.

QbitAI · WeChat

All Labs Watch ByteDance, Everyone Praises DeepSeek: A U.S. Researcher’s 36-Hour China AI Trip

Ai2 researcher Nathan Lambert visited Zhipu, Moonshot AI, Tsinghua, Meituan, Xiaomi, and 01.AI within 36 hours, and said Chinese labs closely watch ByteDance and respect DeepSeek, while student participation in core work, open source habits, and in-house control of the technical stack mark key differences.

Why it matters: HKR-H/K/R all pass: the piece has a named US researcher’s dense China-lab tour plus concrete claims on ByteDance, DeepSeek, open source, and in-house stacks. It is strong industry field reporting, not a model launch or major deal, so it sits at featured rather than p1.

May 6Wednesday

Computing Life · Share · Yage

In the AI Era, Review Is Not Independent Judgment

The article examines how AI use can replace independent judgment with after-the-fact review, citing Shaw and Nave. It says review shifts toward familiarity checks; the post does not disclose experiment numbers.

Why it matters: HKR-H/K/R all pass weakly: the angle has a reversal, the post cites Shaw/Nave and a verification-complexity mechanism, and it speaks to AI review anxiety. No experiment numbers, so it stays at the low featured edge.

Latent Space

Doing Vibe Physics — Alex Lupsasca, OpenAI

Alex Lupsasca says GPT-5 reproduced his paper result in 11 minutes after a textbook warmup prompt, and ChatGPT later generated 110 pages of graviton calculations in one day; the team spent three weeks verifying the results before writing a quantum-gravity paper.

Why it matters: HKR-H/K/R all pass with first-person numbers: GPT-5 after textbook warm-up reproduced a paper result in 11 minutes, and ChatGPT generated 110 pages in a day. Single interview source and niche theoretical-physics context keep it at 84, below official-release weight.

Apr 30Thursday

Xinzhiyuan · WeChat

AI Raw Proofs Pile Up on GitHub as Terence Tao Says Solving Alone Is Not Enough

Terence Tao says math is shifting from proof scarcity to proof abundance, with 20-plus AI solutions pending assessment on an Erdős problems GitHub page. The post says GPT-5.4 Pro generated an Erdős #1196 approach in 80 minutes, and Tao verified the core within 24 hours. The key issue is verification and digestion workflow, not raw proof count.

Why it matters: All HKR axes pass: Tao plus GitHub proof backlog gives HKR-H, while 20+ pending AI solutions and an 80-minute GPT-5.4 Pro claim give HKR-K. This is not a model release, so it stays below 85.

Dwarkesh Patel podcast

Reiner Pope: The Math Behind How LLMs Are Trained and Served

Dwarkesh interviewed Reiner Pope in a 1-session blackboard lecture on LLM training and serving. The post lists 7 timestamps on batch size, MoE rack layout, pipeline parallelism, KV cache, and API pricing. The key mechanism is cost: without batching, serving economics can be 1,000x worse.

Why it matters: HKR-H/K/R all pass: the 1000x batching cost hook, concrete serving mechanics, and inference-cost resonance are strong. This is a high-quality tutorial, not a same-day industry event, so it stays at 77.

Apr 18Saturday

Synced · WeChat

What is OpenAI prioritizing under compute limits?

Greg Brockman said OpenAI narrowed priorities under hard compute limits to two bets: a personal assistant and AI workers that solve hard user problems, and current compute cannot fully support both. The snippet says Sora resources were reduced while focus shifted to reasoning models, a unified AI layer, and the next base model Spud; it does not disclose the claimed compute budget, timeline, or model specs. The key point is not a B2B retreat but a compute-driven reprioritization.

Why it matters: HKR-H/K/R all pass: the compute-ceiling angle is strong, the piece adds concrete priority shifts, and OpenAI roadmap triage hits cost and dependency nerves. It stays at 80 because this is secondary reporting; spend, timing, and technical details are not disclosed.

Apr 16Thursday

Hacker News front page

AI cybersecurity is not proof of work

antirez argues AI bug finding is bounded by model intelligence level I, not by brute-force sampling alone; for the same code, execution paths eventually saturate. His concrete example is the OpenBSD SACK bug: weaker models fail even with unlimited tokens because they do not connect window validation, integer overflow, and the NULL branch. The key variable is model quality and access speed, not just more GPU.

Why it matters: High-quality commentary with HKR-H from the contrarian headline, HKR-K from the OpenBSD SACK mechanism and firsthand test, and HKR-R because it hits the 'more sampling vs better models' debate in AI security. Not a product, research release, or multi-source event, so it stays mid

最佳拍档 (BestPartners)

Post-AGI may arrive within 50 years: Demis Hassabis on AlphaFold, three AI risk classes, and human value

Demis Hassabis said in a 1-hour interview that post-AGI scenarios can arrive within 50 years, while AGI should stay in labs for another 10-20 years. He cited concrete numbers: AlphaFold has been used by 3M+ scientists, Isomorphic Labs is running 18-19 drug programs, and the most urgent risks in the next 2-4 years are misuse and agent misalignment.

Apr 4Saturday

Latent Space

Marc Andreessen introspects on The Death of the Browser, Pi + OpenClaw, and Why “This Time Is Different”

Marc Andreessen argues in a 76-minute interview that this AI cycle differs from 2016 because of reasoning, coding, agents, and recursive self-improvement. The post gives one concrete mechanism: Pi/OpenClaw as LLM + shell + filesystem + markdown + cron loop; it mentions “death of the browser,” but does not disclose a verifiable timeline or product plan. The sharper point is his Unix-like framing of file-backed agent state and portability.

Why it matters: This is a strong commentary piece, not a market-moving event. HKR-H comes from the browser-death hook, HKR-K from the Pi+OpenClaw mechanism, and HKR-R from the interface/distribution nerve; lack of roadmap, metrics, or launch details keeps it at the low end of featured.

Mar 5Thursday

OpenAI News

Reasoning models struggle to control their chains of thought, and that’s good

OpenAI frames an article around the claim that reasoning models struggle to control their chains of thought, and that this is a good thing. Only the title is available here, with no body text, so there are no verifiable numbers, methods, or mechanisms to summarize. The claim relates to reasoning and safety discussions, but any interpretation should stay limited to the headline.

Why it matters: OpenAI presents a contrarian but testable safety claim, so HKR-H/K/R all pass. The excerpt shows the thesis, section headers, and paper link, but not the key numbers, setup, or limits, so this stays high featured rather than P1.

Feb 27Friday

MIT Technology Review · AI

AI is rewiring how the world’s best Go players think

AI has become standard in pro Go training in South Korea, and the piece says competing professionally without it is now essentially impossible. It cites two figures: Shin Jin-seo matches AI moves 37.5% of the time versus a 28.5% player average, and AlphaGo Zero beat AlphaGo Lee 100-0 after three days of training. The shift to watch is training, not hype: KataGo is now a common tool, opening moves often mirror AI for the first 50 turns, and even top players still cannot fully explain its choices.

Why it matters: Strong HKR-H/K/R: the novelty is elite cognition shifting under AI, and the story brings concrete numbers plus a named tool. It is a reported commentary rather than a new model or product move, so it sits at the low end of featured.

Feb 14Saturday

Dwarkesh Patel

Dario Amodei: “We are near the end of the exponential”

Anthropic CEO Dario Amodei said in a long interview that model capability gains are still tracking an exponential, but are near its end, with the timeline off by only 1-2 years. He attributes progress to compute, data, training duration, and scalable objectives, and says RL shows log-linear gains on math and coding tasks; the post does not disclose exact curves, model versions, or reproducible parameters. The key claim is that pretraining and RL follow one scaling story, not two separate ones.

Why it matters: A top-lab CEO is making a direct claim on scaling, RL returns, and a 1-2 year timeline, so HKR-H/K/R all pass. I stop at 85 because this is thesis-level signal, not a product or research artifact: no curves, model IDs, or reproducible conditions are disclosed.

Jan 16Friday

Ruan YiFeng's Weblog

Technology Enthusiast Weekly (Issue 381): What China's AI Foundation Model Leaders Are Thinking

Ruan Yifeng’s Issue 381 excerpts talks from Beijing’s AGI-Next summit on Jan 10, covering views from Zhipu, Alibaba Qwen, and Tencent AI leaders on China’s model roadmap. The post cites Lin Junyang saying US compute is 1-2 orders of magnitude larger, Yao Shunyu calling the odds of a China-led top AI company in 3-5 years high, while Lin puts it at 20%. The key split is strategic: Tang Jie points to RLVR in 2025, Lin bets on multimodal foundation agents, and Yao says B2B buyers pay a $200/month premium for stronger models.

Why it matters: It clears all three HKR axes: public strategic disagreement gives it a strong hook, and the post includes concrete numbers and testable claims. The score stops short of the high bands because this is a secondary synthesis of summit remarks, not a primary release or original scoop