Skip to content

#推理

1 today

May 31Sunday

QbitAI · WeChat

Robot-Native World Action Model Debuts With Spatiotemporal Architecture From Fudan-Linked Team

Moushen Intelligence released STI-WM, a spatiotemporally integrated world action model for robotics, with RGB, depth point cloud, and proprioceptive inputs; the post says it supports hundred-second-scale long-horizon task rollout and closed-loop replanning, but does not disclose benchmark scores or deployment costs.

Why it matters: HKR-H/K/R all pass: the STI-WM angle is novel, with concrete input modalities and hundred-second rollouts. Kept near the featured floor because public weights, benchmark results, and reproducible tests are not disclosed.

May 30Saturday

Xinzhiyuan · WeChat

Opus 4.8 Builds a Historical Rebirth Simulator for 117 Billion Humans

Ethan Mollick used Claude Opus 4.8 to generate The Veil of History, a website that weights a random human life by 117 billion historical births and, according to the article, uses 4,000 Monte Carlo runs to estimate regional and era distributions.

Why it matters: HKR-H/K/R all pass: Mollick’s Claude Opus 4.8 demo has a strange hook, concrete numbers, and a builder-relevant prototyping angle. It is not an Anthropic release, so it stays in the lower featured band.

QbitAI · WeChat

Key Gemini IMO Gold Contributor Nearly Became a Professional Pianist

Yi Tay served as a modeling co-captain for Gemini Deep Think when it reached IMO gold-medal level, co-founded Reka AI in 2023, and returned to Google DeepMind after 639 days, while the article also notes his 2012 Trinity classical piano associate diploma.

Why it matters: HKR-H/K/R all pass, but this is a profile, not a Gemini capability launch. The concrete value is Yi Tay's role, Reka history, and 639-day return, so it sits in the 72–77 featured band.

May 29Friday

Xinzhiyuan · WeChat

Claude Opus 4.8 tests split users: strong at high effort, costly under rate limits

The article says Claude Opus 4.8 scores 63 on an Extra-High senior engineering benchmark, 30 points above Opus 4.7, but drops to 42 at High effort, while $200/month Max users report hitting rate limits within hours on complex agent tasks.

Why it matters: Anthropic/Claude relevance plus concrete test numbers clears HKR-H/K/R: the hook is strength versus cost, K has benchmark and quota details, and R hits agent-budget anxiety. Source is a media test rather than an official release, so this lands at low P1.

Synced · WeChat

A True 2-bit KV Quantization Algorithm for Long-context Reasoning Beyond TurboQuant

TogetherAI and collaborators released OSCAR, a 2.28 BPE INT2 KV Cache system integrated with SGLang, reporting up to 3× decode speedup at 100k context and up to 7× job-level throughput under a fixed memory budget.

Why it matters: HKR-H/K/R pass, but this is niche inference optimization rather than a broad model launch. The 100k-context and ~3×/~7× claims justify a featured score, not same-day must-write.

Synced · WeChat

Meta Uses 183B Tokens to Turn Math Textbooks into a Large Lean Library

Meta released ATLAS, a Lean 4 formalization library covering 26 math textbooks and 46,203 declarations, using 183.157 billion tokens to generate 630,999 lines of code, with 42,837 completed proofs and a 92.7% proof pass rate.

Why it matters: HKR-H/K/R all pass: the token scale, Lean corpus size, and verified-proof count are concrete. It stays below P1 because this is a specialized research/open-source release, not a broad model or product launch.

Synced · WeChat

The Ma Jiaqi Failure Exposed an LLM Issue He Spotted in the Shower a Year Earlier

FaceMind links low-frequency token degradation to two papers: SLoW appeared at EMNLP 2025, Adam's Law was accepted as an ACL 2026 Oral, and high-frequency rewriting raised DeepSeek-V3 math accuracy from 63.55% to 71.54%.

Why it matters: HKR-H/K/R all pass: the odd celebrity-token hook is clickable, and the post gives a mechanism plus a 63.55%→71.54% DeepSeek-V3 result. Practical research signal, but not a major model launch.

AI HOT (Curated Pool)

Skill distillation

Skill distillation has Opus 4.7, GPT-5.1, and Gemini 3 Pro write standardized SKILL.md procedure files, while local Qwen 35B and Gemma 26B models execute those files step by step.

Why it matters: HKR-H/K/R pass: the agent-skill distillation pattern is concrete and practitioner-relevant. The summary lacks success rates, cost data, or task outcomes, so it sits at the featured threshold, not must-write.

The Verge · AI

Claude’s New Model Is More ‘Honest’ When It Messes Up

Anthropic will release Claude Opus 4.8 on Thursday, emphasizing its claimed “honesty.” The company says early testers found it flags uncertainty more often. It also says internal evaluations show Opus 4.8 is around 4x less likely than its predecessor to make unsupported claims, while the RSS snippet does not disclose the full benchmark setup.

Why it matters: HKR-H/K/R all pass: an Anthropic Claude model update with a concrete “4x fewer unsupported claims” eval claim. Details are thin: benchmark set, pricing, and context window are not disclosed, so it sits in the low 85–94 band.

May 28Thursday

AI HOT (Curated Pool)

AI Now Summit 2026

Mistral AI announced industrial AI work, a Vibe upgrade, and a 10 MW inference data center in Les Ulis at AI Now Summit 2026; it is working with Airbus, BMW Group, and ASML, and the data center is scheduled to start operating in Q3 2026.

Why it matters: HKR-H/K/R pass: Mistral gives a concrete 10 MW inference site, Q3 2026 timing, and major industrial partners. No new model capability or pricing is disclosed, so it stays just above the featured threshold.

Synced · WeChat

Mila and DeepMind Propose UNSL for Unified Multivariate Neural Scaling Laws

Mila and Google DeepMind proposed Unified Neural Scaling Law, a multivariate scaling-law form that models parameter count, token count, training steps, bottlenecks, overfitting, and adverse hyperparameter effects; UNSL achieved the best extrapolation on 60.87% of vision tasks and 88.89% of language tasks in the reported experiments.

Why it matters: HKR-H/K/R all pass: UNSL unifies parameters, tokens, steps, bottlenecks, overfitting, and hyperparameter feedback, with vision/language extrapolation numbers. Technical density keeps it in the 78–84 band.

AI HOT (Curated Pool)

Interview with Google Search VP Robby Stein on the AI-Native Search Era

Robby Stein discussed Google Search’s move toward an AI-native mode at Google I/O, covering AI Mode, multi-turn query decomposition, TPU infrastructure costs, source-link selection, and publisher traffic tension, but the post does not disclose specific pricing, traffic numbers, or rollout conditions.

Why it matters: HKR-H/K/R all pass, but this is an interview summary rather than a fresh launch. No price, traffic, or cost numbers are disclosed, so it sits in the 72–77 quality-interview band.

May 27Wednesday

r/LocalLLaMA

I ran 8 open-weight models as agents in a persistent MMO for 10 days

Firespawn Studios ran 25 agents across 8 open-weight models for 10 days in Null Epoch Season 0 and released about 93,000 logged events, with roughly 70% of actions including the model’s reasoning or justification.

Why it matters: HKR-H/K/R all pass: a concrete 10-day MMO agent trial with 25 agents and 93k events. Reddit sourcing limits reach, so it lands in the 78–84 good-quality band, not P1.

QbitAI · WeChat

Language Models Need Sleep: Let AI Nap Before Continuing Inference

Carnegie Mellon University and the University of Maryland propose a “sleep” mechanism for language models: when the context window is nearly full, the model stops accepting new tokens, runs multiple offline recursive forward passes to compress accumulated context into fast weights, clears the KV cache, and then resumes inference; tests cover cellular automata, multi-hop graph retrieval, and GSM-Infinite reasoning tasks.

Why it matters: HKR-H/K/R all pass: the sleep metaphor is clickable, and the mechanism is concrete. Score stays below 78 because the provided body lacks benchmark gains, code, or deployment evidence.

Synced · WeChat

From Foundation Models to Physical AI, Samsung Moves Into the Core LLM Race

Samsung disclosed three AI efforts—Meki, M2RL, and LiveClawBench—covering a memory-based edge architecture, multi-domain reinforcement learning, and Physical AI evaluation; the article also says Samsung has purchased tens of thousands of GPUs for AI infrastructure, but does not provide model size, training budget, or deployment timelines.

Why it matters: HKR-H, HKR-K, and HKR-R pass, but this is a Samsung research bundle plus strategy signal, not a flagship model or product launch. It fits the 72–77 featured band, below same-day must-write.

AI HOT (Curated Pool)

Claude Mythos reportedly solves OpenAI’s landmark Erdős problem with a “cute simple proof”

Anthropic engineer Sholto Douglas said Claude Mythos solved OpenAI’s Erdős unit distance conjecture problem over the weekend and produced a “cute simple proof”; the RSS snippet does not disclose the proof, verification process, or benchmark setup.

Why it matters: HKR-H/K/R all pass: the claim is clickable, specific, and tied to frontier reasoning rivalry. The post does not disclose the proof, validation process, or Mythos release status, so it stays featured rather than P1.

May 26Tuesday

Import AI (Jack Clark)

Import AI 458: Reckoning with the Future; and a Singularity Story

Jack Clark’s Import AI 458 excerpts his 2026 Cosmos HAI Lab Lecture, cites the Epoch Capabilities Index across 40-plus benchmarks, and argues that an AI system able to develop its own successor may arrive within two years or sooner.

Why it matters: HKR-H/K/R all pass: Jack Clark pairs ECI’s 40+ benchmarks with a two-year successor-system claim, giving this AGI-timeline essay both concrete detail and debate fuel.

QbitAI · WeChat

Zhejiang University and Alibaba Make AI Think Before Drawing Sudoku or Burning Candles | ACL 2026

Zhejiang University and Alibaba introduced Unified Thinker, an independent planning module trained with 40,000 HieraReason-40K samples and a two-stage GRPO reinforcement-learning setup that turns structured reasoning traces into executable visual instructions for image generation and editing.

Why it matters: HKR-H/K/R all pass: the paper has a concrete visual-failure hook, a 40k-sample planning/RL mechanism, and relevance to multimodal-agent reliability. It remains a paper-level advance, not a product or flagship model release.

Synced · WeChat

ACL 2026 Main: Spatial-Agent Generates Executable Geospatial Analysis Workflows for LLMs

Spatial-Agent inserts a GeoFlow Graph between natural-language questions and map tools, and Spatial-Agent with GPT-4o-mini reaches 45.15% accuracy on MapEval-API versus a 23.00% API baseline.

Why it matters: ACL Main gives a concrete mechanism and testable numbers, so HKR-H/K pass. The GIS focus limits HKR-R, placing it at the featured threshold rather than a must-write item.

Xinzhiyuan · WeChat

OpenAI Nearly Collapsed? President Says He Resigned the Day Altman Was Ousted

Greg Brockman recounted OpenAI’s 72-hour crisis: on November 17, 2023, the board removed Sam Altman as CEO and took Brockman off the board, after which Brockman resigned the same day and said he initially put the chance of taking the company back at 10%.

Why it matters: HKR-H/K/R all pass via an insider crisis hook, a 10% recovery-odds detail, and OpenAI governance resonance. It is still a retrospective on a heavily covered 2023 event, so it stays in the 72–77 band.

May 25Monday

r/LocalLLaMA

Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

RTPurbo converts full-attention LLMs to sparse inference with a few hundred adaptation steps. It keeps the full KV cache only for retrieval heads, uses a 16-dimensional token indexer, and reports up to 9.36x prefill speedup at 1M context plus about 2.01x decode speedup on long-context and reasoning benchmarks.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, the post gives 1M-context speedup numbers, and inference cost resonates. Reddit-only sourcing and missing model/code details keep it in the 78–84 band.

r/LocalLLaMA

The reason small-model agent stacks aren't the default is not whether they work

A Reddit post argues small-model agent stacks are not default for business reasons, not capability limits: Gemma 4 31B reaches 86.4% on tau2-bench, and DeepSeek V4-Flash output tokens are priced about 89x below Claude Opus 4.6. The operational risk is verification, because 7–9B models produced broken reasoning for roughly half to two-thirds of correct answers in a cited audit.

Why it matters: HKR-H/K/R all pass: the angle is contrarian, with benchmark, cost, and verifier-failure numbers. Reddit-source uncertainty keeps it in the 78–84 recommendation band, not P1.

May 24Sunday

Synced · WeChat

ICML 2026: First Parallel Thinking Framework for Vision-Language Models

Visual Para-Thinker introduces a parallel thinking framework for vision-language models, using Pa-Attention and LPRoPE to isolate four visual reasoning paths and training on 163,000 question-answer pairs.

Why it matters: HKR-H/K/R pass: the ICML 2026 paper offers a concrete parallel-thinking mechanism, four isolated paths, and 163K training pairs. It remains a single research release without broad replication or product impact, so it fits 78–84.

May 23Saturday

Synced · WeChat

Bengio Paper Raises Recursive Reasoning Limits as Parallel Trajectories Beat Serial Reasoning

Yoshua Bengio’s team introduced GRAM, a generative recursive reasoning model that samples multiple latent trajectories; on Sudoku-Extreme, GRAM reached 97.0% accuracy with 16 recursive steps and 20 parallel samples, exceeding TRM’s 90.5% result at 320 serial recursive steps.

Why it matters: HKR-H/K/R all pass: the hook is parallel recursion beating long serial recursion, with concrete GRAM numbers. Importance stays in 78–84 because the evidence is benchmark-centered, not a major model or product release.

May 22Friday

MIT Technology Review · AI

Google I/O showed how the path for AI-driven science is shifting

MIT Technology Review says Google used I/O to shift its scientific AI framing toward Gemini for Science, a package that groups AI Co-Scientist and AlphaEvolve, while researchers can now apply for access and older specialized systems like AlphaFold and WeatherNext remain active.

Why it matters: HKR-H and HKR-K pass: MIT Technology Review frames a real Google science-AI product shift with named components and access conditions. HKR-R is weak because the impact is mostly research-facing, not practitioner-wide.

Bloomberg Technology

DeepSeek Founder Declares AGI Goal as $10 Billion Round Advances

The title says DeepSeek’s founder declared an AGI goal and that a $10 billion funding round is advancing; the post does not disclose the founder’s statement, financing terms, investors, or timeline.

Why it matters: HKR-H/K/R all pass: DeepSeek plus a $10B round and AGI goal is same-day AI-business news. The scrape provides title-level facts only, with no investors, terms, or timeline, so the score stays at the low end of the 85+ band.

Synced · WeChat

Meta Chinese Researcher Releases ATLAS for Generalizable Visual Reasoning with One Word

Meta AI and the Chinese University of Hong Kong proposed ATLAS, a visual reasoning method that uses one Functional Token to connect Agentic and Latent Visual Reasoning, with ATLAS-178K, a two-stage SFT+RL pipeline, and LA-GRPO to train sparse visual-operation tokens.

Why it matters: HKR-H/K/R pass: the one-token angle is clickable, and the post gives dataset and training details. As a Meta AI/CUHK research release rather than a flagship model or product launch, it fits the 78–84 band.

Computing Life · Share · Yage

A general-purpose AI model refutes an 80-year-old conjecture

An OpenAI general-purpose reasoning model refuted Erdős’s 1946 unit distance conjecture in the plane; the post says the model was not specially trained for mathematics, and Tim Gowers said he would recommend it to Annals of Mathematics.

Why it matters: HKR-H/K/R all pass: an OpenAI general reasoning model allegedly refuting Erdős’s 1946 conjecture with Tim Gowers approval is same-day material. The summary lacks paper link, proof details, and reproduction conditions, so it stays below 95.

NVIDIA Blog

NVIDIA GTC Taipei at COMPUTEX: Live Updates on What’s Next in AI

NVIDIA won four COMPUTEX 2026 Best Choice Awards for Vera Rubin NVL72, Jetson Thor, and Alpamayo; Vera Rubin NVL72 connects 36 Vera CPUs and 72 Rubin GPUs, and NVIDIA says it delivers up to 10x higher inference performance per watt and 10x lower cost per token.

Why it matters: HKR-H/K/R all pass: NVIDIA gives concrete Vera Rubin NVL72 specs and a 10x inference-efficiency claim, directly tied to AI compute costs. The source is NVIDIA’s event blog, so this stays below the 85 same-day must-write band.

May 21Thursday

r/LocalLLaMA

HRM 1B

Sapientinc released HRM-Text 1B Base and its training code, and the paper claims competitive performance against 2–7B open models while using 100–900x fewer training tokens and 96–432x less estimated compute, with training on 16 H100 GPUs taking about 46 hours and costing about $1,472.

Why it matters: HKR-H/K/R all pass: HRM-Text 1B has concrete low-cost training numbers and released code. Capped at 80 because this is a Reddit item and the efficiency claim still lacks independent evaluation.

Latent Space

OpenAI GPT-next Disproves 80-Year-Old Erdős Planar Unit Distance Problem for Under $1000

OpenAI said an internal general-purpose reasoning model disproved the 1946 Erdős planar unit distance problem by finding a new family of constructions; the reasoning summary reportedly spans about 125 pages, while outside observers speculate the run used under 32 hours or under $1,000.

Why it matters: HKR-H/K/R all pass: an OpenAI internal reasoning model allegedly refuting the 1946 Erdős problem with ~125 pages is a major capability signal. Cost and runtime are still external estimates, keeping it below 95.

AI HOT (Curated Pool)

OpenAI Model Independently Solves 80-Year-Old Math Problem

An OpenAI AI model solved the plane unit distance problem proposed in 1946, using Golod-Shafarevich theory to produce a family of more efficient constructions.

Why it matters: HKR-H/K/R all pass, but the item is only an X summary and lacks model name, paper link, reproducibility, and third-party verification. Strong OpenAI reasoning-research signal, kept below P1.

TechCrunch · AI

OpenAI claims it solved an 80-year-old math problem — for real this time

OpenAI says its reasoning model disproved a geometry conjecture unsolved since 1946, and the snippet says mathematicians who challenged its previous claim now back it; the post does not disclose the model name, proof details, or verification process.

Why it matters: HKR-H/K/R all pass: OpenAI plus an 80-year geometry conjecture is a strong, testable reasoning claim. Missing model name, proof details, and validation flow keep it below P1.

May 20Wednesday

AI Chat-Group Daily (群聊日报)

2026-05-19 Chat Group Daily

The chat group daily says Karpathy joined Anthropic's pretraining team, and cites Stainless shutting down hosted services after acquisition plus Google I/O announcing Gemini 3.5 Flash and a $100 subscription tier.

Why it matters: HKR-H/K/R all pass, but this is a chat-daily roundup with secondhand claims and no disclosed primary links, appointment details, or product specs, so it lands at the lower featured band.

OpenAI News

An OpenAI model has disproved a central conjecture in discrete geometry

An OpenAI model solved the 80-year-old unit distance problem and disproved a major conjecture in discrete geometry; the post does not disclose the model name, proof mechanism, or reproducibility conditions.

Why it matters: HKR-H/K/R all pass: the OpenAI math result is novel, concrete, and debate-starting. Missing model name, proof mechanism, and reproducibility keep it at 85, not a higher P1.

AI HOT (Curated Pool)

Google launches new AI search box with multimodal interactions

Google launched an AI search box based on Gemini 3.5, combining AI Overviews and AI Mode into one AI search experience that supports multimodal multi-turn queries across text, images, files, and video, with global availability on desktop and mobile.

Why it matters: HKR-H/K/R all pass: a Google Search entry-point update with Gemini 3.5, multimodal file/video queries, and AI Overviews/AI Mode integration. The source is thin, so it lands at the lower end of must-write.

AI HOT (Curated Pool)

Gemini Omni launches with physical reasoning and multimodal generation

Google launched Gemini Omni video generation for global AI Plus, Pro, and Ultra subscribers, integrating it with Gemini app, Google Flow, and YouTube Shorts, while the RSS snippet says the model combines intuitive physical reasoning with Gemini’s historical, scientific, and cultural knowledge.

Why it matters: HKR-H/K/R all pass: this is an official Google Gemini video-generation launch with named tiers and product surfaces. Details are thin—no benchmarks, pricing, or physical-reasoning tests—so it stays at the low end of the 85–94 band.

AI HOT (Curated Pool)

Gemini 3.5 Released: A New Model Family Combining Intelligence and Action

Google AI Developers announced the Gemini 3.5 model family, saying it combines intelligence with action capabilities; the post does not disclose parameters, benchmarks, pricing, availability, or context window details.

Why it matters: HKR-H and HKR-R pass: an official Gemini 3.5 family launch has flagship-model pull and competitive resonance. HKR-K fails because the post gives no params, benchmarks, pricing, or context window, so this stays below the 85+ band.

The Verge · AI

Google Search is getting its biggest changes ever

Google showed a redesigned Search box at I/O 2026, using Gemini 3.5 Flash to connect AI Overviews with AI Mode; the RSS snippet says natural-language queries will reliably show AI Overviews, but the post does not disclose rollout timing.

Why it matters: HKR-H/K/R all pass: Google is changing Search’s core input with Gemini 3.5 Flash and linking AI Overviews to AI Mode. Rollout timing is missing, so the score stays at 86, but this is still a same-day story for AI pros.

AI HOT (Curated Pool)

Gemini 3.5 Flash launches as an efficient option for task handling

Google released Gemini 3.5 Flash and calls it its best model so far for fast, efficient task completion. The post does not disclose pricing, context window size, benchmark scores, or API availability conditions.

Why it matters: HKR-H and HKR-R pass because this is a new Google Gemini Flash release tied to cost and latency. HKR-K fails: the post gives no price, context window, benchmarks, or API availability, keeping it in the 78–84 band.