Skip to content

Reasoning

Progress in model reasoning: chain of thought, reasoning models, math and logic benchmarks and the debates around them.

Latest picks

341–360 of 585

May 29Friday

Synced · WeChat

Meta Uses 183B Tokens to Turn Math Textbooks into a Large Lean Library

Meta released ATLAS, a Lean 4 formalization library covering 26 math textbooks and 46,203 declarations, using 183.157 billion tokens to generate 630,999 lines of code, with 42,837 completed proofs and a 92.7% proof pass rate.

Why it matters: HKR-H/K/R all pass: the token scale, Lean corpus size, and verified-proof count are concrete. It stays below P1 because this is a specialized research/open-source release, not a broad model or product launch.

Synced · WeChat

The Ma Jiaqi Failure Exposed an LLM Issue He Spotted in the Shower a Year Earlier

FaceMind links low-frequency token degradation to two papers: SLoW appeared at EMNLP 2025, Adam's Law was accepted as an ACL 2026 Oral, and high-frequency rewriting raised DeepSeek-V3 math accuracy from 63.55% to 71.54%.

Why it matters: HKR-H/K/R all pass: the odd celebrity-token hook is clickable, and the post gives a mechanism plus a 63.55%→71.54% DeepSeek-V3 result. Practical research signal, but not a major model launch.

AI HOT (Curated Pool)

Skill distillation

Skill distillation has Opus 4.7, GPT-5.1, and Gemini 3 Pro write standardized SKILL.md procedure files, while local Qwen 35B and Gemma 26B models execute those files step by step.

Why it matters: HKR-H/K/R pass: the agent-skill distillation pattern is concrete and practitioner-relevant. The summary lacks success rates, cost data, or task outcomes, so it sits at the featured threshold, not must-write.

The Verge · AI

Claude’s New Model Is More ‘Honest’ When It Messes Up

Anthropic will release Claude Opus 4.8 on Thursday, emphasizing its claimed “honesty.” The company says early testers found it flags uncertainty more often. It also says internal evaluations show Opus 4.8 is around 4x less likely than its predecessor to make unsupported claims, while the RSS snippet does not disclose the full benchmark setup.

Why it matters: HKR-H/K/R all pass: an Anthropic Claude model update with a concrete “4x fewer unsupported claims” eval claim. Details are thin: benchmark set, pricing, and context window are not disclosed, so it sits in the low 85–94 band.

May 28Thursday

AI HOT (Curated Pool)

AI Now Summit 2026

Mistral AI announced industrial AI work, a Vibe upgrade, and a 10 MW inference data center in Les Ulis at AI Now Summit 2026; it is working with Airbus, BMW Group, and ASML, and the data center is scheduled to start operating in Q3 2026.

Why it matters: HKR-H/K/R pass: Mistral gives a concrete 10 MW inference site, Q3 2026 timing, and major industrial partners. No new model capability or pricing is disclosed, so it stays just above the featured threshold.

Synced · WeChat

Mila and DeepMind Propose UNSL for Unified Multivariate Neural Scaling Laws

Mila and Google DeepMind proposed Unified Neural Scaling Law, a multivariate scaling-law form that models parameter count, token count, training steps, bottlenecks, overfitting, and adverse hyperparameter effects; UNSL achieved the best extrapolation on 60.87% of vision tasks and 88.89% of language tasks in the reported experiments.

Why it matters: HKR-H/K/R all pass: UNSL unifies parameters, tokens, steps, bottlenecks, overfitting, and hyperparameter feedback, with vision/language extrapolation numbers. Technical density keeps it in the 78–84 band.

AI HOT (Curated Pool)

Interview with Google Search VP Robby Stein on the AI-Native Search Era

Robby Stein discussed Google Search’s move toward an AI-native mode at Google I/O, covering AI Mode, multi-turn query decomposition, TPU infrastructure costs, source-link selection, and publisher traffic tension, but the post does not disclose specific pricing, traffic numbers, or rollout conditions.

Why it matters: HKR-H/K/R all pass, but this is an interview summary rather than a fresh launch. No price, traffic, or cost numbers are disclosed, so it sits in the 72–77 quality-interview band.

May 27Wednesday

r/LocalLLaMA

I ran 8 open-weight models as agents in a persistent MMO for 10 days

Firespawn Studios ran 25 agents across 8 open-weight models for 10 days in Null Epoch Season 0 and released about 93,000 logged events, with roughly 70% of actions including the model’s reasoning or justification.

Why it matters: HKR-H/K/R all pass: a concrete 10-day MMO agent trial with 25 agents and 93k events. Reddit sourcing limits reach, so it lands in the 78–84 good-quality band, not P1.

QbitAI · WeChat

Language Models Need Sleep: Let AI Nap Before Continuing Inference

Carnegie Mellon University and the University of Maryland propose a “sleep” mechanism for language models: when the context window is nearly full, the model stops accepting new tokens, runs multiple offline recursive forward passes to compress accumulated context into fast weights, clears the KV cache, and then resumes inference; tests cover cellular automata, multi-hop graph retrieval, and GSM-Infinite reasoning tasks.

Why it matters: HKR-H/K/R all pass: the sleep metaphor is clickable, and the mechanism is concrete. Score stays below 78 because the provided body lacks benchmark gains, code, or deployment evidence.

Synced · WeChat

From Foundation Models to Physical AI, Samsung Moves Into the Core LLM Race

Samsung disclosed three AI efforts—Meki, M2RL, and LiveClawBench—covering a memory-based edge architecture, multi-domain reinforcement learning, and Physical AI evaluation; the article also says Samsung has purchased tens of thousands of GPUs for AI infrastructure, but does not provide model size, training budget, or deployment timelines.

Why it matters: HKR-H, HKR-K, and HKR-R pass, but this is a Samsung research bundle plus strategy signal, not a flagship model or product launch. It fits the 72–77 featured band, below same-day must-write.

AI HOT (Curated Pool)

Claude Mythos reportedly solves OpenAI’s landmark Erdős problem with a “cute simple proof”

Anthropic engineer Sholto Douglas said Claude Mythos solved OpenAI’s Erdős unit distance conjecture problem over the weekend and produced a “cute simple proof”; the RSS snippet does not disclose the proof, verification process, or benchmark setup.

Why it matters: HKR-H/K/R all pass: the claim is clickable, specific, and tied to frontier reasoning rivalry. The post does not disclose the proof, validation process, or Mythos release status, so it stays featured rather than P1.

May 26Tuesday

Import AI (Jack Clark)

Import AI 458: Reckoning with the Future; and a Singularity Story

Jack Clark’s Import AI 458 excerpts his 2026 Cosmos HAI Lab Lecture, cites the Epoch Capabilities Index across 40-plus benchmarks, and argues that an AI system able to develop its own successor may arrive within two years or sooner.

Why it matters: HKR-H/K/R all pass: Jack Clark pairs ECI’s 40+ benchmarks with a two-year successor-system claim, giving this AGI-timeline essay both concrete detail and debate fuel.

QbitAI · WeChat

Zhejiang University and Alibaba Make AI Think Before Drawing Sudoku or Burning Candles | ACL 2026

Zhejiang University and Alibaba introduced Unified Thinker, an independent planning module trained with 40,000 HieraReason-40K samples and a two-stage GRPO reinforcement-learning setup that turns structured reasoning traces into executable visual instructions for image generation and editing.

Why it matters: HKR-H/K/R all pass: the paper has a concrete visual-failure hook, a 40k-sample planning/RL mechanism, and relevance to multimodal-agent reliability. It remains a paper-level advance, not a product or flagship model release.

Synced · WeChat

ACL 2026 Main: Spatial-Agent Generates Executable Geospatial Analysis Workflows for LLMs

Spatial-Agent inserts a GeoFlow Graph between natural-language questions and map tools, and Spatial-Agent with GPT-4o-mini reaches 45.15% accuracy on MapEval-API versus a 23.00% API baseline.

Why it matters: ACL Main gives a concrete mechanism and testable numbers, so HKR-H/K pass. The GIS focus limits HKR-R, placing it at the featured threshold rather than a must-write item.

Xinzhiyuan · WeChat

OpenAI Nearly Collapsed? President Says He Resigned the Day Altman Was Ousted

Greg Brockman recounted OpenAI’s 72-hour crisis: on November 17, 2023, the board removed Sam Altman as CEO and took Brockman off the board, after which Brockman resigned the same day and said he initially put the chance of taking the company back at 10%.

Why it matters: HKR-H/K/R all pass via an insider crisis hook, a 10% recovery-odds detail, and OpenAI governance resonance. It is still a retrospective on a heavily covered 2023 event, so it stays in the 72–77 band.

May 25Monday

r/LocalLLaMA

Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

RTPurbo converts full-attention LLMs to sparse inference with a few hundred adaptation steps. It keeps the full KV cache only for retrieval heads, uses a 16-dimensional token indexer, and reports up to 9.36x prefill speedup at 1M context plus about 2.01x decode speedup on long-context and reasoning benchmarks.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, the post gives 1M-context speedup numbers, and inference cost resonates. Reddit-only sourcing and missing model/code details keep it in the 78–84 band.

r/LocalLLaMA

The reason small-model agent stacks aren't the default is not whether they work

A Reddit post argues small-model agent stacks are not default for business reasons, not capability limits: Gemma 4 31B reaches 86.4% on tau2-bench, and DeepSeek V4-Flash output tokens are priced about 89x below Claude Opus 4.6. The operational risk is verification, because 7–9B models produced broken reasoning for roughly half to two-thirds of correct answers in a cited audit.

Why it matters: HKR-H/K/R all pass: the angle is contrarian, with benchmark, cost, and verifier-failure numbers. Reddit-source uncertainty keeps it in the 78–84 recommendation band, not P1.

May 24Sunday

Synced · WeChat

ICML 2026: First Parallel Thinking Framework for Vision-Language Models

Visual Para-Thinker introduces a parallel thinking framework for vision-language models, using Pa-Attention and LPRoPE to isolate four visual reasoning paths and training on 163,000 question-answer pairs.

Why it matters: HKR-H/K/R pass: the ICML 2026 paper offers a concrete parallel-thinking mechanism, four isolated paths, and 163K training pairs. It remains a single research release without broad replication or product impact, so it fits 78–84.

May 23Saturday

Synced · WeChat

Bengio Paper Raises Recursive Reasoning Limits as Parallel Trajectories Beat Serial Reasoning

Yoshua Bengio’s team introduced GRAM, a generative recursive reasoning model that samples multiple latent trajectories; on Sudoku-Extreme, GRAM reached 97.0% accuracy with 16 recursive steps and 20 parallel samples, exceeding TRM’s 90.5% result at 320 serial recursive steps.

Why it matters: HKR-H/K/R all pass: the hook is parallel recursion beating long serial recursion, with concrete GRAM numbers. Importance stays in 78–84 because the evidence is benchmark-centered, not a major model or product release.

May 22Friday

MIT Technology Review · AI

Google I/O showed how the path for AI-driven science is shifting

MIT Technology Review says Google used I/O to shift its scientific AI framing toward Gemini for Science, a package that groups AI Co-Scientist and AlphaEvolve, while researchers can now apply for access and older specialized systems like AlphaFold and WeatherNext remain active.

Why it matters: HKR-H and HKR-K pass: MIT Technology Review frames a real Google science-AI product shift with named components and access conditions. HKR-R is weak because the impact is mostly research-facing, not practitioner-wide.