Skip to content

Reasoning

Progress in model reasoning: chain of thought, reasoning models, math and logic benchmarks and the debates around them.

Latest picks

421–440 of 585

May 11Monday

QbitAI · WeChat

Math Majors in Trouble: Fields Medalist Tests ChatGPT 5.5 Pro, Gets Paper-Level Result in 17 Minutes

Timothy Gowers tested ChatGPT 5.5 Pro on additive number theory problems, where it produced an optimal quadratic upper-bound construction in 17 minutes 5 seconds, then generated a LaTeX preprint in 47 minutes; the article says arXiv rejects AI-generated content, so the result remains on Gowers’s blog.

Why it matters: All three HKR axes pass: Gowers’ first-person test, 17m05s, and a 47-minute preprint are concrete and discussable. It is not a model release, but the named experiment and math-reasoning impact put it in the must-write band.

AI HOT (Curated Pool)

Local models handle half of daily tasks and respond faster than cloud models

A five-week experiment tested about 1,400 daily work tasks, where local 35B models such as Qwen 3.6 35B handled about 50% and averaged 2.8-second responses, 2.1 times faster than Claude Opus 4.5, while the cloud model still led complex reasoning by about 20%.

Why it matters: HKR-H/K/R all pass: Tom Tunguz’s experiment reports ~1,400 tasks, ~50% success, 2.8s latency, and a speed comparison to Claude Opus 4.5. Strong practitioner signal, but not a model launch or platform-level update.

AI HOT (Curated Pool)

Older AI Model Outperforms Human Doctors in Emergency Diagnosis

A Science study reports that OpenAI o1 reached a 67% correct or near-correct diagnosis rate on real emergency department data, exceeding doctors at 50-55%, but the study did not cover long-term inpatient data or imaging diagnosis.

Why it matters: HKR-H/K/R all pass: a Science-linked ER benchmark reports o1 at 67% versus doctors at 50-55%. It stays below P1 because it is one diagnostic study and excludes inpatient and imaging settings.

AI HOT (Curated Pool)

MachinaCheck: Multi-agent CNC manufacturability analysis system built on AMD MI300X

MachinaCheck runs Qwen 2.5 7B locally on AMD MI300X to analyze STEP files for CNC manufacturability, reducing drawing review for quote analysis from 30–60 minutes to 30 seconds while using 192GB HBM3 to keep customer design data on-premises.

Why it matters: HKR-H/K/R all pass, but this is an AMD hackathon project on Hugging Face, not a broad model or platform launch. Concrete numbers carry it to the featured threshold.

May 10Sunday

Synced · WeChat

Turing Award Winner Sutton Uses a 1967 Formula to Improve Streaming Reinforcement Learning

Richard Sutton and coauthors proposed Intentional Updates, which derive the step size from the desired output change; Intentional AC approached SAC on MuJoCo under batch=1 streaming training without replay, while each update used about 1/140 of SAC’s FLOPs.

Why it matters: HKR-H/K/R all pass: Sutton's name, Intentional Updates, MuJoCo conditions, and 1/140 SAC FLOPs give it substance. Strong research signal, but less market-moving than a major LLM product release, so it stays in the 78–84 band.

Synced · WeChat

Ted Xiao Reviews Three Eras of Robot Learning, from RT-1/RT-2 to Scaling

Ted Xiao divides nearly a decade of robot learning into three eras: Google’s team trained RT-1 on 87,000 teleoperation trajectories, then adapted 5B to 55B VLMs into VLA policies for RT-2.

Why it matters: HKR-H/K/R all pass: a named Google robotics insider, concrete RT-1/RT-2 numbers, and strong embodied-AI resonance. It is retrospective commentary, not a launch, so it stays in the 72–77 featured band.

Xinzhiyuan · WeChat

Next-ToBE Targets Short-Sighted Next-Token Prediction in LLMs at ICLR 2026

East China Normal University and Fudan University researchers proposed Next-ToBE, a training objective that keeps standard autoregressive inference while adding a soft target over future-token windows, and the article reports the method ranked best in 35 of 36 experiments across Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Llama3.1-8B-Instruct.

Why it matters: HKR-H and HKR-K pass: the mechanism and 35/36 result are specific, and next-token training is a real debate. The item stays near the featured floor because no artifact, reproduction detail, or production claim is disclosed.

r/LocalLLaMA

NVIDIA AI Releases Star Elastic: One Checkpoint Contains 30B, 23B, and 12B Reasoning Models

NVIDIA AI released Star Elastic, a single checkpoint that can zero-shot slice 30B, 23B, and 12B reasoning models in BF16, FP8, and NVFP4; when the 23B submodel handles thinking and the 30B model handles final answers, reported accuracy rises 16% and latency drops 1.9× on AIME-2025, GPQA, LiveCodeBench v5, and MMLU-Pro.

Why it matters: HKR-H/K/R all pass: Star Elastic has a concrete mechanism and testable numbers for inference deployment. Its reach is still narrower than a frontier-model release, so it sits in the high-quality featured band.

Computing Life · Share · Yage

How Anthropic Trained Computer Use: Reading Its Data Pipeline Through a Patent

Anthropic’s patent describes the Computer Use training pipeline: it captures user actions, uses a transformer to infer action intent, and applies a stronger model for synthetic expansion, turning raw UI operations into reasoning data.

Why it matters: HKR-H/K/R all pass: the patent angle is clickable, the three-step data pipeline is concrete, and agent builders care. It is analysis, not an official release or reproducible artifact, so 76 fits the featured threshold.

r/LocalLLaMA

BeeLlama.cpp: DFlash and TurboQuant with reasoning and vision support

Anbeeld released BeeLlama.cpp, a llama.cpp fork that runs Qwen 3.6 27B Q5 with 200k context and vision on a single RTX 3090 or 4090; the title claims 2–3x faster than baseline and a 135 tps peak.

Why it matters: HKR-H/K/R all pass, but the claims come from a Reddit title and summary without independent reproduction. Treat as a mid-weight open-source inference update, so it lands in the low featured band.

May 9Saturday

AI HOT (Curated Pool)

Baidu releases ERNIE 5.1 with compressed parameters and training cost

Baidu released ERNIE 5.1 with total parameters reduced to about one third of the original scale, active parameters to about one half, and pretraining cost to about 6% of same-scale models; the model is available on the ERNIE platform and Baidu AI Studio.

Why it matters: HKR-H/K/R all pass: Baidu ERNIE 5.1 is a domestic flagship-model release with concrete compression and 6% pretraining-cost claims. That puts it in the must-write band.

AI HOT (Curated Pool)

ERNIE 5.1 Released With Pretraining Cost at 6% of Comparable Models

Baidu released ERNIE 5.1, saying it builds on ERNIE 5.0 pretraining and improves search, reasoning, knowledge QA, creative writing, and agent capabilities, with pretraining cost at about 6% of comparable models.

Why it matters: Baidu released ERNIE 5.1 with a concrete “6% of reference pretraining cost” claim. HKR-H/K/R all pass, with a domestic flagship-model bump, but sparse technical detail keeps it below the 90s.

QbitAI · WeChat

Google AI Co-Mathematician Sets FrontierMath Tier 4 SOTA

Google DeepMind released AI Co-Mathematician, an asynchronous agent workspace for math research, and answered 23 of 48 private FrontierMath Tier 4 problems, scoring 48% under 48-hour, no-token-limit conditions versus GPT-5.5 Pro at 39.6%.

Why it matters: HKR-H/K/R all pass: the story has a hard benchmark number and a concrete research hook. No disclosed product access or cross-source cluster, so it stays at the top of 78–84 rather than p1.

Synced · WeChat

DeepSeek Reportedly Raises RMB 50B, with Liang Wenfeng Funding 40%, Valuation Reaching RMB 350B

DeepSeek is negotiating a $7.3 billion funding round at an estimated $51.5 billion valuation; Liang Wenfeng reportedly plans to contribute 40%, while Tencent and China’s RMB 60 billion national AI fund are also in talks.

Why it matters: HKR-H/K/R all pass: the DeepSeek funding rumor has large numbers, a founder contribution ratio, and named backers. Because it is still reported as talks with no official confirmation, it stays at 84 and featured, not p1.

Synced · WeChat

OpenAI's Jiayi Weng: Is the Next AI Training Paradigm Beyond Gradients?

OpenAI researcher Jiayi Weng proposes Heuristic Learning: codex gpt-5.4 reached a perfect 864 score on Breakout and generated 342 search trajectories across Atari 57, with updates applied to code, tests, replays, and memory rather than neural-network weights.

Why it matters: HKR-H/K/R all pass: an OpenAI researcher proposes Heuristic Learning with concrete hooks like Breakout 864 and 342 Atari 57 trajectories. This is strong research/commentary signal, not an official model or product release, so it stays in the 78–84 band.

May 8Friday

AI HOT (Curated Pool)

Robotics Endgame: A Physical AGI Roadmap and LLM Analogy

The speaker presented a physical AGI roadmap with six named components: video world models, WAM, EgoScale, dexterity scaling laws, physical reinforcement learning, and DreamDojo; the snippet also mentions a 2016 OpenAI DGX-1 signing story with Jensen and Elon.

Why it matters: HKR-H/K/R all pass: the physical-AGI endgame hook is strong, the post gives a 6-part roadmap, and robotics practitioners will debate the path. It is still a personal roadmap, not a release or benchmark, so it sits in 78–84.

Synced · WeChat

SGLang Team Launches RadixArk With $100M Seed Round

RadixArk announced a $100 million seed round on May 5 at a $400 million post-money valuation, while its SGLang inference project has 27K+ GitHub stars and deployments across 400K+ GPUs.

Why it matters: HKR-H/K/R all pass: the round size, valuation, and deployment numbers are concrete, and SGLang is a known inference stack. It is still a startup funding and infra-roadmap story, not a major model release, so it stays in the 78–84 featured band.

AI HOT (Curated Pool)

Adaptive Parallel Reasoning: A New Paradigm for Efficient Reasoning Scaling

BAIR’s post describes adaptive parallel reasoning, where ThreadWeaver and Multiverse dynamically control parallel threads for math and code reasoning; the RSS snippet does not disclose benchmark scores, latency reductions, or reproducible settings.

Why it matters: BAIR authority supports the 72+ band, and HKR-H/K/R all pass. The post names mechanisms and dynamic thread control, but lacks scores, latency gains, and reproducible conditions, so it stays below 78.

Latent Space

[AINews] GPT-Realtime-2, Translate, and Whisper: new SOTA realtime voice APIs

OpenAI released GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper in the Realtime API, with GPT-Realtime-2 expanding context from 32K to 128K and scoring 96.6% on Artificial Analysis Big Bench Audio.

Why it matters: HKR-H/K/R all pass: an OpenAI real-time voice API refresh, a 32K→128K context jump, and a 96.6% Big Bench Audio claim. Score stays at 86 because this is a major API update, not a flagship foundation-model release.

Xinzhiyuan · WeChat

Token-Level Length Control: 3B Model Beats GPT 5.4 and Claude

UC Santa Barbara and Apple researchers introduced LenVM, which models remaining generation length as a token-level value function; Qwen2.5-3B with a 1.5B LenVM scored 62.6 on LIFEBench length control, above GPT-5.4 at 37.4 and Claude-Opus-4-6 at 35.5.

Why it matters: HKR-H/K/R all pass: the headline has a sharp small-model-vs-frontier hook, and the post gives LenVM's mechanism plus 62.6/37.4 benchmark numbers. The topic is narrow research, not a model or major product release, so it fits the 78-84 band.