Skip to content

Reasoning

Progress in model reasoning: chain of thought, reasoning models, math and logic benchmarks and the debates around them.

Latest picks

401–420 of 585

May 15Friday

AI HOT (Curated Pool)

Anthropic's Mythos AI helped find and exploit two unknown macOS kernel vulnerabilities in five days

Anthropic’s Mythos AI helped researchers find two previously unknown macOS kernel vulnerabilities in five days and chain them into a privilege-escalation exploit that bypassed Apple’s memory integrity protection, according to the Wall Street Journal snippet.

Why it matters: HKR-H/K/R all pass, and Anthropic-linked AI security work is high-signal. The score stays in 78–84 because the source is a social post and lacks paper details, reproducible conditions, or exploit mechanics.

r/LocalLLaMA

I Let a Small Model Train on Its Own Mistakes; It Reached 80% on HumanEval and Beat GPT-3.5 on Math

The author fine-tuned Qwen 2.5 7B base on self-mined mistake-correction pairs, raising HumanEval from 25/164 to 112/164; Qwen 2.5 14B used 100 pairs and a 95-minute H100 run costing $3.50.

Why it matters: HKR-H/K/R pass: the hook is strong and the post gives samples, H100 time, cost, and HumanEval deltas. Kept at 78 because it is a single Reddit post and the 80% claim differs from 112/164.

TechCrunch · AI

What Happens When AI Starts Building Itself?

Richard Socher’s new $650 million startup plans to build an AI system that can research and improve itself indefinitely, and the RSS snippet says it will ship products; the post does not disclose the technical mechanism, launch timeline, or product format.

Why it matters: HKR-H/K/R all pass, but the post lacks mechanism, timeline, and product form, keeping it in the 72–77 threshold band. TechCrunch authority, Socher’s name, and the $650M figure support featured.

r/LocalLLaMA

MOOSE-Star (ICML 2026): 7B Model and 108K-Paper Dataset for Scientific Hypothesis Discovery

MiroMind researchers released the MOOSE-Star collection with three 7B models and TOMATO-Star, a dataset of 108,717 NCBI papers. MS-IR-7B reaches 54.37% inspiration-retrieval accuracy, uses DeepSeek-R1-Distill-Qwen-7B as its base, runs at about 14GB fp16, and supports llama.cpp, vLLM, and SGLang.

Why it matters: HKR-H/K/R all pass via the local 7B research-agent hook and concrete dataset metrics. Single Reddit source and limited lab gravity keep it below the must-write band.

r/LocalLLaMA

inclusionAI/Ring-2.6-1T on Hugging Face

inclusionAI released Ring-2.6-1T, a 1T-parameter reasoning model on Hugging Face; it supports high and xhigh reasoning effort levels, targets agent workflows and long-horizon tasks, and uses Async RL with the IcePop algorithm for reinforcement-learning training stability.

Why it matters: HKR-H/K/R pass: a 1T HF model with two reasoning modes and named training methods is real signal. Benchmarks, license, and inference cost are not disclosed, so this stays at the lower edge of featured.

May 14Thursday

AI HOT (Curated Pool)

SenseNova U1 technical report released with MoE-based open model weights

Li Mu’s team released the SenseNova U1 technical report and MoE-based weights; the snippet says it covers architecture and training methods, but the post does not disclose parameter size, license terms, or benchmark results.

Why it matters: HKR-H/K/R pass: SenseNova U1 combines a named Li Mu team release, MoE weights, and practical open-weight relevance. Missing model size, license, and evaluations keep it at 75, below the 78+ band.

AI HOT (Curated Pool)

MiMo V2.5 Pro Places Third on DesignArena

MiMo V2.5 Pro placed third on the DesignArena overall leaderboard; its Thinking version rose 8 spots over MiMo-V2.5 and matched Claude Sonnet 4.6 performance on frontend coding tasks.

Why it matters: HKR-H/K/R all pass, but the facts come from one official X post with no methodology, access, or pricing. This fits a mid-weight benchmark/product update, not a same-day must-write.

Xinzhiyuan · WeChat

Yuandong Tian and Seven Co-Founders Launch Recursive Superintelligence at $4.65B Valuation

Recursive Superintelligence, founded by Yuandong Tian and seven other AI researchers, has a 25-person team, $650 million in funding, and a $4.65 billion valuation, with a stated goal to automate evaluation, data filtering, training, post-training, and research-direction selection.

Why it matters: All three HKR axes pass: a $650M raise at a $4.65B valuation for a 25-person recursive-improvement startup is not routine funding. The stated target spans evals, data selection, training, post-training, and research selection.

Synced · WeChat

ACL 2026: Alibaba DAMO I²B-LPO Improves RLVR Exploration

Alibaba DAMO Academy introduced I²B-LPO, an RLVR post-training framework that branches rollouts at high-entropy nodes and filters them with an information-bottleneck self-reward, reporting up to 5.3% accuracy gains and 7.4% semantic-diversity gains on math benchmarks using Qwen2.5-7B and Qwen3-14B.

Why it matters: HKR-H/K/R all pass: the ACL 2026 DAMO paper has a clear RLVR exploration hook, concrete I²B-LPO mechanics, and benchmark gains. It is still a training-method paper, not a major model or product release, so 78 fits the lower good-quality band.

r/LocalLLaMA

sensenova/SenseNova-U1-A3B-MoT · Hugging Face

SenseNova published SenseNova-U1-A3B-MoT on Hugging Face; the post lists A3B MoT, 8B MoT, and 0.4B LoRA weight links, and says the NEO-unify architecture unifies multimodal understanding, reasoning, and generation in one model family.

Why it matters: HKR-H/K/R all pass: an open multimodal model release with multiple weight sizes and a named NEO-unify mechanism. Source authority and missing benchmarks/license details keep it in the lower featured band.

May 13Wednesday

r/LocalLLaMA

AIDC-AI/Ovis2.6-80B-A3B on Hugging Face

AIDC-AI released Ovis2.6-80B-A3B, a multimodal MoE model with 80B total parameters and about 3B active parameters at inference, supporting a 64K-token context window and images up to 2880×2880 resolution.

Why it matters: HKR-H/K/R pass: the open multimodal MoE has concrete specs and a real efficiency hook. Score stays near the featured floor because the post gives no benchmarks, license details, or hands-on results.

AI HOT (Curated Pool)

Build long-running AI agents that pause, resume, and never lose context with ADK

Google Developers describes using ADK to build long-running agents for enterprise workflows lasting days or weeks, such as HR onboarding, with a persistent state machine, persistent session storage, event-driven webhooks, and multi-agent delegation to pause during idle time and resume after restarts without losing context.

Why it matters: HKR-H/K/R all pass: the ADK tutorial gives concrete persistence mechanisms for long-running agents. It is useful engineering guidance from Google Developers, not a major model or platform release, so it sits at the 72–77 featured threshold.

May 12Tuesday

AI HOT (Curated Pool)

Install the official Codex plugin in Claude Code

The author describes installing OpenAI’s official Codex plugin in Claude Code via the plugin marketplace: add the repository, install the plugin, reload, and configure it, then use it to build a Skill where Claude Code handles reasoning and Codex acts as moderator.

Why it matters: HKR-H/K/R all pass: cross-stack plugin use is clickable, the install path is concrete, and it matters to AI dev workflows. It stays low-featured because this is a tutorial-style tip, not a model or platform release.

Xinzhiyuan · WeChat

OpenAI releases GPT-Realtime-2, described as a GPT-5-level reasoning audio model

OpenAI released GPT-Realtime-2 alongside Realtime-Translate and Realtime-Whisper, with a 128K context window, five reasoning-effort levels, and API pricing of $32 per million input tokens and $64 per million output tokens.

Why it matters: HKR-H/K/R all pass: realtime audio reasoning is a strong hook; 128K context, five reasoning levels, and $32/$64 per 1M tokens add substance; voice-agent cost and stack choices hit practitioners. This is a same-day OpenAI product update.

QbitAI · WeChat

Shanghai AI Lab Study: SFT Generalizes Under Three Conditions

Shanghai AI Lab, Shanghai Jiao Tong University, and USTC tested Long-CoT SFT on Qwen3-14B-Base and found that cross-domain performance recovered and improved after 8 epochs, with generalization conditioned on optimization depth, data quality and structure, and base-model capability.

Why it matters: HKR-H/K/R all pass: the SFT-generalization claim has a clear hook, Qwen3-14B-Base plus an 8-epoch finding, and direct relevance to fine-tuning teams. It lacks deployment impact or full benchmark detail, so it stays in the mid-featured band.

May 11Monday

AI HOT (Curated Pool)

Fields Medalist Tests ChatGPT 5.5 Pro: Paper-Level Result in 17 Minutes

Timothy Gowers tested ChatGPT 5.5 Pro and said it independently solved an open additive number theory problem in 17 minutes with only a simple prompt, producing PhD thesis-level work; he warned that this pace threatens mathematics research training, while Terence Tao said human value lies in digesting and deeply understanding proofs.

Why it matters: HKR-H/K/R all pass: a named mathematician, a 17-minute result, and a PhD-training warning. The exact problem, prompt, and verification path are not disclosed, keeping it below P1.

AI HOT (Curated Pool)

AntLingAGI Releases Trillion-Parameter Ring-2.6-1T Model

AntLingAGI released Ring-2.6-1T, a trillion-parameter thinking model available for free on OpenRouter until May 15, with adjustable thinking intensity, agent-oriented multi-step execution, tool calling, and tasks covering math logic and scientific research.

Why it matters: HKR-H/K/R all pass, but the post is thin: no benchmarks, pricing, architecture, or training details. Treat as a mid-weight model launch on OpenRouter, not a same-day must-write.

AI HOT (Curated Pool)

Tencent Hunyuan Hy3 Preview Released for Complex Agent Tasks

Tencent Hunyuan opened early access to the Hy3 preview, which uses a 256K context window and a mixture-of-experts architecture with fast and slow thinking for complex agent tasks.

Why it matters: HKR-H/K/R all pass: Tencent Hunyuan Hy3 preview names 256K context and a fast/slow-thinking MoE for complex agents. Benchmarks, pricing, and access scope are not disclosed, keeping it in the 78–84 band.

Xinzhiyuan · WeChat

Largest IPO Nears, Topping SpaceX; 2028 AI Self-Iteration Countdown

Xinzhiyuan says Anthropic is considering a near-$1 trillion valuation, with ARR rising to $45 billion in five months; Jack Clark predicts a greater than 50% chance that AI systems can autonomously build better versions of themselves by the end of 2028, while the article cites a 72% Kalshi probability of an IPO announcement before November 1.

Why it matters: HKR-H/K/R all pass: the hook is sharp and the post gives valuation, ARR, and 2028 odds. Source is secondary and IPO/ARR claims lack official confirmation, so it stays in 78-84.

Synced · WeChat

ICML 2026: PRISM Brings Efficient Test-Time Scaling to dLLMs

PRISM raises LLaDA-8B-Instruct on GSM8K from 67.58% to 85.30% by combining hierarchical trajectory search, partial remasking, and self-verified feedback, reducing dLLM test-time scaling cost from O(NT) toward O(N+KT) under a final candidate width K.

Why it matters: HKR-H/K/R all pass: the hook rejects brute-force scaling, the post gives GSM8K and complexity numbers, and it speaks to inference cost. Still an ICML framework paper, not a mainstream product release, so it sits in 78–84.