Skip to content

Reasoning

Progress in model reasoning: chain of thought, reasoning models, math and logic benchmarks and the debates around them.

Latest picks

321–340 of 585

Jun 3Wednesday

AI HOT (Curated Pool)

DeepSeek Reportedly Seeks RMB 50 Billion in First Funding Round with Tencent and CATL

DeepSeek plans to raise about RMB 50 billion in its first funding round, with post-money valuation expected at RMB 350 billion to RMB 400 billion; Liang Wenfeng, Tencent, and CATL plan to invest RMB 20 billion, RMB 10 billion, and RMB 5 billion respectively.

Why it matters: HKR-H/K/R all pass: DeepSeek's rumored RMB 50B first round includes a RMB 350B-400B valuation and named checks from Tencent and CATL. The rumor status keeps it at 88, below confirmed industry-shaking funding news.

r/LocalLLaMA

Microsoft Aion 1.0 Instruct and Aion 1.0 Plan models

Microsoft announced two on-device Aion 1.0 models at Build 2026. Aion 1.0 Plan is a 14B-parameter reasoning and tool-calling model with 32K context, shipping in-box with Windows on capable devices, while Aion 1.0 Instruct targets summarization, rewriting, intents, accessibility, Edge integration, and open-weight availability.

Why it matters: Microsoft announced Aion 1.0 Instruct and Plan at Build 2026, with Plan listed as a 14B, 32K-context model for eligible Windows devices. HKR-H/K/R all pass, but licensing, benchmarks, and hardware requirements are not disclosed, so it stays in the 78–84 band.

Computing Life · Share · Yage

Microsoft AI's MAI-Thinking-1: Getting Models to Think Is Easy, Sustained Thinking Is Hard

Microsoft AI says MAI-Thinking-1 uses three mechanisms—thermostat, circuit breaker, and self-distillation—to keep RL training stable for several thousand steps; the RSS snippet contrasts MAI’s discipline with DeepSeek’s efficiency and GLM’s endurance.

Why it matters: HKR-H/K/R all pass: the hook is training persistence, the new facts are three stability mechanisms and thousand-step RL runs, and the audience cares about reasoning-model stability. Not a major model launch, so it stays below 85.

AI HOT (Curated Pool)

Microsoft releases MAI-Thinking-1 model

Microsoft released MAI-Thinking-1, an MoE model with 35B active parameters and 1T total parameters, pretrained from scratch on 30T tokens without third-party model distillation.

Why it matters: HKR-H/K/R all pass: Microsoft released MAI-Thinking-1 with concrete MoE scale and training-token figures. Benchmarks, access, and pricing are not disclosed, so it stays in the 78–84 band rather than P1.

Hacker News front page

MAI-Thinking-1

The title names MAI-Thinking-1, and the RSS snippet says Microsoft is launching seven MAI models; the post does not disclose parameters, capabilities, benchmarks, pricing, or rollout timing.

Why it matters: HKR-H/K/R pass because Microsoft names a Thinking model and seven MAI models, touching the OpenAI-dependence nerve. Sparse specs, evals, and roadmap keep it in the 72–77 featured-threshold band.

AI HOT (Curated Pool)

Microsoft releases its first advanced reasoning AI model, MAI-Thinking-1

Microsoft released MAI-Thinking-1 at Build 2026, describing it as a medium-sized reasoning model that matches leading models on key software engineering benchmarks.

Why it matters: HKR-H/K/R all pass: Microsoft released its first advanced reasoning model with a mid-sized design and SWE benchmark claim. Exact scores, access, and pricing are not disclosed, so it stays below 85.

AI HOT (Curated Pool)

Google DeepMind releases Gemini multi-agent research system

Google DeepMind introduced Co-Scientist, a Gemini-based multi-agent system that generates, debates, and evolves scientific hypotheses; the post does not disclose the Gemini version, benchmark results, access model, or release timeline.

Why it matters: HKR-H/K/R all pass, but model version, eval results, and availability are not disclosed. This fits a strong research/product release, not the 85+ must-write band.

Jun 2Tuesday

r/LocalLLaMA

Replaced Claude with local Qwen3.6-27B in my multi-agent orchestrator for 2 weeks

The author ran Qwen3.6-27B on one RTX 3090 across 47 multi-step coding workflows. Plan generation reached about 95% schema validity, but tool-call formatting errors were about 12%, and practical long-context use degraded past about 12k tokens.

Why it matters: HKR-H/K/R all pass: a named first-person local-vs-Claude experiment with concrete numbers. The single Reddit source and 47-workflow scope keep it below the 78–84 band.

Synced · WeChat

Turing Award Winner Sutton’s New Paper Argues AI Should Move Toward Enactive Cognition

Banafsheh Rafiee and Richard S. Sutton propose an enactive cognition framework for AI, naming four pillars: experience, perception-action inseparability, autonomy, and embodiment.

Why it matters: HKR-H/K/R all pass, but the article centers on a conceptual framework and does not disclose experiments, code, or reproducible tests. Sutton’s name and the four pillars put it in the 78–84 research-commentary band.

AI HOT (Curated Pool)

StepFun releases Step 3.7 Flash for efficient inference

StepFun released Step 3.7 Flash with a 196B MoE architecture, using multi-matrix factorized attention to cut KV-cache cost to about 22% of DeepSeek models.

Why it matters: HKR-H/K/R all pass: Step 3.7 Flash has concrete specs, not just launch copy, with 196B MoE and ~22% KV-cache cost versus DeepSeek. It is below top-lab flagship weight, so 78 featured.

Jun 1Monday

AI HOT (Curated Pool)

NVIDIA Open-Sources Cosmos 3, Its First Generalist Model for Physical AI

NVIDIA open-sourced Cosmos 3 at GTC Taipei, releasing two variants, Super 32B and Nano 8B, with model weights, code, and datasets made available.

Why it matters: HKR-H/K/R all pass: the concrete hook is NVIDIA opening Cosmos 3 with 32B/8B variants and released artifacts. The post is sparse and single-source, with no benchmarks or license details, so it stays in the 78–84 band.

r/LocalLLaMA

Deepseek V4 Flash performance on DGX Spark

A Reddit user ran DeepSeek-V4-Flash with vLLM on two ASUS GX10 DGX Spark nodes and reported 1,680 prefill tokens/s plus 39.8 decode tokens/s at a 256K context with MTP=2; the setup uses TP=2 over RoCE, fp8 KV cache, and fits about 1M tokens safely in KV cache.

Why it matters: This is not broad industry news, but it is a first-person benchmark with reproducible details: TP=2, RoCE, fp8 KV cache, 256K context, and ~1M KV. HKR-H/K/R all pass, so it lands at low featured.

AI HOT (Curated Pool)

Cosmos 3 Released: First Open Physical AI Generalist Model

NVIDIA released Cosmos 3 as an open physical AI generalist model with native visual reasoning, world generation, and action generation, offering two variants: Super at 32B parameters and Nano at 8B parameters.

Why it matters: HKR-H/K/R all pass: NVIDIA names two Cosmos 3 variants and concrete physical-AI capabilities. Source is a single launch post with no benchmark or license detail, so it stays in the 78–84 band.

r/LocalLLaMA

I bolted an 8-arm reasoning MoE onto a frozen 1.4B Mamba backbone on a single RTX 3060

The author trained Mamba-Titan-1.4B-Reasoning on a 12GB RTX 3060: a frozen 1.4B Mamba-1 backbone with 8 trainable MoE arms, 2.54B total parameters, Top-2 routing at layers 24/25, and about 50% math accuracy.

Why it matters: HKR-H/K/R all pass via a numbered first-person experiment, but it is a single Reddit post with no independent replication and a fairly technical setup, so it stays in the low featured band.

May 31Sunday

r/LocalLLaMA

13 abliterated Gemma 4 E2B variants, 44 GPU hours, benchmark and comparison

Abliterlitics tested 13 abliterated Gemma 4 E2B variants using 44 RTX 5090 GPU hours, and HarmBench ASR rose from the base model’s 32.2% to 82%–100%, while coder3101 scored 84.8% on GSM8K versus the base model’s 83.5%.

Why it matters: HKR-H/K/R all pass, with a named first-person benchmark and concrete numbers. Scope stays narrow around abliterated Gemma 4 E2B variants, so it lands at the featured threshold rather than a must-write item.

QbitAI · WeChat

Robot-Native World Action Model Debuts With Spatiotemporal Architecture From Fudan-Linked Team

Moushen Intelligence released STI-WM, a spatiotemporally integrated world action model for robotics, with RGB, depth point cloud, and proprioceptive inputs; the post says it supports hundred-second-scale long-horizon task rollout and closed-loop replanning, but does not disclose benchmark scores or deployment costs.

Why it matters: HKR-H/K/R all pass: the STI-WM angle is novel, with concrete input modalities and hundred-second rollouts. Kept near the featured floor because public weights, benchmark results, and reproducible tests are not disclosed.

May 30Saturday

Xinzhiyuan · WeChat

Opus 4.8 Builds a Historical Rebirth Simulator for 117 Billion Humans

Ethan Mollick used Claude Opus 4.8 to generate The Veil of History, a website that weights a random human life by 117 billion historical births and, according to the article, uses 4,000 Monte Carlo runs to estimate regional and era distributions.

Why it matters: HKR-H/K/R all pass: Mollick’s Claude Opus 4.8 demo has a strange hook, concrete numbers, and a builder-relevant prototyping angle. It is not an Anthropic release, so it stays in the lower featured band.

QbitAI · WeChat

Key Gemini IMO Gold Contributor Nearly Became a Professional Pianist

Yi Tay served as a modeling co-captain for Gemini Deep Think when it reached IMO gold-medal level, co-founded Reka AI in 2023, and returned to Google DeepMind after 639 days, while the article also notes his 2012 Trinity classical piano associate diploma.

Why it matters: HKR-H/K/R all pass, but this is a profile, not a Gemini capability launch. The concrete value is Yi Tay's role, Reka history, and 639-day return, so it sits in the 72–77 featured band.

May 29Friday

Xinzhiyuan · WeChat

Claude Opus 4.8 tests split users: strong at high effort, costly under rate limits

The article says Claude Opus 4.8 scores 63 on an Extra-High senior engineering benchmark, 30 points above Opus 4.7, but drops to 42 at High effort, while $200/month Max users report hitting rate limits within hours on complex agent tasks.

Why it matters: Anthropic/Claude relevance plus concrete test numbers clears HKR-H/K/R: the hook is strength versus cost, K has benchmark and quota details, and R hits agent-budget anxiety. Source is a media test rather than an official release, so this lands at low P1.

Synced · WeChat

A True 2-bit KV Quantization Algorithm for Long-context Reasoning Beyond TurboQuant

TogetherAI and collaborators released OSCAR, a 2.28 BPE INT2 KV Cache system integrated with SGLang, reporting up to 3× decode speedup at 100k context and up to 7× job-level throughput under a fixed memory budget.

Why it matters: HKR-H/K/R pass, but this is niche inference optimization rather than a broad model launch. The 100k-context and ~3×/~7× claims justify a featured score, not same-day must-write.