Skip to content

Reasoning

Progress in model reasoning: chain of thought, reasoning models, math and logic benchmarks and the debates around them.

Latest picks

361–380 of 585

May 22Friday

Bloomberg Technology

DeepSeek Founder Declares AGI Goal as $10 Billion Round Advances

The title says DeepSeek’s founder declared an AGI goal and that a $10 billion funding round is advancing; the post does not disclose the founder’s statement, financing terms, investors, or timeline.

Why it matters: HKR-H/K/R all pass: DeepSeek plus a $10B round and AGI goal is same-day AI-business news. The scrape provides title-level facts only, with no investors, terms, or timeline, so the score stays at the low end of the 85+ band.

Synced · WeChat

Meta Chinese Researcher Releases ATLAS for Generalizable Visual Reasoning with One Word

Meta AI and the Chinese University of Hong Kong proposed ATLAS, a visual reasoning method that uses one Functional Token to connect Agentic and Latent Visual Reasoning, with ATLAS-178K, a two-stage SFT+RL pipeline, and LA-GRPO to train sparse visual-operation tokens.

Why it matters: HKR-H/K/R pass: the one-token angle is clickable, and the post gives dataset and training details. As a Meta AI/CUHK research release rather than a flagship model or product launch, it fits the 78–84 band.

Computing Life · Share · Yage

A general-purpose AI model refutes an 80-year-old conjecture

An OpenAI general-purpose reasoning model refuted Erdős’s 1946 unit distance conjecture in the plane; the post says the model was not specially trained for mathematics, and Tim Gowers said he would recommend it to Annals of Mathematics.

Why it matters: HKR-H/K/R all pass: an OpenAI general reasoning model allegedly refuting Erdős’s 1946 conjecture with Tim Gowers approval is same-day material. The summary lacks paper link, proof details, and reproduction conditions, so it stays below 95.

NVIDIA Blog

NVIDIA GTC Taipei at COMPUTEX: Live Updates on What’s Next in AI

NVIDIA won four COMPUTEX 2026 Best Choice Awards for Vera Rubin NVL72, Jetson Thor, and Alpamayo; Vera Rubin NVL72 connects 36 Vera CPUs and 72 Rubin GPUs, and NVIDIA says it delivers up to 10x higher inference performance per watt and 10x lower cost per token.

Why it matters: HKR-H/K/R all pass: NVIDIA gives concrete Vera Rubin NVL72 specs and a 10x inference-efficiency claim, directly tied to AI compute costs. The source is NVIDIA’s event blog, so this stays below the 85 same-day must-write band.

May 21Thursday

r/LocalLLaMA

HRM 1B

Sapientinc released HRM-Text 1B Base and its training code, and the paper claims competitive performance against 2–7B open models while using 100–900x fewer training tokens and 96–432x less estimated compute, with training on 16 H100 GPUs taking about 46 hours and costing about $1,472.

Why it matters: HKR-H/K/R all pass: HRM-Text 1B has concrete low-cost training numbers and released code. Capped at 80 because this is a Reddit item and the efficiency claim still lacks independent evaluation.

Latent Space

OpenAI GPT-next Disproves 80-Year-Old Erdős Planar Unit Distance Problem for Under $1000

OpenAI said an internal general-purpose reasoning model disproved the 1946 Erdős planar unit distance problem by finding a new family of constructions; the reasoning summary reportedly spans about 125 pages, while outside observers speculate the run used under 32 hours or under $1,000.

Why it matters: HKR-H/K/R all pass: an OpenAI internal reasoning model allegedly refuting the 1946 Erdős problem with ~125 pages is a major capability signal. Cost and runtime are still external estimates, keeping it below 95.

AI HOT (Curated Pool)

OpenAI Model Independently Solves 80-Year-Old Math Problem

An OpenAI AI model solved the plane unit distance problem proposed in 1946, using Golod-Shafarevich theory to produce a family of more efficient constructions.

Why it matters: HKR-H/K/R all pass, but the item is only an X summary and lacks model name, paper link, reproducibility, and third-party verification. Strong OpenAI reasoning-research signal, kept below P1.

TechCrunch · AI

OpenAI claims it solved an 80-year-old math problem — for real this time

OpenAI says its reasoning model disproved a geometry conjecture unsolved since 1946, and the snippet says mathematicians who challenged its previous claim now back it; the post does not disclose the model name, proof details, or verification process.

Why it matters: HKR-H/K/R all pass: OpenAI plus an 80-year geometry conjecture is a strong, testable reasoning claim. Missing model name, proof details, and validation flow keep it below P1.

May 20Wednesday

AI Chat-Group Daily (群聊日报)

2026-05-19 Chat Group Daily

The chat group daily says Karpathy joined Anthropic's pretraining team, and cites Stainless shutting down hosted services after acquisition plus Google I/O announcing Gemini 3.5 Flash and a $100 subscription tier.

Why it matters: HKR-H/K/R all pass, but this is a chat-daily roundup with secondhand claims and no disclosed primary links, appointment details, or product specs, so it lands at the lower featured band.

OpenAI News

An OpenAI model has disproved a central conjecture in discrete geometry

An OpenAI model solved the 80-year-old unit distance problem and disproved a major conjecture in discrete geometry; the post does not disclose the model name, proof mechanism, or reproducibility conditions.

Why it matters: HKR-H/K/R all pass: the OpenAI math result is novel, concrete, and debate-starting. Missing model name, proof mechanism, and reproducibility keep it at 85, not a higher P1.

AI HOT (Curated Pool)

Google launches new AI search box with multimodal interactions

Google launched an AI search box based on Gemini 3.5, combining AI Overviews and AI Mode into one AI search experience that supports multimodal multi-turn queries across text, images, files, and video, with global availability on desktop and mobile.

Why it matters: HKR-H/K/R all pass: a Google Search entry-point update with Gemini 3.5, multimodal file/video queries, and AI Overviews/AI Mode integration. The source is thin, so it lands at the lower end of must-write.

AI HOT (Curated Pool)

Gemini Omni launches with physical reasoning and multimodal generation

Google launched Gemini Omni video generation for global AI Plus, Pro, and Ultra subscribers, integrating it with Gemini app, Google Flow, and YouTube Shorts, while the RSS snippet says the model combines intuitive physical reasoning with Gemini’s historical, scientific, and cultural knowledge.

Why it matters: HKR-H/K/R all pass: this is an official Google Gemini video-generation launch with named tiers and product surfaces. Details are thin—no benchmarks, pricing, or physical-reasoning tests—so it stays at the low end of the 85–94 band.

AI HOT (Curated Pool)

Gemini 3.5 Released: A New Model Family Combining Intelligence and Action

Google AI Developers announced the Gemini 3.5 model family, saying it combines intelligence with action capabilities; the post does not disclose parameters, benchmarks, pricing, availability, or context window details.

Why it matters: HKR-H and HKR-R pass: an official Gemini 3.5 family launch has flagship-model pull and competitive resonance. HKR-K fails because the post gives no params, benchmarks, pricing, or context window, so this stays below the 85+ band.

The Verge · AI

Google Search is getting its biggest changes ever

Google showed a redesigned Search box at I/O 2026, using Gemini 3.5 Flash to connect AI Overviews with AI Mode; the RSS snippet says natural-language queries will reliably show AI Overviews, but the post does not disclose rollout timing.

Why it matters: HKR-H/K/R all pass: Google is changing Search’s core input with Gemini 3.5 Flash and linking AI Overviews to AI Mode. Rollout timing is missing, so the score stays at 86, but this is still a same-day story for AI pros.

AI HOT (Curated Pool)

Gemini 3.5 Flash launches as an efficient option for task handling

Google released Gemini 3.5 Flash and calls it its best model so far for fast, efficient task completion. The post does not disclose pricing, context window size, benchmark scores, or API availability conditions.

Why it matters: HKR-H and HKR-R pass because this is a new Google Gemini Flash release tied to cost and latency. HKR-K fails: the post gives no price, context window, benchmarks, or API availability, keeping it in the 78–84 band.

May 19Tuesday

TechCrunch · AI

OpenAI co-founder Andrej Karpathy joins Anthropic’s pre-training team

Andrej Karpathy joined Anthropic’s pre-training team, which runs large-scale training for Claude’s core knowledge and capabilities; the RSS snippet does not disclose his title, reporting line, or start date.

Why it matters: HKR-H/K/R all pass: Karpathy’s OpenAI identity, Anthropic pre-training role, and Claude-scale training work make this a same-day talent-war story, even though level, reporting line, and start date are undisclosed.

AI HOT (Curated Pool)

Former OpenAI core member Andrej Karpathy chooses Anthropic to return to frontier LLM research

Andrej Karpathy has joined Anthropic to return to frontline LLM research; the post identifies him as a former OpenAI core team member and Tesla Autopilot architect, but does not disclose his team, title, or specific research projects.

Why it matters: HKR-H/K/R all pass: Karpathy joining Anthropic is a high-signal personnel move in frontier labs. Team, title, and project are undisclosed, so the score stays at the low end of the 85 band.

r/LocalLLaMA

Sapient Intelligence releases HRM-Text 1B: 40B tokens, ~$1k pretrain

Sapient Intelligence released HRM-Text 1B, a 1B-parameter model trained from scratch on 16 GPUs for 1.9 days with 40B tokens and a reported ~$1,000 budget; its self-reported chart shows MATH 56.2 and DROP 82.2, while independent evaluation remains pending.

Why it matters: HKR-H/K/R all pass: low-cost pretraining plus a smaller model beating a larger one is clickable, with concrete training and benchmark numbers. Independent eval is unfinished, so this stays at 78, not 85.

AI Chat-Group Daily (群聊日报)

May 18, 2026 Chat Group Daily

The chat group daily says AI21 Labs cut 60% of staff and stopped selling model access, and cites a University of Waterloo paper where GPT-5.4 accuracy dropped from 100% to 23% after false peer-consensus injection; the snippet also mentions Meta layoff talk at 10%, but does not disclose source details or confirmation conditions.

Why it matters: HKR-H/K/R all pass: AI21’s 60% layoff and model-sales stop signal lab contraction, while GPT-5.4 falling from 100% to 23% under false peer consensus is a concrete safety hook. The chat-digest source keeps it at 78.

QbitAI · WeChat

JD and CAS IIE Publish Three Papers Defining Self-Taught RLVR

JD and CAS IIE released three Self-Taught RLVR papers covering RLSD, NPO, and CoPD; RLSD reports that 200 training steps on Qwen3-VL-8B-Instruct exceed GRPO at 400 steps across 8 benchmarks.

Why it matters: HKR-H/K/R pass: self-taught RLVR is a clear hook; RLSD reports 8 benchmarks and a 200-vs-400-step GRPO comparison; it hits reasoning fine-tuning cost. Not a top-lab model launch and replication heat is undisclosed, so it stays low featured.