Skip to content

Reasoning

Progress in model reasoning: chain of thought, reasoning models, math and logic benchmarks and the debates around them.

Latest picks

441–460 of 585

May 8Friday

QbitAI · WeChat

HIT and Huawei propose Dynamic-dLLM, a training-free acceleration framework with 4.48x speedup

HIT Shenzhen, Huawei, and Shenzhen Hetao College proposed Dynamic-dLLM, a training-free dLLM acceleration framework that raises LLaDA-8B-Instruct throughput on GSM8k from 8.32 TPS to 37.29 TPS with almost no accuracy loss.

Why it matters: HKR-H/K/R all pass: the 4.48x speedup is clickable, and GSM8k TPS figures add concrete substance. It is inference-optimization research, not a mainstream model launch, so it fits the 78–84 band.

QbitAI · WeChat

OpenAI releases three realtime voice models for reasoning, translation, and transcription

OpenAI launched GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper as API models, covering 128K-context voice reasoning, streaming translation from more than 70 input languages into 13 output languages, and realtime transcription priced at $0.017 per minute.

Why it matters: OpenAI shipped three realtime voice APIs across reasoning, translation, and transcription, hitting HKR-H/K/R. The 128K context, 70+ languages, and $0.017/min price make this a same-day must-write item.

QbitAI · WeChat

All Labs Watch ByteDance, Everyone Praises DeepSeek: A U.S. Researcher’s 36-Hour China AI Trip

Ai2 researcher Nathan Lambert visited Zhipu, Moonshot AI, Tsinghua, Meituan, Xiaomi, and 01.AI within 36 hours, and said Chinese labs closely watch ByteDance and respect DeepSeek, while student participation in core work, open source habits, and in-house control of the technical stack mark key differences.

Why it matters: HKR-H/K/R all pass: the piece has a named US researcher’s dense China-lab tour plus concrete claims on ByteDance, DeepSeek, open source, and in-house stacks. It is strong industry field reporting, not a model launch or major deal, so it sits at featured rather than p1.

r/LocalLLaMA

11.67% ARC-AGI-2 Local Eval on a Single 4090: The TOPAS Recursive Architecture

Doug_Bitterbot says TOPAS scored 11.67% on ARC-AGI-2 using one RTX 4090 after about 14 days of training. The 100M-parameter checkpoint hit 36% locally, but recursive TTT caused null outputs on nearly half of Kaggle puzzles. The key detail is time management: the author expects 20% after threshold tuning and 3-5 more weeks of training.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post with unstable Kaggle submissions. It clears featured, not the higher research-release band.

May 7Thursday

AI HOT (Curated Pool)

Trillion-parameter instruction model Ling-2.6-1T released

inclusionAI says Ling-2.6-1T is now live on OpenRouter. The trillion-parameter instruction model uses “fast thinking” and claims top AIME26 and SWE-bench Verified results with about 75% lower cost. The post does not disclose pricing, context length, or full benchmark scores.

Why it matters: HKR-H/K/R all pass: a 1T instruction model on OpenRouter with fast thinking, AIME26/SWE-bench claims, and ~75% cost reduction. Missing price, context window, and full scores keep it in the 78–84 band.

r/LocalLLaMA

Qwen/WebWorld 32B/14B/8B (Qwen3 finetune)

Qwen released WebWorld 32B/14B/8B, Qwen3 finetunes for training and evaluating web agents. It uses 1M+ real web trajectories and supports 30+ step simulation plus A11y Tree, HTML, XML, Markdown, and natural-language states. Agents trained on its synthetic trajectories gain 9.9% on MiniWob++ and 10.9% on WebArena.

Why it matters: HKR-H/K/R all pass: WebWorld has an agent hook, concrete scale, and benchmark gains. It is a useful Qwen research release for agent builders, but limited source detail keeps it below the 85 must-write band.

OpenAI News

Advancing Voice Intelligence with New Models in the API

OpenAI introduced new realtime voice models in its API for voice intelligence. The RSS snippet says they reason, translate, and transcribe speech; the post does not disclose counts, pricing, or limits.

Why it matters: OpenAI’s official voice API update hits HKR-H/K/R, but the available body gives capability direction only. Model count, pricing, latency, and context limits are not disclosed, so it stays at the top of 78–84.

TechCrunch · AI

DeepSeek could hit $45B valuation from its first investment round

DeepSeek could reach a $45B valuation in its first investment round, according to the title. The snippet says it rose in early 2025 after training an LLM with far less compute and cost; the post does not disclose round size, investors, or terms.

Why it matters: HKR-H/K/R all pass: DeepSeek’s first round targeting $45B is a strong valuation story. Missing investors, amount, and terms keep it in the lower 78–84 band, not P1.

May 6Wednesday

r/LocalLLaMA

Qwen3.6 27B NVFP4 + MTP on a Single RTX 5090: 200k Context in vLLM

A Reddit user ran Qwen3.6 27B NVFP4 on one RTX 5090 32GB and validated 200k context in vLLM. The setup used fp8_e4m3 KV cache, FlashInfer, and MTP with 3 speculative tokens; a 10-run 200k pass completed with 73.6 tok/s mean generation and 70.2s TTFT. The key constraint is 32GB VRAM: logs showed 8.3GiB KV cache and about 30478MiB total GPU use.

Why it matters: HKR-H/K/R all pass: the hook is single-GPU 200k context, with concrete vLLM settings and 10-run stability data. Reddit sourcing keeps it in the 78–84 band, not P1.

Xinzhiyuan · WeChat

GPT-5.5 Instant becomes ChatGPT’s free default model

OpenAI made GPT-5.5 Instant the default ChatGPT model, rolling it out free to all users. AIME 2025 rose from 65.4% to 81.2%, responses are 30.2% shorter, and hallucinations fell 52.5% versus GPT-5.3 Instant on high-risk prompts. Plus and Pro web users get chat, file, and Gmail personalization first; the API model ID is chat-latest.

Why it matters: HKR-H/K/R all pass: a free default ChatGPT model switch, concrete benchmark and behavior deltas, and direct impact on daily OpenAI workflows. This fits the 85–94 must-write band.

Computing Life · Share · Yage

In the AI Era, Review Is Not Independent Judgment

The article examines how AI use can replace independent judgment with after-the-fact review, citing Shaw and Nave. It says review shifts toward familiarity checks; the post does not disclose experiment numbers.

Why it matters: HKR-H/K/R all pass weakly: the angle has a reversal, the post cites Shaw/Nave and a verification-complexity mechanism, and it speaks to AI review anxiety. No experiment numbers, so it stays at the low featured edge.

Latent Space

Doing Vibe Physics — Alex Lupsasca, OpenAI

Alex Lupsasca says GPT-5 reproduced his paper result in 11 minutes after a textbook warmup prompt, and ChatGPT later generated 110 pages of graviton calculations in one day; the team spent three weeks verifying the results before writing a quantum-gravity paper.

Why it matters: HKR-H/K/R all pass with first-person numbers: GPT-5 after textbook warm-up reproduced a paper result in 11 minutes, and ChatGPT generated 110 pages in a day. Single interview source and niche theoretical-physics context keep it at 84, below official-release weight.

The Verge · AI

OpenAI claims ChatGPT’s new default model hallucinates way less

OpenAI says ChatGPT’s default GPT-5.5 Instant reduced hallucinations in internal evaluations. Versus GPT-5.3 Instant, hallucinated claims fell 52.5% on high-stakes prompts. Inaccurate claims fell 37.3% on flagged hard chats; the post does not disclose full eval size.

Why it matters: OpenAI changed ChatGPT’s default model and gave two hallucination-reduction figures, satisfying HKR-H/K/R. Internal evals lack set size and reproduction details, but a default ChatGPT model change is same-day material.

TechCrunch · AI

OpenAI releases GPT-5.5 Instant, a new default model for ChatGPT

OpenAI released GPT-5.5 Instant as ChatGPT’s new default model. The company says it reduces hallucinations in law, medicine, and finance while keeping prior low latency; the post does not disclose benchmarks, rollout scope, or pricing.

Why it matters: HKR-H/K/R all pass: a new ChatGPT default model, testable reliability claims, and direct workflow impact. Missing eval numbers, rollout scope, and pricing keep it in the mid 85–94 band.

May 5Tuesday

r/LocalLLaMA

Interactive Guide from Hugging Face Comparing RL Environments Across Frameworks

Hugging Face’s post-training team published an interactive guide comparing RL environment frameworks. The team spent one month building environments in verifiers, OpenEnv, Nemo-Gym, OpenRewards, and others, then trained models to study scaling. The post does not disclose benchmark scores, model sizes, or training costs.

Why it matters: HKR-H/K/R pass through the HF comparison hook, one-month hands-on setup, and post-training cost nerve. Missing benchmark scores, model sizes, and training costs keep it at the low featured band.

OpenAI News

GPT-5.5 Instant: smarter, clearer, and more personalized

OpenAI updated ChatGPT’s default model to GPT-5.5 Instant for default chat use. The RSS snippet says answers are more accurate, hallucinations are reduced, and personalization controls improved; the post does not disclose metrics, pricing, or context window.

Why it matters: HKR-H/K/R all pass: OpenAI changed ChatGPT’s default model to GPT-5.5 Instant. The post lacks evals, pricing, and context window details, so it stays at the low end of the 85–94 band.

Synced · WeChat

Agent-World Scales Real-World Environment Synthesis for Evolving General Agents

Agent-World builds 1,978 environments and 19,822 tools to train agents on long-horizon tasks. It combines web mining, tool generation, verifiable task synthesis, and GRPO training, with tasks averaging over 15 turns. The key signal is the scaling link among environment count, self-evolution rounds, and 23 benchmarks.

Why it matters: HKR-H/K/R all pass: Agent-World reports 1,978 environments, 19,822 tools, 15+ average turns, and 23 benchmarks. It is a strong agent research release, not a same-day must-write product launch.

May 4Monday

Synced · WeChat

ACL 2026: PolyU Open-Sources SignThought for Gloss-Free Sign Language Translation

PolyU and Sichuan University introduced SignThought, accepted to ACL 2026 Main and slated for oral recommendation. It uses latent thoughts, plan-then-ground, and dual-stream decoding, reaching top gloss-free BLEU-4 on five SLT benchmarks. The team also built LC-HKSLT with 1,311 hours, 432K clips, and 14 signers.

Why it matters: ACL 2026 Main, an open model, and a new dataset satisfy HKR-H/K/R, with concrete mechanisms and five benchmarks. The niche sign-language focus keeps it below broader model or developer-tool releases.

Xinzhiyuan · WeChat

Top AI wrote dozens of pages of derivation before reviewers found the problem was wrong

Xinzhiyuan says Google DeepMind used Aletheia on 700 Erdős problems and got 13 original answers. The pipeline had Gemini Deep Think produce 200 candidates, then a verifier reduced them to 63. The post says Erdős-75 had a wrong premise, yet Aletheia wrote dozens of proof pages.

Why it matters: HKR-H/K/R all pass: the mistaken Erdős-75 setup gives a sharp hook, while the 700/13/200/63 pipeline adds substance. This is strong research coverage, not a GPT-scale product release, so it fits 78–84.

最佳拍档 (BestPartners)

Why Claude Code Got Worse: Anthropic’s Review of Three Bugs

The title says Anthropic reviewed Claude Code regressions involving three bugs. It names reasoning-strength changes, a cache optimization error, and a system-prompt length limit; the post does not disclose repro steps, timeline, or fix status. The key point is AI reviewing AI code under engineering constraints.

Why it matters: HKR-H/K/R all pass, but the post gives three cause categories without repro steps, timeline, or fix status. Claude Code relevance is high, so this sits in the 72–77 band.