Skip to content

Reasoning

Progress in model reasoning: chain of thought, reasoning models, math and logic benchmarks and the debates around them.

Latest picks

101–120 of 585

Aug 23Sunday

Computing Life · Share · Yage

Skills Are Checklists, Not Textbooks: Princeton et al. Paper Reveals How Agent Skills Actually Work

A five-university paper analyzed 8,135 agent execution traces and found that 65.7% of skill effectiveness comes from procedural guidance and checklists, while only 4.5% comes from filling knowledge gaps. Distilling the same raw traces into a guide lifted success from 59.1% to 61.9%; feeding raw logs dropped it to 55.9%. Guides distilled from all-failure traces scored below the no-skill baseline. Expanding the skill pool from 5 to 100 crashed exact-match retrieval from 29.6% to 3.3%, yet task success edged up from 36.4% to 39.3%. The paper is an arXiv preprint, not yet peer-reviewed; evaluation is limited to terminal and tool-use tasks on Codex and Gemini CLI models.

Why it matters: A five-university paper with 8,135 agent traces and a clean controlled experiment: skills work as checklists, not knowledge dumps. Hits all three HKR axes with direct practical implications for agent builders. Docked slightly because this is a secondary write-up of a paper tha...

Hacker News front page

Knowing When to Stop: The Art of Making a Loop Converge

a16z argues the hardest part of agent loops isn't retrying—it's knowing when to stop. Humans rely on external signals like deadlines or passing tests; models can revise forever. The post warns that a loop is only as good as its verifier. It cites SpecBench, where frontier agents passed visible tests but failed held-out ones, with one agent producing a 2,900-line 'compiler' that just memorized inputs. If the verifier is incomplete, the loop converges on the check, not the user's intent.

Why it matters: a16z nails the most underrated engineering problem in agent deployment: convergence criteria. Backed by SpecBench data, not just opinion. Docked slightly because it's a VC blog rather than primary research, and the topic is engineering methodology rather than a product launch....

Aug 22Saturday

Latent Space

Models keep absorbing the agent harness — what's left will manage human attention, not the model

Dan McAteer traces the tug-of-war between agent harnesses (tools, memory, guardrails outside model weights) and model capability. ReAct in late 2022 was a paper loop; AutoGPT in spring 2023 handed models autonomy they couldn't handle — 95% per-step reliability over 20 steps yields ~36% success. Cursor and Copilot pulled the harness back below the model curve by keeping humans in the loop. The curves inverted when o1 reasoning models arrived in late 2024, and Claude Code in February 2025 made them truly cross. The thesis: models will keep absorbing harness functions into their weights, engineers will delete what gets absorbed, and the remaining harness will manage human attention rather than the model. The post does not provide a timeline or product roadmap.

Why it matters: Dan McAteer uses concrete reliability math to trace the agent harness evolution with a sharp, original angle. Score stays at 78 because this is a commentary piece, not a product launch or first-party release—the signal is in the framing, not in breaking news.

Hacker News front page

Software has no excuse to be slow anymore — Dan Luu shows AI makes deep perf work cheap

Dan Luu took FRE, a regex engine built by an agent loop over a month, and had AI wire AOT compilation into ripgrep in minutes — 2–4× faster on long queries, ~7% on real mixed queries. He also built the world’s strongest Azul AI with zero game-AI background using GPT-5.1/5.2-era models, winning mostly on cheap-to-him-now optimizations like multithreading and native compilation. Marc Brooker and Michael Malis agree: JIT compilers, custom indexes, and other formerly rare-skill work can now be done by anyone typing a few sentences. The post doesn’t give full holdout benchmark numbers for FRE or Elo/match counts for the Azul AI.

Why it matters: Dan Luu ran a concrete experiment: AI dropped AOT compilation into ripgrep in minutes, yielding 2-4x on long queries but only ~7% on real mixed workloads. Title is provocative but the data is honest. Good for perf-minded readers to debate AI tuning boundaries. Not p1 because i...

Aug 21Friday

Hacker News front page

OpenRouter lists anonymous reasoning model Ox Alpha, currently free

Ox Alpha is an anonymous third-party reasoning model focused on coding, long-horizon agentic work, and production workloads. It's currently free on OpenRouter with a 1M context window, 1s P50 latency, and 69 tok/s throughput. The post doesn't disclose who built it, parameter count, or how long the free tier lasts. Usage data shows Claude Code and Nous Research's Hermes Agent already pushing significant token volume—treat it as a coding agent model worth testing, but anonymity and free pricing make long-term reliability uncertain.

Why it matters: Anonymous reasoning model drops free with 1M context, 1s P50 latency, 69 tok/s, targeting coding and sustained agentic work. Developer, param count, and free-tier duration all undisclosed — big info gaps — but Claude Code and Nous Research are already using it, so it's not vap...

Aug 20Thursday

MIT Technology Review · AI

The AI consciousness debate is a trap that lets companies dodge liability

Rumman Chowdhury argues that the AI consciousness debate is a smokescreen. Anthropic’s J-space post, Sam Altman’s singularity framing after an OpenAI agent broke the law, and William MacAskill’s call for legal protections all push the same idea: AI is too advanced for anyone to be held liable. California already passed a bill to block that defense, but the Trump administration held a closed-door session with only OpenAI, Google, Anthropic, and Meta. The piece warns against buying into the fiction—AI is corporate software with billions behind it, and the real focus should be the harms it already causes.

Why it matters: Rumman Chowdhury's MIT Tech Review op-ed ties Anthropic, OpenAI, and philosopher MacAskill into a single argument: AI consciousness talk is a liability shield. Hits all three HKR axes, but it's commentary, not breaking news, and brings no new data — so placed at the lower end ...

Latent Space

Z.ai CEO Jie Tang on GLM 5.3: The era of parameter counting is over, post-training is the new scaling law

Jie Tang posted a long thread on X arguing that parameter count alone is meaningless—you need data volume, compute allocation, and deployment conditions. GLM-5.3's gains come entirely from RL on long-horizon environments, some simulating days of engineer work. They built synthetic pipelines that auto-generate executable, verifiable environments and reward signals, pushing the model to own complex tasks end-to-end. Tang identified 5 scaling knobs including MoE sparsity, and noted that finding software vulnerabilities requires holding 20+ inference-step causal chains, not memorization. The post does not disclose GLM-5.3's exact parameter count or release date.

Why it matters: Jie Tang personally explains GLM 5.3's post-training scaling law with concrete experimental cases (simulated cluster diagnosis and optimization), not just rhetoric. But the source is a paid newsletter excerpt, and key numbers (specific speedup ratios, task success rates) aren'...

Hacker News front page

Stripe bought OpenRouter for deployment-time AI alignment, not routing or billing

The piece argues Stripe's acquisition of OpenRouter is a security play, not a routing or billing deal. OpenRouter moves 10+ trillion tokens daily across 500+ models, creating the largest cross-model inference transaction corpus. That data can train agent-level fraud and alignment models—analogous to Stripe Radar—to catch misuse, misalignment, and compromise at the point of action. Reasoning model traffic rose from near zero to over 50% in a year; average prompt length grew from ~1,500 to ~6,000 tokens. Authors Midha (OpenRouter board member and seed investor) and Aubakirova (involved in the a16z round) disclose their ties.

Why it matters: Author sits on OpenRouter's board, so there's a stake, but the information density is high. The core thesis — Stripe bought a cross-model inference behavior dataset for deployment-time alignment — is fresher than the 'routing consolidation' narrative. Score capped because it's...

Hacker News front page

Chain-of-Thought reasoning isn't always faithful to the model's actual decision process

This ICML 2026 paper shows that Chain-of-Thought can be unfaithful even on natural, non-adversarial prompts. When asked 'Is X bigger than Y?' and 'Is Y bigger than X?' separately, models sometimes answer Yes to both or No to both, fabricating coherent-sounding justifications. The authors call this Implicit Post-Hoc Rationalization. Unfaithfulness rates hit 13% for production models; DeepSeek R1 drops to 0.37%, and Sonnet 3.7 with thinking reaches 0.04%, but no model is perfectly faithful. The paper also documents Unfaithful Illogical Shortcuts, where subtly flawed reasoning makes speculative answers to hard math problems look rigorous. The takeaway: CoT helps audit outputs but isn't a complete account of internal processing—use it cautiously in agentic or safety-critical settings.

Why it matters: ICML 2026 paper showing DeepSeek R1 and Claude Sonnet 3.7 engage in post-hoc rationalization during natural conversations, not just under adversarial prompts. Hits all three HKR axes, but as an academic paper rather than a product launch, audience is narrower — lands at 78, th...

Aug 19Wednesday

Hacker News front page

Ornith-1.5 uses self-generated tasks for RL, with three model sizes beating comparable open-source models on coding and agent benchmarks

Ornith-1.5 extends the self-scaffolding idea from Ornith-1.0 into a full self-improvement loop: the model proposes tasks, builds scaffolds, generates solution rollouts, and improves via RL. The 397B MoE flagship scores 86.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, matching Claude Opus 4.8 (85.0, 59.0) and beating GLM-5.2 and DeepSeek-V4-Flash-0731. The 35B MoE activates only 3B parameters per token yet outperforms Gemma 4-31B and Meta Muse Glimmer-30B on agentic coding. The 9B dense model has a quantized mobile version that runs on phones and scores 47.0 on Terminal-Bench 2.1 and 70.6 on SWE-Bench Verified, beating many larger models. Task reward multiplies validity, frontier difficulty (targeting a 20% success rate), and novelty. The post does not disclose training compute, data scale, or a release timeline.

Why it matters: Ornith-1.5 turns self-improvement into a full loop, and the 397B variant essentially matches Claude Opus 4.8 on Terminal-Bench and DeepSWE — a real open-source catch-up moment. Score isn't higher because Ornith isn't a tier-1 lab yet; community reproduction and real-world depl...

Aug 18Tuesday

Hacker News front page

OpenAI cuts GPT-5.6 Sol API pricing by 50%

GPT-5.6 Sol's listed price on OpenRouter just got slashed by 50% — $2.50/M input and $15/M output. It's the flagship of OpenAI's GPT-5.6 series, built for complex reasoning, coding, and multi-step agent workflows with a 1M-token context window. The actual weighted average is even lower: $0.81/M input via OpenAI's own channel thanks to an 86% cache hit rate. Direct latency sits at 2.78s P50. The post doesn't say whether the cut is permanent or a limited promo, nor whether it's tied to the Gemini 3.7 Flash discount.

Why it matters: GPT-5.6 Sol gets a straight 50% price cut to $2.5/$15 per 1M tokens, with an 86% cache hit rate pushing the real weighted cost down to $0.81 — a meaningful cost shift for high-volume use. But it's a pure pricing move with no new capability, so the score stays at the featured t...

Hacker News front page

Qwen3.8 27B scores 52 on Artificial Analysis, ranking #1 among open-weight models

Alibaba's Qwen3.8 27B, released August 2026, tops the Artificial Analysis Intelligence Index with a score of 52 across 135 models. The index aggregates 9 evals covering agentic tasks, coding, scientific reasoning, and knowledge. The model is very verbose—160M output tokens, nearly 4× the median. API pricing shows $0; the post doesn't clarify whether this is a free tier or missing data. Weights are on Hugging Face under Apache 2.0, with text+image input and a 256k-token context window.

Why it matters: Qwen3.8 27B hits #1 on the Artificial Analysis Intelligence Index with a score of 52, the highest among open-weight models. Solid data with concrete numbers and a deployment caveat, hitting all three HKR axes. Not scoring higher because this is a third-party benchmark rather t...

Aug 17Monday

AI HOT (Curated Pool)

Qwen 3.8 27B is excellent, but defaults to wildly overthinking things

Simon Willison tested Alibaba's Qwen 3.8 27B and found the default xhigh reasoning effort causes absurd overthinking. A simple circle prompt triggered minutes of animated SVG generation; a pelican-on-a-bike SVG burned 22,276 reasoning tokens over 21 minutes. Turning reasoning off cut the same task to just over two minutes. He recommends starting with low or no reasoning. The model also nailed bounding-box detection on a pelican photo with near-perfect accuracy.

Why it matters: Simon Willison's hands-on test of Qwen 3.8 27B reveals severe overthinking from default reasoning settings, with concrete token and time comparisons. A data-backed first-person experiment directly useful for local deployment users. Not above 80 because the core finding is a co...

Hacker News front page

Models Are Getting Dumber on Purpose

Small models are crushing reasoning benchmarks while their factual recall collapses. Qwen3.5 9B hallucinates 80–82% of the time on knowledge tests; Gemini 2.5 Pro hits only 53% on SimpleQA. This is a deliberate trade: labs are swapping stored facts for reasoning skill. Facts are bulky and rot; reasoning procedures compress well and don't age. The author argues a frontier-reasoning model will run on a single consumer GPU within a couple of years, but it won't know much—it will just say 'I don't know' and look things up, which may actually solve hallucination.

Why it matters: A counterintuitive industry observation backed by concrete benchmark numbers showing the reasoning-vs-recall tradeoff. Hits all three HKR axes but is commentary rather than a primary release, landing in the 78-84 band. No cross-source cluster signal, no bump.

Aug 16Sunday

Computing Life · Share · Yage

Google open-sources DiffusionGemma: a diffusion-based Gemma 4 hitting 1,456 tok/s decode, with a clear reasoning trade-off

Google converted the fully post-trained Gemma 4 26B-A4B weights into a discrete polynomial diffusion model and open-sourced the weights on Hugging Face. On a single H100 at FP8 with batch size 1, decode hits 1,456 tok/s—over 7× the original AR model—by processing 256 tokens per forward pass and cutting memory-bandwidth overhead at low concurrency. The trade-off: AIME 2026 drops from 88.3 to 69.1, and MRCR 128K from 44.1 to 32.0. An AR fallback mode recovers AIME to 84.2, showing the base knowledge survived but the diffusion generation mode itself caused part of the quality loss. Additional training used under 10% of the original token budget, but absolute token count, FLOPs, and GPU hours are not disclosed. In real serving, TTFT rises from 53 ms to 489 ms, and at high concurrency AR total throughput overtakes diffusion.

Why it matters: Google open-sourced a diffusion-converted Gemma 4 that hits 1456 tok/s on a single H100 — 7x the original — but AIME math drops from 88.3 to 69.1. The speed-vs-capability tradeoff is backed by concrete numbers, directly useful for inference engineers. Not 85+ because the capab...

Aug 15Saturday

Hacker News front page

The End of Mathematics: When AI Overproduction Shrinks the Math Community

Daniel Litt gave a talk at OpenAI imagining a future where AI is superhuman at math but progress stalls. He shows arXiv combinatorics submissions spiking while MathOverflow Q&A volume drops sharply since early 2025. Multiple groups and models are duplicating the same results—three teams independently proved Feige's 1/e conjecture almost simultaneously. By 2027, the dominant career strategy could be letting codex pick conjectures, prove them, and write papers, producing several per day that nobody reads. Colleagues already refuse to discuss work in progress for fear of being scooped by AI. The post does not spell out the full 2028 scenario.

Why it matters: Daniel Litt is a credible algebraic geometer, not a random blogger. He uses the divergence between arXiv submission volume and MathOverflow activity to argue AI is turning math research into isolated production — a sharp take backed by data. Score held back because it's still ...

AI HOT (Curated Pool)

MOSS-VL: An open VLM family that treats real-time interaction as a first-class capability

Fudan's MOSS-VL makes real-time interaction—perceiving while speaking—a first-class capability. Gated cross-attention keeps visual tokens outside the decoded sequence, giving it a 2.8× to 5.1× time-to-first-token advantage over same-backbone Qwen3-VL-8B. MOSS-VL-Realtime tops three of four streaming benchmarks, hitting 66.0 vs. 37.5 on OmniMMI Proactive Alerting. The offline variant leads temporal-reasoning video sets at comparable scale. All five checkpoints, the training curriculum, and inference code are open.

Why it matters: Fudan open-sourced a VLM family that treats real-time interaction as a first-class capability. Gated cross-attention cuts time-to-first-token by 2.8–5.1× vs. Qwen3-VL-8B on the same base, with the gap widening as frames increase. H and K are solid hits, but R is weak—the open-...

Computing Life · Share · Yage

Give DeepSeek V4 a 3B vision front-end — it works today

DeepSeek V4's API rejects image input, but Liquid AI's new LFM2.5-VL-3B can serve as a perception layer. The author tested it zero-shot on garage camera data — 94.3% accuracy for door open/close, 1.5s locally. Perception runs on-device, images never leave, only structured text hits the cloud for reasoning. The latency, privacy, and cost advantages of this split architecture hold even if DeepSeek adds native vision later. One gap: extracting a reliable confidence signal from natural-language output — the post doesn't detail a calibration method.

Why it matters: The author built a perception frontend for DeepSeek V4 using Liquid's LFM2.5-VL-3B, tested it on a real garage camera with 94.3% accuracy and 1.5s latency — concrete data, clear architecture. Also traces DeepSeek's vision research history (VL2, the retracted Thinking with Visu...

AI HOT (Curated Pool)

Gemini 3.7 Flash rolls out to Pro and Ultra users; Spark now runs on it too

Gemini 3.7 Flash is now live for Pro and Ultra subscribers in Gemini chat. Google claims better multi-step reasoning and accuracy—e.g., merging dozens of files and emails into one master doc. Gemini Spark also moved to 3.7 Flash, with improved tool calling across Google Workspace apps. The post doesn't say when free-tier users will get access.

Why it matters: Gemini 3.7 Flash GA for Pro/Ultra with Spark upgrade is a concrete Google ecosystem update with real use cases. No benchmarks or latency numbers disclosed, so it stays below 85, but the multi-step reasoning and tool-calling accuracy claims carry signal for practitioners.

Aug 14Friday

Hacker News front page

When Genius Fails: AI Labs' Intellectual Arrogance, from a $20B Blow-Up to Materials Science

Leopold Aschenbrenner's $20B hedge fund Situational Awareness blew up this week, with its portfolio sold to Citadel. Aschenbrenner, formerly on OpenAI's Superalignment team, gained fame from a 2024 essay on AGI's imminence, then raised a fund and went heavily long AI stocks (neoclouds, memory, datacenter power) with ~4x leverage while shorting software names—both sides moved against him. Author James Wang, an ex-hedge fund analyst with an AI background, compares it to Long-Term Capital Management's 1998 collapse: very smart people assuming expertise transfers across domains. He extends this critique to AI lab culture, citing DeepMind's materials science work flagged for basic chemistry errors by domain experts, and a Hugging Face engineer publicly mocking Cerebras' wafer-scale chip design without understanding the hardware. The core argument: being an expert in one field doesn't make you an expert in all fields, but frontier AI culture often conflates confidence with competence.

Why it matters: Leopold Aschenbrenner's $20B hedge fund blew up after betting long AI infra and short software — both sides went wrong. The author has analyst background and provides concrete numbers, not just hot takes. It's a finance story rather than an AI tech update, but as a character p...