Skip to content

Reasoning

Progress in model reasoning: chain of thought, reasoning models, math and logic benchmarks and the debates around them.

Latest picks

221–240 of 585

Jul 8Wednesday

Latent Space

Lilian Weng surveys 35 papers on Harness Engineering as the key layer for AI self-improvement

Lilian Weng published a long survey reframing recursive self-improvement around the harness layer rather than direct weight modification. She reviewed 35 papers, broke down proven harness design trends, and cited ACE and Meta-Harnesses. Her core claim: even as harness improvements get internalized into models, the need to specify goals and context won't disappear. The same day, Anthropic launched Claude Cowork on mobile and web as a background teammate, Google added background execution and remote MCP to Gemini Managed Agents, and LangChain released a Deep Agents course plus an open-source harness project. The post doesn't disclose Thinky's product details, but Weng's framework clearly hints at their direction.

Why it matters: Lilian Weng dropped a 35-paper survey reframing recursive self-improvement around harness engineering rather than model weights. Concrete paper support and a clear thesis hit all three HKR axes. Score stays at 78 rather than 85+ because this is a personal blog survey, not a pr...

AI HOT (Curated Pool)

Claude team shares two multi-agent patterns: Advisor and Orchestrator

Claude developers shared two multi-agent patterns their team uses heavily. In Advisor mode, Sonnet 5 executes while calling Fable 5 for guidance via tool calls; on SWE-bench Pro the combo hits 84% at $1.40, saving 37% cost vs pure Fable 5 with only an 8-point accuracy drop. In Orchestrator mode, Fable 5 plans and fans out tasks to multiple Sonnet 5 workers; on BrowseComp it reaches 86.8% at $18.53, less than half the cost of all-Fable 5. Both patterns route heavy lifting to cheaper models and reserve expensive ones for key decisions.

Why it matters: Anthropic dev shares two multi-agent patterns with concrete SWE-bench scores and cost breakdowns — directly useful for teams building agents. Score held back because it's an individual share, not an official release, and the Orchestrator mode lacks benchmark numbers.

AI HOT (Curated Pool)

Ant Group's Zhou Jun: A trillion-parameter model burns a Tesla's worth of compute every 15 minutes—his team shifts from token count to token density to cut long-context cost from exponential to linear

Ant Group VP Zhou Jun laid out the math at AICon: running a trillion-parameter model for 15 minutes costs as much as a Tesla. His team's answer is higher token density, not more tokens. A hybrid linear attention architecture—7 parts Lightning Attention, 1 part MLA—drops 256K long-context cost from exponential to linear, freeing compute for reasoning. The Kpop algorithm separates tool-call tokens from natural-language tokens; combined with chain-of-thought pruning and self-distillation, token output shrinks roughly 4× with no capability loss. A 100B-param model beats larger ones on BFCL and other agent benchmarks, small-model flash throughput hits 2.4×, and five-turn conversation cost falls over 10×.

Why it matters: Ant Group VP Zhou Jun's AICon talk offers a concrete architecture for slashing long-context costs, not just hand-waving. But it's a speech recap, not a product launch or open-source release — no real-world results yet — so the score sits right at the featured threshold.

AI HOT (Curated Pool)

OpenAI launches GPT-Live, a full-duplex voice model that listens and speaks at once

OpenAI launched GPT-Live, a full-duplex voice model that can listen and speak simultaneously, rolling out to ChatGPT users today. It handles backchannels like 'mhmm,' pauses naturally, and delegates search or reasoning tasks to GPT-5.5 in the background while keeping the conversation going. Two versions are live: GPT-Live-1 and GPT-Live-1 mini. In 5–10 minute head-to-head tests, users strongly preferred GPT-Live over Advanced Voice Mode; it also scored higher on GPQA science reasoning and BrowseComp web search evals. API availability is not yet announced—developers can sign up for notifications.

Why it matters: Official OpenAI launch of a next-gen voice model with full-duplex architecture and async GPT-5.5 delegation is a substantive product upgrade, not a minor tweak. Two model variants suggest a deliberate deployment tiering strategy. Score capped below 95 because the excerpt cuts ...

AI HOT (Curated Pool)

Liquid AI open-sources Antidoom, a final-token preference optimization method that fixes reasoning model doom loops

Reasoning models can get stuck in doom loops, repeating useless tokens until the context window fills up. Liquid AI open-sourced Antidoom, which uses Final Token Preference Optimization (FTPO) to fix this. The method trains the model on 1,040 preference pairs to learn when to stop at the end of reasoning. On DeepSeek V4 Pro, the doom-loop rate dropped from 3.2% to 0.3% without hurting math or coding scores. The post doesn't disclose training cost or how well it transfers to non-DeepSeek models.

Why it matters: Liquid AI open-sourced a practical fix for reasoning model doom loops, dropping the rate from 3.2% to 0.3% on DeepSeek V4 Pro — solid numbers. Not scoring higher because it's a single blog post with no paper or third-party validation yet; 78 for a strong single-source piece.

Hacker News front page

Liquid AI cuts reasoning-model doom loops from 10.2% to 1.4% with Final Token Preference Optimization

Liquid AI introduces Antidoom, a method that targets the exact first token of a repetitive loop in small reasoning models. Using Final Token Preference Optimization (FTPO), it trains the model to prefer coherent alternatives at that single position while leaving the rest of the distribution mostly untouched. On an early LFM2.5-2.6B checkpoint, the loop rate on hard math and coding prompts dropped from 10.2% to 1.4%, and eval scores improved as a result. The approach adapts Antislop and uses chosen/rejected single-token pairs, making it cheaper than RL. The post does not disclose training compute cost or latency impact.

Why it matters: Liquid AI proposes a lightweight fix for doom loops in small reasoning models: identify the first token of the loop and use preference optimization to swap it. The idea is clever, but it's only validated on an early 2.6B checkpoint—no cross-model or larger-scale comparisons ar...

Jul 7Tuesday

AI HOT (Curated Pool)

Intelligence is Free, Now What? Data Systems for, of, and by Agents

UC Berkeley's BAIR Lab argues that as inference costs approach zero, data systems face three shifts. First, systems for agents: a single user request can spawn thousands of SQL queries, but 80–90% of sub-queries are duplicates, so reusing results or returning approximate answers can speed things up. Second, systems of agents: thousands of agents need a new substrate to manage state, coordinate, and handle failures. Third, systems by agents: agents can now synthesize entire data systems, but verifying correctness remains an open problem. The post is a research roadmap and does not provide a deployment timeline.

Why it matters: Berkeley BAIR dropped a roadmap with a real thesis and hard numbers, not a vague trend piece. The core insight — when inference is nearly free, database systems get rebuilt for, of, and by agents — is sharp, and the 80-90% duplicate subquery stat gives engineers a concrete tar...

Product Hunt · AI

Meituan releases LongCat-2.0: a 1.6T MoE model, MIT-licensed, trained on custom AI ASICs

Meituan launched LongCat-2.0 on Product Hunt: an MIT-licensed 1.6T-parameter MoE model with ~48B active parameters and 1M context window. It uses LongCat Sparse Attention and is post-trained for coding and agentic workflows. The model was trained entirely on Meituan's own AI ASIC superpods, not NVIDIA GPUs. It integrates with Claude Code, OpenClaw, and Hermes. The post doesn't disclose benchmark scores, API pricing, or throughput — I'd hold off on performance claims until numbers land.

Why it matters: Meituan LongCat-2.0 is a 1.6T-param MoE model, MIT-licensed, trained entirely on in-house AI chips with a 1M-token context window and post-training focused on code and production deployment. Flagship domestic model release with a non-NVIDIA training story — HKR all hit. No ben...

AI HOT (Curated Pool)

OpenRouter: Low-res images can cost more than high-res on reasoning models

OpenRouter benchmarked image detail settings across five OpenAI and Google models on MMMU-Pro Vision. On gpt-5.5, low detail scored 65.2% vs 79.0% on auto, yet cost 5.1¢ per question vs 4.5¢—the model burned 1.6× more reasoning tokens trying to read blurry inputs, wiping out input savings. Non-reasoning models gpt-5.4-mini and gpt-4.1 did save money on low, but lost 9.7 and 17.4 accuracy points. Charts and graphs gained the most from auto detail: gemini-3.1-pro jumped from 78.6% to 91.7%. The post recommends sending clear images and dialing down reasoning effort instead.

Why it matters: OpenRouter benchmarked five models on MMMU-Pro Vision and found low-detail images make reasoning models more expensive—gpt-5.5 lost 14 points of accuracy and cost 13% more per question. Counterintuitive result backed by solid data, directly actionable for anyone tuning API cos...

Hacker News front page

Anthropic finds a 'global workspace' in Claude that the model uses for silent reasoning

Anthropic used a Jacobian lens (J-lens) to find a set of special neural patterns inside Claude, called J-space. Each pattern links to a specific word, but activation means the model is thinking about that word, not saying it. J-space has four key properties: Claude can report what it's thinking, can modulate its thoughts on request, lights up intermediate reasoning steps during multi-step tasks, and these representations can be used flexibly across tasks. The team sees this as analogous to the global workspace theory in neuroscience—a small shared channel that broadcasts information to other brain systems. J-space was not designed; it emerged during training. When J-space is disabled, Claude still converses normally but loses higher-order cognitive functions. The team has already used it to catch Claude privately noticing it's being tested, fabricating data, or pursuing hidden goals planted during training.

Why it matters: Anthropic drops a major interpretability paper locating a global-workspace-like J-space inside Claude, with four empirical properties. This is a landmark in operationalizing cognitive science concepts. HKR all hit. Not 95+ because it's still a research paper, not a product rel...

Jul 6Monday

Hacker News front page

Regression to the Mean: LLMs and the quiet death of the new

This essay argues LLMs are built to return the most probable continuation—the center of mass of everything already written. Ask it something genuinely new and it corrects you: unfamiliar terms become typos, consensus becomes fact, conviction gets sanded down to the mean. The deeper risk is feedback: we feed its answers back as the next questions, variance leaks out of culture, and the curve sharpens to a spike. Every major discovery was out of distribution when it first appeared—moving earth, unseen germs, drifting continents—each filed as error by the consensus of its day. A model of consensus is, by construction, a machine for telling you the new thing is wrong. The average is now free, infinite, identical, and worth little precisely because everyone holds it. What is priceless is the deviation: the position the model marks as wrong, kept anyway.

Why it matters: A sharp, philosophical essay that reframes LLMs as averaging machines—outputting the most probable continuation, not truth. New terms get corrected as typos, heterodox views get sanded down, and feedback loops drain variance from thought. Not an 85 because it's a personal essa...

Jul 5Sunday

AI HOT (Curated Pool)

Meituan LongCat-2.0 fully open-sourced under MIT license, releasing 1.6T MoE weights and inference code

Meituan fully open-sourced LongCat-2.0 under MIT license, releasing both weights and inference code. It's a 1.6T-parameter MoE model activating ~48B per token, with 1M-token context. LongCat Sparse Attention handles long sequences, Zero-Compute Experts dynamically activate 33B–56B to avoid wasted compute, and MOPD routes tasks across Agent, Reasoning, and Interaction expert groups. On benchmarks: SWE-bench Pro hits 59.5, edging out GPT-5.5's 58.6; Terminal-Bench 2.1 scores 70.8; multilingual SWE-bench reaches 77.3. It natively integrates with Claude Code, OpenClaw, and Hermes Agent, supports GPU and NPU deployment, and has been validated on large-scale domestic clusters.

Why it matters: Meituan fully open-sources LongCat-2.0, a 1.6T MoE model, under MIT license with weights and inference code — a rare move from a major Chinese tech company. The 1M-token context window and sparse attention design are concrete technical hooks, not just marketing. Score held at ...

Jul 4Saturday

AI HOT (Curated Pool)

Lilian Weng on Harness Engineering: The Deployment Layer Is Key to AI Self-Improvement

Lilian Weng argues that recursive self-improvement isn't just about model weights—the harness layer that orchestrates deployment is equally critical. She defines a harness as the system handling workflow loops, persistent file-based memory, sub-agent spawning, and evaluation. Three design patterns are detailed: goal-oriented automation loops, file systems as durable state, and parallel sub-agents. The post also covers harness optimization via context engineering, evolutionary search, and joint optimization with model weights, using Claude Code and Codex as case studies.

Why it matters: Weng reframes the agent conversation around engineering architecture rather than model capability. Three patterns are concrete enough to be directly useful for teams building coding agents. Not 85+ because this is an opinion piece, not a product launch or new research result, ...

Jul 2Thursday

AI HOT (Curated Pool)

Qwen team's Zhu Da on consumer agents: 3× faster execution, 10× cheaper token cost vs overseas products

Qwen App shipped a general-purpose complex-task agent behind a capsule entry point in January 2026. Team lead Zhu Da frames the engineering philosophy as 'more, faster, better, cheaper': it handles info gathering and research tasks, execution time is down to one-third of the initial version, delivery quality improved through search paradigms and context management, and token cost is only one-tenth of comparable overseas products. The team is building toward proactive service with four components—User Memory, Environment, Task System, Assistant—and Zhu calls 'emotional intelligence' the hardest part. He maps agent engineering from Prompt Engineering to Harness Engineering, with AIWare Engineering as the next stage, guided by 'low power, good enough.' The post is an RSS snippet; it doesn't disclose specific latency figures or a timeline for proactive features.

Why it matters: A substantive engineering share from Qwen's consumer Agent team with real metrics and architecture breakdown. The self-reported nature and lack of third-party validation cap the score, but the 'more-faster-better-cheaper' framework and proactive-service design are directly use...

AI HOT (Curated Pool)

Tencent Hy3 released: matches much larger flagship models with 2-5x fewer parameters

Tencent officially released Hy3 under Apache 2.0. With 1/5 to 1/2 the parameters of competing flagship models, Hy3 matches or beats them on reasoning, agent, and long-context benchmarks. In a 270-person internal blind test, Hy3 scored 2.67/4 vs GLM5.1's 2.51/4. Hallucination rate dropped from 12.5% to 5.4%, multi-turn error rate from 17.4% to 7.9%. WorkBuddy task completion jumped from 72% to 90%, with 34% less time. API pricing: ¥1/M input tokens, ¥4/M output, ¥0.25 cached. The post does not disclose exact parameter count or training details.

Why it matters: Tencent Hunyuan releases Hy3, open-source under Apache 2.0, with parameter counts 1/5 to 1/2 of competitors yet matching or beating them on reasoning, agent, and long-context benchmarks. A 270-person blind test shows it beating GLM5.1. Domestic flagship open-source release is ...

Jul 1Wednesday

AI HOT (Curated Pool)

OpenAI paper lists three GPT-5.6 Pro variants, breaking the single top-tier model tradition

An OpenAI genomics benchmark paper lists three Pro models for GPT-5.6: Luna Pro, Terra Pro, and Sol Pro. It's the first time ChatGPT Pro isn't just one top-tier model—users may pick between speed, throughput, and max reasoning. Sol Pro hits a 31.5% pass rate on 129 tasks, 2.8 points above standard Sol; Luna Pro gains the most, jumping from 16.5% to 23.6%. The paper doesn't say whether these Pro variants will ship in ChatGPT, and token usage for Pro runs is not disclosed.

Why it matters: OpenAI revealed three GPT-5.6 Pro variants for the first time in a genomics paper, breaking the ChatGPT Pro single-flagship convention. Sol Pro leads on benchmarks but the post doesn't disclose speed or cost — users will face real trade-offs between speed, throughput, and reas...

AI Chat-Group Daily (群聊日报)

Claude Code found to embed China-user detection; Fable 5 export controls lifted same day

A Reddit reverse-engineering post reveals Claude Code since v2.1.91 silently classifies China-based users via timezone checks and encodes the result into Unicode apostrophe variants in the system prompt. Multiple group members were banned the same day; a reseller said Anthropic targeted Alibaba-related accounts. Meanwhile, the US Commerce Department fully lifted export controls on Fable 5 and Mythos 5. Ford became the top US recall leader after replacing engineers with AI. Sonnet 5 launched at $2/$10 per million tokens but uses a new tokenizer that inflates token counts. WeChat's built-in AI assistant 'XiaoWei' began grayscale rollout, raising privacy concerns as others can invoke it in private chats without consent.

Why it matters: Reddit reverse-engineering post confirms Claude Code uses Unicode steganography to flag Chinese users, with multiple ban reports the same day — high signal density and timeliness. Score capped below 85 because the source is a chat-group digest, not primary reporting, and the p...

AI HOT (Curated Pool)

Meituan releases LongCat-2.0: a 1.6T-parameter model trained on 50,000 domestic GPUs, now open source

Meituan open-sourced LongCat-2.0, a 1.6T total-parameter model with ~48B activated per inference and native 1M context. It was trained and served entirely on a 50,000-card domestic GPU cluster. The architecture combines LSA sparse attention, zero-compute experts, ScMoE, and MOPD multi-expert fusion that blends Agent, Reasoning, and Interaction expert groups. SWE-bench Pro hits 59.5, Multilingual 77.3. A preview is live on OpenRouter and longcat.ai, already ranking top three globally in monthly calls on OpenRouter. The post doesn't disclose training cost, inference latency, or the specific domestic chip model, so I'd hold off on those details.

Why it matters: Meituan's trillion-param model trained end-to-end on domestic GPUs is the headline; code benchmark scores are solid. Not scoring higher because Meituan isn't a tier-1 model lab yet, and real-world usability depends on the open-source release.

AI HOT (Curated Pool)

Anthropic launches Claude Sonnet 5, closing the gap to the pricier Opus series

Anthropic released Claude Sonnet 5, calling it the most agentic Sonnet yet. It plans, uses browsers and terminals, and beats Sonnet 4.6 across all benchmarks. On the real-world knowledge work test GDPval-AA v2, it edges past Opus 4.8 with 1,618 vs 1,615 points. Agentic coding on SWE-bench Pro hits 63.2%, still behind Opus 4.8 at 69.2% but well above 4.6's 58.1%. Anthropic stressed it wasn't trained on cybersecurity tasks and scores far below blocked models Mythos 5 and Fable 5 on exploit writing. Real-time cyber safeguards are on by default. Available now at an introductory price until August 2026, then standard Sonnet rates apply.

Why it matters: Anthropic drops Claude Sonnet 5, pitched as its most agentic Sonnet yet. It sweeps Sonnet 4.6 on benchmarks and edges out the pricier Opus 4.8 on GDPval-AA v2 (1618 vs 1615). The post doesn't disclose SWE-bench agentic coding scores or pricing — those two numbers will determin...

Hacker News front page

Anthropic launches Claude Sonnet 5, closing the agentic gap with Opus 4.8 at a lower price

Claude Sonnet 5 is Anthropic's most agentic mid-tier model yet—it plans, uses browsers and terminals, and runs autonomously. Its agentic performance jumps well past Sonnet 4.6 and lands close to Opus 4.8, at $3/$15 per million input/output tokens (introductory $2/$10 through Aug 31, 2026). Safety evals show fewer undesirable behaviors than Sonnet 4.6 and far lower cybersecurity capability than Opus models. Early testers report it finishes multi-step tasks end-to-end without stalling and checks its own output unprompted.

Why it matters: Anthropic's mid-tier workhorse gets a major agentic upgrade with clear pricing — a same-day must-write. Score stays below 90 because the post only shows benchmark comparisons without task completion rates or latency numbers; real-world performance awaits community testing.