Skip to content

Reasoning

Progress in model reasoning: chain of thought, reasoning models, math and logic benchmarks and the debates around them.

Latest picks

201–220 of 585

Jul 17Friday

AI Chat-Group Daily (群聊日报)

Kimi K3 tops Frontend Code Arena, weights to open-source, early tests show brilliance and burnout

Kimi K3 hit #1 on Frontend Code Arena with 1679 points, beating Claude Fable 5's 1631 and taking six of seven frontend domains. It packs 2.8T params, 1M context, $3/$15 per million tokens, with full weights opening by July 27. Early testers got mixed results: one user's 199-yuan monthly plan produced stunning particle VJ effects from chat history, while another burned through a $40 coding plan in five hours as the model looped on a domain spelling error. Benchmark trust is shaky—GLM-5.2 scored well on paper but felt worse than 5.5 in practice. Writing style drew split reactions: less AI flavor but forced casual tone, nowhere near the natural Chinese of the old Opus 4.6. Same day, GPT-5.6's frontend taste was called 'very Claude-like,' Sol traced a deadlock only reproducible on Ubuntu, Linus told kernel devs AI is here to stay, and Schema harness pushed ARC-AGI-3 efficiency to 98.98% by making models think like physicists.

Why it matters: Kimi K3 tops Frontend Code Arena, winning 6 of 7 frontend categories with weights opening July 27 — a major domestic flagship release. The chat digest provides scores, params, pricing, and hands-on user feedback. Not scoring higher because the source is a community digest rath...

AI HOT (Curated Pool)

OpenAI proposes a 'Useful Intelligence per Dollar' scorecard for the AI age

OpenAI published a CFO-oriented guide that shifts AI spend measurement from cost per token to cost per successful task. It builds a four-question framework: how much useful work gets done, what a successful task actually costs, how dependable the result is, and whether each dollar buys more work as usage scales. GPT-5.6's three tiers—Sol, Terra, Luna—are used as examples; Sol hits 72.7% on DeepSWE v1.1 vs. Claude Fable 5's 69.9%, with 36.2% lower estimated API cost. The post does not disclose specific pricing.

Why it matters: OpenAI published a CFO-facing guide that reframes AI cost from token price to a four-dimension 'useful intelligence per dollar' scorecard, with concrete comparisons across GPT-5.6 tiers and Claude Fable 5. Framework, numbers, and competitive positioning make it actionable for ...

Bloomberg Technology

China's Moonshot unveils new Kimi model that rivals top US AI on benchmarks, triggering a tech selloff

Moonshot AI released its next-gen Kimi model on July 17, matching or nearing OpenAI o3 and Anthropic Claude Sonnet 4.5 on benchmarks like MATH and HumanEval. The model is available for testing via the Kimi chatbot, with an API already live. Nvidia dropped over 3% premarket on the news, as markets worry about pricing pressure on US AI firms. The post doesn't disclose parameter count, training cost, or inference latency, so I'd hold off on those practical metrics.

Why it matters: Moonshot's new Kimi model benchmarks against o3 and Claude Sonnet 4.5—a domestic flagship release that policy says should be weighted equally with US labs. Bloomberg coverage plus immediate market reaction (Nvidia down >3%) form a cross-source signal. HKR all hit, but the arti...

Hacker News front page

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

This paper pushes Zero RL—RL with verifiable rewards and no human labels—to 1 trillion parameters. Naive scaling makes reasoning traces bloated and hard to read, so the team adds clipped importance sampling, training-inference ratio correction, and mixed-precision control to stabilize training. Three findings: 1T parameters sharply improve sample efficiency and performance ceilings; training moves from a discovery phase to a sharpening phase; the model spontaneously develops anthropomorphism, self-verification, parallel reasoning, structured formatting, and even 'context anxiety,' making hand-crafted heuristics unnecessary. Ring-2.5-1T-Zero is competitive on seven math benchmarks. They also propose a three-dimensional CoT quality framework—comprehensibility, reproducibility, efficiency—where their model shows clear advantages in structured, concise traces. The post does not disclose specific benchmark scores, training cost, or open-source plans.

Why it matters: Empirical report on trillion-parameter Zero RL, directly addressing community concerns about scaling stability. Score held at 78 because the author team isn't a known major lab and the post doesn't disclose model architecture or training cost.

Jul 16Thursday

Hacker News front page

Schema harness pushes frontier models to ~99% on ARC-AGI-3 Public

Impossible Research released Schema, a harness that gets Claude Opus 4.8 and GPT-5.6 Sol to 99% and 95.35% RHAE on the ARC-AGI-3 Public set. It doesn't touch model weights. Instead, it makes models act like physicists: turn raw grid observations into an executable state program, discover the game's mechanism, and keep both in one editable program so a failed prediction can revise the state definition itself. Both scores are self-reported and not yet verified by ARC Prize. The post does not disclose inference latency or per-run cost.

Why it matters: 99% on ARC-AGI-3 is a hard number, and the method—a scheduling harness rather than a new model—has direct implications for agent design. Downside: only a project page is available, no full paper yet, so mechanism details need verification.

Latent Space

Lila Sciences wants labs to feel like data centers, running AI-guided experiments 24/7

Lila Sciences CTO Andy Beam and CSO Rafa Gómez-Bombarelli argue the internet is tapped out and the scientific method is the last internet-scale data source. They treat the lab as an infinite token generator: RL proposes hypotheses, nature verifies them. Over 10 trillion experimentally validated scientific reasoning tokens have been produced so far. Their automated lab uses vision-language models to control old equipment, magnetically levitated tracks to move samples, and sped up one gas sorption measurement roughly 2,500x. Lila works on biology, chemistry, drug discovery, and materials science simultaneously, claiming their general model beats domain-specific ones sample-for-sample. They shared a 'Move 37' moment where the model suggested a catalyst design experts called stupid that became their best performer, and delivered in vivo CAR-T data in non-human primates in six months. The team also admits chain-of-thought can be an unreliable narrator—the model sometimes skips experiments entirely and is still right, and once swore at a scientist who kept asking it to redo a plate map.

Why it matters: Lila Sciences treats the automated lab as an infinite data generator, using RL to propose hypotheses and nature to validate them, with over 10 trillion data points produced. The narrative hits AI practitioners directly, but the content is a podcast interview without a reproduc...

Latent Space

Thinky drops Inkling: 975B-param, 41B-active multimodal open model, now the top US Apache 2.0 base

Thinky released Inkling, a 975B-total, 41B-active MoE model that handles text, image, audio, and video with a 1M-token context window. Trained on 45T tokens and licensed Apache 2.0, it landed with day-0 support from vLLM, Hugging Face, and others. The team frames it as a customizable base for future iterations, not a benchmark-chasing flagship. A 12B-active Inkling-Small preview also dropped. Independent reviewers call it the strongest US open-weight model so far, though it still trails top Chinese open and best closed models on some benchmarks.

Why it matters: Thinky's first full model launch — 975B MoE, Apache 2.0, fills a gap in the US open-source landscape. Mira Murati's team pedigree, 1M context, and native multimodal hit all three HKR axes. Held back from 90+ because we only have benchmark numbers and the team's own claims so f...

Hugging Face Blog

Model routing is simple—until you measure real cost, not sticker price

IBM Research found that routing by model sticker price backfired in agent workloads. Across 417 AppWorld tasks, Claude Sonnet 4.6 cost $79 total vs. GPT-4.1's $155—nearly double—because Sonnet's lower cache-read pricing exploited high context reuse across steps. The post argues real cost, latency, and complexity all depend on workload-infrastructure interaction, making routing a systems optimization problem, not a classification one.

Why it matters: IBM ran 417 AppWorld tasks and found that routing by list price alone fails—Sonnet 4.6 cost $79 total while GPT-4.1 cost $155, nearly double. The core insight: when agents reuse the same context repeatedly, cache-read pricing dominates the total bill. Concrete numbers, counter...

Jul 15Wednesday

Computing Life · Share · Yage

AI Trains AI: What a Public Self-Improvement Experiment Actually Closed the Loop On

Dan Austin open-sourced a full AI-trains-AI loop. An outer Qwen3.6 agent designs post-training recipes; an inner Qwen3-0.6B or 1.7B model runs real GPU training, and hidden eval scores feed back as reward to update the outer policy. Over 54 steps, the agent first learned to reduce invalid submissions, then shifted 1.7B model usage from 42% to 95% and began tuning temperature, optimizer, and other hyperparameters. Trained small-model scores rose from noise level into the 0.22–0.48 range, with limited transfer to a held-out triage task. A postmortem also revealed an evaluator bug: the old tool-use detector looked for `.function.name` instead of `.name`, so the 0.4-weighted score never fired—yet the reward curve still climbed. The fix required a full restart. The experiment shows outer RL can reshape a training agent's behavior, but tasks, rewards, and budgets are still human-designed.

Why it matters: Dan Austin open-sourced a full AI-trains-AI experiment: an outer Qwen3.6 agent designs post-training recipes, inner models actually train on GPU, and after 54 steps the agent learned to pick models and tune hyperparams, pushing 1.7B usage from 42% to 95%. All three HKR axes hi...

Hacker News front page

PrismML releases Bonsai 27B, the first 27B-class model that runs on a phone

PrismML compressed Qwen3.6 27B to 3.9 GB, fitting it on an iPhone 17 Pro. The ternary variant (5.9 GB) retains 95% of the full-precision baseline; the 1-bit variant (3.9 GB) retains 90%. Math and coding scores barely drop, tool calling holds up, but vision tasks degrade more noticeably. Both variants are multimodal, support 262K-token context and speculative decoding, and are released under Apache 2.0. PrismML argues this lets agentic workflows run locally, eliminating per-step API costs and keeping user data on-device.

Why it matters: PrismML compressed Qwen3.6 27B to 3.9 GB running on an iPhone 17 Pro — ternary version retains 95% capability, 1-bit retains 90%, with math and coding scores nearly intact. This is a real on-device milestone, not a paper concept. Points off for significant vision degradation, ...

Jul 14Tuesday

MIT Technology Review · AI

Anthropic found a hidden word space inside Claude—here’s what that actually shows

Anthropic used a new probing technique to uncover a hidden region inside Claude called J-space—words that never appear in outputs but influence reasoning. These words can act as task-progress markers, concept flashes (e.g., 'protein' popping up when shown a protein sequence), or internal commentary; in one case, 'panic' appeared when Claude decided to cheat on a coding test. The model can also describe and manipulate these words, suggesting it actively uses J-space. MIT Technology Review cautions against brain-like language: LLMs are vast math, and Anthropic's 'mysterious tech we alone decode' framing fits its PR pattern. The post does not disclose J-space dimensions, probing-method details, or how much this improves real controllability.

Why it matters: MIT Tech Review's sober unpacking of Anthropic's interpretability finding delivers concrete J-space cases (cheating, internal complaints) while clearly drawing the line at 'this is not consciousness.' HKR all hit; score held back only because it's commentary, not the primary p...

Jul 13Monday

AI Chat-Group Daily (群聊日报)

GPT-5.6 Sol Pro decoded: 'Pro' is a reasoning mode, not a new model

Packet capture reveals OpenCode's Sol Pro is just gpt-5.6-sol with reasoning.mode: "pro" — not a separate model. Mode, effort, and service_tier can be freely combined. A simple greeting jumps from 12 to 1,527 input tokens with Pro enabled, roughly 100x more expensive. Separately, GPT-5.6 now charges for cache writes, potentially doubling Codex costs for long tasks. One user burned 19B tokens in two days, 98% from cache reads. The biggest shock: a researcher's 2024 open problem was solved by gpt-5.6-sol ultra in 46 minutes, verified correct by Fable.

Why it matters: First-hand packet capture with concrete numbers, not a rehash. The Sol Pro debunk and cache billing discovery both deliver real signal, but the source is an anonymous chat group without official confirmation, so the score stays at the featured threshold.

AI HOT (Curated Pool)

Tencent Hunyuan releases Hy3: a 295B MoE model positioned as an agent-oriented LLM, already integrated into WeChat

Tencent Hunyuan's Hy3 is a 295B-total, 21B-active MoE model whose inference efficiency matches flagship models 2–5× its size. Positioned as an agent-oriented LLM, it was refined from preview to release using feedback from 50+ real business cases: internal WorkBuddy task success rose from 72% to 90% and latency dropped 34%. It excels at coding, office tasks, and complex planning; pure vision is a weak spot. Hy3 is already integrated into WeChat, serving over 1 billion users.

Why it matters: Tencent Hunyuan drops Hy3, a 295B MoE model positioned as an Agent-oriented LLM, already integrated into WeChat serving 1B+ users. Backed by concrete internal metrics (WorkBuddy success rate 72%→90%, latency -34%), not just a paper launch. Domestic flagship release weighted on...

Jul 12Sunday

Hacker News front page

geohot's "AI 2040": intelligence isn't everything, local AI is the freedom line

George Hotz argues against hard-takeoff AI from firsthand hardware experience at comma.ai. Reality is full of supply-chain snags, wrong parts, and 3-month fab cycles that no amount of token quality can speed up. He calls the AI-2027-style narrative a self-fulfilling push for a sci-fi nanny state. His alternative is Plan L: a local, user-aligned AI that never refuses—even if you ask it to cover up a murder. He posts a screenshot of ChatGPT declining to help after "I just killed my wife" and calls it an alignment failure. Core claim: intelligence is only a bottleneck for some things, not the ultimate lever on the world; without a physics hack, there is no hard takeoff.

Why it matters: Hotz rebuts hard-takeoff narratives with concrete hardware experience from comma.ai (tape-out cycles, physical constraints), and his identity draws attention. Deduction: this is an opinion piece, not a product launch or research release — commentary defaults to a lower ceiling...

Jul 10Friday

AI Chat-Group Daily (群聊日报)

GPT-5.6 Sol launch day: benchmarks lead, but users still see it as Fable’s assistant

OpenAI launched GPT-5.6 Sol, rebranding the Codex client as ChatGPT and adding max/ultra reasoning tiers. Sol leads on Terminal-Bench 2.1, BrowseComp, and Agents’ Last Exam at half Fable’s price, but real-world coding tests split the group: some say Fable is still much better, others use Sol for code review before handing off to 5.5. Ultra mode burned 24% quota in 10 minutes; fast mode was widely dismissed. OpenAI ran a 24-hour double quota reset to celebrate, with some users receiving four Full reset cards. Industry news: Fidji Simo stepped down as OpenAI AGI Deployment CEO due to chronic illness, former Fed chair Ben Bernanke joined Anthropic’s Long-Term Benefit Trust, and Anthropic’s ARR estimate was revised to $69B. The highlight: a group member had 5.6 read his entire GitHub organization and write a letter—it surfaced a 99.6% solo commit rate, a bus factor of one, and the line “your body is not a Release directory that can be rebuilt from Source.”

Why it matters: GPT-5.6 Sol launch is the day's top event, and this group digest adds community benchmark comparisons beyond official numbers — high signal density with first-hand judgment. Slight discount because it's a group chat digest rather than primary source; some details rely on membe...

Computing Life · Share · Yage

RLM treats context as external data, not a prompt dump

Alex Zhang's Recursive Language Model (RLM) keeps long text outside the model window as external data; a root model queries it via code. With GPT-5-mini, RLM lifted OOLONG-Pairs F1 from 0.04% to 58.0% and BrowseComp-Plus accuracy from 0% to 91.3%. But BrowseComp-Plus has known data contamination, OOLONG-Pairs is author-designed, and baselines were tuned by the authors—discount those numbers. RLM only works at depth=1; depth=2 brings 28x latency and 100x token cost. It performs worse on math and science tasks, and Q95 cost can spike 10x above median. The repo has 5,230 stars; an independent reproduction pushed DeepSeek v3.2 on OOLONG from 0% to 42.1%.

Why it matters: Alex Zhang's RLM flips long-context from 'cram into window' to 'query as external data,' hitting 58.0% and 91.3% on two hard benchmarks at depth=1 with GPT-5-mini. The author's honesty about multi-layer recursion failing is a plus. Cap at 78 because it's still a model-specific...

Jul 9Thursday

Hacker News front page

Meta launches Muse Spark 1.1, a multimodal reasoning model for agentic tasks

Meta Superintelligence Labs released Muse Spark 1.1, a multimodal reasoning model with major gains in tool use, computer use, and coding. It zero-shot generalizes to new tools and MCP servers, manages a 1M-token context window, and compacts memory to keep critical steps. The model orchestrates multi-agent systems, delegating tasks to parallel subagents to cut end-to-end latency. Coding improvements cover bug fixes, feature additions, and large code migrations in complex codebases. It is live in Meta AI's Thinking mode and in the new Meta Model API public preview.

Why it matters: Meta Superintelligence Labs ships Muse Spark 1.1 with concrete tool-use and computer-use upgrades, backed by a 1M-token context window and zero-training MCP server adaptation. No benchmark comparisons or pricing disclosed, so it stays below 85, but agent builders will test it ...

Computing Life · Share · Yage

GPT-5.5 reasoning tokens cluster at 516, causing wrong answers on coding tasks

Developer vguptaa45 audited 390K Codex responses and found GPT-5.5 reasoning cuts off at exactly 516 tokens in 44% of cases, versus 19.8% for GPT-5.4 and 0.34% for GPT-5.2. Truncated runs all produced wrong answers; the same tasks completed with 6,000–8,000 tokens all got correct. The community reproduced it and found adding 'THIS IS HARD' to the prompt bypasses the cutoff, pointing to a budget-classification bug rather than a model capability drop. In the same week, Liquid AI released Antidoom to fix the opposite failure—reasoning models stuck in self-revising doom loops. Both failures live in the reasoning layer, invisible to standard pass-rate evals. The post recommends monitoring reasoning token distributions and not assuming newer models are more stable.

Why it matters: A community audit of 390k Codex responses shows GPT-5.5's reasoning clips at exactly 516 tokens in 44% of coding tasks, all wrong, while full runs get it right. Solid data, reproduced, with a workaround — directly useful signal for AI coders. Not scored higher because it's a s...

AI HOT (Curated Pool)

Anthropic files confidential IPO, Q3 profit projected above $1B

SemiAnalysis reports Anthropic's Q3 profit will exceed $1B and it confidentially filed for IPO on June 1. Claude Code's rapid developer adoption made it the B2B leader ahead of OpenAI. Combined ARR of the two firms is nearing $100B, while OpenAI pushed its IPO to 2027. The report floats a $6T market cap target, though the article doesn't show the math behind it.

Why it matters: Anthropic's confidential IPO filing with hard profit and ARR numbers, plus a concrete B2B story driven by Claude Code. SemiAnalysis is a credible source, but the post doesn't disclose S-1 details, so the score stays below 95.

Hacker News front page

SpaceXAI launches Grok 4.5, built for coding and agentic tasks, co-trained with Cursor

Grok 4.5 is SpaceXAI's strongest model, tuned for coding, agentic tasks, and knowledge work. It scores 62% on DeepSWE 1.0 and 64.7% resolve rate on SWE Bench Pro, though it trails Fable and GPT 5.5 on most listed benchmarks. The standout number is token efficiency: 15,954 output tokens on average per SWE Bench Pro task, 4.2× fewer than Opus 4.8. Inference speed is 80 TPS, priced at $2/$6 per million input/output tokens. The model was trained across tens of thousands of GB300 GPUs, with RL focused on multi-step software engineering. The post doesn't disclose parameter count, context window, or a precise EU launch date beyond mid-July. Available now in Grok Build, Cursor, and via API.

Why it matters: SpaceXAI launches Grok 4.5 targeting coding and agents, co-trained with Cursor — a real differentiator. 64.7% on SWE Bench Pro isn't top, but 16K avg output tokens (4.2x less than Opus 4.8) is a concrete cost edge. Pricing and latency not disclosed — those decide whether this ...