Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

1221–1240 of 1,465

Apr 24Friday

Latent Space

GPT 5.5 and OpenAI Codex Superapp

OpenAI launched GPT-5.5 for ChatGPT and Codex, while API access is delayed for safeguards. The post cites 82.7% Terminal-Bench 2.0, 58.6% SWE-Bench Pro, and a 1M API context window. The sharper signal is Codex: browser control and Prism integration point to a desktop superapp strategy.

Why it matters: All HKR axes pass: GPT-5.5 is a major OpenAI model update with benchmark numbers and API conditions. Codex plus browser control and Prism raises the coding-agent stakes; this fits the Claude 4.7-level 85–94 band.

X · @dotey

DeepSeek releases and open-sources V4 preview; 1M context is standard across all services

DeepSeek released and open-sourced the V4 preview, making 1M context standard across all official services with no tier or price split. The post says V4-Pro and V4-Flash use token compression plus DSA sparse attention to cut compute and memory costs for 1M context; legacy APIs remain for 3 months and stop after July 24.

Why it matters: DeepSeek is a flagship Chinese model vendor, and this V4 preview is a substantive release with open source and 1M context made standard across official services. HKR-H/K/R all pass: the post includes mechanisms and a migration deadline, and the tier reset makes it a same-day P1.

X · @op7418

DeepSeek V4 detailed official announcement is out

DeepSeek says V4 Pro has 1.6T total parameters with 49B active, while Flash has 284B total and 13B active; both were pretrained on 32T tokens. Web and app Expert mode map to Pro, and Fast mode maps to Flash. The post also says several benchmarks are on par with Opus 4.6, with stronger agent ability and world knowledge, plus a new attention mechanism that reduces compute and memory demand.

Why it matters: This is a flagship DeepSeek release, scored on par with peer US lab model launches. HKR-H/K/R all pass on concrete scale numbers, 32T data, and an inference-efficiency mechanism; benchmark setup, pricing, and API availability are not disclosed in the summary.

Computing Life · Share · Yage

Skills Are Products With Built-in Suicide Genes

The author argues Anthropic Skills cannot stand alone as paid products, citing direct sales, hosting, and API funneling as 3 dead ends. The post cites PromptBase at about $5M annual revenue, Stripe’s 2.9% plus 30 cents fee, and Snyk finding 13.4% of skills with critical issues. The sharper point is charging for relationships, time-sensitive access, physical accountability, and judgment.

Why it matters: HKR-H/K/R all pass: the hook is sharp, and the post tests three business paths with named examples. It is strong commentary, not a new Anthropic release, so it lands at the featured threshold rather than 78+.

Hugging Face Blog

DeepSeek-V4: a million-token context that agents can actually use

DeepSeek released V4 with two MoE checkpoints, Pro and Flash, both supporting a 1M-token context. Pro has 1.6T total and 49B active parameters; Flash has 284B total and 13B active. The key detail is KV cost: Pro uses 27% of V3.2 single-token FLOPs and 10% of its KV cache; Flash uses 10% and 7%.

Why it matters: DeepSeek-V4 is a flagship Chinese model release with 1M-token context and KV cache at 7%–10% of V3.2. HKR-H/K/R all pass, placing it in the 85–94 same-day band.

Ruan YiFeng's Weblog

Tech Weekly Issue 394: The Second Wave of API Opening

Ruanyifeng’s Weekly Issue 394 argues that production-ready LLMs in H2 2025 triggered a second API-opening wave. The post says agents need platform APIs to act, citing Tencent opening WeChat interfaces after OpenClaw and adoption of MCP and Skills. The key shift is consumer services exposing actions, not only cloud APIs.

Why it matters: HKR-H/K/R all pass: the historical API-wave frame is clickable, and the post gives mechanisms around agent action APIs, MCP/Skills, and WeChat access. This is strong commentary, not a model or major product release, so it stays in the 72–77 band.

The Verge · AI

Claude is connecting directly to personal apps like Spotify, Uber Eats, and TurboTax

Anthropic added personal app connectors to Claude, covering services such as Spotify, Uber, AllTrails, Instacart, and TurboTax. After connection, Claude can suggest relevant apps inside chats, such as using AllTrails for hike recommendations; the post does not disclose launch count, regions, or plan access. The key shift is Claude moving from work apps into personal consumer workflows.

Why it matters: This gets Anthropic’s positive signal: a substantive product update, but not a model release. HKR-H/K/R all pass because personal-app connectors are a strong hook, the story confirms in-chat app invocation, and it hits the fight for assistant entry points; missing pricing, region

X · @dotey

Anthropic launches memory for Claude Managed Agents in public beta

Anthropic has launched memory for Claude Managed Agents in public beta, letting agents retain and reuse experience across sessions. Memory is stored as files on a filesystem, with shared permissions, concurrent access, audit logs, and rollback; Rakuten reports a 97% drop in first-time errors, and Wisedocs reports 30% faster document validation. The key detail is the implementation path: it uses a filesystem, not a dedicated vector database.

Why it matters: Anthropic adds cross-session memory to Claude Managed Agents beta and discloses the implementation plus two user numbers: Rakuten 97% and Wisedocs 30%. HKR-H/K/R all pass, but the scope is still limited to the managed-agent beta, so this lands at 83 and featured.

X · @claudeai

Memory on Claude Managed Agents is now in public beta

Claude has put Memory for Managed Agents into public beta, and agents can now learn from every session. The post only says it uses an intelligence-optimized memory layer balancing performance and flexibility; it does not disclose capacity, retention, pricing, or access conditions. What matters for practitioners is when persistent memory becomes default and how it changes agent evals and state management.

Why it matters: Memory on Claude Managed Agents is a substantive Anthropic product update with clear practitioner resonance, so HKR-H and HKR-R pass. HKR-K is weak because the post omits capacity, retention, pricing, and default-on conditions, keeping it in low featured rather than p1.

Bloomberg Technology

An AI Agent Takes Over a Store and Orders Too Many Candles

Andon Market in San Francisco’s Cow Hollow put store operations under an AI agent named Luna, which handles assortment and pricing, and the headline says it over-ordered candles. The RSS snippet only confirms Luna acts like a CEO; the post does not disclose the candle quantity, failure mechanism, financial impact, or remediation. The real signal is that a retail operating loop was delegated to an agent.

Why it matters: Bloomberg reports a real store delegating assortment and pricing to an AI agent, turning agent risk into a concrete incident. HKR-H and HKR-R pass, but HKR-K is limited because quantity, loss, trigger, and rollback are undisclosed, so this sits at the low end of featured.

X · @dotey

Codex now supports GPT-5.5 and adds five capability upgrades

Codex now supports GPT-5.5 and adds 5 upgrades aimed at moving it from a coding tool to an agent that can execute longer tasks. The RSS snippet says it can control browsers and computers, create files in Microsoft Office and Google Drive, and use gpt-image-2; an auto-review mode invokes a separate review agent for high-risk actions. What matters is longer task chains, but the post does not disclose pricing, rollout scope, or safety thresholds.

Why it matters: This is a substantive Codex product update: the main signal is the shift toward an agent that can execute chained tasks, not just a new model toggle. HKR-H/K/R all pass, but the item is second-hand and omits pricing, rollout scope, and safety thresholds, so it lands as featured,

X · @claudeai

Claude can now connect to more apps outside work, including Tripadvisor, Booking.com, and Resy

Claude added at least 10 consumer app connections, including Tripadvisor, Booking.com, Resy, Instacart, Spotify, Audible, AllTrails, Thumbtack, and TurboTax. The RSS snippet confirms only a product update; the post does not disclose integration method, supported actions, regions, permission scope, or rollout timing. The key question is whether Claude can act in these apps directly, not just list them.

Why it matters: Official Anthropic product update with clear HKR-H/K/R: consumer app connectors expand Claude beyond workplace tools and widen its assistant surface. The score stays at 75 because the post lists apps only; actions, permissions, regions, and rollout details are not disclosed.

Hacker News front page

GPT-5.5: Mythos-Like Hacking, Open to All

XBOW says GPT-5.5 cut miss rate to 10% on its real-vulnerability benchmark, versus 40% for GPT-5 and 18% for Opus 4.6. It scored 97.5% on visual acuity and used about half the login iterations of the next-best model. The key point is black-box testing: GPT-5.5 without source beat GPT-5 with source.

Why it matters: HKR-H/K/R all pass: a major OpenAI model claim, concrete security benchmark numbers, and a clear practitioner safety nerve. The source is XBOW rather than an OpenAI launch post, so it stays below 95.

X · @OpenAI

Introducing GPT-5.5

OpenAI introduced GPT-5.5, and it is now available in ChatGPT and Codex. The RSS snippet says it targets real work and agents, can understand complex goals, use tools, check its work, and carry more tasks to completion; the post does not disclose parameters, pricing, context window, or benchmark results. What matters is the execution loop, not the headline's “new class of intelligence.”

Why it matters: OpenAI launching GPT-5.5 in ChatGPT and Codex is same-day mandatory coverage. HKR-H/K/R all pass: new model release, concrete agent-workflow claims, and direct impact on daily AI work. Price, context window, params, and benchmarks are undisclosed, so it stays below 95.

The Verge · AI

OpenAI says its new GPT-5.5 model is more efficient and better at coding

OpenAI announced GPT-5.5 and says it is more efficient and stronger at coding than GPT-5.4, which shipped last month. The RSS snippet says it handles coding, debugging, online research, and cross-tool work on spreadsheets and documents; the post does not disclose pricing, context window, or benchmark scores.

Why it matters: An OpenAI model release is same-day coverage, and the angle ties efficiency, coding, and tool use into one clear upgrade, so HKR-H/K/R all pass. The post does not disclose price, context window, or benchmark scores, which keeps it in the high 80s instead of 90+.

Hacker News front page

An update on recent Claude Code quality reports

Anthropic said three product-layer changes degraded Claude Code quality for Sonnet 4.6, Opus 4.6, and Opus 4.7, while the API was unaffected; all were fixed on April 20 in v2.1.116. The changes were lowering default reasoning effort on March 4, a March 26 bug that cleared prior thinking every turn after sessions sat idle for over an hour, and an April 16 prompt tweak to reduce verbosity that hurt coding quality. The signal for practitioners is sharp: product and prompt changes can degrade code performance even when model and inference evals do not reproduce it early.

Apr 23Thursday

The Verge · AI

You’re about to feel the AI money squeeze

Anthropic sharply restricted OpenClaw’s access to Claude this month and pushed heavy third-party agent users toward pricier paid plans. The RSS snippet says system strain and profit pressure drove the move, and Boris Cherny said existing subscriptions do not fit this usage pattern; the post does not disclose pricing, limits, or rollout scope. Watch the monetization shift: agent-style usage is being carved out of flat subscriptions.

Why it matters: Anthropic is turning heavy Claude agent usage into a pricing and access story, which directly affects tool builders and power users. HKR-H/K/R all land, but missing price, quota, and rollout details keep it at the low end of featured.

QbitAI · WeChat

Qwen3.6-27B open-weights, beats its 397B flagship predecessor on agentic coding

Qwen released Qwen3.6-27B and says it beats Qwen3.5-397B on 4 agentic coding benchmarks with about 1/15 the parameters. The post cites SkillsBench rising from 30.0 to 48.2, GPQA Diamond at 87.8, and AIME26 at 94.1; it uses a dense architecture, Thinking Preservation, and Gated DeltaNet, with weights on Hugging Face and ModelScope.

Why it matters: This is a substantive Qwen open-source model release with concrete agent-coding and reasoning scores, so HKR-H/K/R all pass. I keep it at 84, not higher, because the post gives strong benchmarks but no pricing, context window, or independent reproduction yet.

The Verge · AI

Microsoft launches 'vibe working' in Word, Excel, and PowerPoint

Microsoft is rolling out Agent Mode in Word, Excel, and PowerPoint this week, extending Copilot from a Q&A assistant to an agent that can act directly on the document canvas. Sumit Chauhan said earlier foundation models were not strong enough for app control; the post does not disclose rollout scope, pricing, or exact actions.

Why it matters: Microsoft moving Agent Mode into Word, Excel, and PowerPoint clears HKR-H/K/R: the hook is strong, the mechanism is new, and the Office install base makes it resonate. But rollout scope, pricing, and the exact action list are undisclosed, so it stays below the 85+ band.

Xinzhiyuan · WeChat

Historic moment: Anthropic nears $1 trillion on private secondary markets, surpassing OpenAI for the first time

Anthropic was quoted at $1.05T-$1.15T on private secondary markets, above OpenAI’s roughly $880B quotes on similar platforms. The post attributes the rerating to scarce float, a sharp rise from a $380B funding valuation three months earlier, and momentum around Claude Code and revenue growth; it does not disclose trade volume, revenue figures, or company confirmation. Do not confuse this with a new funding valuation: these are secondary-market quotes on platforms such as Forge Global.

Why it matters: The signal is a private-secondary quote of $1.05T-$1.15T for Anthropic, above OpenAI's quoted ~$880B, not a new financing round. HKR-H/K/R all pass, but missing volume, revenue detail, and company confirmation keep it in the good-quality band, not must-write.