Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

881–900 of 1,465

May 20Wednesday

r/LocalLLaMA

Public repository Codegraph claims 94% fewer Claude, Cursor, Codex, and OpenCode tool calls locally

Codegraph uses a pre-indexed knowledge graph for symbol relationships, call graphs, and code structure. In the VS Code test, it reduced tool calls from 52 to 3 and runtime from 1m37s to 17s.

Why it matters: All HKR axes pass, but evidence is a Reddit/public-repo self-test without independent replication. The 94% reduction and 52→3 call count clear featured, not p1.

AI HOT (Curated Pool)

Gemini 3.5 launches with frontier intelligence and real-world action

Google DeepMind introduced Gemini 3.5 with 3.5 Flash as the first release; the post says it is the company’s strongest model for agents and coding, but does not disclose parameters, pricing, or context window.

Why it matters: HKR-H/K/R all pass because Google DeepMind shipped a major Gemini 3.5 update with 3.5 Flash first and agent/coding claims. Missing price, parameters, and context window keep it mid-85–94, not higher.

r/LocalLLaMA

Cursor and Claude Code Are Not Getting Dumber; Agent Loops Are Suffocating Context

A Reddit user says an API-log audit showed Cursor and Claude Code recursively grep about 40 files in 10k-plus-line repositories, sometimes load 2k-line files for 5-line edits, and spend roughly 30k tokens on tool definitions and logs before generating code.

Why it matters: HKR-H/K/R all pass: the hook is contrarian, the API-log numbers are concrete, and coding-agent context waste is a live practitioner pain. Reddit single-post sourcing and no shared logs keep it at the featured threshold.

May 19Tuesday

AI HOT (Curated Pool)

OpenRouter Tool-Calling Models Can Now Run Web Search Autonomously

OpenRouter now lets any tool-calling model on its platform autonomously invoke web search and webpage scraping, with the model deciding when to search, what to query, and how many searches to run; OpenRouter also added @p0 as a web search provider.

Why it matters: HKR-H/K/R pass: OpenRouter lets tool-calling models decide search timing, queries, and frequency. The source is tweet-thin and lacks pricing, limits, or evals, so it lands near the featured threshold.

AI HOT (Curated Pool)

Claude Managed Agents add two safety features

Claude Managed Agents added two safety improvements: self-hosted sandboxes keep agent execution environments in the user’s infrastructure or hosted sandbox provider, while MCP tunnels let agents connect to services inside the user’s security boundary.

Why it matters: HKR-K and HKR-R pass: the post names two agent-safety mechanisms and a concrete execution-boundary change. HKR-H is weak, and this is not a model release, so it sits in the low featured band.

AI HOT (Curated Pool)

Membrane launches single-skill API integration for AI agents

Membrane launched a universal skill that lets Claude Code, ChatGPT, and Cursor call more than 100,000 APIs with one instruction, covering services from Stripe payments to NASA Mars rover data.

Why it matters: HKR-H/K/R pass: one skill for 100K+ APIs is a strong agent-tooling hook. Source is a social post summary with no pricing, auth model, safety boundary, or live case, so this stays in the mid-weight product-update band.

Hacker News front page

Show HN: Forge takes an 8B model from 53% to 99% on agentic tasks

Forge adds five guardrail layers to self-hosted LLM tool calling, raising Ministral 8B to 99.3% across 18 multi-step agentic scenarios, with the accepted ACM CAIS ’26 paper covering 97 model/backend configurations and 50 runs per scenario.

Why it matters: HKR-H/K/R all pass: the 53%→99.3% jump is clickable, the test setup has concrete numbers, and self-hosted agent reliability is a live practitioner pain. Single-source Show HN/GitHub evidence keeps it in the 78–84 open-source-tool band, not P1.

AI HOT (Curated Pool)

Former executive says Microsoft’s AI strategy faltered, with Copilot paid usage below 3%

Former Microsoft executive Matt Veloso said Microsoft generated about $30 billion from its AI partnership between 2023 and 2025, while related costs reached $100 billion; he also said actual usage among paid Copilot users is below 3%.

Why it matters: HKR-H/K/R all pass: a former executive gives concrete Microsoft AI cost, revenue, and Copilot usage numbers. Kept at 80 because this is a single former-exec claim, not an official Microsoft disclosure.

AI HOT (Curated Pool)

Claude Managed Agents Add Self-Hosted Sandboxes and MCP Tunnels

Anthropic added two updates to the Claude managed agents platform: self-hosted sandboxes are in public beta, and MCP tunnels are in research preview for private network database and API access.

Why it matters: HKR-H/K/R all pass: this is an official Anthropic Claude agent-platform update with two concrete mechanisms. It is below model-release weight, but strong enough for featured agent-infra coverage.

AI HOT (Curated Pool)

Claude launches self-hosted sandboxes and MCP tunnels

Claude launched self-hosted sandboxes in public beta and MCP tunnels in research preview for Claude Managed Agents, letting agents run inside a user’s own security boundary with the user’s security controls applied by default.

Why it matters: HKR-H/K/R all pass: this is an official Claude agent-infra update with concrete self-hosted sandbox and MCP tunnel mechanisms, tied to enterprise security boundaries. It is beta/preview scope, not a model release, so it stays in the 78–84 band.

Latent Space

[AINews] How to Land a Job at a Frontier Lab (on Pretraining)

Latent Space says Vlad Feinberg’s pretraining job-prep notes reduce frontier-lab readiness to kernel-level performance work: derive Chinchilla laws, compare dense and MoE architectures, code the solution in JAX, then write a Pallas kernel that beats jax.lax.ragged_dot for F > D by fusing up/down projections.

Why it matters: HKR-H/K/R all pass: the career hook is strong and the prep list is concrete. It is not a model release or major product update, and the kernel-heavy angle keeps it at the lower featured band.

QbitAI · WeChat

World model supports multiplayer FPS gameplay before Fei-Fei Li

Odyssey released Agora-1, a world model that supports up to four human and AI players fighting in the same generated FPS world in real time. The system decouples simulation from rendering and trains on GoldenEye internal game states.

Why it matters: HKR-H/K/R all pass: Agora-1 moves world models from solo demos to up to 4-player real-time FPS, with decoupled simulation/rendering and training-data clues. The lab is not a top-tier foundation-model vendor, so this stays in the 78–84 band.

Xinzhiyuan · WeChat

CUHK and Zhejiang University Question Whether AI Agent Memory Is Just a Memo

CUHK and Zhejiang University researchers argue that mainstream Agent memory is retrieval-based memo storage, not true memory, citing an Ω(k²) case requirement for compositional tasks and a PoisonedRAG result where 5 adversarial texts reached a 90% attack success rate.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the summary gives Ω(k²) and 90% attack success, and the issue matters to agent-memory and RAG-security builders. Strong research signal, not a same-day model-release event.

Xinzhiyuan · WeChat

AI startup annualized revenue hits $80B, with OpenAI and Anthropic taking 89%

The Information says 34 leading AI startups reached about $80 billion in annualized revenue, with OpenAI and Anthropic taking 89%, while Anthropic exceeded $30 billion in April 2026 and surpassed OpenAI’s reported $25 billion.

Why it matters: HKR-H/K/R all pass: the story has a sharp Anthropic-vs-OpenAI hook, concrete revenue-concentration numbers, and startup-economics resonance. It is secondary financial reporting, not a model release, so it stays in 78–84.

Synced · WeChat

Recent LLM Architecture Changes: From Gemma 4 to DeepSeek V4

Jiqizhixin translated Sebastian Raschka’s blog on recent LLM architecture changes, covering long-context cost reductions in Gemma 4, Laguna XS.2, and ZAYA1-8B; the article states that Gemma 4 E2B saves about 2.7GB of KV cache at 128K context with bfloat16 precision.

Why it matters: HKR-H/K/R pass: notable model names, a concrete 128K bf16 KV-cache saving, and inference-cost relevance. As a translated survey rather than a release, it stays in the 72–77 featured band.

Synced · WeChat

From Selling Tokens to Selling Outcomes: AI Companies Start Taking KPI Risk

Sierra raised $950 million in May at a valuation above $15 billion, while Lingxi says it reached scaled profitability and positive cash flow in 2025; the article uses both companies to frame RaaS as charging for measurable business outcomes rather than tokens or subscriptions.

Why it matters: HKR-H/K/R all pass: the KPI hook is clickable, Sierra’s $950M raise and RaaS pricing add concrete facts, and the angle hits agent monetization. This is strong business-model signal, not a model-release-level event.

AI HOT (Curated Pool)

Cursor releases Composer 2.5, calling it its strongest model yet

Cursor released Composer 2.5, claiming a 10x efficiency gain at comparable capability, with larger training scale, more complex reinforcement-learning environments, and a text-feedback mechanism.

Why it matters: Cursor Composer 2.5 is a substantive model update for a front-line AI coding tool, with HKR-H/K/R from the 10x efficiency and RL-training details. The single social-source summary lacks benchmarks, pricing, and reproducible tests, keeping it in the 78–84 band.

AI HOT (Curated Pool)

First real-time multi-agent world model released, humans interact with AI on the same screen

Odyssey Labs released Agora-1, described as the first real-time multi-agent world model, using a GoldenEye deathmatch demo where multiple humans and AI agents interact in the same simulated world; the post says a playable research preview is available now, but does not disclose model architecture or latency figures.

Why it matters: HKR-H/K/R all pass: Agora-1 combines multi-agent world modeling with a live human-AI preview. Sparse details on architecture, latency, cost, and benchmarks keep it in the 78–84 band.

Hacker News front page

Mexican Government Breached by Solo User with Claude, 150 GB Exfiltrated

The title says a solo user used Claude to breach the Mexican government and exfiltrate 150 GB of data; the RSS body does not disclose the attack mechanism, timeline, affected systems, or confirmation source.

Why it matters: HKR-H/K/R all pass: a solo Claude-assisted government breach with 150 GB allegedly exfiltrated is a strong security story. Source details are thin—no attack path, timeline, or affected systems—so it stays below P1.

Bloomberg Technology

Self-Improving AI Startup Recursive AI Valued at $4.65B

Recursive came out of stealth at a $4.65 billion valuation, building AI that runs experiments on safe self-improvement, with backers including Google Ventures, Greycroft, Nvidia, and AMD Ventures.

Why it matters: HKR-H/K/R all pass: Bloomberg gives a $4.65B valuation and named backers, with a self-improving AI safety angle. No model capability, experiment result, or product path is disclosed, so it stays below 85.