Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

801–820 of 1,465

May 24Sunday

Computing Life · Share · Yage

You May Have Coded for 10 Years, but You Are Still a Beginner with AI

The article discusses the debate sparked by Armin Ronacher using Pi to develop Pi, citing issue tracker data to argue that experienced programmers can still be misled by confident but wrong AI outputs.

Why it matters: HKR-H/K/R all pass, but this is commentary around the Armin Ronacher debate, not a model or product launch. The issue-tracker evidence lifts it to the featured threshold.

r/LocalLLaMA

llama.cpp server has built-in native tools: exec_shell, edit_file, and more

llama.cpp server exposes an experimental --tools flag with 8 native tools, including file reads, grep search, shell execution, file edits, diffs, and datetime; the post says file operations are relative to the server launch directory and no command whitelist or strict sandbox is provided yet.

Why it matters: HKR-H/K/R all pass: llama.cpp adding native shell and file tools is a concrete agent-runtime shift with safety stakes. Reddit sourcing and experimental status keep it in the lower featured band.

AI HOT (Curated Pool)

StepAudio 2.5 Realtime Voice Released with Paralinguistic Awareness and Persona Interaction

StepFun released StepAudio 2.5 Realtime with Chinese and English real-time voice support, API-based custom personas, more than 10,000 native persona options, millions of composable traits, and 5 built-in preset personas.

Why it matters: HKR-H/K/R all pass, but the source is an official X post and lacks latency, pricing, benchmarks, and rollout scope. This fits the low featured band for a mid-weight product update.

May 23Saturday

AI HOT (Curated Pool)

Microsoft Says AI Use Can Cost More Than Human Wages

Microsoft says AI use costs more than human wages in specific work scenarios, with its report comparing token- and agent-based usage costs against the cost of hiring people for the same tasks.

Why it matters: HKR-H/K/R all pass, but the disclosed facts stop at a broad Microsoft cost claim; jobs, amounts, and methodology are not given. Strong featured cost signal, not a major release.

Computing Life · Share · Yage

AI Is Splitting Into Two Markets: Which Side Do You Choose?

Token prices fall 10x per year, but enterprise AI bills keep expanding; the post says Chinese open-source models push the low-cost tier toward zero, while enterprise lock-in and agent workloads raise the premium tier, creating a 300x price gap.

Why it matters: HKR-H/K/R all pass: the hook is the pricing paradox, the facts include 10x annual drops and a 300x spread, and the nerve is cost plus lock-in. It remains a single commentary piece, not a release or first-person test.

AI HOT (Curated Pool)

Gemini update: over 900 million users and new agent features

Google announced that the Gemini app has surpassed 900 million monthly active users and introduced two agent features: Daily Brief for personalized daily summaries and Gemini Spark, a 24/7 personal agent that manages tasks under user authorization.

Why it matters: HKR-H/K/R all pass: Google gives a 900M MAU number and two agent features for Gemini. This is an entry-point product update with competitive weight, not a routine small feature.

AI HOT (Curated Pool)

v2.1.149 release summary

Claude Code v2.1.149 adds categorized /usage reporting, an enterprise allowAllClaudeAiMcps setting for cloud MCP connectors, and fixes three security issues involving PowerShell permission bypass, Git worktree sandbox allowlist overflow, and otelHeadersHelper failures when script paths contain spaces.

Why it matters: Official Claude Code point release with concrete changes but limited blast radius: /usage categories, an enterprise MCP allow switch, and PowerShell bypass fixes hit developer security and governance needs.

AI HOT (Curated Pool)

Claude Auto Mode Adds Pro Plan and Model Support

Claude Auto Mode is now available on the Pro plan and supports Sonnet 4.6 and Opus 4.7; users can start it with Shift+Tab, while the post does not disclose pricing changes or rollout scope.

Why it matters: HKR-H/K/R all pass: official Claude dev channel gives Pro access, two supported models, and a shortcut. This is a mid-weight Claude product update, not a major model or capability release.

AI HOT (Curated Pool)

Project Glasswing: Initial Update

Anthropic says Project Glasswing used Claude Mythos Preview with about 50 partners to find more than 10,000 high or critical vulnerabilities in global critical systems, with independently verified accuracy of 90.6%.

Why it matters: HKR-H/K/R all pass: Anthropic gives concrete numbers—~50 partners, 10,000+ high/critical bugs, 90.6% validation—and the story hits AI-agent security automation and critical-system risk.

AI HOT (Curated Pool)

Project Glasswing Collaborative AI Cybersecurity Project Reports Results

Anthropic says Project Glasswing and its partners found more than 10,000 high or critical vulnerabilities in key software since the initiative launched last month; the post does not disclose the vulnerability list, reproduction conditions, or remediation status.

Why it matters: HKR-H/K/R all pass: Anthropic ties AI security work to 10,000+ severe flaws. Missing vulnerability lists, reproduction details, and fix status keep it in the featured-threshold band, not p1.

r/LocalLLaMA

How small can the orchestration model in an agent be? Separating it from code generation

HomoAgens1 runs a local ReAct orchestration loop on Qwen3.6-35B-A3B, with about 3B active parameters, a 12GB GPU, 30 expert offload, and 40 tokens/s prompt generation; smaller dense models fail first on tool-call discipline, inventing arguments or repeating bad calls, while reasoning is not identified as the first break point.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit experiment rather than a formal release. The VRAM, speed, and failure-mode details put it at the 72 featured threshold.

AI HOT (Curated Pool)

Kakuna: An AI Agent Tool for Automated Codebase Hardening

Kakuna hardens prototype codebases with built-in checklists and a plan-goal workflow; one roughly 16-hour run can generate hundreds of commits while preserving functionality.

Why it matters: HKR-H/K/R all pass: the post has a 16-hour run, hundreds of commits, and a workflow mechanism tied to coding-agent pain. Single X source and a non-major vendor keep it at the featured threshold.

AI HOT (Curated Pool)

Google I/O Releases AI Agent Development Toolchain

Google announced an AI agent development and deployment toolchain at I/O, including Antigravity 2.0, managed agent services in the Gemini API, WebMCP in Chrome 149, and Chrome DevTools access for automated agent debugging.

Why it matters: HKR-H/K/R all pass: Google is shipping a named agent stack across tooling, managed services, WebMCP, and Chrome. Single-source social summary lacks pricing, API details, and demos, so it stays in the 78–84 band.

AI HOT (Curated Pool)

Agent Workloads Quietly Reshape Inference Economics

SemiAnalysis analyzed 432,000 real coding-agent requests and found a median input length of 96,000 tokens, not 32,000 or 64,000. The post does not disclose the model mix, cost curve, sampling method, or time window.

Why it matters: HKR-H/K/R all pass: SemiAnalysis adds a 432k coding-agent request dataset and 96k-token median input. Missing models, cost curves, and sampling keep it in the strong-data-point band, not must-write.

May 22Friday

Hacker News front page

Launch HN: Superset (YC P26) – IDE for the agents era

Superset launched an open-source agentic IDE that runs coding agents such as Claude Code, Codex, and OpenCode in parallel through git worktrees, and the team added Remote Workspaces in beta for running agents on remote machines while managing work from the desktop app.

Why it matters: HKR-H/K/R all pass, but Superset is still a new YC launch and the post lacks usage, pricing, or performance data. The git-worktree agent workflow clears the featured bar, not the must-write band.

Mistral AI

Mistral launches Connectors in Studio with built-in and custom MCP

Mistral launched Connectors in Studio. All built-in connectors and custom MCP are now callable through the API/SDK by every model and agent. New features include direct tool calling, human-in-the-loop approval flows, and programmatic access to create, modify, list and delete connectors.

Why it matters: The original gives the API usage and code examples for Connectors, enough to judge how enterprise MCP integration gets built.

AI HOT (Curated Pool)

Alibaba Qianwen App, PC, and Web Add Qwen3.7-Max

Alibaba added Qwen3.7-Max to the Qianwen app, PC client, and web client, with free access after updating the app to version 6.9.7 or later, and the official test reports a 35-hour autonomous kernel optimization run with more than 1,000 tool calls.

Why it matters: HKR-H/K/R all pass: Alibaba ships Qwen3.7-Max across three Qianwen clients, with v6.9.7+ free access and a 35-hour, 1,000+ tool-call claim. Benchmarks, context window, and API pricing are not disclosed, so it stays below 90.

MIT Technology Review · AI

Google I/O showed how the path for AI-driven science is shifting

MIT Technology Review says Google used I/O to shift its scientific AI framing toward Gemini for Science, a package that groups AI Co-Scientist and AlphaEvolve, while researchers can now apply for access and older specialized systems like AlphaFold and WeatherNext remain active.

Why it matters: HKR-H and HKR-K pass: MIT Technology Review frames a real Google science-AI product shift with named components and access conditions. HKR-R is weak because the impact is mostly research-facing, not practitioner-wide.

Latent Space

[AINews] New AI Infra Unicorns: Exa, Modal, TurboPuffer

Latent Space summarized AI News for May 20-21, 2026, confirming TurboPuffer reached $100 million ARR and profitability, Exa raised a $250 million Series C at a $2.2 billion valuation, and Modal raised a $355 million Series C at a $4.7 billion valuation.

Why it matters: HKR-H/K/R all pass because the roundup gives concrete AI-infra funding and ARR numbers. It stays below 78 because it is market aggregation, not a new model, product capability, or technical release.

Xinzhiyuan · WeChat

OpenClaw Case: Routine Chats Can Poison an Agent’s Long-Term Memory

Researchers from The Hong Kong Polytechnic University and HKUST (Guangzhou) introduced ULSPB with 350 settings; routine conversations can poison an agent’s long-term state without malicious prompts, while StateGuard audits state diffs before persistence and reduces Harm Score to near zero in Targeted-Ensemble settings.

Why it matters: HKR-H/K/R all pass: the story has a non-malicious agent corruption hook and a concrete ULSPB benchmark with 350 settings. It is useful agent-safety research, not a top-lab product release.