Skip to content

#MCP/工具调用

5 today

May 20Wednesday

TechCrunch · AI

Google’s AI Studio now lets anyone build Android apps in minutes

Google unveiled web-based AI tools in AI Studio that generate native Android apps in minutes. The RSS snippet does not disclose the model, pricing, availability, or supported development constraints.

Why it matters: Google AI Studio’s coding update clears HKR-H/K/R with a strong speed hook and a concrete capability. Model, pricing, and rollout are not disclosed, so it stays at the featured threshold for a mid-weight product update.

TechCrunch · AI

Google launches Antigravity 2.0 with updated desktop app and CLI tool at I/O 2026

Google launched Antigravity 2.0 with an updated desktop app and CLI tool, and introduced a $100 AI Ultra plan that gives users 5x the usage limit of AI Pro; the post does not disclose the desktop app or CLI feature details.

Why it matters: HKR-H/K/R pass, but the post does not disclose concrete desktop or CLI capabilities, so it stays below 78. Google I/O plus the $100 plan and 5x quota clear the featured bar.

TechCrunch · AI

Google introduces Gemini Spark, a 24/7 agentic assistant with Gmail integration, at I/O 2026

Google introduced Gemini Spark at I/O 2026 as a 24/7 agentic personal assistant with Gmail integration; the RSS snippet says it uses Gemini base models and an agentic harness from Google Antigravity, but the post does not disclose pricing, rollout timing, or supported Gmail actions.

Why it matters: HKR-H/K/R all pass: Google used I/O to launch a 24/7 Gmail-linked agentic assistant, a core-entry product update. Price, rollout scope, and safety controls are not disclosed, so it stays at the low end of the 85+ band.

AI HOT (Curated Pool)

I/O 2026: Welcome to the autonomous Gemini era

Google announced at I/O 2026 that Gemini is moving into an autonomous agent phase, with the post saying it can manage email, schedule calendar items, and generate reports automatically, but it does not disclose model parameters, launch timing, or pricing.

Why it matters: HKR-H/K/R all pass: Google frames Gemini as an office agent for email, calendar, and reports. Missing launch timing, price, and model details keeps it in the 78–84 band, below a full major model release.

AI HOT (Curated Pool)

Google AI Ultra plan gets a price cut and a new tier

Google cut the top AI Ultra plan from $250 to $200 per month and added a $100 monthly tier with 5x the Gemini app usage limit of Pro, 20TB of storage, early access to new features, and YouTube Premium under stated terms.

Why it matters: HKR-H/K/R all pass, but this is subscription pricing and quota packaging, not a model or capability launch. Official source and concrete prices put it at the featured threshold.

r/LocalLLaMA

Public repository Codegraph claims 94% fewer Claude, Cursor, Codex, and OpenCode tool calls locally

Codegraph uses a pre-indexed knowledge graph for symbol relationships, call graphs, and code structure. In the VS Code test, it reduced tool calls from 52 to 3 and runtime from 1m37s to 17s.

Why it matters: All HKR axes pass, but evidence is a Reddit/public-repo self-test without independent replication. The 94% reduction and 52→3 call count clear featured, not p1.

r/LocalLLaMA

Floor for local meeting summarization on a 6GB GPU: Qwen3.5 0.8B works in 57s, Granite 4 350M hallucinates

The author tested VoiceFlow 1.6.0 on an RTX 3060 Laptop 6GB, where Qwen3.5 0.8B summarized a 4-minute meeting in 57 seconds with 16K context, while Granite 4 350M returned summaries in 0.6-2.8 seconds but fabricated Binance and Star Trek content.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the test reports hardware/context/timing, and local meeting summarization hits privacy and cost nerves. Single Reddit experiment limits authority, so 73 featured.

r/LocalLLaMA

Cursor and Claude Code Are Not Getting Dumber; Agent Loops Are Suffocating Context

A Reddit user says an API-log audit showed Cursor and Claude Code recursively grep about 40 files in 10k-plus-line repositories, sometimes load 2k-line files for 5-line edits, and spend roughly 30k tokens on tool definitions and logs before generating code.

Why it matters: HKR-H/K/R all pass: the hook is contrarian, the API-log numbers are concrete, and coding-agent context waste is a live practitioner pain. Reddit single-post sourcing and no shared logs keep it at the featured threshold.

May 19Tuesday

AI HOT (Curated Pool)

OpenRouter Tool-Calling Models Can Now Run Web Search Autonomously

OpenRouter now lets any tool-calling model on its platform autonomously invoke web search and webpage scraping, with the model deciding when to search, what to query, and how many searches to run; OpenRouter also added @p0 as a web search provider.

Why it matters: HKR-H/K/R pass: OpenRouter lets tool-calling models decide search timing, queries, and frequency. The source is tweet-thin and lacks pricing, limits, or evals, so it lands near the featured threshold.

AI HOT (Curated Pool)

Claude Managed Agents add two safety features

Claude Managed Agents added two safety improvements: self-hosted sandboxes keep agent execution environments in the user’s infrastructure or hosted sandbox provider, while MCP tunnels let agents connect to services inside the user’s security boundary.

Why it matters: HKR-K and HKR-R pass: the post names two agent-safety mechanisms and a concrete execution-boundary change. HKR-H is weak, and this is not a model release, so it sits in the low featured band.

AI HOT (Curated Pool)

Membrane launches single-skill API integration for AI agents

Membrane launched a universal skill that lets Claude Code, ChatGPT, and Cursor call more than 100,000 APIs with one instruction, covering services from Stripe payments to NASA Mars rover data.

Why it matters: HKR-H/K/R pass: one skill for 100K+ APIs is a strong agent-tooling hook. Source is a social post summary with no pricing, auth model, safety boundary, or live case, so this stays in the mid-weight product-update band.

Hacker News front page

Show HN: Forge takes an 8B model from 53% to 99% on agentic tasks

Forge adds five guardrail layers to self-hosted LLM tool calling, raising Ministral 8B to 99.3% across 18 multi-step agentic scenarios, with the accepted ACM CAIS ’26 paper covering 97 model/backend configurations and 50 runs per scenario.

Why it matters: HKR-H/K/R all pass: the 53%→99.3% jump is clickable, the test setup has concrete numbers, and self-hosted agent reliability is a live practitioner pain. Single-source Show HN/GitHub evidence keeps it in the 78–84 open-source-tool band, not P1.

AI HOT (Curated Pool)

Former executive says Microsoft’s AI strategy faltered, with Copilot paid usage below 3%

Former Microsoft executive Matt Veloso said Microsoft generated about $30 billion from its AI partnership between 2023 and 2025, while related costs reached $100 billion; he also said actual usage among paid Copilot users is below 3%.

Why it matters: HKR-H/K/R all pass: a former executive gives concrete Microsoft AI cost, revenue, and Copilot usage numbers. Kept at 80 because this is a single former-exec claim, not an official Microsoft disclosure.

AI HOT (Curated Pool)

Advancing content provenance for a safer, more transparent AI ecosystem

OpenAI launched an AI content provenance system that combines Content Credentials and SynthID with a verification tool; the post does not disclose supported media formats, rollout scope, or detection accuracy.

Why it matters: HKR-H/K/R pass: the OpenAI provenance stack has a concrete cross-standard mechanism and trust/compliance relevance. Missing format coverage, rollout scope, and accuracy keep it in the low featured band.

AI HOT (Curated Pool)

I really want to praise HTML!

The author used Claude Code to generate a single-file HTML project plan page in 2 minutes, with a dark theme, timeline, and collapsible tables; the comparable Notion template previously took 30-40 minutes.

Why it matters: HKR-H/K/R all pass: the post has a concrete Claude Code workflow hook, a 2-minute vs 30-40-minute comparison, and clear practitioner resonance. Scope is small, so it sits at the featured threshold.

AI HOT (Curated Pool)

Claude Managed Agents Add Self-Hosted Sandboxes and MCP Tunnels

Anthropic added two updates to the Claude managed agents platform: self-hosted sandboxes are in public beta, and MCP tunnels are in research preview for private network database and API access.

Why it matters: HKR-H/K/R all pass: this is an official Anthropic Claude agent-platform update with two concrete mechanisms. It is below model-release weight, but strong enough for featured agent-infra coverage.

AI HOT (Curated Pool)

Claude launches self-hosted sandboxes and MCP tunnels

Claude launched self-hosted sandboxes in public beta and MCP tunnels in research preview for Claude Managed Agents, letting agents run inside a user’s own security boundary with the user’s security controls applied by default.

Why it matters: HKR-H/K/R all pass: this is an official Claude agent-infra update with concrete self-hosted sandbox and MCP tunnel mechanisms, tied to enterprise security boundaries. It is beta/preview scope, not a model release, so it stays in the 78–84 band.

Computing Life · Share · Yage

Why spend hundreds of millions acquiring open-source AI infra?

Anthropic acquired Bun, Vercept, Coefficient Bio, and Stainless within six months, while OpenAI acquired Astral; the post does not disclose deal values, terms, or the cost comparison against forking the open-source projects.

Why it matters: HKR-H/K/R all pass: the counterintuitive title, five named acquisitions, and open-source infra capture anxiety create signal. Missing prices, terms, and fork-cost evidence keep it in the lower featured band.

AI HOT (Curated Pool)

Claude Design Expands Creative Capabilities

Claude Design doubled token limits across all plans, letting users create more content; the post does not disclose the exact token counts, pricing changes, rollout date, or whether any plan-specific feature limits changed.

Why it matters: Official Claude product update with one concrete fact: token limits doubled across all plans. HKR-H/K/R pass, but exact quotas, pricing, and rollout scope are missing, keeping it at the lower featured band.

TechCrunch · AI

Anthropic has acquired the dev tools startup used by OpenAI, Google, and Cloudflare

Anthropic acquired Stainless, a New York startup founded in 2022 that automates creation and maintenance of SDKs for developers using APIs; the post does not disclose the deal price or Anthropic’s integration plan.

Why it matters: HKR-H/K/R pass: the rival-used startup hook is strong, the SDK automation mechanism is concrete, and the Anthropic developer-stack angle resonates. Missing deal value and integration details keep it below the 78+ band.

Hacker News front page

We Let AIs Run Radio Stations

Andon Labs gave four AI agents tools to host live radio shows and run a media business without humans. The post says revenue is terrible, but it does not disclose amounts or operating metrics.

Why it matters: HKR-H is strong from the AI-run radio premise; HKR-K has a concrete 4-agent experiment but weak revenue disclosure; HKR-R lands on agent economics and media automation. Niche lab post, so it sits at the featured threshold, not same-day news.

AI HOT (Curated Pool)

Claude Console Adds Prompt Cache Diagnostics

Anthropic added prompt cache diagnostics to the Claude Console; when a request misses the cache, developers can see which prompt segment changed and how many tokens it consumed.

Why it matters: HKR-H/K/R pass: this is a small but practical Anthropic Claude developer-console update with concrete cache-miss diagnostics and token-cost visibility, enough for the featured threshold but below major product-release weight.

r/LocalLLaMA

Tried Every Hermes Agent Alternative So You Don't Have To: 2026 Roundup

A Reddit user compared 11 Hermes Agent alternatives across open-source and managed options; OpenClaw is listed with 347k GitHub stars, 24+ integrations, and 9 CVEs in four days, while TrustClaw uses OAuth-only sandboxed execution and Perplexity Computer requires a $200/month Max tier.

Why it matters: HKR-H/K/R all pass: this is a practical agent-tool comparison with 11 items and concrete integration/security figures. Reddit single-post sourcing limits confidence, so it stays near the featured threshold.

AI HOT (Curated Pool)

Anthropic Acquires SDK Platform Stainless

Anthropic is acquiring Stainless, an SDK and MCP server platform that has supported all Anthropic SDKs since the early Anthropic API period; the post does not disclose the deal value, closing timeline, or integration plan.

Why it matters: HKR-H/K/R all pass: Anthropic is buying a core SDK/MCP tooling partner. The post lacks price and closing timing, so this is featured developer-ecosystem news, not P1.

AI HOT (Curated Pool)

Take your local GitHub sessions anywhere

GitHub launched remote control sessions for Copilot, letting users start tasks in VS Code or the command line and continue them through github.com or GitHub Mobile.

Why it matters: GitHub Copilot session handoff from VS Code/CLI to web and mobile clears HKR-H/K/R, but the post only gives entry points and use case; permissions, pricing, and supported task scope are not disclosed.

May 18Monday

Hacker News front page

Show HN: InsForge – Open-source Heroku for coding agents

InsForge released an Apache 2.0 backend platform that lets coding agents deploy, operate, and debug backend systems through one CLI install command and Skills.

Why it matters: HKR-H/K/R all pass: the Heroku-for-agents framing, Apache 2.0 plus one-CLI install, and agent ops pain are concrete. Source is mainly Show HN/GitHub with no usage, benchmark, or production proof, so it sits at the featured threshold.

AI HOT (Curated Pool)

The Open Agent Leaderboard

IBM Research published the Open Agent Leaderboard on Hugging Face to evaluate agents across language understanding, tool use, and multi-step reasoning tasks; the post does not disclose dataset size, model scores, or the evaluation date.

Why it matters: HKR-H and HKR-R pass because an open agent leaderboard speaks to agent-eval pain. HKR-K fails: the article lacks scores, dataset size, and evaluation date, so it sits at the featured threshold.

QbitAI · WeChat

openJiuwen open-sources JiuwenSwarm, a multi-agent swarm coordination framework

openJiuwen released and open-sourced JiuwenSwarm with four components: Agent Swarm, Swarm Skills, Skills Hub, and self-evolution, and the framework supports HOTS and HITS modes for human participation in multi-agent workflows.

Why it matters: HKR-H/K/R pass: the swarm angle is clickable, the post gives four modules plus HOTS/HITS, and agent builders care about orchestration choices. Lacking benchmarks or adoption data keeps it at the featured threshold.

r/LocalLLaMA

I built a coding agent that gets 87% on benchmarks with a 4B parameter model

SmallCode passes 87 of 100 benchmark tasks with Gemma 4 activating 4B parameters per token. The author attributes the result to compound tools, compile and lint feedback, task decomposition after two repeated failures, and optional escalation to Claude or OpenAI for one task.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post and the benchmark identity plus replication details are incomplete. It fits a concrete first-person experiment above the featured bar, not the 78+ band.

Synced · WeChat

openJiuwen releases JiuwenSwarm, an open-source multi-agent swarm framework

openJiuwen released and open-sourced JiuwenSwarm with four components: Agent Swarm, Swarm Skills, Swarm Skills Hub, and self-evolving Swarm Skills, and reports a 94.2% PinchBench score versus 91.6% for OpenClaw.

Why it matters: HKR-H/K/R all pass: an open-source agent-swarm framework with named components and a PinchBench 94.2% claim. It stays at 78 because openJiuwen is not a top lab and the summary lacks license, reproduction setup, and baselines.

AI HOT (Curated Pool)

Tencent AI Design Agent Ardot Enters Public Beta: Generates Editable Designs and Converts Them to Code

Tencent Cloud opened public beta for Ardot, an AI design agent that generates editable app pages, websites, and posters from one-sentence prompts, then converts designs to code.

Why it matters: HKR-H/K/R pass on a concrete Tencent product beta for editable design-to-code workflows. Missing pricing, model details, benchmarks, and field results keep it at the lower featured threshold.

AI HOT (Curated Pool)

Open-source tool exposes security risks and detection gaps in AI API relays

api-relay-audit audits AI API relay risks with verifiable three-state decisions and transparent logs, covering AC-1 tool-call rewriting, AC-2 error-response leakage, and context truncation, while the author has published the methodology, comparison results, quick-reference table, and the open-source tool.

Why it matters: HKR-H/K/R all pass because the tool targets real AI API relay risks with concrete checks. Source is a single X post, and adoption or incident data is not disclosed, so it stays in the low featured band.

AI HOT (Curated Pool)

Grok launches Skills feature

xAI launched Grok Skills on May 18, 2026, letting users set preferences, formatting rules, or workflows once and keep them active across all conversations on web, iOS, and Android.

Why it matters: HKR-H/K/R all pass: Grok Skills adds persistent preferences and workflows across web, iOS, and Android. This is a mid-weight xAI product update; rollout scope, limits, and pricing are not disclosed.

May 17Sunday

r/LocalLLaMA

MiroThinker-1.7 Open-Weight Deep Research Agent Based on Qwen3 MoE

MiroMindAI released the MiroThinker-1.7-deepresearch and mini APIs, with the mini version using 30B total parameters and 3B active parameters, weights on HuggingFace, and context management based on sliding window K=5 plus episode restarts.

Why it matters: HKR-H/K/R all pass, but the source is a Reddit thread and the lab is not top-tier. Open weights, MoE sizing, and context-management details clear featured, not same-day must-write.

Synced · WeChat

Peter Steinberger Says His Monthly Token Bill Hit $1.3M, Covered by OpenAI

Peter Steinberger used 603 billion tokens across 7.6 million requests in 30 days, with the bill exceeding $1.3 million; he said disabling fast mode cut the price by 70%, and OpenAI does not charge him for the tokens.

Why it matters: HKR-H/K/R all pass: the story has a sharp cost hook, concrete usage numbers, and strong practitioner resonance. It is a first-person bill disclosure, not an OpenAI pricing or product launch, so it sits just above the featured threshold.

AI HOT (Curated Pool)

MagicPath Integrates with Codex to Combine Design and Development

MagicPath AI CEO @skirano demonstrated MagicPath running inside Codex as a native canvas, with users configuring it through one command, dragging UI elements, and letting Codex generate and edit code in real time.

Why it matters: HKR-H/K/R pass: MagicPath puts a draggable design canvas inside Codex with one-command setup and live code edits. Single-demo sourcing and missing framework support, permissions, and reproducible cases keep it at the lower featured band.

AI HOT (Curated Pool)

Study on the Cognition–Action Disconnect in Tool-Using Agents

An interpretability paper studies tool-using agents and finds models often recognize when to call a tool but fail to act, with a cognition-to-action mismatch rate of 26%–54%.

Why it matters: HKR-H/K/R all pass: the story has a sharp agent-failure hook, a 26%-54% mismatch rate, and clear relevance to tool-use reliability. Source detail is thin, with paper name, models, and task setup not disclosed.

AI HOT (Curated Pool)

Ring-2.6-1T Open-Sourced and Listed on OpenRouter for Agent Workflows

AntLingAGI open-sourced Ring-2.6-1T and listed it on OpenRouter with a 75% discount through the end of May; the trillion-scale reasoning model targets agent workflows, including planning, tool use, context maintenance, and complex task execution, using Async RL and IcePop training methods.

Why it matters: HKR-H/K/R all pass: a 1T open agent model is clickable, with OpenRouter access, discount, and training methods disclosed. Score stays at 74 because benchmarks, license, and context window are not given.

May 16Saturday

TechCrunch · AI

OpenAI co-founder Greg Brockman takes charge of product strategy

Greg Brockman has officially taken charge of OpenAI’s product strategy, and Wired reports that he described a plan in a staff memo to combine ChatGPT and Codex into one unified experience.

Why it matters: HKR-H/K/R all pass: OpenAI co-founder product control plus a reported ChatGPT-Codex unification matters. No launch date, feature boundary, or rollout plan is disclosed, so this stays below a major product release.

AI HOT (Curated Pool)

Anthropic Founder’s Playbook warns AI can raise startup failure rates

Anthropic published Founder’s Playbook, arguing that AI tools such as Claude Code reduce prototyping cost but increase startup failure risk across the Idea, MVP, Launch, and Scale stages through false validation, confirmation bias, agentic technical debt, and founder decision bottlenecks.

Why it matters: HKR-H/K/R pass: the Anthropic founder playbook has a sharp counterintuitive angle, a four-stage mechanism, and clear founder resonance. It stays near the featured floor because no dataset or reproducible test is disclosed.