Skip to content

MCP & tool use

How models connect to the outside world: the MCP ecosystem, function calling and tool integrations.

760 picksRelated topicsAgentsAI codingOpen source

Latest picks

201–220 of 760

May 21Thursday

Alibaba Technology · WeChat

Building an Agent from 0 to 1: Principles and Personal Assistant Practice

Zhan Xupeng published a roughly 50-minute article on Agent theory and a personal assistant implementation, covering memory, ReAct planning, progressive skill loading, subagents, and harness-level fault recovery.

Why it matters: HKR-K/R pass via concrete agent mechanisms and practitioner reliability pain; HKR-H is weak because the headline is a standard tutorial frame. This fits the quality-tutorial threshold, not the 78+ news band.

Xinzhiyuan · WeChat

Anthropic Acquires SDK Toolmaker Stainless, Leaving OpenAI and Others to Maintain SDKs

Anthropic has completed its acquisition of Stainless, an SDK generation company used by OpenAI, Anthropic, Meta, Cloudflare, and other infrastructure vendors; Stainless says prior SDK ownership remains with customers, but it will shut down hosted products including SDK generator and stop providing ongoing support.

Why it matters: HKR-H/K/R all pass: the deal targets API SDK generation, names OpenAI/Meta/Cloudflare as customers, and says hosted products will shut down. Anthropic bump applies, but this is not a model or core capability release, so it fits 78–84.

r/LocalLLaMA

Moved from prompt-based output validation to schema-enforced execution, with significant reliability gains

A Reddit user tested Claude structured outputs and reported 90–95%+ first-pass parse rates with tool_use, typed schemas, enum constraints, and stepwise validation, versus 65–70% for prompt instructions followed by regex or JSON parsing and retries.

Why it matters: HKR-H/K/R all pass: the post has a clear reliability contrast and concrete 90–95%+ vs 65–70% numbers. Source authority is limited to one Reddit experiment, with sample and task details not disclosed, so it stays at low featured.

AI HOT (Curated Pool)

Equipping AI with a Scientific Toolkit to Accelerate Research Workflows

Google DeepMind released Science Skills for Google Antigravity, integrating insights from more than 30 life-science sources, including the UniProt and AlphaFold databases.

Why it matters: Google DeepMind has a concrete product update: Science Skills adds 30+ life-science sources to Antigravity, hitting HKR-H/K. It is a vertical toolkit rather than a model release, so it sits at the low featured band.

AI HOT (Curated Pool)

Tencent Launches OS-Level AI Assistant Mavis on Windows, Mac, and Android

Tencent launched the OS-level AI assistant Mavis on May 21 across Windows, Mac, and Android, with document parsing, image recognition, system maintenance, partial offline use, model dispatching, and desktop control of mobile apps listed as supported functions.

Why it matters: HKR-H/K/R all pass: Tencent’s OS-level assistant spans Windows, Mac, and Android with concrete tool abilities. Model, pricing, and permission design are not disclosed, so it stays at the lower featured band.

Latent Space

Railway: The Agent-Native Cloud — Jake Cooper

Railway serves 3 million users with a 35-person team, adds about 100,000 signups per week, has raised $124 million, and has moved most workloads to its own bare-metal data centers with a reported three-month payback versus rented cloud capacity.

Why it matters: HKR-H/K/R all pass: the Railway interview has concrete growth, funding, and bare-metal details tied to agent infrastructure. It is a strong practitioner story, not a core model or major AI product release, so 74 fits the featured threshold.

AI HOT (Curated Pool)

Google Stitch update: AI design assistant supports end-to-end building

Google updated its AI design partner Stitch with real-time streaming design builds, direct edits and feedback, codebase or Design.md imports, dynamic UI generation, shareable URL exports, and global availability.

Why it matters: HKR-H/K/R pass: Google Stitch adds streaming builds, codebase/Design.md import, and global access. It stays at the featured threshold because model details, pricing, and measured output quality are not disclosed.

AI HOT (Curated Pool)

ChatGPT mobile app adds Codex support for cross-device collaboration

OpenAI Devs says the ChatGPT mobile app now supports Codex, letting users ask questions on mobile and continue the same conversation on desktop; the post does not disclose supported platforms, app versions, or rollout scope.

Why it matters: OpenAI Devs is authoritative and HKR-H/K/R pass, but the post only confirms mobile Codex access and handoff; platform, version, and rollout scope are not disclosed, so it sits at the featured threshold.

The Verge · AI

Google Search’s AI Evolution Includes More Ads

Google is adding Gemini-generated product explanations to Search ads for product queries, including Sponsored Product placements and some ads with built-in chatbots; the snippet cites a compact espresso pod machine example, but the post does not disclose rollout scope or pricing.

Why it matters: HKR-H/K/R all pass: Google is inserting Sponsored Product units and ad chatbots into Gemini shopping results. Scope, pricing, and performance data are not disclosed, so this stays in low featured rather than a major product-release band.

May 20Wednesday

The Verge · AI

If Google Can’t Make AI Agents Useful, Maybe No One Can

The Verge says Google announced multiple AI agents at I/O 2026 for information gathering, event planning, and inbox or calendar summarization; the RSS snippet says the agents can run continuously in the background, but the post does not disclose launch timing, pricing, or evaluation results.

Why it matters: HKR-H and HKR-R are strong because the Verge frames Google agents as a sector test; HKR-K is limited to background-running agents. Missing launch timing, pricing, and evals keeps it in the lower featured band.

TechCrunch · AI

Figma adds an AI assistant to its collaborative canvas

Figma added an AI agent to its collaborative canvas, letting users use natural-language prompts to create new designs, edit existing ones, or automate tasks such as generating design iterations.

Why it matters: HKR-K and HKR-R pass: Figma puts an AI agent into a core collaborative design surface. Details stop at capability scope, with no model, pricing, or rollout timing, so this sits at the featured threshold.

AI Chat-Group Daily (群聊日报)

2026-05-19 Chat Group Daily

The chat group daily says Karpathy joined Anthropic's pretraining team, and cites Stainless shutting down hosted services after acquisition plus Google I/O announcing Gemini 3.5 Flash and a $100 subscription tier.

Why it matters: HKR-H/K/R all pass, but this is a chat-daily roundup with secondhand claims and no disclosed primary links, appointment details, or product specs, so it lands at the lower featured band.

AI HOT (Curated Pool)

OpenAI offers $2 million API investment to every YC startup

OpenAI is offering each startup in Y Combinator’s current batch $2 million in API credits in exchange for equity; the post does not disclose the equity stake, credit expiration, or usage limits.

Why it matters: HKR-H/K/R all pass: $2M per current YC startup in API credits for equity is concrete and talkable. Missing equity %, term and usage caps keep it at 78, below the 85+ must-write band.

AI HOT (Curated Pool)

Microsoft reportedly warns internally that GitHub faces existential risk as AI coding tools reduce hosting need

Microsoft internally warned that GitHub faces an existential risk from AI coding assistants such as Cursor and Claude Code, and told some teams to stop using Claude Code by the end of June 2026 and move to GitHub Copilot CLI.

Why it matters: HKR-H/K/R all pass: the angle is sharp, the summary gives a Claude Code-to-Copilot CLI deadline, and the workflow stakes are real. Single-source “reported” framing and no Microsoft response keep it below the 85 must-write band.

AI HOT (Curated Pool)

Qwen3.7: Agent Frontier

Qwen Studio released Qwen3.7 with chatbots, image and video understanding, and image generation. It also covers document processing, web search integration, tool calling, and artifact generation. The RSS snippet frames it as an agent-focused model, but the post does not disclose context length. It also omits benchmark scores, pricing, API limits, release schedule, and reproducible evaluation conditions.

Why it matters: HKR-H/K/R all pass: this is a Qwen flagship-model update with concrete capability coverage. Lack of benchmarks, pricing, and context-window details keeps it at the low end of the 85–94 band.

AI HOT (Curated Pool)

Gemini 3.5 Flash price rises sharply as Google plans broad rollout

Google released Gemini 3.5 Flash at I/O with $1.50 per million input tokens and $9 per million output tokens, making it 3x and 6x the previous model’s pricing, while adding roughly 1 million input tokens and about 65,000 maximum output tokens.

Why it matters: HKR-H/K/R all pass: a Google model update with a sharp pricing twist and concrete token costs. It stays below p1 because the body only gives price and rollout intent, not capability deltas, benchmarks, or context window.

AI HOT (Curated Pool)

Gemini launches personal AI agent and Daily Brief

Gemini added Gemini Spark and Daily Brief: Spark acts as an always-on personal AI agent across Gmail, Google Docs, and Slides after user authorization, while Daily Brief is available to Google AI subscribers in the U.S. aged 18 or older.

Why it matters: HKR-H/K/R all pass: Google is adding Gemini Spark’s authorized actions across Gmail, Docs, and Slides, plus Daily Brief eligibility for US 18+ AI subscribers. This is a same-day Google agent product update.

AI HOT (Curated Pool)

Claude Code’s HTML Output: The Unreasonable Effectiveness of HTML

The Claude Code team is shifting its primary output format from Markdown to HTML, and the post names four mechanisms: tables, CSS styling, SVG charts, and JavaScript interactions.

Why it matters: Official Claude Code post with a concrete shift from Markdown to HTML and 4 output mechanisms; strong practitioner utility, but not a major product launch, so it sits in the featured threshold band.

TechCrunch · AI

You Can Now Talk to Your Gmail Inbox, as Seen at Google I/O 2026

Google expanded Gmail’s AI Inbox with conversational voice search, letting users ask Gemini to find details buried in email. The RSS snippet does not disclose rollout scope, supported languages, pricing, latency, or the retrieval mechanism behind Gmail search.

Why it matters: HKR-H/K pass: a Google-scale Gmail voice inbox feature is concrete and clickable. HKR-R is weak because rollout, language support, pricing, and retrieval mechanics are not disclosed.

The Verge · AI

Google’s AI Future Demands Trust — and Your Personal Data

Google presented Gemini Spark, Daily Brief, and expanded Gmail AI inbox access at I/O 2026; the Verge snippet says these tools depend on large amounts of personal information, but the post does not disclose detailed data-handling terms.

Why it matters: HKR-H/K/R all pass, but the body gives product names and a personal-data dependency without data-handling details. Google I/O makes it featured, not a same-day must-write.