Skip to content

MCP & tool use

How models connect to the outside world: the MCP ecosystem, function calling and tool integrations.

760 picksRelated topicsAgentsAI codingOpen source

Latest picks

321–340 of 760

May 13Wednesday

AI HOT (Curated Pool)

Codex Enables Background Multitasking Across Apps

OpenAI Devs says computer use lets Codex click, type, and keep working across Mac apps in the background; the post does not disclose release timing, permission design, or availability scope.

Why it matters: HKR-H/K/R all pass: OpenAI Devs gives a concrete Codex mechanism for background Mac cross-app actions, with clear developer relevance. Release timing, permissions, and availability are missing, so it stays at the lower featured band.

AI HOT (Curated Pool)

How Anthropic's Cybersecurity Team Uses Claude Code to Build a Threat Detection Platform

Anthropic’s detection platform engineering team used Claude Code to build the CLUE threat detection and response platform, completing a proof of concept in one day and delivery in one week while reducing analyst log investigation from hours to minutes.

Why it matters: HKR-H/K/R all pass, but this is an Anthropic internal dogfooding case rather than a Claude Code capability launch. CLUE and the timing metrics keep it just above the featured threshold.

AI HOT (Curated Pool)

Claude Opus 4.7 Fast Mode Opens Research Preview

Claude Opus 4.7 Fast Mode is now available as a research preview in the API and Claude Code. The post does not disclose model parameters, pricing, rate limits, or a general availability date.

Why it matters: HKR-H/K/R pass because this is a Claude fast-mode preview in API and Claude Code, directly tied to developer latency and workflows. Thin disclosure on pricing, limits, parameters, and GA timing keeps it at the featured threshold, not 78+.

Hacker News front page

Show HN: Needle Distills Gemini Tool Calling into a 26M Model

Cactus open-sourced Needle, a 26M-parameter tool-calling model that reaches 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, with MIT-licensed weights released on Hugging Face.

Why it matters: HKR-H/K/R all pass: the tiny Gemini-style tool-calling angle is clickable, with concrete speed and license claims. Source is still Show HN/GitHub self-reporting, not an independent benchmark or major lab release, so it stays below the 78–84 band.

r/LocalLLaMA

Needle: We Distilled Gemini Tool Calling Into a 26M Model

Cactus Compute open-sourced Needle, a 26M-parameter tool-calling model that reaches 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, using an attention-and-gating architecture with no MLPs.

Why it matters: HKR-H/K/R all pass: a 26M tool-calling model has a strong hook and concrete speed/design claims. Single Reddit source and a less-known team keep it in the lower 78–84 band.

AI HOT (Curated Pool)

Claude Enters the Legal Industry

Anthropic released more than 20 MCP connectors and 12 legal plugins, letting Claude work inside Word and Outlook for contract drafting, revision, clause comparison, and routine legal workflows.

Why it matters: HKR-H/K/R all pass: a substantive Anthropic vertical product update with 20+ MCP connectors and Office workflows. It is not a model release or platform-wide capability, so it stays in the 72–77 band.

AI HOT (Curated Pool)

Google Launches New Android Smart Assistant

Google introduced Android Intelligence at Android Show 2026, with multi-step automation across Android apps, browser-use features for Gemini in Chrome, automatic form filling, Rambler voice-note transcription, and custom Gen UI widgets; the post does not disclose rollout timing, supported devices, or pricing.

Why it matters: HKR-H/K/R all pass: the hook is Android-level agent control, the new facts are concrete automation surfaces, and the resonance is the mobile AI platform fight. Thin source detail keeps it at the low end of the 85-94 band.

Hacker News front page

Show HN: Agentic Interface for Mainframes and COBOL

Hypercubic launched Hopper, an agentic development environment that combines a real TN3270 terminal, z/OS-aware panels for datasets, jobs, and spool output, and an AI agent; sensitive operations require approval, and the terminal remains visible during agent actions.

Why it matters: HKR-H/K/R all pass: the mainframe-agent angle is novel, with concrete TN3270, z/OS, and approval mechanics. Small-vendor Show HN status and missing customer/pricing/results data keep it at the featured floor.

TechCrunch · AI

Everything Google announced at its Android Show, from Googlebooks to vibe-coded widgets

Google announced AI-first Googlebooks laptops, more agentic Gemini features, vibe-coded Android widgets, Gemini in Chrome, and refreshed Android Auto ahead of I/O; the RSS snippet does not disclose specs, pricing, availability, or rollout timelines.

Why it matters: HKR-H/K/R all pass because Google bundled several Gemini/Android AI entry points with named product hooks. Missing parameters, pricing, rollout dates, and testable performance keeps it in the mid-weight product-update band.

The Verge · AI

Gemini’s Latest Updates Are All About Controlling Your Phone

Google announced Gemini Intelligence at its pre-I/O Android showcase, placing Gemini in Chrome on Android, autofill suggestions, and in-app actions; the RSS snippet does not disclose supported device models, rollout timing, or pricing.

Why it matters: HKR-H/K/R all pass because Gemini is moving into Android phone control across Chrome, autofill, and app actions. Missing device list, launch timing, and pricing keep it in the 72–77 product-update band.

TechCrunch · AI

Google brings agentic AI and vibe-coded widgets to Android

Google is adding Gemini Intelligence to Android, and the RSS snippet only discloses Gboard-based dictation and form-filling capabilities; the post does not disclose launch timing, supported devices, pricing, or technical implementation details.

Why it matters: HKR-H/K/R pass because this is a Google Android platform AI update with named features. Missing rollout timing, device scope, and pricing keep it at the low featured band, not a must-write release.

TechCrunch · AI

The AI legal services industry is heating up — Anthropic is getting in on the action

Anthropic introduced tools for law firms that cover five clerical workflows: document search and review, case law resources, deposition preparation, document drafting, and related tasks; the RSS snippet does not disclose pricing, launch timing, or model details.

Why it matters: HKR-H/K/R pass: Anthropic is moving into a high-value legal workflow with five named use cases. No pricing, customer scale, or new model capability is disclosed, so this stays just above the featured threshold.

AI HOT (Curated Pool)

Code w/ Claude SF 2026: Building on Exponential AI Growth

Anthropic expanded developer tooling at Code w/ Claude SF 2026: Claude Code rate limits doubled, Claude Opus API limits increased, and hosted agents on the Claude platform added four functions, including memory review, multi-agent delegation, output criteria, and webhooks.

Why it matters: Anthropic ships a substantive Claude Code update with concrete numbers and feature additions; HKR-H/K/R all pass. This is strong dev-tool news, not a flagship model release, so it fits the 78–84 band.

May 12Tuesday

AI HOT (Curated Pool)

Install the official Codex plugin in Claude Code

The author describes installing OpenAI’s official Codex plugin in Claude Code via the plugin marketplace: add the repository, install the plugin, reload, and configure it, then use it to build a Skill where Claude Code handles reasoning and Codex acts as moderator.

Why it matters: HKR-H/K/R all pass: cross-stack plugin use is clickable, the install path is concrete, and it matters to AI dev workflows. It stays low-featured because this is a tutorial-style tip, not a model or platform release.

AI HOT (Curated Pool)

Dungeons & Desktops: Building a Procedurally Generated Roguelike with GitHub Copilot CLI

A GitHub employee used GitHub Copilot CLI to build an extension that parses any codebase into one Roguelike-style dungeon layout, with procedural level generation used as the core mechanism for a creative coding and game prototyping demo.

Why it matters: HKR-H and HKR-K pass: an official GitHub tutorial has a novel demo and a clear mechanism. It is not a major Copilot capability release, and lacks production metrics, pricing, or benchmark data, so it sits at the tutorial-featured floor.

Hacker News front page

Show HN: Statewright – Visual State Machines for More Reliable AI Agents

Statewright uses a Rust state-machine engine to constrain Claude Code tool access, iterations, transitions, and guards; the post says 13–20B models improved consistently on real SWE-bench tasks, but it does not disclose benchmark scores, sample size, or the exact evaluation protocol.

Why it matters: HKR-H/K/R all pass: the state-machine constraint is a clear agent-reliability hook with a testable SWE-bench claim. Exact scores and reproduction details are not disclosed, so it stays just above the featured threshold.

AI HOT (Curated Pool)

Exporting consumer data for personalized AI Agent services

The author surveyed export methods across 5 consumer platforms: Taobao supports exports, JD.com needs a Codex-built Chrome extension, Ele.me can provide Excel exports by request, Meituan Waimai has no method, and JD.com plus Dianping tools are open sourced.

Why it matters: HKR-H/K/R pass: the post has a concrete personal-agent hook, 5-platform data, and open-source JD/Dianping tools. It stays at the featured floor because this is a practitioner experiment, not a platform release.

Financial Times · Technology

Amazon staff use AI tool for unnecessary tasks to inflate usage scores

Amazon’s in-house MeshClaw tool lets employees delegate work to AI agents and raise their position on the company’s AI leaderboard; the post does not disclose the number of staff involved, the scoring rules, or the specific unnecessary tasks.

Why it matters: FT gives a concrete Amazon AI-adoption gaming story, clearing HKR-H/K/R. Missing participant count, scoring rules, and task examples keep it at the featured threshold, not a must-write item.

QbitAI · WeChat

OpenClaw quietly updates with Peekaboo v3 for Mac computer use

OpenClaw-related Peekaboo v3 adds Mac agent capabilities for pixel-level screenshots, UI position reading, clicks, text input, hotkeys, scrolling, and drag-and-drop, with MCP server integration for Cursor, Claude Code, and Codex.

Why it matters: HKR-H/K/R all pass: Peekaboo v3 adds Mac GUI perception and action primitives plus MCP access for Cursor, Claude Code, and Codex. This is a useful open-source agent-tooling update, not a model-level event, so it sits in the low featured band.

QbitAI · WeChat

Markdown Is Fading? Karpathy Also Backs HTML

Anthropic engineer Thariq argued for using HTML instead of Markdown and gave 5 reasons; the post says HTML generation takes about 2 to 4 times longer than Markdown.

Why it matters: HKR-H/K/R all pass, but this is a developer format debate rather than a model or product launch. Named Anthropic/Karpathy context and the 2-4x time figure clear the featured threshold at the low end.