Skip to content

MCP & tool use

How models connect to the outside world: the MCP ecosystem, function calling and tool integrations.

760 picksRelated topicsAgentsAI codingOpen source

Latest picks

261–280 of 760

May 18Monday

r/LocalLLaMA

I built a coding agent that gets 87% on benchmarks with a 4B parameter model

SmallCode passes 87 of 100 benchmark tasks with Gemma 4 activating 4B parameters per token. The author attributes the result to compound tools, compile and lint feedback, task decomposition after two repeated failures, and optional escalation to Claude or OpenAI for one task.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post and the benchmark identity plus replication details are incomplete. It fits a concrete first-person experiment above the featured bar, not the 78+ band.

Synced · WeChat

openJiuwen releases JiuwenSwarm, an open-source multi-agent swarm framework

openJiuwen released and open-sourced JiuwenSwarm with four components: Agent Swarm, Swarm Skills, Swarm Skills Hub, and self-evolving Swarm Skills, and reports a 94.2% PinchBench score versus 91.6% for OpenClaw.

Why it matters: HKR-H/K/R all pass: an open-source agent-swarm framework with named components and a PinchBench 94.2% claim. It stays at 78 because openJiuwen is not a top lab and the summary lacks license, reproduction setup, and baselines.

AI HOT (Curated Pool)

Tencent AI Design Agent Ardot Enters Public Beta: Generates Editable Designs and Converts Them to Code

Tencent Cloud opened public beta for Ardot, an AI design agent that generates editable app pages, websites, and posters from one-sentence prompts, then converts designs to code.

Why it matters: HKR-H/K/R pass on a concrete Tencent product beta for editable design-to-code workflows. Missing pricing, model details, benchmarks, and field results keep it at the lower featured threshold.

AI HOT (Curated Pool)

Open-source tool exposes security risks and detection gaps in AI API relays

api-relay-audit audits AI API relay risks with verifiable three-state decisions and transparent logs, covering AC-1 tool-call rewriting, AC-2 error-response leakage, and context truncation, while the author has published the methodology, comparison results, quick-reference table, and the open-source tool.

Why it matters: HKR-H/K/R all pass because the tool targets real AI API relay risks with concrete checks. Source is a single X post, and adoption or incident data is not disclosed, so it stays in the low featured band.

AI HOT (Curated Pool)

Grok launches Skills feature

xAI launched Grok Skills on May 18, 2026, letting users set preferences, formatting rules, or workflows once and keep them active across all conversations on web, iOS, and Android.

Why it matters: HKR-H/K/R all pass: Grok Skills adds persistent preferences and workflows across web, iOS, and Android. This is a mid-weight xAI product update; rollout scope, limits, and pricing are not disclosed.

May 17Sunday

r/LocalLLaMA

MiroThinker-1.7 Open-Weight Deep Research Agent Based on Qwen3 MoE

MiroMindAI released the MiroThinker-1.7-deepresearch and mini APIs, with the mini version using 30B total parameters and 3B active parameters, weights on HuggingFace, and context management based on sliding window K=5 plus episode restarts.

Why it matters: HKR-H/K/R all pass, but the source is a Reddit thread and the lab is not top-tier. Open weights, MoE sizing, and context-management details clear featured, not same-day must-write.

Synced · WeChat

Peter Steinberger Says His Monthly Token Bill Hit $1.3M, Covered by OpenAI

Peter Steinberger used 603 billion tokens across 7.6 million requests in 30 days, with the bill exceeding $1.3 million; he said disabling fast mode cut the price by 70%, and OpenAI does not charge him for the tokens.

Why it matters: HKR-H/K/R all pass: the story has a sharp cost hook, concrete usage numbers, and strong practitioner resonance. It is a first-person bill disclosure, not an OpenAI pricing or product launch, so it sits just above the featured threshold.

AI HOT (Curated Pool)

MagicPath Integrates with Codex to Combine Design and Development

MagicPath AI CEO @skirano demonstrated MagicPath running inside Codex as a native canvas, with users configuring it through one command, dragging UI elements, and letting Codex generate and edit code in real time.

Why it matters: HKR-H/K/R pass: MagicPath puts a draggable design canvas inside Codex with one-command setup and live code edits. Single-demo sourcing and missing framework support, permissions, and reproducible cases keep it at the lower featured band.

AI HOT (Curated Pool)

Study on the Cognition–Action Disconnect in Tool-Using Agents

An interpretability paper studies tool-using agents and finds models often recognize when to call a tool but fail to act, with a cognition-to-action mismatch rate of 26%–54%.

Why it matters: HKR-H/K/R all pass: the story has a sharp agent-failure hook, a 26%-54% mismatch rate, and clear relevance to tool-use reliability. Source detail is thin, with paper name, models, and task setup not disclosed.

AI HOT (Curated Pool)

Ring-2.6-1T Open-Sourced and Listed on OpenRouter for Agent Workflows

AntLingAGI open-sourced Ring-2.6-1T and listed it on OpenRouter with a 75% discount through the end of May; the trillion-scale reasoning model targets agent workflows, including planning, tool use, context maintenance, and complex task execution, using Async RL and IcePop training methods.

Why it matters: HKR-H/K/R all pass: a 1T open agent model is clickable, with OpenRouter access, discount, and training methods disclosed. Score stays at 74 because benchmarks, license, and context window are not given.

May 16Saturday

TechCrunch · AI

OpenAI co-founder Greg Brockman takes charge of product strategy

Greg Brockman has officially taken charge of OpenAI’s product strategy, and Wired reports that he described a plan in a staff memo to combine ChatGPT and Codex into one unified experience.

Why it matters: HKR-H/K/R all pass: OpenAI co-founder product control plus a reported ChatGPT-Codex unification matters. No launch date, feature boundary, or rollout plan is disclosed, so this stays below a major product release.

AI HOT (Curated Pool)

Anthropic Founder’s Playbook warns AI can raise startup failure rates

Anthropic published Founder’s Playbook, arguing that AI tools such as Claude Code reduce prototyping cost but increase startup failure risk across the Idea, MVP, Launch, and Scale stages through false validation, confirmation bias, agentic technical debt, and founder decision bottlenecks.

Why it matters: HKR-H/K/R pass: the Anthropic founder playbook has a sharp counterintuitive angle, a four-stage mechanism, and clear founder resonance. It stays near the featured floor because no dataset or reproducible test is disclosed.

AI HOT (Curated Pool)

Codex adds multi-device remote control and shared context

Codex controls multiple devices through ChatGPT, switches by project to access each device’s context and files, and supports remote SSH setup for other VMs.

Why it matters: HKR-H/K/R all pass, but the item is a thin X-post summary with no official release note, pricing, permission model, or reproducible demo. Treat it as a mid-weight coding-agent product update at the featured threshold.

AI HOT (Curated Pool)

OpenAI Restructures as Brockman Takes Over Product Strategy

OpenAI merged ChatGPT, Codex, and API into one product organization, with Greg Brockman taking over product strategy; the post says Anthropic’s valuation reached $900 billion, but it does not disclose the restructuring timeline.

Why it matters: HKR-H/K/R all pass: this is an OpenAI top-level product reorg covering ChatGPT, Codex, and API. Single-source summary keeps it below the highest band, but it is same-day must-write news.

QbitAI · WeChat

Codex Integrates HeyGen for Prompt-Based Video Generation and Editing

Codex integrates the HeyGen plugin to run image generation, talking-avatar video, subtitles, and edits from natural-language prompts; the article tests roughly one-minute avatar generation, trimming content after 10 seconds, and deleting a blink at the eighth second.

Why it matters: HKR-H/K/R all pass, backed by a numbered hands-on test. The scope is still one Codex-to-HeyGen plugin workflow, not a model or platform release, so it lands in the 72-77 featured band.

Computing Life · Share · Yage

OpenAI Reaches Into Your Bank Account

OpenAI uses Plaid to let ChatGPT connect to bank accounts; the post does not disclose launch timing, authorization flow, or the exact data scope ChatGPT can access.

Why it matters: HKR-H/R are strong and HKR-K passes via the Plaid integration mechanism. Missing launch timing, authorization flow, and data scope keep it at the featured threshold rather than a higher OpenAI product-update score.

AI HOT (Curated Pool)

Ignoring Token Costs, Using 100 AI Instances to Automate an Open Source Project

The OpenClaw team runs about 100 Codex instances to handle code review, security analysis, issue deduplication, test reproduction, task creation from meetings, spam filtering, and performance regression monitoring.

Why it matters: HKR-H/K/R all pass: 100 Codex instances running open-source maintenance is a strong operational anecdote with concrete task types. Single X post, no cost, outcome metrics, or reproducible setup, so it stays in the lower featured band.

The Verge · AI

OpenAI now wants ChatGPT to access your bank accounts

OpenAI previewed a ChatGPT feature that lets users connect financial accounts through Plaid, which links to 12,000 institutions. OpenAI says more than 200 million people ask ChatGPT finance questions each month; the post does not disclose a general release date.

Why it matters: HKR-H/K/R all pass: OpenAI is moving ChatGPT toward real financial-account access, with Plaid’s 12,000 institutions and 200M monthly finance askers as concrete facts. It stays below 85 because launch timing is not disclosed.

TechCrunch · AI

OpenAI launches ChatGPT for personal finance, will let users connect bank accounts

OpenAI launched ChatGPT for personal finance, and connected users can view portfolio performance, spending, subscriptions, and upcoming payments; the RSS snippet does not disclose supported banks, launch regions, pricing, or account-security terms.

Why it matters: HKR-H is strong because ChatGPT connects to bank accounts; HKR-K has concrete finance features; HKR-R hits privacy and fintech competition. Banks, regions, and pricing are undisclosed, so this stays in the low P1 band.

May 15Friday

AI HOT (Curated Pool)

X open-sources the “For You” feed recommendation algorithm

X open-sourced the For You recommendation pipeline on GitHub, using a Grok-based Phoenix Transformer to score candidate posts and predict engagement probabilities such as likes, replies, and reposts.

Why it matters: HKR-H/K/R all pass, but the item only gives the open-source claim and Phoenix Transformer ranking mechanism; repo details, license, and reproducible tests are not disclosed, so it stays low-featured.