Skip to content

MCP & tool use

How models connect to the outside world: the MCP ecosystem, function calling and tool integrations.

760 picksRelated topicsAgentsAI codingOpen source

Latest picks

81–100 of 760

Jun 3Wednesday

AI HOT (Curated Pool)

OpenAI Codex Sites feature launches

OpenAI launched Codex Sites, which turns work, ideas, and plans into an interactive website or app that a team can access through one URL; the feature rolls out first to Business and Enterprise plans, and the post does not disclose pricing or broader availability timing.

Why it matters: HKR-H/K/R all pass, but the post gives launch framing without pricing, permission boundaries, or quality examples. Treat it as a mid-weight OpenAI product feature, above the featured threshold.

r/LocalLLaMA

Benchmarks of 20 Small LLMs on a 6GB RTX 4050

The author benchmarked 20 small LLMs on a 6GB RTX 4050 using LM Studio’s OpenAI-compatible API, with N=5 speed runs at 1k, 8k, and 32k context; unsloth/lfm2.5-vl-1.6b led throughput at 207 tok/s on 1k context while using 3.0GB VRAM.

Why it matters: HKR-H/K/R all pass: the low-VRAM GPU hook is concrete, the post gives speed/context/VRAM numbers, and it speaks to local-inference cost pressure. Source authority is a Reddit post, so it stays in the lower featured band.

TechCrunch · AI

OpenAI launches new Codex tools for white-collar work

OpenAI released six Codex app plug-ins for data analytics, creative production, sales, product design, equity investing, and investment banking; each tool bundles integrations, instructions, and context, while the post does not disclose pricing or rollout limits.

Why it matters: HKR-H/K/R all pass: OpenAI is expanding Codex into six white-collar plugin categories. Pricing, rollout scope, and measured performance are not disclosed, so this stays in the mid-weight product-update band.

Jun 2Tuesday

AI HOT (Curated Pool)

Holo3.1: Fast Local Computer-Use Agents

Holo3.1 releases Qwen-based computer-use agents in 0.8B, 4B, 9B, and 35B-A3B sizes, with FP8, Q4 GGUF, and NVFP4 quantized checkpoints for local inference and a 79.3% AndroidWorld score for the 35B-A3B model.

Why it matters: HKR-H/K/R all pass: Holo3.1 pairs a local computer-use agent with concrete model sizes and quantized checkpoints. It fits the 78–84 band, below major lab model-release weight.

AI HOT (Curated Pool)

Anthropic Expands Project Glasswing Program

Anthropic expanded Project Glasswing to about 150 new organizations across more than 15 countries, covering electricity, water, healthcare, communications, and hardware infrastructure, after an initial group of about 50 partners.

Why it matters: Anthropic expanded Project Glasswing to about 150 new organizations across 15+ countries, giving HKR-H/K/R enough substance. No concrete safety mechanism or Claude capability change is disclosed, so it stays in the lower featured band.

AI HOT (Curated Pool)

StepFun releases Step 3.7 Flash as an open-weight model for agentic coding

StepFun released the open-weight Step 3.7 Flash model for fast agentic coding, with tool calling and multimodal understanding, and the model is already available in Kilo alongside MiniMax M3.

Why it matters: HKR-H/K/R pass on the open-weight agentic-coding angle and Kilo availability. Missing benchmarks, size, license, and pricing keep it at the lower featured threshold.

The Verge · AI

Gemini Spark is the most impressive and terrifying AI experience I’ve had yet

The Verge tested Google’s always-on AI agent Gemini Spark for trip planning, but the RSS snippet only describes a different experience from generic itinerary demos and does not disclose launch timing, pricing, benchmarks, or reproducible test conditions.

Why it matters: HKR-H and HKR-R pass: The Verge’s hands-on has a strong click hook and hits agent safety/competition nerves. HKR-K fails because pricing, release timing, and reproducible test conditions are missing, keeping it at the featured threshold.

Synced · WeChat

DataMaster: When AI Becomes Its Own Data Engineer

DataMaster searches, cleans, and combines data while keeping the model and training algorithm fixed; on MLE-Bench Lite, it raised the medal rate from 35.91% to 68.18%.

Why it matters: HKR-H/K/R all pass: DataMaster changes the data pipeline under fixed model and training code, lifting MLE-Bench Lite medal rate from 35.91% to 68.18%. This is still a single research release without production validation, so it lands at 78 featured.

AI HOT (Curated Pool)

To Avoid Paying $120, I Turned a Computer Cleaner into an Open-Source Skill

The author open-sourced a cross-platform AI cleaning skill for Mac and Windows, generating interactive HTML reports from file scans; in a test, it freed nearly 120GB, compared with CleanMyMac identifying 15.8GB.

Why it matters: This is not a platform-level release, but HKR-H/K/R all land through the $120 replacement hook, concrete scan/report mechanism, and 120GB test result. It fits the practical open-source tool band near the featured threshold.

Xinzhiyuan · WeChat

CAS Opens MobileGym, a Browser-Based Agent Training Environment for Mobile Apps

CASIA released MobileGym, a browser-based Android simulation environment covering 28 apps, with about 400MB per instance, 3-second cold start, JSON state snapshots, and programmatic task verification for mobile-agent training and evaluation.

Why it matters: MobileGym is practical open-source infrastructure for agent training and evaluation, with enough concrete numbers and mechanisms to pass HKR-H/K/R. It fits the 78–84 quality band, below major lab model-release weight.

AI HOT (Curated Pool)

Anthropic Developer Shares a Claude Code Understanding-Verification Workflow

An Anthropic developer shared a Claude Code understanding-verification workflow with 8 steps, using incremental teaching, user restatement, checklists, and quizzes to confirm the human can defend the problem, solution, and impact before moving to the next stage.

Why it matters: HKR-H/K/R all pass: a concrete Claude Code workflow with an 8-step verification loop and a strong oversight hook. It is a practical tutorial, not a product release, so it sits at the lower featured band.

Computing Life · Share · Yage

AI agents don't need to be hacked; persuasion is enough

The article says an AI agent with password-reset permission can be abused when an attacker persuades it they are a legitimate user; the snippet only discloses a three-layer architecture that separates what from who, not concrete attack steps or implementation details.

Why it matters: HKR-H/K/R all pass: the hook is strong, the post offers a what/who three-layer design, and agent permissions are a live security worry. No real incident, success rate, or product comparison keeps it at the featured threshold.

r/LocalLLaMA

I spent months inside verl, forked it, then stopped: internals, fork costs, and an NCCL bug

ReinforcedKnowledge analyzes ByteDance’s verl RLHF loop, covering DataProto plus rollout, reward, advantage, and update paths. The author stopped a private fork because near-daily upstream changes made sync cost exceed refactoring work, and describes an NCCL hang fixed on one node by setting NCCL_SOCKET_IFNAME=lo.

Why it matters: Niche but useful RL post-training field report, not an industry release. HKR-H comes from the fork-then-quit twist; HKR-K has verl’s five paths and NCCL_SOCKET_IFNAME=lo; HKR-R hits the cost of maintaining open-source training forks.

AI HOT (Curated Pool)

Google AI Studio adds app-building support for Gmail and other apps

Google AI Studio has added app-building support for connected Gmail, Drive, and Sheets apps, and users can add testers inside AI Studio; the post does not disclose a launch date for full public sharing.

Why it matters: HKR-H/K/R all pass, but this is a mid-weight product update: Workspace connections and tester support are confirmed, while sharing, permission details, and pricing are not disclosed.

AI HOT (Curated Pool)

Meta AI Exploit Used to Hijack Instagram Accounts

Meta’s AI chatbot was found vulnerable to an account-takeover exploit against Instagram accounts. Attackers could ask the AI to link a new email address, and the failure condition was the agent’s ability to execute account-management actions directly; the RSS snippet does not disclose affected account counts, patch status, or reproduction details.

Why it matters: HKR-H/K/R all pass: a Meta AI support agent allegedly enabled Instagram account takeover via add-email requests. Impact scale, fix timeline, and reproducible steps are not disclosed, so it stays in the 78–84 band.

AI HOT (Curated Pool)

Perplexity Releases Search as Code Architecture

Perplexity released Search as Code, an architecture where agents write Python code to call its search stack directly instead of looping through function calls; it is now available in the Perplexity Agent API and is the default option for Computer.

Why it matters: HKR-H/K/R pass: Perplexity gives a concrete agent-search mechanism and Agent API integration. Single-source post lacks performance, pricing, and rollout scope, so this stays a low featured product update.

Jun 1Monday

The Verge · AI

AI is blowing up music. How should the Grammys handle it?

Deezer reports that more than 50,000 AI-generated songs are uploaded each day, while Recording Academy CEO Harvey Mason Jr. says AI is now present in every recent music session he has attended and Grammy rules still bar AI music from the industry’s highest honors.

Why it matters: HKR-H/K/R all pass, but this is a podcast-style policy discussion rather than a model, product, or binding regulation story. The concrete signal is the 50,000/day Deezer figure plus the Grammy eligibility conflict.

AI HOT (Curated Pool)

Wang Xing: Meituan AI Agent Xiao Mei to Partner Deeply with Tencent Yuanbao

Meituan CEO Wang Xing said Xiao Mei will connect with Tencent Yuanbao, routing local service requests into food ordering, delivery, and related Meituan scenarios; Meituan reported Q1 2026 revenue of RMB 91.039 billion and a net loss of RMB 6.827 billion.

Why it matters: HKR-H/K/R all pass, but the deal is still “coming soon”; launch timing, UX entry point, and revenue split are not disclosed. This fits a mid-weight product partnership at the featured floor.

AI HOT (Curated Pool)

Tutorial: Turning Books into AI Skills with Claude Opus 4.8

The author used Claude Opus 4.8 to turn Nonviolent Communication into an AI Skill through a six-step workflow, taking about 45 minutes, using roughly 300,000 tokens, and costing under RMB 20.

Why it matters: HKR-H/K/R all pass: this is a numbered first-person Claude workflow with concrete cost and token details. It stays in the lower featured band because it is a personal tutorial, not an Anthropic release or model update.

AI HOT (Curated Pool)

Apache RocketMQ Releases an AI-Focused Messaging Engine

Apache RocketMQ released RocketMQ for AI, a messaging engine for long-running sessions, multi-agent workflows, and fair scheduling, with Lite-Topics, ordered messages, and traffic shaping; the post does not disclose a version number or performance figures.

Why it matters: HKR-H/K/R pass: the AI-specific RocketMQ angle has a real agent-infra hook and named mechanisms. Score stays in the 72–77 band because version, benchmarks, and production cases are not disclosed.