Skip to content

#MCP/工具调用

5 today

Jun 3Wednesday

AI HOT (Curated Pool)

Claude Platform Adds CLI Tool

Claude Platform added a CLI that runs every API endpoint from the terminal, calls the Messages API, launches Claude-hosted agents, and pipes results directly into the shell.

Why it matters: Claude Platform CLI clears HKR-H/K/R as a practical developer-tooling update, but the post only gives capability scope; install flow, permissions, safety limits, and pricing are not disclosed.

TechCrunch · AI

Microsoft Offers Developers a Better Way to Control AI Agent Behavior

Microsoft released an agent policy specification that lets developer, compliance, and security teams define behavior rules in portable policy files; the post does not disclose the version, license, supported frameworks, or rollout timeline.

Why it matters: HKR-H/K/R pass: the portable-policy mechanism is concrete and the safety/compliance nerve is real for agent builders. Missing version, license, and framework support keeps it at the featured threshold, not a same-day must-write.

AI HOT (Curated Pool)

Microsoft Scout: A New OpenClaw-Based AI Personal Assistant

Microsoft launched Microsoft Scout, an OpenClaw-based personal assistant that can run persistently inside Outlook, OneDrive, and Teams, and enterprises can assign it to employees for calendar management, expense processing, and email drafting.

Why it matters: HKR-H/K/R all pass, but the body is thin: it gives integrations and task scope, not pricing, launch timing, or technical depth. Treat it as a Microsoft workplace-agent product update at the low featured band.

r/LocalLLaMA

Using Gemma 4 E4B with LiteRT: about 2.4× faster text generation than Q4 GGUF

The author tested Gemma 4 E4B on an RTX 4060 Ti 16GB, where LiteRT averaged 157.2 tok/s for text generation versus 66.3 tok/s for llama.cpp Q4 GGUF; image captioning on 111 full-resolution images improved only 1.1×, at about 72 seconds versus 80 seconds.

Why it matters: HKR-H/K/R all pass, with a first-person benchmark including hardware, throughput, and sample count. Source authority is limited to one Reddit test, so it sits at the featured threshold rather than the 78+ band.

Bloomberg Technology

Uber Caps Usage of AI Tools Like Claude Code to Manage Costs

Uber Technologies set usage caps on staff AI tools including Claude Code after the company exceeded its AI budget earlier this year; the post does not disclose the cap size, affected teams, or budget amount.

Why it matters: HKR-H/K/R all pass: the Bloomberg item gives a named enterprise cost-control case for Claude Code-like tools. Budget size, cap rules, and affected headcount are not disclosed, keeping it at the featured threshold.

Latent Space

GitHub's Plan for Agents — Kyle Daigle, GitHub

GitHub COO Kyle Daigle said AI-driven code commits grew 14x in 2026, and the interview covers Copilot, Actions, MCP, WorkIQ, cloud agents, and the infrastructure availability pressure created when code review, CI/CD, and open-source contribution volume scale beyond human-speed workflows.

Why it matters: HKR-H/K/R all pass: a GitHub executive gives a 14x AI code-submission figure and ties Copilot, Actions, MCP, WorkIQ, and cloud agents into one roadmap. Not a major release, so it stays at 80.

AI HOT (Curated Pool)

Claude Code Team Practice: How Agentic Coding Changes Engineering Organizations and Processes

The Claude Code engineering team described process changes after making agentic coding the default at Code w/ Claude SF 2026: JIT planning, asking Claude first for context collection, Claude handling style and tests in code review, and humans focusing on legal and safety judgments.

Why it matters: First-party Claude Code workflow post with concrete engineering mechanisms and strong HKR-H/K/R fit. It is not a model or major product release, so it stays in the 78–84 band.

AI HOT (Curated Pool)

OpenAI Codex releases Python SDK for direct app integration

OpenAI Codex released a Python SDK with the install command pip install openai-codex, and the snippet says it can reuse the Codex login state; the post does not disclose API pricing, model versions, or rate-limit conditions.

Why it matters: HKR-H/K/R pass: a Codex SDK for embedded app use is practical and discussable. Sparse sourcing keeps it in the mid-weight product-update band: package and auth are given, but price, model, and rate limits are not.

AI HOT (Curated Pool)

OpenAI Codex Sites feature launches

OpenAI launched Codex Sites, which turns work, ideas, and plans into an interactive website or app that a team can access through one URL; the feature rolls out first to Business and Enterprise plans, and the post does not disclose pricing or broader availability timing.

Why it matters: HKR-H/K/R all pass, but the post gives launch framing without pricing, permission boundaries, or quality examples. Treat it as a mid-weight OpenAI product feature, above the featured threshold.

r/LocalLLaMA

Benchmarks of 20 Small LLMs on a 6GB RTX 4050

The author benchmarked 20 small LLMs on a 6GB RTX 4050 using LM Studio’s OpenAI-compatible API, with N=5 speed runs at 1k, 8k, and 32k context; unsloth/lfm2.5-vl-1.6b led throughput at 207 tok/s on 1k context while using 3.0GB VRAM.

Why it matters: HKR-H/K/R all pass: the low-VRAM GPU hook is concrete, the post gives speed/context/VRAM numbers, and it speaks to local-inference cost pressure. Source authority is a Reddit post, so it stays in the lower featured band.

TechCrunch · AI

OpenAI launches new Codex tools for white-collar work

OpenAI released six Codex app plug-ins for data analytics, creative production, sales, product design, equity investing, and investment banking; each tool bundles integrations, instructions, and context, while the post does not disclose pricing or rollout limits.

Why it matters: HKR-H/K/R all pass: OpenAI is expanding Codex into six white-collar plugin categories. Pricing, rollout scope, and measured performance are not disclosed, so this stays in the mid-weight product-update band.

Jun 2Tuesday

AI HOT (Curated Pool)

Holo3.1: Fast Local Computer-Use Agents

Holo3.1 releases Qwen-based computer-use agents in 0.8B, 4B, 9B, and 35B-A3B sizes, with FP8, Q4 GGUF, and NVFP4 quantized checkpoints for local inference and a 79.3% AndroidWorld score for the 35B-A3B model.

Why it matters: HKR-H/K/R all pass: Holo3.1 pairs a local computer-use agent with concrete model sizes and quantized checkpoints. It fits the 78–84 band, below major lab model-release weight.

AI HOT (Curated Pool)

Anthropic Expands Project Glasswing Program

Anthropic expanded Project Glasswing to about 150 new organizations across more than 15 countries, covering electricity, water, healthcare, communications, and hardware infrastructure, after an initial group of about 50 partners.

Why it matters: Anthropic expanded Project Glasswing to about 150 new organizations across 15+ countries, giving HKR-H/K/R enough substance. No concrete safety mechanism or Claude capability change is disclosed, so it stays in the lower featured band.

AI HOT (Curated Pool)

StepFun releases Step 3.7 Flash as an open-weight model for agentic coding

StepFun released the open-weight Step 3.7 Flash model for fast agentic coding, with tool calling and multimodal understanding, and the model is already available in Kilo alongside MiniMax M3.

Why it matters: HKR-H/K/R pass on the open-weight agentic-coding angle and Kilo availability. Missing benchmarks, size, license, and pricing keep it at the lower featured threshold.

The Verge · AI

Gemini Spark is the most impressive and terrifying AI experience I’ve had yet

The Verge tested Google’s always-on AI agent Gemini Spark for trip planning, but the RSS snippet only describes a different experience from generic itinerary demos and does not disclose launch timing, pricing, benchmarks, or reproducible test conditions.

Why it matters: HKR-H and HKR-R pass: The Verge’s hands-on has a strong click hook and hits agent safety/competition nerves. HKR-K fails because pricing, release timing, and reproducible test conditions are missing, keeping it at the featured threshold.

Synced · WeChat

DataMaster: When AI Becomes Its Own Data Engineer

DataMaster searches, cleans, and combines data while keeping the model and training algorithm fixed; on MLE-Bench Lite, it raised the medal rate from 35.91% to 68.18%.

Why it matters: HKR-H/K/R all pass: DataMaster changes the data pipeline under fixed model and training code, lifting MLE-Bench Lite medal rate from 35.91% to 68.18%. This is still a single research release without production validation, so it lands at 78 featured.

AI HOT (Curated Pool)

To Avoid Paying $120, I Turned a Computer Cleaner into an Open-Source Skill

The author open-sourced a cross-platform AI cleaning skill for Mac and Windows, generating interactive HTML reports from file scans; in a test, it freed nearly 120GB, compared with CleanMyMac identifying 15.8GB.

Why it matters: This is not a platform-level release, but HKR-H/K/R all land through the $120 replacement hook, concrete scan/report mechanism, and 120GB test result. It fits the practical open-source tool band near the featured threshold.

Xinzhiyuan · WeChat

CAS Opens MobileGym, a Browser-Based Agent Training Environment for Mobile Apps

CASIA released MobileGym, a browser-based Android simulation environment covering 28 apps, with about 400MB per instance, 3-second cold start, JSON state snapshots, and programmatic task verification for mobile-agent training and evaluation.

Why it matters: MobileGym is practical open-source infrastructure for agent training and evaluation, with enough concrete numbers and mechanisms to pass HKR-H/K/R. It fits the 78–84 quality band, below major lab model-release weight.

AI HOT (Curated Pool)

Anthropic Developer Shares a Claude Code Understanding-Verification Workflow

An Anthropic developer shared a Claude Code understanding-verification workflow with 8 steps, using incremental teaching, user restatement, checklists, and quizzes to confirm the human can defend the problem, solution, and impact before moving to the next stage.

Why it matters: HKR-H/K/R all pass: a concrete Claude Code workflow with an 8-step verification loop and a strong oversight hook. It is a practical tutorial, not a product release, so it sits at the lower featured band.

Computing Life · Share · Yage

AI agents don't need to be hacked; persuasion is enough

The article says an AI agent with password-reset permission can be abused when an attacker persuades it they are a legitimate user; the snippet only discloses a three-layer architecture that separates what from who, not concrete attack steps or implementation details.

Why it matters: HKR-H/K/R all pass: the hook is strong, the post offers a what/who three-layer design, and agent permissions are a live security worry. No real incident, success rate, or product comparison keeps it at the featured threshold.

r/LocalLLaMA

I spent months inside verl, forked it, then stopped: internals, fork costs, and an NCCL bug

ReinforcedKnowledge analyzes ByteDance’s verl RLHF loop, covering DataProto plus rollout, reward, advantage, and update paths. The author stopped a private fork because near-daily upstream changes made sync cost exceed refactoring work, and describes an NCCL hang fixed on one node by setting NCCL_SOCKET_IFNAME=lo.

Why it matters: Niche but useful RL post-training field report, not an industry release. HKR-H comes from the fork-then-quit twist; HKR-K has verl’s five paths and NCCL_SOCKET_IFNAME=lo; HKR-R hits the cost of maintaining open-source training forks.

AI HOT (Curated Pool)

Google AI Studio adds app-building support for Gmail and other apps

Google AI Studio has added app-building support for connected Gmail, Drive, and Sheets apps, and users can add testers inside AI Studio; the post does not disclose a launch date for full public sharing.

Why it matters: HKR-H/K/R all pass, but this is a mid-weight product update: Workspace connections and tester support are confirmed, while sharing, permission details, and pricing are not disclosed.

AI HOT (Curated Pool)

Meta AI Exploit Used to Hijack Instagram Accounts

Meta’s AI chatbot was found vulnerable to an account-takeover exploit against Instagram accounts. Attackers could ask the AI to link a new email address, and the failure condition was the agent’s ability to execute account-management actions directly; the RSS snippet does not disclose affected account counts, patch status, or reproduction details.

Why it matters: HKR-H/K/R all pass: a Meta AI support agent allegedly enabled Instagram account takeover via add-email requests. Impact scale, fix timeline, and reproducible steps are not disclosed, so it stays in the 78–84 band.

AI HOT (Curated Pool)

Perplexity Releases Search as Code Architecture

Perplexity released Search as Code, an architecture where agents write Python code to call its search stack directly instead of looping through function calls; it is now available in the Perplexity Agent API and is the default option for Computer.

Why it matters: HKR-H/K/R pass: Perplexity gives a concrete agent-search mechanism and Agent API integration. Single-source post lacks performance, pricing, and rollout scope, so this stays a low featured product update.

Jun 1Monday

The Verge · AI

AI is blowing up music. How should the Grammys handle it?

Deezer reports that more than 50,000 AI-generated songs are uploaded each day, while Recording Academy CEO Harvey Mason Jr. says AI is now present in every recent music session he has attended and Grammy rules still bar AI music from the industry’s highest honors.

Why it matters: HKR-H/K/R all pass, but this is a podcast-style policy discussion rather than a model, product, or binding regulation story. The concrete signal is the 50,000/day Deezer figure plus the Grammy eligibility conflict.

AI HOT (Curated Pool)

Wang Xing: Meituan AI Agent Xiao Mei to Partner Deeply with Tencent Yuanbao

Meituan CEO Wang Xing said Xiao Mei will connect with Tencent Yuanbao, routing local service requests into food ordering, delivery, and related Meituan scenarios; Meituan reported Q1 2026 revenue of RMB 91.039 billion and a net loss of RMB 6.827 billion.

Why it matters: HKR-H/K/R all pass, but the deal is still “coming soon”; launch timing, UX entry point, and revenue split are not disclosed. This fits a mid-weight product partnership at the featured floor.

AI HOT (Curated Pool)

Tutorial: Turning Books into AI Skills with Claude Opus 4.8

The author used Claude Opus 4.8 to turn Nonviolent Communication into an AI Skill through a six-step workflow, taking about 45 minutes, using roughly 300,000 tokens, and costing under RMB 20.

Why it matters: HKR-H/K/R all pass: this is a numbered first-person Claude workflow with concrete cost and token details. It stays in the lower featured band because it is a personal tutorial, not an Anthropic release or model update.

AI HOT (Curated Pool)

Apache RocketMQ Releases an AI-Focused Messaging Engine

Apache RocketMQ released RocketMQ for AI, a messaging engine for long-running sessions, multi-agent workflows, and fair scheduling, with Lite-Topics, ordered messages, and traffic shaping; the post does not disclose a version number or performance figures.

Why it matters: HKR-H/K/R pass: the AI-specific RocketMQ angle has a real agent-infra hook and named mechanisms. Score stays in the 72–77 band because version, benchmarks, and production cases are not disclosed.

Xinzhiyuan · WeChat

400 tokens/s: StepFun Step 3.7 Flash cuts Agent task costs

StepFun released Step 3.7 Flash, a sparse MoE model with 196B parameters plus a 1.8B ViT, activating 11B parameters per inference and reaching up to 400 tokens per second.

Why it matters: HKR-H/K/R all pass with concrete speed and parameter numbers. The feed does not disclose pricing, benchmark setup, or open-source terms, so this stays in the 78–84 quality update band.

QbitAI · WeChat

How Cloud Models Reach the Physical World: CMG Lion Rock AI Lab Uses LiOS for Embodied AI

CMG Lion Rock AI Lab released the LiOS edge-cloud architecture for embodied robotics, reporting about 30 ms one-way latency from local camera to cloud GPU memory in cross-machine tests, and open-sourced the low-latency video transmission module plus the LeFold laundry-folding dataset.

Why it matters: HKR-H/K/R pass: LiOS offers a concrete latency claim and open artifacts for embodied AI. Impact stays mid-tier because the lab is not a top platform vendor and no cross-source cluster is shown.

r/LocalLLaMA

Deepseek V4 Flash performance on DGX Spark

A Reddit user ran DeepSeek-V4-Flash with vLLM on two ASUS GX10 DGX Spark nodes and reported 1,680 prefill tokens/s plus 39.8 decode tokens/s at a 256K context with MTP=2; the setup uses TP=2 over RoCE, fp8 KV cache, and fits about 1M tokens safely in KV cache.

Why it matters: This is not broad industry news, but it is a first-person benchmark with reproducible details: TP=2, RoCE, fp8 KV cache, 256K context, and ~1M KV. HKR-H/K/R all pass, so it lands at low featured.

AI HOT (Curated Pool)

NVIDIA Releases FOX Factory Operations Blueprint for Autonomous Factory Management Agents

NVIDIA released the FOX factory operations blueprint at GTC Taipei, and Foxconn used it to build the MoMClaw multi-agent system with an expected 80% reduction in root-cause analysis time.

Why it matters: HKR-H/K/R pass: NVIDIA is pushing an agent blueprint into factory ops, with Foxconn’s MoMClaw and an expected 80% RCA time cut. Kept at the featured floor because the source is a vendor blog and the result is projected.

AI HOT (Curated Pool)

NVIDIA Releases RTX Spark and Local AI Agent Security and Performance Updates

NVIDIA released RTX Spark, a Windows PC for local AI agents with 1 petaflops of AI compute and 128GB of unified memory. OpenShell uses new Windows security primitives with Microsoft, while llama.cpp optimizations raise Qwen 27B throughput by up to 2x.

Why it matters: HKR-H/K/R all pass: NVIDIA frames RTX Spark for local agents and gives hard specs: 1 petaflops, 128GB, and up to 2x llama.cpp throughput. Vendor-blog framing keeps it in the low 78–84 band.

AI HOT (Curated Pool)

Qwen3.7-Plus: Multimodal Agent Intelligence

Qwen Studio lists seven capability areas: chatbots, image and video understanding, image generation, document processing, web search integration, tool use, and artifact generation; the post does not disclose Qwen3.7-Plus parameters, pricing, or release timing.

Why it matters: HKR-H/K/R pass, but the facts are thin: 7 capability categories, no params, pricing, benchmarks, or launch terms. A Qwen flagship update clears featured, not p1.

r/LocalLLaMA

I ported NVIDIA Parakeet speech-to-text to ggml: same output as NeMo, faster, GGUF-quantized, no Python

mudler_it ported NVIDIA Parakeet speech-to-text models to C++/ggml with no Python or PyTorch, reporting byte-for-byte NeMo parity on f32/f16, up to about 5x GPU speedups on larger TDT and hybrid models, and GGUF quantization across f16, q8_0, q6_k, q5_k, and q4_k.

Why it matters: HKR-H/K/R all pass: the port has a concrete local-inference hook, byte-parity and speed claims, and clear practitioner resonance. Source scope keeps it at the low featured band, not P1.

May 31Sunday

AI HOT (Curated Pool)

Apple WWDC AI Upgrade: Gemini-Distilled Model Runs Locally, With Heavy External Dependencies

Apple will present Siri and on-device AI upgrades at next month’s WWDC, with iPhones running a smaller Gemini-distilled model locally while complex queries route to Google Cloud using Nvidia confidential computing.

Why it matters: HKR-H/K/R all pass: the Apple-Google-Nvidia stack is a strong WWDC AI hook with a concrete routing mechanism and clear industry tension. Capped at 82 because this is a single X-sourced claim with no model size, latency, pricing, or contract terms disclosed.

r/LocalLLaMA

Use any model and provider with the official OpenAI Codex Desktop App without modifying its code

Reddit user thibautrey describes a 3-step setup: edit Codex Desktop config.toml, store an API key, and use a multicodex proxy alias to map gpt-5.3-codex to MiniMax-Latest. The post lists a local base_url of 127.0.0.1:1455 and says the proxy disguises returned model names as gpt-5.3-codex.

Why it matters: This is a reproducible developer workflow trick, not an official release. HKR-H comes from the lock-in workaround, HKR-K has concrete config details, and HKR-R hits cost and model-choice pressure, placing it at the tutorial featured threshold.

Synced · WeChat

Microsoft open-sources SkillOpt for training Agent skill documents, reaching 3.3k stars in a week

Microsoft open-sourced SkillOpt, a text-space optimization framework that trains Agent skill documents without changing model weights; the paper reports best or tied-best results across 52 combinations covering 7 target models, 6 benchmarks, and 3 execution environments.

Why it matters: Microsoft’s open-source SkillOpt is a strong Agent tooling and research release. HKR-H has the 3.3k-star/trainable-skill hook, HKR-K has the text-parameter mechanism and 52 eval setups, and HKR-R hits agent engineering pain, so it lands in featured at 82.

QbitAI · WeChat

Fudan and Tongyi introduce ToolCUA for GUI-Tool path selection in agents

Fudan University and Tongyi Lab introduced ToolCUA-8B, which reaches 46.85% accuracy on OSWorld-MCP after training with about 4k synthetic tools and 180k interleaved GUI-Tool trajectory steps.

Why it matters: HKR-H/K/R all pass: the tool-selection failure hook is concrete, with OSWorld-MCP 46.85% and 180k steps. It stays in the 78–84 band because this is a research release, not a major model or product launch.

AI HOT (Curated Pool)

Run Python ASGI Apps in the Browser with Pyodide and Service Workers

Simon Willison demonstrated running Python ASGI apps in the browser with Pyodide and Service Workers, with Claude Opus 4.8 assisting development, and showed two working demos: a basic ASGI FastCGI demo and Datasette 1.0a31.

Why it matters: HKR-H/K/R all pass: the post has a surprising browser-runtime hook, concrete mechanisms, and developer resonance. Impact stays in the 72–77 band because this is a developer experiment, not a model or platform launch.