Skip to content

#MCP/工具调用

5 today

May 27Wednesday

AI HOT (Curated Pool)

Reachy Mini enables fully local voice interaction

Reachy Mini implements local voice interaction through the speech-to-speech library, using a cascaded pipeline with a Realtime API-compatible WebSocket interface and default components including Silero VAD, Parakeet-TDT, and Qwen3-TTS.

Why it matters: HKR-H/K/R all pass: the post has a clear local-robot voice hook, concrete stack details, and edge-agent resonance. Scope stays limited to Reachy Mini voice interaction, so it sits at the featured threshold.

AI HOT (Curated Pool)

Shipping a Trillion Parameters With a Hub Bucket: Delta Weight Sync in TRL

Hugging Face merged TRL PR 5417 for delta weight sync, sending only changed weights as sparse safetensors via a Hugging Face Bucket; on Qwen3-0.6B, the per-step payload falls from 1.2GB to 20–35MB.

Why it matters: HKR-H/K/R all pass: TRL gets delta weight sync with a concrete sparse-safetensors mechanism and a 1.2GB to 20–35MB example. Scope is training infra, so it stays below must-write.

Computing Life · Yage

Using AI Better, Step Two: Write the Skill Before Execution

The author proposes writing a Skill before asking AI to execute a task; each Skill should include three elements—success criteria, observed pitfalls, and deterministic tools—and can be organized through index.md plus AGENTS.md or CLAUDE.md for reuse.

Why it matters: HKR-H/K/R pass via a concrete Skill-first workflow and reusable agent practice. No model release, product capability, or experiment numbers, so it sits at the featured threshold.

Computing Life · Yage

Step Two to Using AI Well: Write the Skill Before You Execute

Yage argues that users should externalize work before execution by writing reusable Skills for Claude Code, Codex, and Cursor. The post gives an Outlook email example: spend about 30 minutes documenting username, phone approval, and client choice, then have AI read that file on later runs.

Why it matters: HKR-H/K/R all pass, but this is a workflow tutorial rather than a product or model release. The concrete Skill mechanism and Outlook example clear the featured floor; weak source authority keeps it at 72.

AI HOT (Curated Pool)

How we contain Claude across different products

Anthropic describes three mechanisms for containing Claude agent deployment risks across products: sandboxing or VMs, network egress controls, system-prompt and training constraints, and fine-grained permissions for MCP servers and third-party plugins.

Why it matters: Anthropic discloses a concrete containment stack for Claude agents, stronger than a routine product note. HKR-H/K/R all pass, but this is not a model launch or major capability release, so it stays in the 78–84 band.

May 26Tuesday

r/LocalLLaMA

[OSS] dlmserve: First Serving Engine for Diffusion Language Models

dlmserve released an MIT-licensed serving engine for diffusion language models, with LLaDA-8B-Instruct support and 2.5x HF throughput at batch=4. It exposes an OpenAI-compatible /v1/chat/completions API, batches at the denoising-step level, runs in 12GB VRAM, and adds about 1.8x throughput with optional LocalLeap acceleration.

Why it matters: HKR-H/K/R all pass: an open-source DLM serving engine with concrete throughput and VRAM claims. Single Reddit source and an early ecosystem keep it in low featured, not 78+.

AI HOT (Curated Pool)

Sundar Pichai on AI, the Future of Search, and Changes to the Web

Sundar Pichai said after Google I/O that Google is integrating Gemini into a new smart search box and the Gemini Spark agent platform; the post does not disclose model parameters, launch dates, or traffic impact numbers.

Why it matters: HKR-H and HKR-R pass: Pichai’s interview touches Google Search as an AI entry point and web traffic allocation. HKR-K is weak because the article gives Gemini-in-Search and Spark, but no rollout timing or technical detail.

r/LocalLLaMA

SkillOpt treats markdown skill files as trainable parameters with proper optimization machinery

SkillOpt uses a frontier model to propose add, delete, and replace edits to markdown skill files, then accepts only strict gains on a held-out validation set; the best skills usually converge after 1 to 4 accepted edits.

Why it matters: HKR-H/K/R all pass: the hook is trainable markdown skills, with held-out validation and 1-4 accepted edits. Single Reddit/project source and no broad adoption data keep it at 78, featured not p1.

AI HOT (Curated Pool)

Qwen3.7-Max Becomes the World’s No. 2 AI Coding Model

Qwen3.7-Max scored 1541 on Code Arena and ranked behind Claude; the post says it can run 35-hour tasks and perform more than 1,000 tool calls.

Why it matters: HKR-H/K/R all pass, but the source is a single Alibaba Cloud post and the evidence is benchmark plus vendor claims. This fits a strong product/benchmark update, not P1 without independent validation.

Synced · WeChat

ACL 2026 Main: Spatial-Agent Generates Executable Geospatial Analysis Workflows for LLMs

Spatial-Agent inserts a GeoFlow Graph between natural-language questions and map tools, and Spatial-Agent with GPT-4o-mini reaches 45.15% accuracy on MapEval-API versus a 23.00% API baseline.

Why it matters: ACL Main gives a concrete mechanism and testable numbers, so HKR-H/K pass. The GIS focus limits HKR-R, placing it at the featured threshold rather than a must-write item.

Xinzhiyuan · WeChat

Chinese agent SkyClaw targets Opus 4.6-level performance with free trial

Kunlun Tech released SkyClaw-v1.0 and SkyClaw-v1.0-lite with a 2-4 week free trial, claiming SkyClaw-v1.0 input costs are 1/24 of DeepSeek V4 Pro and about 1/43 of Sonnet 4.6.

Why it matters: HKR-H/K/R all pass: SkyClaw-v1.0 has a sharp cost hook, concrete trial and pricing ratios, and budget resonance. Source facts remain vendor claims, so it stays at the low featured band.

AI HOT (Curated Pool)

OpenAI GPT-5.6 Reportedly Set for Next Month With 1.5M-Token Context

Developers found an unannounced OpenAI GPT-5.6 entry in Codex backend logs under the codename iris-alpha, with a 1.5 million-token context window, about 43% higher than GPT-5.5’s 1.05 million-token limit.

Why it matters: HKR-H/K/R all pass: the Codex-log leak, 1.5M-token window, and 43% increase are concrete and practitioner-relevant. It stays below 85 because this is not an official GPT-5.6 launch.

AI HOT (Curated Pool)

Grok Build Beta Opens to SuperGrok Users

xAI opened Grok Build Beta to all SuperGrok and X Premium+ users, with Plan Mode, Imagine-based image and video creation, and a CLI for automation or orchestrator workflows at x.ai/cli.

Why it matters: HKR-H/K/R all pass: xAI opened a paid beta with named workflow features. The score stays at the featured floor because the post lacks capability limits, pricing detail, and test results.

May 25Monday

r/LocalLLaMA

NuExtract3 released: open-weight 4B VLM for Markdown, OCR and structured extraction

Numind released NuExtract3, a 4B open-weight VLM based on Qwen3.5-4B under Apache-2.0, supporting image and text to Markdown, OCR, and JSON-template extraction, with self-hosting from 4GB VRAM and weights in Safetensors, GGUF, and MLX formats.

Why it matters: HKR-H/K/R all pass: NuExtract3 packages OCR, Markdown, and structured extraction into a 4B open-weight VLM with a 4GB self-hosting condition. Source and lab reach keep it in the low featured band.

r/LocalLLaMA

Computer-use sandbox framework for Codex on headless Linux

superSmitty9999 released ai-sandbox-manager as a PoC that uses LXC templates to give Codex sudo access, browser use, Docker, and shared GPU access, with a hook that blocks git push while the agent works inside isolated copies.

Why it matters: HKR-H/K/R all pass, but this is a Reddit personal PoC with mechanisms only, not adoption, benchmarks or maturity evidence. It fits the featured floor for practical agent-sandbox work.

AI HOT (Curated Pool)

Harness, Scaffold, and AI Agent Terminology Explained

Hugging Face’s post frames an agent as three layers: Model, Scaffolding, and Harness; Scaffolding defines behavior through prompts and tool descriptions, while Harness runs model calls, tool calls, and control loops.

Why it matters: HKR-H/K/R pass: the Hugging Face post gives a concrete agent-stack taxonomy. It clears featured on practitioner relevance, but lacks a release, benchmark, or deployment case, so it stays at the threshold.

May 24Sunday

r/LocalLLaMA

Using llama.cpp native tools for web RAG inside llama-server WebUI

A Reddit user describes using llama.cpp native tools for web RAG inside llama-server WebUI with a 7-step setup: enable get_datetime and exec_shell_command, then run wget through firejail, a separate Linux user, and an Alpine OCI VM sandbox.

Why it matters: HKR-H/K/R all pass: the post gives a concrete local web-RAG recipe with sandboxing. It is a community tutorial, not a model or product launch, so the narrow reach and source authority keep it at the low featured band.

Xinzhiyuan · WeChat

AI Agent Completes Chip Design from 219 Words to 7nm GDSII Without Engineer Input

Verkor’s Design Conductor generated an ASAP7 7nm GDSII layout for the VerCore RISC-V CPU from a 219-word English spec in 12 hours, with no engineer in the design loop; the reported result scored 3,261 CoreMark at 1.48GHz, but it has not been fabricated and lacks cache implementation.

Why it matters: HKR-H/K/R all pass, but VerCore is not taped out and lacks cache, so the claim stays at demo-and-benchmark level. Concrete numbers and test conditions put it in the 78–84 recommendation band.

r/LocalLLaMA

llama.cpp server has built-in native tools: exec_shell, edit_file, and more

llama.cpp server exposes an experimental --tools flag with 8 native tools, including file reads, grep search, shell execution, file edits, diffs, and datetime; the post says file operations are relative to the server launch directory and no command whitelist or strict sandbox is provided yet.

Why it matters: HKR-H/K/R all pass: llama.cpp adding native shell and file tools is a concrete agent-runtime shift with safety stakes. Reddit sourcing and experimental status keep it in the lower featured band.

May 23Saturday

AI HOT (Curated Pool)

v2.1.149 release summary

Claude Code v2.1.149 adds categorized /usage reporting, an enterprise allowAllClaudeAiMcps setting for cloud MCP connectors, and fixes three security issues involving PowerShell permission bypass, Git worktree sandbox allowlist overflow, and otelHeadersHelper failures when script paths contain spaces.

Why it matters: Official Claude Code point release with concrete changes but limited blast radius: /usage categories, an enterprise MCP allow switch, and PowerShell bypass fixes hit developer security and governance needs.

AI HOT (Curated Pool)

Claude Auto Mode Adds Pro Plan and Model Support

Claude Auto Mode is now available on the Pro plan and supports Sonnet 4.6 and Opus 4.7; users can start it with Shift+Tab, while the post does not disclose pricing changes or rollout scope.

Why it matters: HKR-H/K/R all pass: official Claude dev channel gives Pro access, two supported models, and a shortcut. This is a mid-weight Claude product update, not a major model or capability release.

r/LocalLLaMA

How small can the orchestration model in an agent be? Separating it from code generation

HomoAgens1 runs a local ReAct orchestration loop on Qwen3.6-35B-A3B, with about 3B active parameters, a 12GB GPU, 30 expert offload, and 40 tokens/s prompt generation; smaller dense models fail first on tool-call discipline, inventing arguments or repeating bad calls, while reasoning is not identified as the first break point.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit experiment rather than a formal release. The VRAM, speed, and failure-mode details put it at the 72 featured threshold.

AI HOT (Curated Pool)

Kakuna: An AI Agent Tool for Automated Codebase Hardening

Kakuna hardens prototype codebases with built-in checklists and a plan-goal workflow; one roughly 16-hour run can generate hundreds of commits while preserving functionality.

Why it matters: HKR-H/K/R all pass: the post has a 16-hour run, hundreds of commits, and a workflow mechanism tied to coding-agent pain. Single X source and a non-major vendor keep it at the featured threshold.

AI HOT (Curated Pool)

Google I/O Releases AI Agent Development Toolchain

Google announced an AI agent development and deployment toolchain at I/O, including Antigravity 2.0, managed agent services in the Gemini API, WebMCP in Chrome 149, and Chrome DevTools access for automated agent debugging.

Why it matters: HKR-H/K/R all pass: Google is shipping a named agent stack across tooling, managed services, WebMCP, and Chrome. Single-source social summary lacks pricing, API details, and demos, so it stays in the 78–84 band.

r/LocalLLaMA

Experts first llama.cpp

comanderxv published a llama.cpp fork that caches MoE experts in 12GB VRAM; on an RTX 2060 with Qwen3.6-35B-A3B, throughput rose from 19/22 tk/s to 26 tk/s at about a 62% expert-cache hit rate.

Why it matters: HKR-H/K/R all pass: the hook is a 35B MoE speedup on a 12GB RTX 2060, with concrete caching and hit-rate data. Scope stays niche to local inference, so it lands at the featured threshold rather than must-write.

May 22Friday

Hacker News front page

Launch HN: Superset (YC P26) – IDE for the agents era

Superset launched an open-source agentic IDE that runs coding agents such as Claude Code, Codex, and OpenCode in parallel through git worktrees, and the team added Remote Workspaces in beta for running agents on remote machines while managing work from the desktop app.

Why it matters: HKR-H/K/R all pass, but Superset is still a new YC launch and the post lacks usage, pricing, or performance data. The git-worktree agent workflow clears the featured bar, not the must-write band.

Mistral AI

Mistral launches Connectors in Studio with built-in and custom MCP

Mistral launched Connectors in Studio. All built-in connectors and custom MCP are now callable through the API/SDK by every model and agent. New features include direct tool calling, human-in-the-loop approval flows, and programmatic access to create, modify, list and delete connectors.

Why it matters: The original gives the API usage and code examples for Connectors, enough to judge how enterprise MCP integration gets built.

AI HOT (Curated Pool)

Karpathy’s CLAUDE.md Four Rules Raise AI Coding Accuracy to 94%

Karpathy published a 65-line CLAUDE.md with four rules that raised AI coding accuracy from 65% to 94%, and the file received over 220,000 GitHub stars.

Why it matters: HKR-H/K/R all pass: a notable name, a claimed accuracy jump, and a rules-based Claude Code workflow. It stays below 85 because the body only gives summary-level numbers; task set, evaluation method, and the four rules are not disclosed.

AI HOT (Curated Pool)

Alibaba Qianwen App, PC, and Web Add Qwen3.7-Max

Alibaba added Qwen3.7-Max to the Qianwen app, PC client, and web client, with free access after updating the app to version 6.9.7 or later, and the official test reports a 35-hour autonomous kernel optimization run with more than 1,000 tool calls.

Why it matters: HKR-H/K/R all pass: Alibaba ships Qwen3.7-Max across three Qianwen clients, with v6.9.7+ free access and a 35-hour, 1,000+ tool-call claim. Benchmarks, context window, and API pricing are not disclosed, so it stays below 90.

MIT Technology Review · AI

Google I/O showed how the path for AI-driven science is shifting

MIT Technology Review says Google used I/O to shift its scientific AI framing toward Gemini for Science, a package that groups AI Co-Scientist and AlphaEvolve, while researchers can now apply for access and older specialized systems like AlphaFold and WeatherNext remain active.

Why it matters: HKR-H and HKR-K pass: MIT Technology Review frames a real Google science-AI product shift with named components and access conditions. HKR-R is weak because the impact is mostly research-facing, not practitioner-wide.

Xinzhiyuan · WeChat

Microsoft, after investing $13B in OpenAI, saw its engineers run up Claude Code costs

Microsoft plans to end Claude Code subscriptions by the end of June for its Experiences and Devices teams and move nearly 100,000 engineers to GitHub Copilot CLI, with the article attributing the change to external token-based billing costs.

Why it matters: HKR-H/K/R all pass: the OpenAI-Claude contrast hooks, the story gives end-June migration, nearly 100k engineers and token-billing, and it hits enterprise coding-agent cost control. Not a model release or official major launch, so 78–84 fits.

Xinzhiyuan · WeChat

Enterprise Agent Operations Begin? Anthropic Updates Architecture, Chinese Tech Firms Have It Running

Alibaba Cloud JVS Crew splits Agent, Environment, and Session into three layers, with sandboxes, snapshot recovery, RBAC, and usage-based billing. Anthropic added self-hosted sandboxes to Claude Managed Agents on May 19, while the article cites 2-week deployments and 5x or 10x efficiency gains in several Chinese customer cases.

Why it matters: HKR-H/K/R all pass, but the facts are an enterprise agent-infra comparison: Anthropic self-hosted sandboxes and Alibaba Cloud JVS Crew architecture. This is featured-level, not a must-write model release.

AI HOT (Curated Pool)

OpenAI Codex /goal Feature Officially Launches with Usage Guide

OpenAI moved Codex /goal mode from experiment to stable release, letting users set milestones in the Codex app, IDE extension, or CLI and keep tasks running for hours or days with progress checks, direction changes, and pause controls.

Why it matters: HKR-H/K/R all pass: OpenAI Codex /goal is now stable, with milestones across app, IDE extension, and CLI. The article is thin on permissions, safety limits, and tier access, so it stays in the lower featured band.

Hacker News front page

Show HN: Spec-Driven Development Workflow for Claude Code

The sddw author released a Claude Code plugin that splits work into requirements, code analysis, and design specs, then clears context after each step to keep cost and context focused.

Why it matters: HKR-H/K/R all pass for a Claude Code workflow with a concrete spec-and-context mechanism. It stays in the 72–77 featured band because the post lacks benchmarks, adoption data, or an official Anthropic release.

AI HOT (Curated Pool)

Plastic Interfaces: The Future Shape of AI-Driven Software

Salesforce has adopted a headless architecture that lets salespeople update data through AI; the post says MCPs, HTML, audio, and web interfaces can be generated dynamically by context, but it does not disclose implementation metrics or adoption numbers.

Why it matters: HKR-H/K/R all pass, but this is a software-form thesis without user metrics, launch timing, or a reproducible test. It fits the insightful-commentary band, not a must-write release.

AI HOT (Curated Pool)

v2.1.147 Release Update

Claude Code v2.1.147 adds a Workflow tool, disabled by default, for deterministic multi-agent orchestration, and renames /simplify to /code-review with code-correctness reporting and GitHub PR inline-comment generation.

Why it matters: HKR-H/K/R all pass: the official Claude Code release adds a default-off Workflow tool for deterministic multi-agent orchestration. No performance data, pricing, or scope limits are disclosed, so this stays in the mid product-update band.

Latent Space

Giving Agents Computers — Ivan Burazin, Daytona

Daytona provides composable computers for AI agents, with one sandbox starting in about 60 ms, 50,000 sandboxes in about 75 seconds, and its largest customer running roughly 850,000 sandboxes per day.

Why it matters: HKR-H/K/R all pass: the agent-computer framing is clickable, and the sandbox scale numbers are concrete. Still, this is a startup infrastructure story, not a major model or platform release.

AI HOT (Curated Pool)

ChatGPT now supports creating and editing presentations directly in PowerPoint

ChatGPT is testing PowerPoint support for creating and editing presentations directly, including building, updating, understanding, and refining editable slides; the post does not disclose pricing, rollout scope, or availability conditions.

Why it matters: HKR-H/K/R all pass: OpenAI shows ChatGPT creating and editing editable PowerPoint slides. Pricing, rollout scope, and enterprise controls are not disclosed, so this stays featured rather than P1.

AI HOT (Curated Pool)

Datasette Agent

Datasette released Datasette Agent as its first extensible AI assistant, offering conversational data queries, plugin-based chart generation, official plugins for charts, AI image creation, and sandboxed code execution, with support for Gemini 3.1 Flash-Lite cloud models and local open-source models through LM Studio.

Why it matters: HKR-H/K/R all pass: a concrete Datasette agent with chart plugins and LM Studio local execution. The audience is narrower than major lab releases, so it sits in the 72–77 featured band.

AI HOT (Curated Pool)

Codex Enables Secure Cross-Device Mac Control Around the Clock

OpenAI Devs says Codex can use apps on a Mac from a phone while the Mac remains locked and the screen is off; the post does not disclose permission boundaries, pricing, or a release timeline.

Why it matters: HKR-H/K/R all pass: OpenAI Devs disclosed a concrete Codex Mac-control condition. Missing permission boundaries, pricing, and launch timing keep it below the 85+ band.