Skip to content

#MCP/工具调用

4 today

Today · Sep 30Wednesday · 4 items

The Decoder

OpenAI's DevDay updates push ChatGPT from chatbot toward work platform

At DevDay, OpenAI announced a set of ChatGPT updates: an open Plugin Extensions system, shared workspace Space, collaborative Pages and Slides, Slack and Microsoft Teams integrations, automated workflows, and an enterprise marketplace.

Why it matters: It lays out the full set of updates moving ChatGPT from chatbot to work platform, a basis for judging its rivalry with Slack, Notion and similar tools.

TechCrunch · AI

AI-powered app maker Wabi pivots to a messaging experience

AI 应用构建平台 Wabi 本周宣布转型为 AI 即时通讯工具,推出 Wabi 2.0,定位为"为你做事并即时构建所需界面的个人智能体"。用户可在对话中按需生成卡路里追踪、健身记录、家庭日历等应用界面,目前仅通过 X 上发放的邀请码开放使用。

TechCrunch · AI

OpenAI expands ChatGPT’s plugins with app-like interfaces and automations

OpenAI 在 Dev Day 上宣布扩展 ChatGPT 插件,允许开发者在 ChatGPT 侧边栏中构建类应用体验,并提供可交互面板和文件查看器。开发者获得新的 Plugin Creator 工具,可通过重新设计的提交流程将插件提交到插件目录,插件在目录和对话中的排序与推荐方式也已改进。

The Decoder

OpenAI launches always-on agent Dots at DevDay 2026, plus GPT-6.1 Sol

At its DevDay 2026 developer conference, OpenAI launched the always-on agent Dots, which can handle tasks on its own such as fixing bugs reported in Slack or sending forgotten invoices. It also released the cheaper model GPT-6.1 Sol; the high-end GPT-6.1 Astra was held back over safety concerns.

Why it matters: The original details Dots' always-on cloud computer, proactive research and permission boundaries, a basis for judging how always-on agents will actually land.

Yesterday · Sep 29Tuesday

Computing Life · Share · Yage

Amazon 能封住替你购物的 AI 吗?

Amazon 在商城中封锁了 Meta 的 Muse,理由是未经同意的 AI 程序反复访问网站、违反服务条款,而 Meta 此前已在官方安全文档中说明密码进入隔离存储、主模型看不到明文。

Sep 22Tuesday

Simon Willison

llm-typesafe 0.1a0

Simon Willison 发布 LLM 插件 llm-typesafe 0.1a0,为 TypeSafe AI 的新模型 Jev 提供支持,可通过 `llm install llm-typesafe` 安装并用 `llm keys set typesafe` 配置 API key。

Sep 21Monday

Simon Willison

MCP was always a bad idea?

Simon Willison 反驳「MCP 一直是坏主意」的观点,认为该文忽略了 MCP 当下的价值:若运行 Claude Code、Codex、Meta Muse、OpenClaw 等拥有无限制互联网访问的完整终端智能体,确实几乎无需 MCP,直接调用 API 即可。

Sep 18Friday

GitHub Blog · AI & ML

Should you read the code, is RAG dead, and did Skills kill MCP?

GitHub Podcast 最新一期拆解了五个 AI 热门观点:AI 生成的代码仍需阅读和负责,但审查力度应按风险分级;Skills 与 MCP 解决不同问题,前者是打包的团队经验,后者是连接工具与数据的标准,可组合使用;RAG 并未死亡,它为模型提供训练数据之外的相关信息,减少 token 浪费并让回答更有依据。

Jun 10Wednesday

AI HOT (Curated Pool)

Google Gemini 3.5 Live Translate enters public preview with 70+ languages

Google released Gemini 3.5 Live Translate in public preview through the Gemini API, offering low-latency speech-to-speech translation across 70+ languages and 2,000 language pairs.

Why it matters: HKR-H/K/R all pass: Google’s speech-to-speech translation API has a clear developer hook and concrete scale numbers. Single X-source detail and missing price, latency benchmarks, and regions keep it at 78.

AI HOT (Curated Pool)

Claude Managed Agents adds scheduled runs and environment variable storage

Claude Managed Agents added cron-based scheduled runs and vaults environment variable storage in public beta, with real secrets attached only at the network boundary so agents cannot read them directly.

Why it matters: HKR-H/K/R all pass: this first-party Claude update adds concrete agent-ops mechanics with cron scheduling and vault-bound secrets. It is not a model release, so it stays in the lower good-quality band.

AI HOT (Curated Pool)

Claude Code team member Thariq shares 10 tips for improving Claude Code efficiency

Thariq shared 10 Claude Code tips that shift review from checking outputs to steering the right task, with concrete practices including full upfront context, /goal, Workflows for parallel tasks, self-checking, and comparison reports.

Why it matters: This is a strong Claude Code workflow tutorial, with concrete tactics around task calibration, /goal, and Workflows self-checks. It lands in the 72–77 tutorial band; the insider source and all three HKR hits justify featured.

AI HOT (Curated Pool)

OpenRouter Launches Advisor Tool for Low-Cost Models to Consult Stronger Models

OpenRouter released the Advisor server tool, letting GPT-4o Mini consult Claude Fable during generation, but the post does not disclose pricing, latency, or the routing policy.

Why it matters: HKR-H/K/R all pass: OpenRouter turns cheap-model plus strong-model advising into a callable server tool. Price, latency, and call policy are not disclosed, so this stays in the upper mid-weight product-update band.

AI HOT (Curated Pool)

GitHub Copilot CLI Adds Custom AI Agents to Turn One-Off Terminal Prompts into Workflows

GitHub Copilot CLI added custom AI agents that understand a developer’s tech stack and team workflows; the post does not disclose configuration details, rollout scope, or pricing.

Why it matters: Official GitHub product update with HKR-H/R: custom Copilot CLI agents matter for developer workflows. HKR-K is weak because setup, rollout, and pricing are missing, so it sits at the featured threshold.

Jun 9Tuesday

AI HOT (Curated Pool)

GPT-5.5 Replaces OCR as ChinaRxiv Papers Become Freely Available

A developer replaced a complex OCR pipeline with GPT-5.5, making 23,000+ ChinaRxiv papers freely available with more complete English translations.

Why it matters: HKR-H/K/R all pass, but this is a developer use case rather than an OpenAI model launch. The 23,000+ paper corpus and OCR-pipeline replacement put it in the 78–84 recommendation band.

AI HOT (Curated Pool)

How an Agent Chains Two HuggingFace Spaces to Build a 3D Paris Gallery

A coding agent chained ideogram-ai/ideogram4 and VAST-AI/TripoSplat to generate Paris monument images, reconstruct single-image 3D Gaussian splats as .ply files, convert them to .ksplat with about 3× smaller size, and deploy a static Three.js Space using APIs exposed through agents.md.

Why it matters: HKR-H/K/R all pass, but this is a Hugging Face Spaces tutorial-style build, not a model or platform release. The concrete chain and ~3x compression place it in the 72-77 featured band.

AI HOT (Curated Pool)

Qwen3.7-Max Delivers Mobile and Web Apps from Scratch Using One Document

Qwen3.7-Max delivered mobile and web applications from a roughly 150,000-character product research document without design files or backend code; each client took about 4 hours, used staged constraint injection and error feedback, and the web app passed typecheck, build, and 34 reachable routes.

Why it matters: HKR-H/K/R all pass: the coding-agent claim is clickable, quantified, and emotionally relevant to developers. The summary lacks eval setup, failure rate, and human-intervention detail, so it stays in the 78–84 band.

Hacker News front page

Microsoft's Open Source Tools Were Hacked to Steal AI Developers' Passwords

The title says Microsoft's open source tools were hacked to steal passwords from AI developers; the RSS snippet does not disclose the affected tools, attack mechanism, timeline, or victim count.

Why it matters: TechCrunch plus HN front-page placement supports source weight, and the title hits HKR-H and HKR-R. HKR-K fails because tools, mechanism, and victim scale are missing, so the score stays at the featured floor.

AI HOT (Curated Pool)

GitHub 122K-star Skills adds Teach to turn a working directory into a stateful learning space

GitHub’s 122K-star Skills repository added Teach, which turns a working directory into a stateful learning space using MISSION.md, lessons/, learning-records/, and reference/ files to track goals, lessons, learned items, and reusable notes.

Why it matters: HKR-H/K/R pass via a concrete agent-memory workflow and named file structure, but the source is a single X summary with no benchmarks, maintainer detail, or user results, so it sits near the featured threshold.

Financial Times · Technology

Apple unveils “Siri AI” in challenge to rival chatbots

Apple unveiled “Siri AI” as a long-delayed overhaul of Siri, and the title frames it as a challenge to rival chatbots; the RSS snippet only states a user-privacy promise and does not disclose model details, launch timing, or a feature list.

Why it matters: FT authority plus an Apple Siri overhaul clears HKR-H and HKR-R, so it reaches featured. HKR-K fails because the article gives privacy claims but not specs, launch timing, or concrete features.

AI HOT (Curated Pool)

Migrating GitHub CI to Hugging Face Jobs

Hugging Face describes using huggingface/jobs-actions to run GitHub Actions CI as HF Jobs, where the Trackio project cut CPU job time by about 30% and added a GPU test suite using CPU, t4-small, or h200 hardware.

Why it matters: HKR-H/K/R pass via a concrete CI-to-HF Jobs workflow, ~30% speedup, and GPU-test pain point. Scope is ML tooling, not a major platform release, so it sits at the featured threshold.

AI HOT (Curated Pool)

Claude Supports Apple Foundation Models Framework With New Swift Package

Anthropic released a Swift package that lets Apple developers call Claude inside the Foundation Models framework with three lines of code, returning typed Swift values and handing off multi-step reasoning, code generation, web search, and data analysis on iOS 27, macOS 27, and related platforms.

Why it matters: HKR-H/K/R all pass: Anthropic is shipping a concrete Claude Swift package for Apple Foundation Models, but this is a developer integration rather than a model release, so it sits high in the 78–84 featured band.

The Verge · AI

Apple is using AI to fix Safari’s extension problem

Apple demonstrated Safari using Apple Intelligence to generate an extension from a text prompt, with a Recipe Keeper example for saving recipes and notes; the RSS snippet does not disclose release timing, required OS versions, or developer restrictions.

Why it matters: HKR-H/K/R pass, but the post gives only a demo and the Recipe Keeper example; launch timing, OS version, and developer limits are not disclosed. This fits a mid-weight product update at 73, below the 78 band.

TechCrunch · AI

Apple just taught your iPhone to finish your sentences, photos, and workflows

Apple is adding AI-powered features to Safari, Shortcuts, and Passwords, but the post does not disclose release timing, supported iPhone models, or the specific model behind them.

Why it matters: HKR-H/K/R pass: Apple is adding AI to Safari, Shortcuts, and Passwords, a concrete platform-surface update. Missing timing, device scope, and model details keep it at the featured threshold, not a must-write release.

TechCrunch · AI

Apple Will Let You Build Workflows Using AI in Its New Shortcuts App

Apple will add prompt-based workflow creation to its new Shortcuts app; the RSS snippet says users can describe the workflow they want, but the post does not disclose launch timing, OS version, pricing, or the model mechanism.

Why it matters: HKR-H/K/R pass, but the body only says users describe a goal to generate a workflow; launch timing, OS version, and model mechanism are not disclosed. This fits a mid-weight Apple product update.

TechCrunch · AI

Apple’s Long-Awaited AI Siri Overhaul Is Finally Here

Apple announced an AI Siri overhaul that aims to turn the voice-controlled assistant into an AI companion; the RSS snippet does not disclose the model, rollout timeline, pricing, or specific feature list.

Why it matters: HKR-H and HKR-R pass because Apple’s delayed Siri AI overhaul is a high-interest product story. HKR-K fails: the feed gives no model, rollout date, or concrete capability, so it sits near the featured floor.

The Verge · AI

Apple announces Siri AI and its next generation of Apple Intelligence

Apple announced Siri AI and a new Apple Intelligence set at WWDC, with systemwide access, onscreen reading, app interaction, and a customizable voice; the RSS snippet does not disclose launch timing or device eligibility.

Why it matters: HKR-H/K/R all pass: Apple used WWDC to add system-wide access, screen reading, and app actions to Siri, a major on-device agent update. Launch timing is not disclosed, so it lands at 86 rather than higher.

AI HOT (Curated Pool)

ChatGPT adds data chart generation

ChatGPT added data chart generation that turns data and comparisons into charts, and the post says it is available on mobile and web.

Why it matters: HKR-K and HKR-R pass: this is a concrete ChatGPT product update for chart generation across mobile and web. HKR-H is weak, and the post does not disclose formats, limits, or pricing, so it sits at the featured threshold.

AI HOT (Curated Pool)

NotebookLM upgrade adds agent capabilities and advanced reasoning

NotebookLM released an upgrade for Google AI Ultra subscribers, adding in-conversation agent capabilities, advanced reasoning, and new output formats. The post does not disclose the specific formats, pricing, or rollout schedule.

Why it matters: HKR-H/K/R all pass: Google confirms NotebookLM adds in-chat agents, advanced reasoning, and multi-output for AI Ultra users. Missing formats, pricing, and rollout details keep it in the mid-weight product-update band.

The Verge · AI

NotebookLM’s Gemini 3.5 Upgrade Adds a Cloud Computer and Source Discovery

Google is upgrading NotebookLM to Gemini 3.5, letting users start a research project by asking topic questions and use Google Search to find relevant sources, while the RSS snippet does not disclose details about the cloud computer feature.

Why it matters: HKR-H/K/R pass: NotebookLM gains Gemini 3.5, a cloud computer, and Search-based source discovery. This is a mid-weight Google product update, with pricing, rollout scope, and measured quality not disclosed.

Jun 8Monday

r/LocalLLaMA

OpenEnv Is Now Owned by HF, Torch, Prime Intellect, Unsloth, Modal, Mercor, and More

OpenEnv moved to committee coordination with 9 initial members, including Meta-PyTorch, Unsloth, Modal, Prime Intellect, Nvidia, and Mercor, while the post describes it as a tool for creating agent execution environments such as terminals and browsers.

Why it matters: HKR-H/K/R pass, but the post is thin: it gives committee ownership and 9 initial members. This is a mid-weight open-source agent-infra governance update, not a must-write release.

AI HOT (Curated Pool)

AgentScope Java 2.0 Released

Alibaba Cloud released AgentScope Java 2.0 for enterprise AI agent development, with K8s elastic scaling, session recovery, multi-tenant isolation, and Human-in-the-Loop support for JVM production environments.

Why it matters: HKR-K/R pass: AgentScope Java 2.0 names concrete production mechanisms from an Alibaba Cloud source. HKR-H is weak, and no benchmarks, adoption, or pricing are disclosed, so it sits at the featured threshold.

AI HOT (Curated Pool)

WeChat AI Agent Ecosystem Revealed: Mini Program Calls and Phone Maker Partnerships

Tencent is testing a WeChat-embedded AI Agent that opens via a right swipe and uses natural-language commands to call millions of Mini Programs for tasks such as ordering coffee. WeChat also partnered with Huawei, Honor, Xiaomi, OPPO, and vivo on A2A assistant capabilities, and released developer access guidance on June 8.

Why it matters: HKR-H/K/R all pass: WeChat-as-agent-runtime is clickable, concrete, and strategically resonant. Kept below P1 because this is single-source exposure and key details like rollout scope, model stack, and pricing are not disclosed.

AI HOT (Curated Pool)

WeChat AI Enters Internal Testing with Two Access Modes for Developers

WeChat Open Platform confirmed WeChat AI is in internal testing, offering two access modes: automatic mode lets the platform read mini program source code, while developer mode lets developers submit custom skills for review, and both modes can be enabled without affecting existing mini program services.

Why it matters: HKR-H/K/R all pass: WeChat AI is in beta with auto and developer modes that preserve mini-program services. Score stays near the featured floor because model capability, pricing, and rollout timing are not disclosed.

AI HOT (Curated Pool)

Apple Releases Third-Generation Apple Foundation Models (AFM)

Apple released its third-generation AFM family with five models. The RSS snippet says they span on-device use and Private Cloud Compute servers, with Google involved in customization for Apple Intelligence, Siri, and system-level tools.

Why it matters: Official Apple model-family release with 5 models, on-device/PCC deployment, and Google customization clears HKR-H/K/R. Missing benchmark and pricing details keep it at the low end of the 85+ band.

AI HOT (Curated Pool)

Open-source community backs OpenEnv for agentic reinforcement learning

Hugging Face announced broader OpenEnv access, coordinated by a committee from Meta-PyTorch, Reflection, and Unsloth; the project provides Gymnasium-style APIs and first-class MCP support for terminal and browser agent environments.

Why it matters: HKR-H/K/R all pass: this is not a model launch, but OpenEnv ties agent-RL environments, a Gymnasium-style API, and MCP into open governance, making it a solid infra story.

AI HOT (Curated Pool)

ChatGPT Is Set to Become AgentGPT

OpenAI is preparing ChatGPT’s largest redesign since its 2022 launch, shifting it toward an agent platform that integrates Codex, image generation, Canva, and Booking, with web and mobile rollout planned in the coming weeks. ChatGPT has 900 million weekly active users, 50 million paid users, and $2 billion in monthly revenue, but the post says it remains unprofitable.

Why it matters: HKR-H/K/R all pass, but this is a single X post and the body lacks official timing, access scope, and pricing. It sits at the top of 78–84 rather than P1 because the revamp is not yet shipped.

Jun 7Sunday

r/LocalLLaMA

Qwen3.6 35B-A3B on a Laptop: My Zero-to-One Moment

A Reddit user ran Qwen3.6 35B-A3B on an ASUS Zenbook Pro 14 with RTX 4060 8GB VRAM and 64GB RAM, reaching about 27 TPS at 32k context and 18 TPS at 256k context. The setup uses llama.cpp, unsloth’s IQ3_XXS GGUF quantization, and a 262144-token context flag.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit experiment, not an official release or paper. Concrete hardware, quantization, context, and TPS clear the featured bar, but keep it in the 72–77 band.

Computing Life · Share · Yage

How Claude Design Works: Reverse-Engineering an AI Designer from an Open-Source Plugin

The article reverse-engineers Claude Design from Anthropic’s open-source Design plugin and describes a six-layer structure; the snippet only discloses mechanisms such as workflow decomposition, aesthetic injection, evaluation transfer, and connector abstraction.

Why it matters: HKR-H/K/R all pass, but this is third-party reverse engineering rather than an Anthropic launch. It fits the high-quality Claude/agent mechanism analysis band just above featured threshold.

TechCrunch · AI

OpenAI unveils Lockdown Mode to protect sensitive data from prompt injection attacks

OpenAI introduced Lockdown Mode for ChatGPT, disabling live web browsing, web image retrieval and display, deep research, and agent mode for self-serve ChatGPT Business accounts and eligible personal accounts.

Why it matters: HKR-H/K/R all pass: OpenAI turns prompt-injection defense into a visible product switch with four concrete feature limits. Strong safety/product news, below a model release or major capability launch.