Skip to content

#MCP/工具调用

5 today

May 30Saturday

TechCrunch · AI

I put Google’s 24/7 AI assistant Gemini Spark to work, and it’s actually pretty useful

TechCrunch tested Google’s Gemini Spark as a 24/7 AI assistant for inbox summaries and local event planning; the RSS snippet does not disclose pricing, release timing, or why Google made it a separate product.

Why it matters: HKR-H/K/R pass: the hands-on angle is clickable, and inbox plus local-planning automation gives concrete substance. The score stays in the low featured band because price, launch timing, and product positioning are not disclosed.

AI HOT (Curated Pool)

Nano Banana Pro and Nano Banana 2 officially released

Google AI Developers released Nano Banana Pro and Nano Banana 2, mapped to gemini-3-pro-image and gemini-3.1-flash-image. The post says both are production-ready through the Gemini API, but does not disclose pricing, benchmarks, or runtime limits.

Why it matters: HKR-H/K/R all pass: Google names two image models and production Gemini API access. Missing pricing, benchmarks, and invocation limits keep it in the mid product-update band rather than a must-write release.

Xinzhiyuan · WeChat

Claude AI fluency scorecard surfaces, with strong users scoring 7.5

Anthropic is testing a Claude AI Fluency scorecard that analyzes Chat, Cowork, and Claude Code history against 11 observable behaviors, with an 11-point maximum score. The underlying study used 9,830 anonymized multi-turn conversations, and iteration appeared in 85.7% of high-quality conversations.

Why it matters: HKR-H/K/R all land: the angle is clickable, the scorecard has concrete numbers, and Claude users will debate being graded. This is not a model launch or major capability release, so it stays in the 78–84 featured band.

QbitAI · WeChat

RUC and Zhizhi Institute Open-Source Claw Agent Data, Training, and Evaluation Pipeline

Renmin University of China and Zhizhi Institute open-sourced ClawGym, a Claw Agent framework with 13.5K synthetic executable tasks, 200 benchmark tasks, model checkpoints, training data, and training code; ClawGym-30B-A3B scores 56.82 on ClawGym-Bench and exceeds Qwen3-235B-A23B in the reported evaluation.

Why it matters: HKR-H/K/R all pass: ClawGym bundles data, code, checkpoints, and eval tasks rather than just a leaderboard. Its impact is developer-facing, below a major lab model release or market-moving event.

AI HOT (Curated Pool)

Codex Can Manage Conversation Threads and Parallel Tasks

Codex can now create, search, organize, and pin conversation threads inside the Codex interface, and start worktrees for parallel tasks.

Why it matters: HKR-H/K/R pass: Codex gets concrete thread-management and parallel-worktree mechanics that matter to coding-agent users. Scope, pricing, and performance data are not disclosed, so this stays in the lower featured band.

AI HOT (Curated Pool)

Codex now supports computer use on Windows

OpenAI added Windows computer-use support for Codex, letting users start, review, and guide tasks on a Windows PC through the ChatGPT mobile app; the post states this is an early experience and does not disclose pricing or rollout scope.

Why it matters: HKR-H/K/R all pass: OpenAI adds Windows computer use to Codex, controlled through ChatGPT mobile. The post gives the workflow and early-stage condition, but not permissions, pricing, or rollout scope, so this stays at the featured threshold.

AI HOT (Curated Pool)

OpenRouter supports model-generated file patches

OpenRouter now supports apply_patch, a server-side tool that lets any model propose file edits through the Responses API using V4A diffs, covering file creation, updates, and deletion, with OpenRouter validating diff syntax on the server.

Why it matters: HKR-H/K/R pass: the OpenRouter update gives coding agents a concrete cross-model patch path with V4A diffs and server validation. It is useful infra, not a model-level release, so it sits low in the 72–77 band.

May 29Friday

The Verge · AI

Adobe’s Conversational AI Agent Is a Mediocre Design Intern

The Verge tested Adobe Firefly AI Assistant in beta. It can operate Adobe design apps as a conversational middleman, rather than only generating images or video. The post says it explains edit steps clearly, but the results were not impressive. The RSS snippet does not disclose pricing, release timing, or the full list of supported apps.

Why it matters: HKR-H/K/R all pass because this is a Verge hands-on of Adobe’s Firefly AI Assistant beta with a clear negative usability hook. Missing pricing, launch timing, and supported-app details keep it in the 72–77 featured-threshold band.

AI HOT (Curated Pool)

Xiaomi Open-Sources Controllable Video Foley Model ControlFoley

Xiaomi’s large model application team open-sourced ControlFoley, a controllable video Foley model supporting three tasks: text-guided video dubbing, text-controlled video dubbing, and reference-audio-controlled video dubbing, with code, model weights, and an online demo released.

Why it matters: ControlFoley clears HKR-H/K/R with controllable video Foley plus code, weights, and demo. It is a useful multimodal-audio release from Xiaomi, but not a flagship foundation-model launch, so it sits near the featured threshold.

QbitAI · WeChat

Tencent unveils Code Craft, an AI game creation platform for beginners and developers

Tencent Games unveiled Code Craft, an AI game creation platform that turns natural-language prompts into runnable 2D or 3D games, with a planning knowledge base, Skill system, visual tuning panels, and more than 20,000 free cloud assets; the post does not disclose release timing, pricing, model details, or supported engines.

Why it matters: HKR-H/K/R pass on the Tencent game-creation hook, runnable 2D/3D output, and 20,000+ assets. Pricing, access scope, and model limits are not disclosed, so it stays in the lower featured band.

r/LocalLLaMA

Liquid AI releases LFM2.5-8B-A1B

Liquid AI released LFM2.5-8B-A1B with a 128K context window, 38T pre-training tokens, large-scale reinforcement learning, doubled vocabulary for non-Latin tokenization, and availability on Hugging Face.

Why it matters: HKR-H/K/R pass: 8B/A1B, 128K context, and 38T tokens are concrete hooks for local inference. No benchmarks, license, or deployment limits are disclosed, so it stays in the mid featured band.

Latent Space

Anthropic raises $65B Series H, releases Opus 4.8 and Dynamic Workflows

Anthropic announced a $65B Series H at a $965B post-money valuation, disclosed a $47B revenue run rate, and released Claude Opus 4.8 plus Claude Code Dynamic Workflows as a research preview for parallel subagent orchestration.

Why it matters: HKR-H/K/R all pass: this combines a frontier-lab financing event with an Anthropic model and Claude Code workflow release. I score using the summary’s $65B raise and $965B post-money valuation because the title’s dollar figure conflicts with it.

AI HOT (Curated Pool)

Cursor team releases Developer Habits Report

Cursor’s report says developers’ weekly code output rose from about 3.6K to 8.6K lines, while AI agents increased tool calls per session by roughly 30%.

Why it matters: HKR-H/K/R all pass: Cursor’s own report gives concrete 3.6K→8.6K and +30% figures for AI coding work. It is not a product launch or cross-source event, so 78–84 fits better than the must-write band.

Ruan YiFeng's Weblog

Technology Enthusiasts Weekly Issue 398: Token Costs Are Hard to Afford

Peter Steinberger posted one month of usage showing 7.6 million requests and 603 billion tokens, with CodexBar estimating a $1.3 million value under preset rates rather than his actual spend as an OpenAI employee.

Why it matters: HKR-H/K/R all pass: the CodexBar case turns token economics into concrete usage and cost. This is strong practitioner commentary, not a model or platform release, so it fits the 72–77 featured band.

AI HOT (Curated Pool)

StepFun Releases Step 3.7 Flash, Focused on Agent Efficiency

StepFun released the open-source Step 3.7 Flash model with a 198B-parameter MoE architecture, about 11B active parameters, a 256K context window, and a 67.1 score on ClawEval-1.1.

Why it matters: HKR-H/K/R all pass: the release has a clear sparse-model hook, concrete context and benchmark numbers, and practitioner resonance around open agent efficiency. Official-post sourcing and no independent eval keep it in the 78–84 band.

AI HOT (Curated Pool)

Skill distillation

Skill distillation has Opus 4.7, GPT-5.1, and Gemini 3 Pro write standardized SKILL.md procedure files, while local Qwen 35B and Gemma 26B models execute those files step by step.

Why it matters: HKR-H/K/R pass: the agent-skill distillation pattern is concrete and practitioner-relevant. The summary lacks success rates, cost data, or task outcomes, so it sits at the featured threshold, not must-write.

Latent Space

The Age of Async Agents — Cognition's Walden Yan and OpenInspect's Cole Murray

Latent Space discusses async coding agents with Cognition’s Walden Yan and OpenInspect’s Cole Murray, citing Devin’s 7x merged PR growth and an increase from 16% to 80% of commits across Cognition repos.

Why it matters: HKR-H/K/R all pass: the Cognition repo numbers make this more than agent rhetoric. It stays in the 78 band because it is an interview/trend piece, not a major model or product release.

TechCrunch · AI

Anthropic releases Opus 4.8 with new Dynamic Workflows tool

Anthropic released Opus 4.8 with a Dynamic Workflows tool for coordinating swarms of subagents. The RSS snippet does not disclose pricing, context window size, benchmarks, or a rollout schedule.

Why it matters: HKR-H/K/R all pass: an Anthropic model release plus an agent orchestration tool fits the 85–94 same-day band. Missing price, context window, and rollout detail keep it below the top of the band.

May 28Thursday

AI HOT (Curated Pool)

Perplexity Computer Now Integrates with Microsoft Office

Perplexity Computer is now available in Microsoft Excel, Word, PowerPoint, and Outlook, letting users access Computer from the app sidebar to coordinate work, draft documents, model data, create presentations, and handle email.

Why it matters: HKR-H/K/R pass: Perplexity brings Computer into four Office apps with sidebar workflows, a useful product fact and competitive hook. Price, permission model, enterprise rollout, and measured results are not disclosed, so it stays at the featured threshold.

The Verge · AI

CNN sues Perplexity over “verbatim” copycat articles

CNN sued Perplexity in a New York court on Thursday, alleging its AI answer tools generate “verbatim” copies of CNN work and provide users with information locked behind CNN’s subscription wall.

Why it matters: HKR-H/K/R all pass: CNN vs Perplexity has a clear conflict, concrete allegations, and licensing-risk resonance. It is a notable copyright front, but only a lawsuit filing, not a ruling or product change.

QbitAI · WeChat

A New Paradigm for GUI Agent Trajectories: FSMs Generate Trajectories at $0.04 Each

AutoWebWorld synthesized 29 web environments, 875 pages, and 11,663 verified trajectories at about $0.04 per trajectory, using FSM-defined states, preconditions, and transitions to verify GUI agent tasks instead of human labeling or an LLM judge.

Why it matters: HKR-H/K/R all pass: $0.04 per trace, 11,663 verified traces, and FSM state checks give concrete hooks for GUI-agent data and eval cost. The source is not a top lab release, so it stays in the 78–84 research-tool band.

AI HOT (Curated Pool)

Mistral AI Releases Search Toolkit

Mistral AI released the public preview of Search Toolkit, an open source framework that combines data ingestion, retrieval, and evaluation behind shared interfaces for cloud, on-premises, or edge deployment.

Why it matters: HKR-K and HKR-R pass: Mistral combines RAG ingestion, retrieval, and evaluation in an open-source Search Toolkit with cloud, local, and edge deployment. HKR-H is weak, so this sits at the featured threshold for a mid-weight product update.

Mistral AI

Mistral upgrades Le Chat into unified agent Vibe, covering office work and coding

Mistral upgraded Le Chat into a unified AI agent called Vibe, with one license covering both office work and coding. Existing chats, settings and plans all carry over. Work Mode supports enterprise knowledge search, structured data analysis, document and report generation, scheduled multi-step tasks and reusable skills, and connects to Google Workspace, Outlook, SharePoint, Slack, GitHub and more.

Why it matters: It discloses Vibe's Work Mode, coding mode and CLI updates in full, so readers can judge how it plugs into existing workflows.

Alibaba Technology · WeChat

AI-Native Project Management: Two Git Repos Replace Weekly Updates, Insights, and Metrics Reports

Zhou Zhiwei describes a project-management setup that uses two Git repositories, an AI coding assistant, Shell, and Python to replace at least 80% of manual weekly-update chasing, data moving, chart generation, and engineering-metrics reporting.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, the post gives an 80% replacement claim and a two-repo mechanism, and it hits engineering-management toil. This is a strong practical workflow piece, not a model or platform launch.

AI HOT (Curated Pool)

Mistral AI launches physics AI model for industrial engineering

Mistral AI integrated the Emmi AI team and launched a physics AI foundation model for industrial engineering, with the post saying it can learn from geometry, boundary conditions, or measurement data and predict full physical fields on a single GPU in seconds.

Why it matters: HKR-H/K/R pass: a major model lab entering physics simulation with a concrete single-GPU seconds claim. The score stays in the lower featured band because model name, benchmarks, pricing, and access are not disclosed.

The Verge · AI

YouTube Will Let You Ask AI to Make a Custom Video Feed

YouTube is rolling out an AI custom video feed that uses a user-entered prompt or suggested options to build a personalized homepage feed around interests, moods, or topics; the feature supports English first and is available to signed-in users in the US on YouTube’s mobile app and desktop site.

Why it matters: HKR-H/K/R pass: promptable YouTube feeds are a concrete consumer-AI UX shift with rollout conditions. No model, ranking mechanism, or performance data is disclosed, so this stays in the mid-weight product-update band.

Financial Times · Technology

Kirkland & Ellis to Spend $500mn Building Its Own AI Technology

Kirkland & Ellis plans to spend $500mn building its own AI technology platform for the “collective intelligence” of its lawyers; the post does not disclose model architecture, vendors, or launch timing.

Why it matters: FT source authority and a $500mn in-house AI budget give strong HKR-H/K/R for legal AI adoption. Missing model architecture, vendors, and launch timing keep it in the 72–77 enterprise-adoption band.

AI HOT (Curated Pool)

Grok Build 0.1 on API

xAI released Grok Build 0.1 in public beta through the xAI API for agentic coding tasks, with throughput above 100 tokens per second and pricing at $1 per million input tokens and $2 per million output tokens.

Why it matters: HKR-H/K/R all pass, but this is a 0.1 public-beta API and pricing launch; benchmarks, context window, and task success rates are not disclosed. It fits a solid mid-weight product update at 78, featured not p1.

AI HOT (Curated Pool)

Cognition AI raises over $1B, targets 10x software engineering productivity

Cognition AI raised over $1 billion at a $26 billion pre-money valuation, while annualized revenue grew from $37 million to about $492 million in one year, and Devin is positioned as an autonomous junior engineer that can plan, test, and deploy through multi-step workflows.

Why it matters: HKR-H/K/R all pass: the story has hard numbers on funding, valuation, and ARR, plus a direct junior-engineer automation angle. Single-post sourcing keeps it below the 95+ industry-shaking band.

AI HOT (Curated Pool)

OpenAI Products Support Secure Connections to Private MCP Servers

OpenAI supports ChatGPT, Codex, and the Responses API connecting to internal MCP servers through outbound-only HTTPS, while teams keep those servers inside private networks.

Why it matters: HKR-H/K/R pass: OpenAI adds private MCP server support with outbound-only HTTPS, a concrete enterprise agent integration mechanism. Missing permission model, pricing, and rollout details keep it in the lower featured band.

AI HOT (Curated Pool)

Zero-Trust Security Framework for AI Agents

Anthropic published a zero-trust framework for enterprise autonomous AI agents, saying frontier models compress vulnerability exploitation from months to hours; the post outlines a three-tier architecture, an eight-stage rollout process, and threats including prompt injection, tool poisoning, and memory poisoning.

Why it matters: Anthropic’s agent zero-trust framework clears HKR-H/K/R with a concrete exploit-cycle claim, three-layer architecture, and eight-stage process. Strong safety/agent signal, but not a model launch or major product release.

AI HOT (Curated Pool)

Latest Google Pay updates

Google Pay introduced a universal commerce protocol and a new MCP server for AI agents to manage integrations and analyze trends, while Android updates add dynamic callbacks for faster checkout, WebView payments in social apps, cross-device biometric authentication, and new transaction signals.

Why it matters: HKR-H/K/R pass: the MCP payments angle is concrete and relevant to agent commerce. Score stays in the 72–77 band because the post lists features but gives no adoption scale, pricing, or real agent transaction case.

AI HOT (Curated Pool)

Interview with Google Search VP Robby Stein on the AI-Native Search Era

Robby Stein discussed Google Search’s move toward an AI-native mode at Google I/O, covering AI Mode, multi-turn query decomposition, TPU infrastructure costs, source-link selection, and publisher traffic tension, but the post does not disclose specific pricing, traffic numbers, or rollout conditions.

Why it matters: HKR-H/K/R all pass, but this is an interview summary rather than a fresh launch. No price, traffic, or cost numbers are disclosed, so it sits in the 72–77 quality-interview band.

May 27Wednesday

The Verge · AI

Robinhood will let your AI agent trade stocks and make (or lose) lots of money

Robinhood opened its trading platform to AI agents: traders can create a separate account, allocate a specific amount of money, and let the agent buy and sell stocks, while Robinhood warns agentic trading can cause the loss of the entire investment.

Why it matters: HKR-H/K/R all pass: real-money stock trading gives the hook, separate funded accounts add mechanism, and autonomy risk creates resonance. It stays in the 78–84 band because safeguards, rollout scope, and regulatory limits are not disclosed.

AI HOT (Curated Pool)

Runway launches Model Context Protocol server

Runway launched an MCP server that lets compatible agents such as Claude, ChatGPT, and Cursor generate images and videos inside chat interfaces, with access to Gen-4.5, Seedance 2.0, GPT Image 2, Kling 3.0, and Nano Banana Pro.

Why it matters: HKR-H/K/R all pass, but this is a Runway product integration, not an MCP protocol change or model release. It clears featured, with the score kept in the 72–77 band.

TechCrunch · AI

Robinhood now lets your AI agents trade stocks

Robinhood lets AI agents read and analyze users’ portfolios and suggest investments, but order placement is limited to the pre-loaded balance in a dedicated wallet.

Why it matters: HKR-H/K/R all pass: the hook is agents trading real money, the concrete mechanism is portfolio access plus a prefunded wallet, and the resonance is agent safety. Robinhood is not a frontier AI lab, so this stays at the lower featured band.

Alibaba Technology · WeChat

From Language Emergence to Collaborative Emergence: How AI Can Make High-Quality Decisions

Lv Ruofan proposes the Agent Room model: multiple agents share context, a task ledger, Memory, Runtime, and Artifacts, and two software-engineering cases show the system moving from workflow automation toward collaborative judgment rather than predefined task routing.

Why it matters: HKR-H/K/R all pass, but this is a methodology piece rather than a model launch or open-source framework. Concrete Agent Room mechanisms and 2 R&D sites put it in the 72–77 featured band.

New York Times Chinese

How Google Rebounded and Started Winning the AI Race

Google said regular Gemini users more than doubled in one year to 900 million, while ad revenue rose 16% to $77 billion last quarter, and its Siri partnership with Apple will place Gemini inside future iPhone assistant features.

Why it matters: HKR-H/K/R all pass: NYT ties Google’s comeback narrative to 900M Gemini users, ad growth, and a Siri distribution deal. This is strong industry analysis, not a model launch, so it fits the 78–84 band.

Xinzhiyuan · WeChat

OpenRouter processes 100 trillion tokens monthly and raises $113M Series B

OpenRouter raised a $113 million Series B led by CapitalG, lifting its valuation to $1.3 billion; the platform processes 25 trillion tokens per week, about 100 trillion per month, and provides one API for more than 400 models.

Why it matters: HKR-H comes from the 100T-token/month hook; HKR-K has funding, valuation, usage, and model-count numbers; HKR-R maps to routing and API-cost competition. Still, this is infra funding news, not an 85+ must-write release.

AI HOT (Curated Pool)

Code w/ Claude London event: Rethinking the developer experience

Anthropic announced two Claude Managed Agents capabilities at Code w/ Claude London: self-hosted sandboxes in public beta and MCP tunnels in research preview, with Spotify, Base44, and Legora already using them.

Why it matters: Official Anthropic product update with two concrete Claude Managed Agents capabilities. HKR-H/K/R pass, but this is a developer-tooling update rather than a major model release, so it lands at 78.