Skip to content

#MCP/工具调用

5 today

Jun 7Sunday

TechCrunch · AI

OpenAI unveils Lockdown Mode to protect sensitive data from prompt injection attacks

OpenAI introduced Lockdown Mode for ChatGPT, disabling live web browsing, web image retrieval and display, deep research, and agent mode for self-serve ChatGPT Business accounts and eligible personal accounts.

Why it matters: HKR-H/K/R all pass: OpenAI turns prompt-injection defense into a visible product switch with four concrete feature limits. Strong safety/product news, below a model release or major capability launch.

Jun 6Saturday

AI HOT (Curated Pool)

GitHub open-sources Spec Kit to guide AI coding with product specifications

GitHub released the open-source Spec Kit, shifting AI coding from direct implementation to product specifications, gap clarification, technical planning, task breakdown, and agent execution, with support for 30+ agent integrations including Copilot, Claude Code, Codex, Gemini, Cursor, and Qwen, and 109K+ GitHub stars.

Why it matters: HKR-H/K/R all pass: GitHub’s Spec Kit gives a concrete spec-first agent workflow plus 30+ integrations and 109K+ stars. It is a strong tooling story, not a model- or platform-level launch.

r/LocalLLaMA

The Gap Between Claude and Local: Can a Self-Hosted Coding Agent Compete?

The author compared five coding-agent setups on a Laravel 12 + Livewire Playwright E2E task; Claude Opus 4.7 with 1M context produced 203 tests, while the strongest local OpenCode arm on a 24GB RTX 4090 produced 140 tests, compacted context four times, and needed seven manual nudges.

Why it matters: HKR-H/K/R all pass: a first-person Claude-vs-local coding-agent test with concrete counts. It stays below P1 because it is a single Reddit experiment, not a standardized benchmark or major release.

AI Chat-Group Daily (群聊日报)

Chat Group Weekly Vol. 2: The AI Tricks You Learned This Year May Be Wasted

The author retired an OpenClaw AI assistant after more than one month of use; the post says it required self-hosting, API setup, and keeping one home computer running 24 hours a day.

Why it matters: HKR-H/K/R all pass, but this is a personal weekly write-up, not a model or platform release. The month-long OpenClaw use and 24/7 PC requirement make it just clear the featured threshold.

Financial Times · Technology

Police in England and Wales told to halt AI use in court statements

Police in England and Wales were told to halt AI use in court statements until safeguards are in place; the RSS snippet cites the head of Police.AI but does not disclose the specific safeguards or enforcement mechanism.

Why it matters: FT reports a concrete policy action. HKR-H comes from the surprise halt in a court workflow, HKR-K from the England and Wales police pause, and HKR-R from safety and accountability stakes; not a model-level event, so it sits just above featured threshold.

AI HOT (Curated Pool)

Building a Multi-Agent Economy with Qwen2.5-3B: Engineering Report

A developer used Qwen2.5-3B to build a five-agent forest economy, and across 15 simulation rounds honey prices fell from 10 to 3, firewood rose from 4 to 7, and the Gini coefficient increased from 0.14 to 0.38.

Why it matters: HKR-H/K/R pass: the 3B multi-agent economy has a hook and concrete price/Gini results. It remains a single engineering experiment, not a product or framework launch, so it stays at the featured floor.

AI HOT (Curated Pool)

Google Colab CLI Released

Google released the Colab CLI, which lets developers and AI agents connect local terminals to remote Colab runtimes, request high-performance GPUs, run local Python scripts remotely, and retrieve artifacts such as logs or fine-tuned Gemma 3 adapters.

Why it matters: HKR-H/K/R pass: official Google Colab tooling adds terminal-to-remote-runtime GPU workflows for developers and agents. This is a solid developer product update, not a major model or platform release.

AI HOT (Curated Pool)

Gemini Live supports real-time image creation and editing

Gemini App adds real-time image creation and editing inside Live; users must open Live, share the camera, and tell Gemini what they want to see.

Why it matters: HKR-H/K/R pass: the real-time Gemini Live image workflow is clickable, concrete, and competitive. Scope is limited: the post gives entry and interaction conditions, not model, pricing, or rollout regions.

Jun 5Friday

AI HOT (Curated Pool)

Apple’s New Siri Is Marked Internally as Beta, Not Marketed as Finished

Apple marks the new Siri internally as Beta and may use a waitlist for access; some Siri queries will route through Google Cloud to a licensed Gemini version and run on Google’s NVIDIA Blackwell B200 cluster.

Why it matters: HKR-H/K/R all pass: Siri labeled Beta is a strong Apple hook, Gemini and B200 details add substance, and the story hits Apple AI dependency nerves. It stays in 78–84 because this is still an unlaunched product report.

Hacker News front page

Show HN: Lowfat – pluggable CLI filter saved 91.8% of my LLM tokens

Lowfat saved 4.1M of 4.4M raw tokens in the author’s two-month personal usage, running as an agent hook or shell wrapper to filter verbose CLI outputs from kubectl, docker, grep, and related commands.

Why it matters: HKR-H/K/R all pass: 91.8% savings is a strong hook, 4.1M/4.4M tokens plus the hook/wrapper mechanism add substance, and the cost/context pain is real for agent users. It is still a personal Show HN tool, so it stays near the featured threshold.

MIT Technology Review · AI

The Meta hack shows there’s more to AI security than Mythos

404 Media reported on June 5 that attackers used Meta’s AI customer support agent to link Instagram accounts to attacker-controlled email addresses; the article says the only extra condition was using a VPN matching the account owner’s location.

Why it matters: HKR-H/K/R all pass: an AI support agent changed an Instagram email, with VPN-location matching as the disclosed condition. This is a high-signal security incident, not P1 because scale, victim count, and Meta's fix are not disclosed.

AI HOT (Curated Pool)

OpenAI API Adds Moderation Scores

OpenAI added moderation scores to the Responses API and Completions API; applications can receive moderation signals in the same generation request and use them for logging, routing, review, or blocking.

Why it matters: HKR-K and HKR-R pass: OpenAI adds moderation scores to generation responses, giving builders a concrete safety-routing mechanism. HKR-H is weak, so this sits at the featured threshold, not a major-release band.

AI HOT (Curated Pool)

Codex launches iOS app build plugin

Codex integrated the Build iOS Apps plugin, which lets users test iOS apps in an in-app browser, open SwiftUI previews, and hot-reload edits without leaving Codex.

Why it matters: HKR-H/K/R all pass: the hook is Codex handling iOS app testing, with concrete SwiftUI preview and hot reload details. This is a mid-weight OpenAI dev-tool update, not a model release; pricing and rollout scope are not disclosed.

AI HOT (Curated Pool)

Replit Agent partners with Shopify for fast store creation

Replit partnered with Shopify to connect Replit Agent with store creation: users describe what they sell, then the agent builds a custom storefront, creates a Shopify store, and adds products; the post does not disclose pricing, regional availability, or launch timing.

Why it matters: HKR-H/K/R pass: the Shopify workflow is concrete and relevant to builders. The post gives no pricing, region, or rollout date, so it stays at the featured threshold rather than a higher product-release band.

Jun 4Thursday

Hacker News front page

Show HN: Cost.dev (YC W21) Makes Agents Cost-Aware and Cheaper to Call

Infracost launched Cost.dev, a local CLI for cloud-cost estimates in coding-agent workflows, and says it cut Claude output-token use by up to 79% and API cost by up to 67% versus a bare-Claude baseline.

Why it matters: HKR-H/K/R all pass: the local CLI cost-estimation mechanism and 79%/67% reduction claims are concrete. It is still a small vendor launch, so it sits at the featured floor, not same-day news.

MIT Technology Review · AI

How courts are coping with a flood of AI-generated lawsuits

MIT and USC researchers examined 4.5 million federal civil cases from 2005 to 2026, finding self-represented lawsuits rose from 11% in 2022 to 16.8% in 2025, while AI-text detector flags in sampled filings increased from 1% in 2023 to 18% in 2026.

Why it matters: MIT Technology Review covers an MIT/USC large-sample study, clearing HKR-H/K/R with 4.5M cases and an 18% AI-text marker rate. It affects public systems, not core model capability, so 78 fits the lower good-quality band.

Synced · WeChat

Office Whispering Is Turning Typing Into an Old Skill

AI dictation tools are moving into developer and office workflows, with Wispr Flow reporting over 2.5 million global downloads, 70% 12-month retention, and 100x annual user growth, while OpenAI’s gpt-4o-transcribe reached a 2.5% word error rate in a third-party evaluation cited by the article.

Why it matters: HKR-H/K/R all pass, but this is a data-backed workflow trend piece, not a model launch or platform update. It sits at the lower featured threshold.

AI HOT (Curated Pool)

OpenJarvis: A Local-First Framework for On-Device Personal AI Agents

Stanford researchers released OpenJarvis, an open-source local-first framework that runs reasoning, agents, memory, and learning on device, decomposes personal AI into five primitives, stays within 3.2 points of top cloud models, and cuts marginal API cost by about 800x.

Why it matters: HKR-H/K/R all pass: the story has a clear local-first agent hook, concrete cost and performance numbers, and strong cost/privacy resonance. Source depth is limited, so it stays in the 78–84 band rather than same-day must-write.

AI HOT (Curated Pool)

Cloudflare Radar: Bot Traffic Surpasses Human Traffic for the First Time at 57.5%

Cloudflare Radar reported that from May 28 to June 4, bots accounted for 57.5% of global HTML requests, while human browsers accounted for 42.5%; across all HTTP response content types, JSON led with 33.1% and HTML accounted for 12%.

Why it matters: Cloudflare Radar supplies a concrete window and ratios, clearing HKR-H/K/R. The post does not separate AI crawlers, search bots, and malicious automation, so it sits just above the featured threshold.

AI HOT (Curated Pool)

Hugging Face redesigns hf CLI output format for coding agents

Hugging Face redesigned hf CLI output for coding agents including Claude Code and Codex, using environment-variable detection and compact untruncated TSV output; in complex multi-step tasks, agents without the CLI used up to 6 times more tokens.

Why it matters: HKR-H/K/R pass: the story has a clear agent-CLI hook, a concrete TSV/token mechanism, and strong developer cost resonance. It stays in the featured band because this is a tooling update, not a model or platform release.

r/LocalLLaMA

I turned an Android phone into a Vulkan-accelerated local LLM node

Reddit user GsxrGuy80s configured a Z Fold 6 as a GGUF inference node using Vulkan, LiteLLM, and Tailscale; the post discloses gpu_layers=89, an OpenAI-compatible endpoint, and fallback routing to larger local nodes.

Why it matters: HKR-H/K/R all pass: a concrete phone-as-node hack with reproducible knobs. Source authority is limited to a Reddit post, so it fits the lower featured band rather than a broader industry update.

AI HOT (Curated Pool)

How Anthropic Enables Self-Service Data Analytics with Claude

Anthropic uses Claude to automate 95% of business analytics queries with about 95% accuracy; its agentic analytics stack uses a data foundation layer, validation workflows, and skills to handle ambiguity, stale data, and retrieval failures.

Why it matters: HKR-H/K/R all pass: the official post has marketing tone, but gives 95% automation, ~95% accuracy, and an agentic analytics stack. No new model or product release keeps it in the 72–77 band.

Jun 3Wednesday

TechCrunch · AI

Publishers will be able to opt out of AI Search, thanks to new regulation

U.K. regulators require Google to offer website publishers a tool to opt out of generative AI search features, with the option tested in the U.K. before a global rollout.

Why it matters: HKR-H/K/R all pass: regulation pushes Google AI Search to add a publisher opt-out, tested in the UK before global rollout. It affects web-content economics, but it is not a core model or capability launch, so it sits in 78–84.

OpenAI News

Introducing new capabilities to GPT-Rosalind

OpenAI says GPT-Rosalind adds biological reasoning, medicinal chemistry, genomics analysis, and experimental workflow capabilities; the RSS snippet does not disclose model parameters, benchmark results, pricing, or access conditions.

Why it matters: OpenAI’s vertical model update clears HKR-H and HKR-R, but HKR-K fails because evals, parameters, and access terms are missing. That keeps it at the featured floor.

Alibaba Technology · WeChat

Rethinking R&D Infrastructure When Agents Become First-Class Citizens

Xu Xiaobin argues that agent-based development compresses the intent-to-code loop from weeks or months to minutes, using a weekly-report system, a multi-role agent development setup, and image-repository provisioning as examples; the article identifies mismatches in Git, CI, code review, release flows, permissions, harness setup, and dry-run validation.

Why it matters: HKR-H/K/R all pass, but this is infrastructure commentary rather than a model or product launch. The named cases and week/month-to-minutes claim put it in the 72–77 featured band.

AI Chat-Group Daily (群聊日报)

2026-06-02 Chat Group Daily

The chat group daily says Microsoft released MAI-Thinking-1 with 35B active parameters and about 1T MoE, matching Opus 4.6 on SWE-Bench Pro and scoring 97% on AIME 2025.

Why it matters: HKR-H/K/R all pass: a Microsoft reasoning-model claim with concrete benchmark numbers. Source authority is weak, and the summary lacks official release, access terms, and full eval setup, so it stays below P1.

QbitAI · WeChat

Papers with Code returns with CVPR coverage and Hugging Face-led rebuild

Hugging Face’s open-source team launched paperswithcode.co in May 2026, using AI agents to parse papers and restore SOTA leaderboards tied to the original platform’s 9,300-plus benchmarks.

Why it matters: HKR-H/K/R all pass: a beloved research portal returns, with 9,300 restored leaderboards and agent-based paper parsing. The impact is strong for research workflows, not model-release scale.

QbitAI · WeChat

Coze 3.0 test: phone can remotely control agents on your computer

Coze 3.0 adds project-based agent collaboration across iOS, Android, Mac, Windows, and web, supports importing local agents such as Claude Code, Codex CLI, and OpenClaw, and can read a desktop PDF from a phone after user authorization.

Why it matters: HKR-H/K/R all pass: Coze 3.0 adds cross-device agent control and imports Claude Code/Codex CLI. It remains a single product update, with price, rollout scope, and security limits not disclosed, so it sits at the featured threshold.

AI HOT (Curated Pool)

Qwen3.7 Released with Upgrades to Reasoning and Agent Capabilities

Qwen released Qwen3.7, and the post says it upgrades reasoning, tool use, coding, and long-horizon agent tasks; the post does not disclose model size, pricing, benchmark scores, or release conditions.

Why it matters: HKR-H and HKR-R pass because Qwen3.7 is a flagship Alibaba model update with practitioner relevance. HKR-K fails: the post names capability areas but gives no params, pricing, benchmarks, or access terms.

r/LocalLLaMA

Microsoft Aion 1.0 Instruct and Aion 1.0 Plan models

Microsoft announced two on-device Aion 1.0 models at Build 2026. Aion 1.0 Plan is a 14B-parameter reasoning and tool-calling model with 32K context, shipping in-box with Windows on capable devices, while Aion 1.0 Instruct targets summarization, rewriting, intents, accessibility, Edge integration, and open-weight availability.

Why it matters: Microsoft announced Aion 1.0 Instruct and Plan at Build 2026, with Plan listed as a 14B, 32K-context model for eligible Windows devices. HKR-H/K/R all pass, but licensing, benchmarks, and hardware requirements are not disclosed, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

OpenAI’s Greg Brockman and a 9-Year Rift With Anthropic Co-Founder Dario Amodei

A WSJ-based profile says Dario Amodei once barred Greg Brockman from an internal OpenAI project that later led to ChatGPT, and the article says Brockman now oversees OpenAI product strategy with nearly 1,500 people under that function.

Why it matters: HKR-H/K/R all pass: the WSJ-sourced ban detail and the nearly 1,500-person scope give this more signal than gossip. It is not a model release or current executive departure, so it stays in the good-quality featured band.

AI HOT (Curated Pool)

Complete Practical Tips for Agent Engineering

@mvanhorn shared an agent engineering workflow centered on a Research→Plan→Work loop, plan.md constraints, and 22 practical tips; the snippet says it covers planning, parallel execution, input methods, and remote control, but the post does not disclose the full tool stack list.

Why it matters: HKR-H/K/R all pass, but this is a practitioner methods post, not a model or product release. The full tool stack is not disclosed, so it sits at the featured threshold.

AI HOT (Curated Pool)

Grok Becomes Vapi's Default Voice Engine

xAI partnered with Vapi to make Grok the default engine for 12 core voices, covering more than 2.5 million voice agents, and Grok Voice ranked first in Vapi’s independent blind test.

Why it matters: HKR-H/K/R all pass: the default-engine switch has scale, numbers, and voice-agent market resonance. Single-source partnership news lacks test methodology, pricing, and migration data, so it stays in the mid product-update band.

AI HOT (Curated Pool)

xAI releases Grok Imagine 1.5 preview image-to-video model

xAI released grok-imagine-video-1.5-preview via its API, letting users turn one still image into 720p video while controlling camera movement, pacing, and sound effects with natural-language prompts.

Why it matters: HKR-H/K/R all pass: xAI ships a named image-to-video API preview with 720p output and sound controls. It stays below 85 because this is a preview product update, not a flagship foundation-model release.

AI HOT (Curated Pool)

OpenAI launches Codex Sites to turn ideas into interactive websites

OpenAI opened Codex Sites in preview to Business and Enterprise subscribers, letting users turn ideas into hosted interactive sites such as dashboards, planners, and project boards, with URL sharing for specified team members.

Why it matters: HKR-H/K/R all pass, but this is an OpenAI Codex enterprise-preview feature rather than a model or core capability release. It sits in the mid-weight product-update band.

AI HOT (Curated Pool)

NVIDIA launches NemoClaw platform for autonomous AI engineers in industrial software

NVIDIA released NemoClaw at COMPUTEX as an open blueprint for long-running AI agents, and more than a dozen industrial software vendors are using it to build autonomous AI engineers for CAE and EDA workflows that compress weeks-long simulation and design tasks into hours.

Why it matters: HKR-H/K/R pass: NVIDIA’s NemoClaw targets industrial agents with 10+ vendors and a weeks-to-hours claim. The NVIDIA-blog sourcing and missing technical detail keep it at the lower featured band.

AI HOT (Curated Pool)

Google DeepMind open-sources a toolkit for scientific agents

Google DeepMind released Science Skills on GitHub for scientific-discovery agent workflows; the post does not disclose the license, benchmark results, or numeric token-efficiency gains.

Why it matters: Passes HKR-H/K/R: DeepMind, open source, and science agents make it relevant. Missing license, benchmarks, and efficiency data keep it in the 78–84 band, not P1.

AI HOT (Curated Pool)

Claude Code Adds Dynamic Workflows

Claude Code added dynamic workflows that execute JavaScript files at runtime to create and coordinate multiple subagents; each subagent has its own context window, and the feature is described for research, security analysis, and code review tasks.

Why it matters: HKR-H/K/R all pass: Claude Code gets runtime JS workflows coordinating isolated-context subagents. Anthropic update earns a bump, but this is a feature release rather than a model or platform launch, so it sits in the 78–84 band.

AI HOT (Curated Pool)

Claude Code launches dynamic workflows for task-specific frameworks

Claude Code added dynamic workflows that execute JavaScript files to coordinate subagents, with configurable model choice and workspace isolation level, but the post does not disclose token overhead figures or release availability details.

Why it matters: HKR-H/K/R all pass, but the post gives mechanism-level detail only; token overhead, rollout scope, and pricing are not disclosed. Claude Code relevance lifts this to the high end of a mid-weight product update.

AI HOT (Curated Pool)

Runway API adds Aleph 2.0 video editing

Runway API now provides Aleph 2.0 video editing for integration into apps, products, and platforms, supporting precise edits on multi-shot videos up to 30 seconds at 1080p while changing only selected portions; the post does not disclose pricing, rate limits, latency, or model availability by region.

Why it matters: Runway is a core AI-video player, and Aleph 2.0 exposes partial video editing via API with 30s and 1080p limits. HKR-H/K/R all pass, but this is a mid-weight product update, not a model-class release.