Skip to content

#MCP/工具调用

6 today

May 12Tuesday

AI HOT (Curated Pool)

The Evolution of Human-Computer Interfaces: From Text to Interactive Neural Video

Karpathy argues that LLM output is moving from Markdown toward richer HTML, while interactive neural video still has an open problem: how to combine neural generation with precise traditional software.

Why it matters: HKR-H/K/R pass: Karpathy gives a fresh UI frame, a concrete Markdown→HTML→neural-video path, and a builder-facing product question. Single X post with no data keeps it at the featured floor.

May 11Monday

AI HOT (Curated Pool)

Anthropic open-sources full-stack financial AI templates

Anthropic open-sourced a financial services AI template library on GitHub, including 10 end-to-end agents, 7 vertical industry plugins, and MCP connectors for 11 financial data providers, with deployment paths from personal plugins to enterprise APIs and integrations for Microsoft 365 and private cloud.

Why it matters: HKR-H/K/R all pass: Anthropic shipped a reusable finance-agent template library with GitHub artifacts and concrete counts. It is not a model release, so it stays below 85, but the open-source MCP vertical stack clears featured.

AI HOT (Curated Pool)

AntLingAGI Releases Trillion-Parameter Ring-2.6-1T Model

AntLingAGI released Ring-2.6-1T, a trillion-parameter thinking model available for free on OpenRouter until May 15, with adjustable thinking intensity, agent-oriented multi-step execution, tool calling, and tasks covering math logic and scientific research.

Why it matters: HKR-H/K/R all pass, but the post is thin: no benchmarks, pricing, architecture, or training details. Treat as a mid-weight model launch on OpenRouter, not a same-day must-write.

AI HOT (Curated Pool)

OpenAI Launches DeployCo to Help Enterprises Build Businesses Around Intelligence

OpenAI launched DeployCo, an enterprise deployment company focused on moving AI systems into production, while the RSS snippet does not disclose pricing, customer names, deployment scope, or launch timeline.

Why it matters: OpenAI launching DeployCo is a real enterprise strategy signal: HKR-H has a separate-company hook and HKR-R hits deployment competition. HKR-K is weak because pricing, customers, and timing are absent, so it sits at the featured floor.

Xinzhiyuan · WeChat

The Second Half of Agent Evaluation: Why a Live Benchmark Is Needed

Claw-Eval-Live evaluates 13 frontier models on 105 tasks, and the top model stays below a 70% pass rate, while HR tasks average only 6.8% pass rate.

Why it matters: HKR-H/K/R all pass: the live benchmark hook is specific, and the post gives 105 tasks, 13 models, HR at 6.8%. Claw-Eval-Live still lacks proven field impact, so this sits in the lower featured band.

Computing Life · Share · Yage

DeployCo Arrives: OpenAI and Anthropic Form AI Deployment JVs with PE on the Same Day

OpenAI and Anthropic announced AI deployment joint ventures with private equity on May 4, and the snippet cites divergent terms, including a 17.5% guaranteed return versus no guaranteed return.

Why it matters: HKR-H/K/R all pass: the angle has tension, the facts include PE JVs and a 17.5% floor, and the nerve is model-lab commercialization. Single-source commentary keeps it in the 78–84 band, not must-write.

Computing Life · Share · Yage

Google shuts down Project Mariner; Anthropic and OpenAI also hit limits

Google quietly shut down Project Mariner on May 4, and the post says Google, Anthropic, and OpenAI reached the same conclusion: standalone browser agents do not work, while GUI automation still has room outside headless dedicated environments.

Why it matters: HKR-H/K/R all pass: the shutdown date, route-level claim, and Google/OpenAI/Anthropic contrast carry signal. Single-source summary lacks an official notice or failure metrics, so this stays in the low featured band.

AI HOT (Curated Pool)

Codex autonomously completes a security audit and earns a bounty

A user instructed Codex to earn $5; Codex spent about 22 hours finding an open-source security audit bounty, submitting a valid PR, communicating with maintainers, passing GitHub verification, and ultimately receiving a $16.88 payment.

Why it matters: HKR-H/K/R all pass: a Codex agent allegedly closed a bounty loop in 22 hours with concrete money and workflow details. Single social-post evidence lacks reproducible logs, so it stays below P1.

AI HOT (Curated Pool)

MachinaCheck: Multi-agent CNC manufacturability analysis system built on AMD MI300X

MachinaCheck runs Qwen 2.5 7B locally on AMD MI300X to analyze STEP files for CNC manufacturability, reducing drawing review for quote analysis from 30–60 minutes to 30 seconds while using 192GB HBM3 to keep customer design data on-premises.

Why it matters: HKR-H/K/R all pass, but this is an AMD hackathon project on Hugging Face, not a broad model or platform launch. Concrete numbers carry it to the featured threshold.

May 10Sunday

QbitAI · WeChat

Zhejiang University introduces AdaMARP, an AI role-playing framework with scene direction

Zhejiang University and Tencent Youtu proposed AdaMARP for immersive role-playing, using a four-channel message format and a scene manager; its data pipeline includes 81 literary works, 20 synthetic themes, and AdaptiveBench with 100 evaluation seeds.

Why it matters: ACL 2026 role-play agent work brings four-channel messaging, a scene manager, and an 81-book dataset, clearing HKR-H/K. Narrow use cases and missing open-source or production evidence keep it at threshold featured.

Computing Life · Share · Yage

How Anthropic Trained Computer Use: Reading Its Data Pipeline Through a Patent

Anthropic’s patent describes the Computer Use training pipeline: it captures user actions, uses a transformer to infer action intent, and applies a stronger model for synthetic expansion, turning raw UI operations into reasoning data.

Why it matters: HKR-H/K/R all pass: the patent angle is clickable, the three-step data pipeline is concrete, and agent builders care. It is analysis, not an official release or reproducible artifact, so 76 fits the featured threshold.

May 9Saturday

AI HOT (Curated Pool)

YC CEO Open-Sources Personal AI OS GBrain for a Compounding Second Brain

Y Combinator CEO Garry Tan open-sourced GBrain, a personal AI operating system that processed more than 20 books in five months and manages over 100,000 pages of structured knowledge.

Why it matters: HKR-H/K/R pass: Garry Tan’s open-source personal knowledge system has a notable-user hook and three concrete usage numbers. Missing repo activity, architecture detail, and tests keep it at the featured threshold.

AI HOT (Curated Pool)

Peekaboo 3.0 Launches With Action-First macOS Control and UI Detection

Peekaboo 3.0 is now live with action-first macOS control, unified screenshots and UI detection, cleaner JSON exchange between CLI and MCP, and improved snapshots; the post does not disclose pricing, model choices, or release timeline beyond the 3.0 launch.

Why it matters: HKR-H/K/R all pass for a concrete desktop-agent tooling update. Score stays at the featured floor because pricing, model details, and adoption data are not disclosed.

r/LocalLLaMA

80 tok/sec and 128K context on 12GB VRAM with Qwen3.6 35B A3B and llama.cpp MTP

Reddit user janvitos ran Qwen3.6-35B-A3B-MTP-GGUF with a llama.cpp MTP PR on an RTX 4070 Super. The posted benchmark shows 69.2-81.9 tok/s, 0.694-0.947 draft acceptance, 131072 context, and a -fitt 1536 setting that reserves 1536 MB for the draft model and KV cache.

Why it matters: HKR-H/K/R all pass with concrete single-user benchmark data and reproducible settings. Source is one Reddit post, so verification is thin; this lands above featured threshold, not in must-write range.

AI HOT (Curated Pool)

Using Codex to debug and verify fixes in parallel

The author uses Codex in temporary crabbox environments to recreate bug states, verify failures, apply fixes, and re-verify them, while running 10 sessions in parallel to avoid local state pollution and speed loss.

Why it matters: HKR-H/K/R all pass, but this is a single first-person workflow note, not a product release or benchmark. The 10-session Codex/crabbox setup earns featured-level practical signal, near the lower band.

Xinzhiyuan · WeChat

CUHK Open-Sources ArbiterOS Agent Governance Kernel With 92.95% High-Risk Interception

CUHK CURE Lab open-sourced ArbiterOS, an agent runtime governance kernel that intercepts, parses, governs, and observes actions before execution, raising high-risk step interception on OpenClaw tasks from 6.17% to 92.95%.

Why it matters: HKR-H/K/R all pass: the story has a sharp execution-control hook, a concrete 6.17%→92.95% result, and clear agent-safety resonance. It is a strong open-source research tool, not a top-lab model release, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

NVIDIA, AMD, and Intel Back RadixArk’s $100M Seed Round

RadixArk announced a $100 million seed round on May 5 at a $400 million post-money valuation, led by Accel and co-led by Spark Capital, with participation from NVentures, AMD, MediaTek, Databricks, and other investors tied to AI infrastructure.

Why it matters: HKR-H/K/R all pass: a $100M seed round, $400M post-money valuation, and chip/data investors create an AI-infra rivalry angle. It remains a single-company funding item with no product benchmarks or customer data, so it sits near the featured floor.

QbitAI · WeChat

Google AI Co-Mathematician Sets FrontierMath Tier 4 SOTA

Google DeepMind released AI Co-Mathematician, an asynchronous agent workspace for math research, and answered 23 of 48 private FrontierMath Tier 4 problems, scoring 48% under 48-hour, no-token-limit conditions versus GPT-5.5 Pro at 39.6%.

Why it matters: HKR-H/K/R all pass: the story has a hard benchmark number and a concrete research hook. No disclosed product access or cross-source cluster, so it stays at the top of 78–84 rather than p1.

AI HOT (Curated Pool)

Claude Code Practice: The Effectiveness of HTML Output

Thariq Shihipar recommends requesting HTML output from Claude, and the post cites GPT-5.5 generating an interactive Linux vulnerability page with SVG diagrams, interactive components, and in-page navigation.

Why it matters: HKR-H/K/R all pass, but this is a workflow tip rather than a Claude release. As a quality Claude Code tutorial, it sits in the 72–77 band, with Simon Willison’s source authority clearing featured.

May 8Friday

Hacker News front page

Show HN: Git for AI Agents

regent-vcs released the open-source re_gent project for AI-agent version control, currently supporting Claude Code, with workflows for tracking why an agent changed files, rewinding sessions, and bisecting agent actions; the post does not disclose the license, storage format, or installation details.

Why it matters: HKR-H/K/R all pass: the Git analogy is clicky, the mechanism is concrete, and Claude Code rollback pain is real. The post lacks license, storage format, and install details, so it stays at the featured threshold.

Alibaba Technology · WeChat

The AI-Native Era: Where R&D Organizations Go Next

Xu Xiaobin cites internal interviews showing that engineers who use AI heavily cut coding time from 30% to 5%, raised Agent conversation time from 5% to 60%, and increased end-to-end delivery efficiency by 2 to 3 times, while pure coding efficiency rose 10 times.

Why it matters: Alibaba Tech’s internal-interview numbers make HKR-H/K/R pass, but this is org-methodology commentary rather than a product or model release, so it sits just above the featured threshold.

Synced · WeChat

OpenAI launches official CLI for terminal-based model access

OpenAI released the open-source openai-cli, letting developers call Responses, cloud tools, image generation and editing, speech transcription, and TTS from a single terminal command.

Why it matters: HKR-H/K/R all pass: an official OpenAI CLI, open-source packaging, and terminal access to multimodal APIs. This is a useful developer workflow update, not a major model capability release, so it sits in low featured.

Latent Space

[AINews] GPT-Realtime-2, Translate, and Whisper: new SOTA realtime voice APIs

OpenAI released GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper in the Realtime API, with GPT-Realtime-2 expanding context from 32K to 128K and scoring 96.6% on Artificial Analysis Big Bench Audio.

Why it matters: HKR-H/K/R all pass: an OpenAI real-time voice API refresh, a 32K→128K context jump, and a 96.6% Big Bench Audio claim. Score stays at 86 because this is a major API update, not a flagship foundation-model release.

AI HOT (Curated Pool)

China releases L1-L4 AI terminal intelligence standards covering 7 device categories

MIIT and other agencies released AI terminal intelligence standards covering 7 device categories. The framework uses a “2+N” structure with L1 response, L2 tool, L3 assistance, and L4 collaboration; L4 details come later. The post does not disclose concrete test metrics.

Why it matters: HKR-H/K/R pass: the story has a clear L1-L4 standards hook, concrete “2+N” and 7-category details, and compliance impact for device AI teams. Missing L4 rules and test metrics keep it near the featured floor.

AI HOT (Curated Pool)

OpenAI launches official openai-cli for terminal API calls

OpenAI open-sourced openai-cli for direct API calls from the terminal. The Apache 2.0 tool installs via Homebrew or Go and covers Responses API, structured output, image editing, transcription, and key config. The key detail is Agent workflows using cloud tools like web search and code interpreter.

Why it matters: HKR-H/K/R all pass: official OpenAI terminal tooling is clickable, with concrete install/license/API details and workflow resonance. It is still a developer tooling update, not a model or major capability release, so 76 fits the featured threshold.

AI HOT (Curated Pool)

Codex Plugin Now Supports Parallel Runs Across Chrome Tabs

OpenAI says Codex now runs in Chrome on macOS and Windows. The plugin works across tabs in the background without taking browser control; the post does not disclose version, concurrency limits, or enterprise policy.

Why it matters: HKR-H/K/R all pass, but the post gives platform and execution mechanics only; version, concurrency limits, and enterprise controls are not disclosed. Score: 76 as a practical OpenAI Codex product update.

NVIDIA Blog

Powering the Next American Century: Chris Wright and NVIDIA’s Ian Buck on Genesis Mission

The U.S. DOE and NVIDIA are building two AI supercomputers at Argonne; Equinox uses 10,000 Grace Blackwell GPUs. Solstice will use 100,000 Vera Rubin GPUs, which Buck said reach 5,000 exaflops. The key bottleneck is grid work: Wright said AI can cut interconnection studies from years to weeks or hours.

Why it matters: HKR-H/K/R all pass: the GPU counts, DOE-NVIDIA role, and grid bottleneck are concrete. NVIDIA-blog sourcing keeps it below must-write; this fits the 78–84 band.

AI HOT (Curated Pool)

Work with Claude across Excel, PowerPoint, Word, and Outlook

Claude now connects to four Microsoft apps: Excel, PowerPoint, Word, and Outlook. Excel, PowerPoint, and Word are generally available; Outlook is in public beta. Admins can deploy via Microsoft admin center and monitor with OpenTelemetry.

Why it matters: HKR-H/K/R all pass: Claude enters 4 Microsoft 365 apps with rollout status and OpenTelemetry details. This is a strong Anthropic product update, but not a model release or core capability jump, so it stays in the 78–84 band.

AI HOT (Curated Pool)

Perplexity launches Personal Computer app for Mac

Perplexity opened its Personal Computer Mac app to all users. It runs on any Mac and works across local files, native Mac apps, the web, and Perplexity secure servers. The post does not disclose pricing, permission boundaries, or task success rates.

Why it matters: HKR-H/K/R all pass, but the source is a single product post with no pricing, permission model, or task success rate. Score stays in the mid-weight product-update band.

r/LocalLLaMA

WARNING: Open-OSS/privacy-filter Malware

A Reddit user says Hugging Face repo Open-OSS/privacy-filter is an infostealer. It mimics OpenAI's privacy filter, uses loader.py to fetch PowerShell, then downloads an EXE and runs it via Task Scheduler. The author says they reported it to Microsoft and Hugging Face; the post says Linux is unaffected.

Why it matters: HKR-H/K/R all pass: malware disguised as an OpenAI privacy filter has a concrete Windows execution chain. Single Reddit sourcing keeps it at the 72-77 featured threshold.

May 7Thursday

AI HOT (Curated Pool)

Apify mcpc and x402 Give AI Agents an Auto-Payment Wallet

Apify mcpc integrates the x402 payment protocol, letting AI agents auto-sign payments on HTTP 402. x402 compresses paid API settlement into one HTTP round trip plus a signature; mcpc supports Claude Code and USDC-funded wallets. The key point is machine settlement for paid tool calls, not the wallet label.

Why it matters: HKR-H/K/R all pass: the hook is fresh, the mechanism is concrete, and agent payments hit a real practitioner nerve. It is still a mid-weight integration with no usage scale, pricing, or production case disclosed.

QbitAI · WeChat

Vidu Claw Generates Ad Videos From One Prompt and a Hundred-Yuan Budget

Shengshu Technology opened Vidu Claw, which generates ad scripts, voiceover, music, editing, and final videos from one prompt; its Video Plan includes up to 40 minutes of daily generation across video, image, and audio.

Why it matters: HKR-H has a concrete ad-test hook, HKR-K adds the 40-minute daily quota and one-prompt workflow, and HKR-R hits production-cost pressure. No benchmark or pricing detail, so this stays at the featured threshold.

Ben's Bites

Elon Doubled Limits

Ben’s Bites says Anthropic doubled Claude usage on paid plans via SpaceX’s Colossus 1. The issue also lists GPT-5.5 Instant, ChatGPT spreadsheet integration, and three Claude Managed Agents features. The title names Elon, but the post does not disclose exact limits.

Why it matters: HKR-H/K/R pass: the SpaceX Colossus 1 angle, 2x Claude usage, and quota pressure are all concrete. Missing exact caps, pricing, and rollout scope keep it in the low featured band.

OpenAI News

Scaling Trusted Access for Cyber with GPT-5.5 and GPT-5.5-Cyber

OpenAI expanded Trusted Access for Cyber to GPT-5.5 and GPT-5.5-Cyber. The RSS snippet says access is for verified defenders; the post does not disclose criteria, pricing, or benchmark data.

Why it matters: HKR-H/K/R all pass: OpenAI expands trusted cyber access to GPT-5.5 and GPT-5.5-Cyber. Kept below 85 because admission rules, pricing, evals, and reproducible tests are not disclosed.

AI HOT (Curated Pool)

Consistent web search and scraping for all models

OpenRouter released tools for tool-calling models to run web search and page scraping. The post says multiple search and scraping engines are supported, but does not disclose names, pricing, or limits. The key item is cross-model tool interface consistency.

Why it matters: HKR-H/K/R all pass, but engines, pricing, and limits are not disclosed. This is a mid-weight Product update: useful for model-agnostic agent stacks, not a major model or capability release.

r/LocalLLaMA

Running Qwen3.5/Qwen3.6 with NextN MTP in llama.cpp on one RTX 3090 Ti

A Reddit user posted a llama.cpp guide for Qwen3.5/3.6 with NextN MTP on one RTX 3090 Ti. It requires two unmerged PRs, #22400 and #22673; Qwen3.6-35B-A3B-MTP reaches 157 tok/s at 350W, 1700MHz, with q8 KV. The key reproducible detail is nextn=q8_0 quant override; missing it yields “////” output.

Why it matters: HKR-H/K/R all pass: single-GPU 157 tok/s is a strong hook, and the PR/power settings make it testable. Scope stays narrow because it is a Reddit guide using unmerged PRs.

AI HOT (Curated Pool)

China’s First Criminal AI Short-Drama Copyright Case Sentenced Over 1,700 Pirated Works

China’s first criminal AI short-drama copyright case reached a first-instance verdict over 1,700 pirated works. The defendant sold the bundle for 66.66 yuan and received eight months in prison, suspended for 14 months, plus a 6,000 yuan fine. The court held prompt-generated dramas contain original expression protected by copyright.

Why it matters: HKR-H/K/R all pass: first criminal AI short-drama copyright ruling, concrete figures, and direct pressure on gen-content IP compliance. Strong legal signal, but narrower than a major model or platform release.

AI HOT (Curated Pool)

Amp releases Neo CLI as coding agents shift toward long-horizon workflows

Amp released Neo, a CLI tool covering remote orchestration, automatic context compression, and a Plugin API. Neo lets local threads be controlled remotely, allows all operations by default, and moves safety control to plugins; the post does not disclose version, pricing, or performance gains.

Why it matters: HKR-H/K/R all pass: Neo adds remote orchestration, context compression, Plugin API, and default-allow permissions. Amp’s reach and missing price/version/perf data keep it in the 72–77 band.

AI HOT (Curated Pool)

Open Slide lets AI write PPT code

Open Slide builds PPTs with React, using a workflow designed for AI agents. It integrates SVGL with 1,500+ brand logos, supports manual edits, and lets AI read user comments for revisions.

Why it matters: HKR-H/K/R pass: the programmable-slide angle is clickable, with concrete React and 1500+ logo details, and deck work is a real practitioner pain. No usage metrics or hands-on test keeps it at the featured threshold.

The Verge · AI

Google shuts down Project Mariner

Google shut down Project Mariner on May 4, 2026. The experimental web-task agent once supported up to 10 concurrent tasks. Its technology moved into Google products, including Gemini Agent.

Why it matters: HKR-H/K/R all pass, but the disclosed facts are limited to shutdown timing, a 10-task limit, and migration into Gemini Agent. Strong source authority supports low featured, not a major launch.