Skip to content

MCP & tool use

How models connect to the outside world: the MCP ecosystem, function calling and tool integrations.

760 picksRelated topicsAgentsAI codingOpen source

Latest picks

381–400 of 760

May 8Friday

r/LocalLLaMA

WARNING: Open-OSS/privacy-filter Malware

A Reddit user says Hugging Face repo Open-OSS/privacy-filter is an infostealer. It mimics OpenAI's privacy filter, uses loader.py to fetch PowerShell, then downloads an EXE and runs it via Task Scheduler. The author says they reported it to Microsoft and Hugging Face; the post says Linux is unaffected.

Why it matters: HKR-H/K/R all pass: malware disguised as an OpenAI privacy filter has a concrete Windows execution chain. Single Reddit sourcing keeps it at the 72-77 featured threshold.

May 7Thursday

AI HOT (Curated Pool)

Apify mcpc and x402 Give AI Agents an Auto-Payment Wallet

Apify mcpc integrates the x402 payment protocol, letting AI agents auto-sign payments on HTTP 402. x402 compresses paid API settlement into one HTTP round trip plus a signature; mcpc supports Claude Code and USDC-funded wallets. The key point is machine settlement for paid tool calls, not the wallet label.

Why it matters: HKR-H/K/R all pass: the hook is fresh, the mechanism is concrete, and agent payments hit a real practitioner nerve. It is still a mid-weight integration with no usage scale, pricing, or production case disclosed.

QbitAI · WeChat

Vidu Claw Generates Ad Videos From One Prompt and a Hundred-Yuan Budget

Shengshu Technology opened Vidu Claw, which generates ad scripts, voiceover, music, editing, and final videos from one prompt; its Video Plan includes up to 40 minutes of daily generation across video, image, and audio.

Why it matters: HKR-H has a concrete ad-test hook, HKR-K adds the 40-minute daily quota and one-prompt workflow, and HKR-R hits production-cost pressure. No benchmark or pricing detail, so this stays at the featured threshold.

Ben's Bites

Elon Doubled Limits

Ben’s Bites says Anthropic doubled Claude usage on paid plans via SpaceX’s Colossus 1. The issue also lists GPT-5.5 Instant, ChatGPT spreadsheet integration, and three Claude Managed Agents features. The title names Elon, but the post does not disclose exact limits.

Why it matters: HKR-H/K/R pass: the SpaceX Colossus 1 angle, 2x Claude usage, and quota pressure are all concrete. Missing exact caps, pricing, and rollout scope keep it in the low featured band.

OpenAI News

Scaling Trusted Access for Cyber with GPT-5.5 and GPT-5.5-Cyber

OpenAI expanded Trusted Access for Cyber to GPT-5.5 and GPT-5.5-Cyber. The RSS snippet says access is for verified defenders; the post does not disclose criteria, pricing, or benchmark data.

Why it matters: HKR-H/K/R all pass: OpenAI expands trusted cyber access to GPT-5.5 and GPT-5.5-Cyber. Kept below 85 because admission rules, pricing, evals, and reproducible tests are not disclosed.

AI HOT (Curated Pool)

Consistent web search and scraping for all models

OpenRouter released tools for tool-calling models to run web search and page scraping. The post says multiple search and scraping engines are supported, but does not disclose names, pricing, or limits. The key item is cross-model tool interface consistency.

Why it matters: HKR-H/K/R all pass, but engines, pricing, and limits are not disclosed. This is a mid-weight Product update: useful for model-agnostic agent stacks, not a major model or capability release.

r/LocalLLaMA

Running Qwen3.5/Qwen3.6 with NextN MTP in llama.cpp on one RTX 3090 Ti

A Reddit user posted a llama.cpp guide for Qwen3.5/3.6 with NextN MTP on one RTX 3090 Ti. It requires two unmerged PRs, #22400 and #22673; Qwen3.6-35B-A3B-MTP reaches 157 tok/s at 350W, 1700MHz, with q8 KV. The key reproducible detail is nextn=q8_0 quant override; missing it yields “////” output.

Why it matters: HKR-H/K/R all pass: single-GPU 157 tok/s is a strong hook, and the PR/power settings make it testable. Scope stays narrow because it is a Reddit guide using unmerged PRs.

AI HOT (Curated Pool)

China’s First Criminal AI Short-Drama Copyright Case Sentenced Over 1,700 Pirated Works

China’s first criminal AI short-drama copyright case reached a first-instance verdict over 1,700 pirated works. The defendant sold the bundle for 66.66 yuan and received eight months in prison, suspended for 14 months, plus a 6,000 yuan fine. The court held prompt-generated dramas contain original expression protected by copyright.

Why it matters: HKR-H/K/R all pass: first criminal AI short-drama copyright ruling, concrete figures, and direct pressure on gen-content IP compliance. Strong legal signal, but narrower than a major model or platform release.

AI HOT (Curated Pool)

Amp releases Neo CLI as coding agents shift toward long-horizon workflows

Amp released Neo, a CLI tool covering remote orchestration, automatic context compression, and a Plugin API. Neo lets local threads be controlled remotely, allows all operations by default, and moves safety control to plugins; the post does not disclose version, pricing, or performance gains.

Why it matters: HKR-H/K/R all pass: Neo adds remote orchestration, context compression, Plugin API, and default-allow permissions. Amp’s reach and missing price/version/perf data keep it in the 72–77 band.

AI HOT (Curated Pool)

Open Slide lets AI write PPT code

Open Slide builds PPTs with React, using a workflow designed for AI agents. It integrates SVGL with 1,500+ brand logos, supports manual edits, and lets AI read user comments for revisions.

Why it matters: HKR-H/K/R pass: the programmable-slide angle is clickable, with concrete React and 1500+ logo details, and deck work is a real practitioner pain. No usage metrics or hands-on test keeps it at the featured threshold.

The Verge · AI

Google shuts down Project Mariner

Google shut down Project Mariner on May 4, 2026. The experimental web-task agent once supported up to 10 concurrent tasks. Its technology moved into Google products, including Gemini Agent.

Why it matters: HKR-H/K/R all pass, but the disclosed facts are limited to shutdown timing, a 10-task limit, and migration into Gemini Agent. Strong source authority supports low featured, not a major launch.

r/LocalLLaMA

GB10 inference engine Atlas is open source, with Qwen3.6-35B-FP8 over 100 tok/s

Avarok open-sourced Atlas, an inference engine running Qwen3.5-35B at ~111 tok/s sustained on one DGX Spark. It uses Rust+CUDA, a ~2.5GB image, and sub-2-minute cold start; the author claims 3.0–3.3x vLLM in tests. The key details are Blackwell SM120/121 kernels, NVFP4/FP8, and MTP decoding.

Why it matters: HKR-H/K/R pass: open-source inference engine, 35B FP8 at 111 tok/s, and a direct vLLM comparison. Single Reddit sourcing and unreproduced benchmarks keep it at the lower featured band.

r/LocalLLaMA

Analysis of 922 Agentic Task Traces Finds DeepSeek v4’s Cost Edge in Caching

A Reddit user analyzed 922 agentic task traces and reported $0.01 per task for DeepSeek v4 Flash versus $1.52 for Opus 4.7. Both used about 960K tokens per task, but DeepSeek showed a 97% cache hit rate versus 87%, with a 0.02 cache read/write price ratio versus 0.08. The key issue is caching, not headline pricing.

Why it matters: HKR-H/K/R all pass: 922 agent traces tie a large cost gap to cache hit rate and cache read/write pricing. Reddit single-source data and incomplete method detail keep it in the 78–84 band.

May 6Wednesday

r/LocalLLaMA

CopilotKit (MIT): Open-source building blocks for agent apps and generative UI

CopilotKit offers MIT-licensed React components and claims 30k GitHub stars. It covers chat, streaming, tool calls, HITL, and generative UI, with AG-UI support for LangGraph, CrewAI, LlamaIndex, and other backends. The key point is decoupling the UI layer from agent frameworks.

Why it matters: HKR-H/K/R all pass: MIT open source, 30k stars, and AG-UI links to major agent backends. Kept in 72–77 because the post lacks a new version, benchmark, or named production adopter.

r/LocalLLaMA

Qwen3.6 27B NVFP4 + MTP on a Single RTX 5090: 200k Context in vLLM

A Reddit user ran Qwen3.6 27B NVFP4 on one RTX 5090 32GB and validated 200k context in vLLM. The setup used fp8_e4m3 KV cache, FlashInfer, and MTP with 3 speculative tokens; a 10-run 200k pass completed with 73.6 tok/s mean generation and 70.2s TTFT. The key constraint is 32GB VRAM: logs showed 8.3GiB KV cache and about 30478MiB total GPU use.

Why it matters: HKR-H/K/R all pass: the hook is single-GPU 200k context, with concrete vLLM settings and 10-run stability data. Reddit sourcing keeps it in the 78–84 band, not P1.

The Verge · AI

Google’s AI Search Summaries Will Now Quote Reddit

Google updated AI Search to include firsthand views from Reddit, social media, and forums in summaries. The post says a “perspectives” preview links queries to related online discussions; it does not disclose rollout scope or timing. For search teams, the key issue is how AI summaries cite and rank UGC sources.

Why it matters: HKR-H is strong because Google AI summaries quoting Reddit alters the search surface. HKR-K has the perspectives mechanism, and HKR-R hits SEO/UGC traffic concerns; missing rollout scope keeps it in the 72–77 product-update band.

NVIDIA Blog

NVIDIA Spectrum-X AI-Native Ethernet Fabric Adds MRC for Gigascale AI

NVIDIA added MRC support to Spectrum-X Ethernet, letting one RDMA connection spread traffic across multiple paths. MRC ran in Blackwell deployments, with microsecond failure bypass and hardware rerouting. The key detail is the OCP open specification and multiplane support for clusters up to hundreds of thousands of GPUs.

Why it matters: HKR-K/R are solid: MRC stripes one RDMA flow across paths, detects failures in microseconds, and is tied to Blackwell deployments. HKR-H is narrow and the source is vendor-owned, so this stays below major release level.

Synced · WeChat

DeepSeek Version of Claude Code Tops Trending Chart With 8,700 Stars

DeepSeek TUI topped GitHub trending with over 8,700 stars. Hunter Bown built it in Rust for local terminal use with DeepSeek V4, supporting chat, file edits, shell commands, and task management. The key detail is RLM mode: up to 16 V4 Flash subtasks, plus a 1M-token context window and approval gates.

Why it matters: HKR-H/K/R all pass: the 8,700-star hook is strong, RLM adds concrete mechanisms, and coding-agent competition resonates. It is a third-party open-source tool, not an official DeepSeek model release, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

Salesforce plans to hire 1,000 graduates as agent roles expand

Salesforce CEO Marc Benioff said the company will hire 1,000 graduates or interns for Agentforce growth. The post cites Agentforce ARR up 169% to $800 million, with roles covering prompts, evals, agent supervision, and delivery. The key shift is entry roles moving from execution to agent orchestration and output checks.

Why it matters: HKR-H/K/R all pass: 1,000 junior hires, $800M Agentforce ARR, and 169% growth give concrete signal, with a strong jobs angle. This is Salesforce hiring plus Agentforce expansion, not a major model or product release.

TechCrunch · AI

Apple plans to make iOS 27 a Choose Your Own Adventure of AI models

Apple reportedly plans to let iOS 27 users choose third-party AI models for multiple tasks. The RSS snippet does not disclose model names, task scope, launch timing, or integration mechanics.

Why it matters: HKR-H/K/R pass: system-level model choice on iOS has a strong platform hook and distribution stakes. Kept in 72–77 because the RSS summary lacks model names, task scope, launch timing, and API mechanics.