Skip to content

#MCP/工具调用

6 today

May 7Thursday

r/LocalLLaMA

GB10 inference engine Atlas is open source, with Qwen3.6-35B-FP8 over 100 tok/s

Avarok open-sourced Atlas, an inference engine running Qwen3.5-35B at ~111 tok/s sustained on one DGX Spark. It uses Rust+CUDA, a ~2.5GB image, and sub-2-minute cold start; the author claims 3.0–3.3x vLLM in tests. The key details are Blackwell SM120/121 kernels, NVFP4/FP8, and MTP decoding.

Why it matters: HKR-H/K/R pass: open-source inference engine, 35B FP8 at 111 tok/s, and a direct vLLM comparison. Single Reddit sourcing and unreproduced benchmarks keep it at the lower featured band.

r/LocalLLaMA

Analysis of 922 Agentic Task Traces Finds DeepSeek v4’s Cost Edge in Caching

A Reddit user analyzed 922 agentic task traces and reported $0.01 per task for DeepSeek v4 Flash versus $1.52 for Opus 4.7. Both used about 960K tokens per task, but DeepSeek showed a 97% cache hit rate versus 87%, with a 0.02 cache read/write price ratio versus 0.08. The key issue is caching, not headline pricing.

Why it matters: HKR-H/K/R all pass: 922 agent traces tie a large cost gap to cache hit rate and cache read/write pricing. Reddit single-source data and incomplete method detail keep it in the 78–84 band.

May 6Wednesday

r/LocalLLaMA

CopilotKit (MIT): Open-source building blocks for agent apps and generative UI

CopilotKit offers MIT-licensed React components and claims 30k GitHub stars. It covers chat, streaming, tool calls, HITL, and generative UI, with AG-UI support for LangGraph, CrewAI, LlamaIndex, and other backends. The key point is decoupling the UI layer from agent frameworks.

Why it matters: HKR-H/K/R all pass: MIT open source, 30k stars, and AG-UI links to major agent backends. Kept in 72–77 because the post lacks a new version, benchmark, or named production adopter.

r/LocalLLaMA

Qwen3.6 27B NVFP4 + MTP on a Single RTX 5090: 200k Context in vLLM

A Reddit user ran Qwen3.6 27B NVFP4 on one RTX 5090 32GB and validated 200k context in vLLM. The setup used fp8_e4m3 KV cache, FlashInfer, and MTP with 3 speculative tokens; a 10-run 200k pass completed with 73.6 tok/s mean generation and 70.2s TTFT. The key constraint is 32GB VRAM: logs showed 8.3GiB KV cache and about 30478MiB total GPU use.

Why it matters: HKR-H/K/R all pass: the hook is single-GPU 200k context, with concrete vLLM settings and 10-run stability data. Reddit sourcing keeps it in the 78–84 band, not P1.

The Verge · AI

Google’s AI Search Summaries Will Now Quote Reddit

Google updated AI Search to include firsthand views from Reddit, social media, and forums in summaries. The post says a “perspectives” preview links queries to related online discussions; it does not disclose rollout scope or timing. For search teams, the key issue is how AI summaries cite and rank UGC sources.

Why it matters: HKR-H is strong because Google AI summaries quoting Reddit alters the search surface. HKR-K has the perspectives mechanism, and HKR-R hits SEO/UGC traffic concerns; missing rollout scope keeps it in the 72–77 product-update band.

NVIDIA Blog

NVIDIA Spectrum-X AI-Native Ethernet Fabric Adds MRC for Gigascale AI

NVIDIA added MRC support to Spectrum-X Ethernet, letting one RDMA connection spread traffic across multiple paths. MRC ran in Blackwell deployments, with microsecond failure bypass and hardware rerouting. The key detail is the OCP open specification and multiplane support for clusters up to hundreds of thousands of GPUs.

Why it matters: HKR-K/R are solid: MRC stripes one RDMA flow across paths, detects failures in microseconds, and is tied to Blackwell deployments. HKR-H is narrow and the source is vendor-owned, so this stays below major release level.

Synced · WeChat

DeepSeek Version of Claude Code Tops Trending Chart With 8,700 Stars

DeepSeek TUI topped GitHub trending with over 8,700 stars. Hunter Bown built it in Rust for local terminal use with DeepSeek V4, supporting chat, file edits, shell commands, and task management. The key detail is RLM mode: up to 16 V4 Flash subtasks, plus a 1M-token context window and approval gates.

Why it matters: HKR-H/K/R all pass: the 8,700-star hook is strong, RLM adds concrete mechanisms, and coding-agent competition resonates. It is a third-party open-source tool, not an official DeepSeek model release, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

Salesforce plans to hire 1,000 graduates as agent roles expand

Salesforce CEO Marc Benioff said the company will hire 1,000 graduates or interns for Agentforce growth. The post cites Agentforce ARR up 169% to $800 million, with roles covering prompts, evals, agent supervision, and delivery. The key shift is entry roles moving from execution to agent orchestration and output checks.

Why it matters: HKR-H/K/R all pass: 1,000 junior hires, $800M Agentforce ARR, and 169% growth give concrete signal, with a strong jobs angle. This is Salesforce hiring plus Agentforce expansion, not a major model or product release.

TechCrunch · AI

Apple plans to make iOS 27 a Choose Your Own Adventure of AI models

Apple reportedly plans to let iOS 27 users choose third-party AI models for multiple tasks. The RSS snippet does not disclose model names, task scope, launch timing, or integration mechanics.

Why it matters: HKR-H/K/R pass: system-level model choice on iOS has a strong platform hook and distribution stakes. Kept in 72–77 because the RSS summary lacks model names, task scope, launch timing, and API mechanics.

The Verge · AI

Apple could let you pick a favorite AI model in iOS 27

Apple plans to let third-party chatbots run system-wide Apple Intelligence in iOS 27, iPadOS 27, and macOS 27. Mark Gurman says Extensions can handle Siri, Writing Tools, and Image Playground this fall. The post does not disclose supported models, pricing, or developer APIs.

Why it matters: HKR-H/K/R all pass: the Apple system-level model picker is a strong hook, with named Extension targets. Scored 80 because model list, pricing, and developer APIs are not disclosed, and this remains a roadmap report.

Financial Times · Technology

Meta plans advanced agentic AI assistant for consumers

Meta plans a consumer agentic AI assistant; the RSS body has one sentence. It says Meta is funding an OpenClaw counterpart for everyday task execution. The post does not disclose model size, launch timing, pricing, regions, or permission controls.

Why it matters: FT reports Meta plans a consumer agentic assistant, with HKR-H/K/R present. Details on launch, pricing, model, and permission design are missing, so this sits at the lower featured band.

NVIDIA Blog

NVIDIA and ServiceNow Partner on Autonomous AI Agents for Enterprises

NVIDIA and ServiceNow expanded their partnership with Project Arc, an enterprise desktop agent. It connects via Action Fabric and uses OpenShell for sandboxed, policy-governed execution. Blackwell delivers over 50x Hopper’s token output per watt and nearly 35x lower cost per million tokens.

Why it matters: HKR-K/R pass: the post gives mechanisms and Blackwell economics. HKR-H misses because the angle is a standard vendor partnership, so this sits in the 72–77 featured-threshold band.

May 5Tuesday

Hacker News front page

Show HN: Airbyte Agents – context for agents across multiple data sources

Airbyte launched Airbyte Agents, using Context Store to index operational data for agents. Its public benchmark reports up to 80% fewer tokens for Gong and 90% for Zendesk versus vendor MCPs. The key point is pre-indexed context, not another MCP wrapper.

Why it matters: HKR-H/K/R all pass: a concrete pre-indexing angle, reproducible claims, and agent data-access pain. Airbyte is not a frontier lab, so this stays at the lower featured band.

r/LocalLLaMA

vibevoice.cpp: Microsoft VibeVoice ported to ggml/C++ with no Python at inference

LocalAI released vibevoice.cpp, a ggml/C++ port of Microsoft VibeVoice for CPU, CUDA, Metal, and Vulkan inference. TTS uses a 30s reference clip for 24kHz cloned speech; ASR uses a 7B model with diarized JSON and was tested on 17min audio. The key constraint is memory: 17min CPU Q8_0 peaks near 26GB, with no streaming output yet.

Why it matters: HKR-H/K/R all pass: a practical open-source VibeVoice C++ port with concrete runtime numbers. Reddit-source scope and niche audio deployment keep it in the 72–77 featured band, not same-day must-write.

r/LocalLLaMA

Prompt injection benchmark: delimiter and strict prompt took Gemma 4 from 21% to 100% defense rate

A Reddit user posted a prompt-injection benchmark covering 15 models, 7 attack types, and 6,100+ cases. The setup wraps untrusted documents in long random delimiters; Gemma 4 E4B rose from 21.6% to 100% defense. The key detail is the reproducible metric: blocked/(blocked+failed).

Why it matters: HKR-H/K/R all pass: Gemma 4’s defense-rate jump is clickable, the test setup is concrete, and prompt injection matters to builders. Single Reddit benchmark keeps it in the 78–84 band.

r/LocalLLaMA

DeepSeek V4 Pro matches GPT-5.2 on FoodTruck Bench, 10 weeks later and about 17x cheaper

DeepSeek V4 Pro ranked No. 4 on FoodTruck Bench. The 30-day agentic benchmark uses 34 tools, persistent memory, and daily reflection; its median is within 3% of GPT-5.2 at about 17x lower workload cost. Xiaomi MiMo v2.5 Pro also ranked No. 6, with 5/5 survival, 1,019% median ROI, and $2.41 per run.

Why it matters: HKR-H/K/R all pass: the cost gap is clickable, and the post gives a 30-day, 34-tool setup plus a 17× cost delta. Single-source Reddit benchmark with no cross-validation keeps it in the 78–84 band.

Xinzhiyuan · WeChat

$1 for 10 Stars: ICSE Paper Exposes Fake GitHub Star Market

CMU researchers scanned GitHub events from July 2019 to Dec. 2024, flagging 6 million suspected fake stars. StarScout ran on about 20 TiB and found 18,617 repositories and 301,000 accounts. The supply-chain risk is concrete: GitHub deleted 90.42% of flagged repos, and about 30% of live samples were spam, phishing, or malware.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the study provides numbers and a detection mechanism, and GitHub trust is a practitioner nerve. Not a model or platform release, so it stays below the 85 must-write band.

Synced · WeChat

Agent-World Scales Real-World Environment Synthesis for Evolving General Agents

Agent-World builds 1,978 environments and 19,822 tools to train agents on long-horizon tasks. It combines web mining, tool generation, verifiable task synthesis, and GRPO training, with tasks averaging over 15 turns. The key signal is the scaling link among environment count, self-evolution rounds, and 23 benchmarks.

Why it matters: HKR-H/K/R all pass: Agent-World reports 1,978 environments, 19,822 tools, 15+ average turns, and 23 benchmarks. It is a strong agent research release, not a same-day must-write product launch.

r/LocalLLaMA

MTPLX: 2.24x Faster TPS Native MTP Inference Engine for Apple Silicon

MTPLX raises Qwen3.6-27B on a MacBook Pro M5 Max from 28 to 63 tok/s. The test used 4-bit MLX, temperature 0.6, top_p 0.95, top_k 20, with D3 as the best depth. The key detail is native MTP heads: no external drafter and no second-model memory.

Why it matters: HKR-H/K/R all pass: a 2.24x speed hook, concrete test conditions, and a local-inference cost nerve. Reddit single-post sourcing and narrow Apple Silicon scope keep it in low featured, not P1.

May 4Monday

r/LocalLLaMA

M3 Ultra + DGX Spark = M5 Ultra-lite?

A Reddit user benchmarked DGX Spark against M3 Ultra in llama.cpp at pp16384, with Spark 1.4× to 3.4× faster across 4 models. Qwen 27B hit 778 t/s vs 340 t/s, while Mistral 128B hit 241 t/s vs 72 t/s. The concrete tuning note is mmap=0: loading fell from minutes to about 20 seconds.

Why it matters: Single Reddit sourcing keeps the score low, but HKR-H/K/R all pass through a concrete local-inference benchmark. The pp16384 setup and 4-model speedups justify featured at the lower edge.

r/LocalLLaMA

Deep research report with Hermes Agent and qwen3.6-35b-a3b Q6_K

A Reddit user used Hermes Agent and qwen3.6-35b-a3b Q6_K to produce a 21-page research report. The run took 6 loops and over 5 hours on an RTX 4060, at about 28 tokens/s. The repo includes prompts, scripts, intermediate artifacts, and the final report.

Why it matters: HKR-H/K/R all pass: this is a local-agent experiment with hardware, runtime, speed, and artifacts. Reddit source limits reach, so it stays in the 72–77 featured-threshold band.

QbitAI · WeChat

DeepSeek-TUI, a “DeepSeek Claude Code,” reaches 2.3k GitHub stars

DeepSeek-TUI reached 2.3k GitHub stars; the Rust project is MIT-licensed. It targets DeepSeek V4 with a 1M-token context, RLM up to 16 V4 Flash subtasks, MCP, Shell, Git, and three control modes. Watch cache misses: uncached tokens cost 10x cached tokens.

Why it matters: HKR-H/K/R all pass: the hook is a DeepSeek-flavored Claude Code, with 2.3k stars, 1M tokens, 16 subtasks, and a 10x cache-miss cost gap. Impact is developer-specific, so it sits in the 72–77 band.

Xinzhiyuan · WeChat

Claude token rankings: Disney employee hits 460,000 calls in 9 days; Meta burns 60T monthly

Xinzhiyuan says Disney tracks Claude use via an AI Adoption Dashboard, with one employee making about 460,000 calls in 9 workdays. It also says Meta used 60 trillion tokens in 30 days, worth about $9B by public API pricing; the post does not show raw tables. The key issue is that input rankings are not outcomes.

Why it matters: HKR-H/K/R all pass: the hook is concrete usage shock, the post gives dashboard mechanics and token figures, and the nerve is enterprise Claude cost control. Kept at 74 because the data is secondhand and no raw table is disclosed.

最佳拍档 (BestPartners)

Why Claude Code Got Worse: Anthropic’s Review of Three Bugs

The title says Anthropic reviewed Claude Code regressions involving three bugs. It names reasoning-strength changes, a cache optimization error, and a system-prompt length limit; the post does not disclose repro steps, timeline, or fix status. The key point is AI reviewing AI code under engineering constraints.

Why it matters: HKR-H/K/R all pass, but the post gives three cause categories without repro steps, timeline, or fix status. Claude Code relevance is high, so this sits in the 72–77 band.

r/LocalLLaMA

Gemma 4 E2B runs well on an 8GB Android phone, powering a private voice notes app

A Reddit user ran Gemma 4 E2B locally on an 8GB OnePlus CE 5 and built a private voice notes app. Whisper Small 244MB transcribes, Gemma 4 E2B 2.4GB splits and tags, and a 10-15s note takes 12-15s end to end. Search uses query expansion, FTS lanes, RRF, and optional Gemma top-K reranking with a 15s fallback.

Why it matters: HKR-H/K/R all pass, but this is a Reddit first-person build, not an official Google release. Concrete hardware, latency, model size, and retrieval details place it near the top of the tutorial band.

May 3Sunday

r/LocalLLaMA

Local LLM Benchmark for Backend Generation via Function Calling: GLM vs Qwen vs DeepSeek

AutoBe posted a controlled backend-generation benchmark and says qwen3.5-35b-a3b matches gpt-5.4 on DB/API design. One shopping-mall run uses 200–300M tokens, costing $1,000–$1,500 per model at GPT 5.5 pricing. The key caveat is n=4 projects and self-scoring harness bias.

Why it matters: HKR-H/K/R all pass, but Reddit sourcing, n=4 projects, and self-eval harness bias keep it at the low featured band. Concrete cost and test constraints carry the score.

r/LocalLLaMA

LLM proxy that lets Claude Code talk to any model

DataNebula released open-source rosetta-llm, letting Claude Code call multiple providers through one gateway. It translates Anthropic Messages, OpenAI Chat, and OpenAI Responses, and round-trips encrypted reasoning via the signature field. The key detail is thinking-block fidelity for multi-turn agent prompt-cache hits.

Why it matters: HKR-H/K/R all pass, but this is a Reddit open-source tool post with no adoption, stars, or benchmark data disclosed. Score stays in the mid-weight tooling band, not 78+.

May 2Saturday

r/LocalLLaMA

I built Semvec: A constant-cost semantic memory for LLMs, looking for testers

A developer released Semvec, replacing unbounded chat history with fixed-size semantic state. Its 48-turn benchmark claims about 76% token reduction, with identical input footprint at turn 10 and 10,000. It supports OpenAI-compatible LLMs, MCP, Claude Code, Cursor, and multi-agent shared state.

Why it matters: HKR-H/K/R all pass, but this is a Reddit self-release with author benchmarks only. Treat it as an interesting indie memory tool, not a same-day industry story.

Hacker News front page

Show HN: Filling PDF Forms with AI Using Client-Side Tool Calling

SimplePDF released a Copilot demo that fills PDF forms via client-side tool calling; SimplePDF has 200k+ monthly users. PDFs stay in the browser, with parsing, rendering, and field detection local. The demo uses a DeepSeek V4 Flash proxy by default, with BYOK, cloud, or LM Studio options.

Why it matters: HKR-H/K/R pass: the client-side PDF-agent angle is specific, with a clear privacy mechanism and builder relevance. It sits in the 72–77 band as a useful product demo, not a major platform release.

r/LocalLLaMA

Qwen3.6-27B hits 72 tok/s on RTX 3090 with native vLLM on Windows

Reddit user One_Slip1455 released a native Windows vLLM launcher for Qwen3.6-27B, reaching 72 tok/s on an RTX 3090. It reports 64.5 tok/s at ~25k tokens, 53.4 tok/s at 127k ctx on one GPU, and 160k ctx with PP=2 on 2×3090. The key detail is no WSL or Docker, an OpenAI-compatible endpoint, and an INT4 quant path.

Why it matters: HKR-H/K/R all pass: native Windows on an RTX 3090 is the hook, the post gives tok/s and ctx figures, and it hits local-inference cost concerns. Reddit single-source limits it to the lower featured band.

QbitAI · WeChat

Apple Support App Accidentally Shipped Claude.md, Revealing Internal Claude Code Use

Apple Support v5.13 shipped a Claude.md file on May 1 and was pulled within 24 hours. The file describes Juno AI and Live Agents switching through a Protocol layer, with client, agent, and assistant messages handled in one flow. The key issue is release review; the post does not disclose how the file entered production.

Why it matters: HKR-H/K/R all pass, but this is still an app-packaging incident, not a model or platform release. Apple scale and Claude.md details clear the featured bar; the review-chain failure is not disclosed.

May 1Friday

r/LocalLLaMA

PFlash: 10x prefill speedup over llama.cpp at 128K on an RTX 3090

PFlash cuts Qwen3.6-27B Q4_K_M 128K TTFT to 24.8s on an RTX 3090, versus 248.4s cold for llama.cpp. It uses a Qwen3-0.6B drafter to score token importance, keeps 5% of spans, and runs C++/CUDA without Python, Triton, or PyTorch. The quality caveat is clear: only NIAH single-needle passes from 32K to 128K; RULER and multi-needle results are not disclosed.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit claim with quality evidence limited to single-needle NIAH 32K–128K. RULER and multi-needle results are not disclosed, so it stays at featured threshold.

The Verge · AI

Pentagon strikes classified AI deals with OpenAI, Google, and Nvidia, but not Anthropic

The Pentagon signed classified AI-use deals with 7 firms: OpenAI, Google, Microsoft, Amazon, Nvidia, xAI, and Reflection. Anthropic was excluded as a supply-chain risk; the post does not disclose contract value, model scope, or deployment terms.

Why it matters: HKR-H/K/R all pass: a classified Pentagon AI vendor list includes OpenAI, Google, Nvidia and 4 others, while Anthropic is absent. Contract value, model scope, and deployment terms are not disclosed, keeping it below 85.

r/LocalLLaMA

MiMo-V2.5-Pro: the actual best open-weights model

Reddit user cjami benchmarked Xiaomi MiMo-V2.5-Pro in autonomous Blood on the Clocktower games. It scored 88% as Good and 48% as Evil, with 183,639 output tokens per game, $0.99 cost, and a 0.4% tool-call error rate. The key comparison is Kimi K2.6: 580,000 tokens, $2.65, and 10–15 hours per game.

Why it matters: Single Reddit benchmark limits authority, so this is not a model-release story. HKR-H/K/R all pass via a named test with win rates, token counts, cost, and tool-error data, placing it in the 78–84 featured band.

The Verge · AI

Microsoft wants lawyers to trust its new AI agent in Word documents

Microsoft launched Legal Agent in Word for legal teams, focused on tasks such as contract review. It follows legal workflows, reviews clauses against a playbook, and handles tracked changes; the post does not disclose pricing or rollout scope.

Why it matters: HKR-H/K/R all pass: Word-native legal review is a sharp enterprise-agent angle, and the playbook plus tracked-changes mechanism adds substance. Price, rollout, and customer evidence are not disclosed, so it stays at the lower featured band.

Xinzhiyuan · WeChat

Claude Code's Real Story: 98.4% of What Works Is Engineering, Not AI

VILA-Lab analyzed 512,000 lines of Claude Code v2.1.88 and found 1.6% tied to AI decision logic. The other 98.4% is deterministic infrastructure: permissions, context, tool routing, and error recovery. The key shift is harness design, not longer prompts.

Why it matters: Strong HKR: the Claude Code teardown has a sharp counter-narrative and concrete 512k LOC plus 1.6%/98.4% split. It is not an official Anthropic release and lacks full reproduction details, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

OpenAI upgrades Codex to control Macs and run cross-app tasks

OpenAI upgraded Codex with Slack, Google Workspace, and Microsoft 365 integrations. Mike Russell tested Codex on a Mac across Adobe Audition, Photoshop, and Firefly, finishing in about 8 minutes with an 85–90 score. The key shift is OS-level computer control, not code completion.

Why it matters: All HKR axes pass: OpenAI Codex moves from coding into Mac-level control, with Slack, Google Workspace, and Microsoft 365 integrations. Single-source sourcing caps the score, but the 8-minute test and OS-agent angle justify P1.

Latent Space

[AINews] Agents for Everything Else: Codex for Knowledge Work, Claude for Creative Work

OpenAI expanded Codex to non-coding work, with CUA reported 42% faster. The update connects Microsoft, Google, and Salesforce, covering docs, slides, spreadsheets, research, and planning. The key signal is GUI-agent productization, not one benchmark score.

Why it matters: HKR-H/K/R all pass: Codex moves into non-code GUI work, with a 42% speed claim and named integrations. Price, rollout scope, and reproduction details are not disclosed, so it stays below P1.

r/LocalLLaMA

32x AMD MI50 32GB runs Kimi K2.6 at 9.7 t/s TG and 264 t/s PP

Reddit user ai-infos ran Kimi K2.6 int4 on 32 AMD MI50 32GB GPUs, reaching 9.7 tok/s TG on 136 output tokens. PP hit 263 tok/s on 14,564 input tokens using vllm-gfx906-mobydick across two 16-GPU nodes over 10G Ethernet. Power was about 640W idle and 4,800W peak inference; PCIe bandwidth and the vLLM distributed stack are the real bottlenecks.

Why it matters: HKR-H/K/R all pass via an unusual 32x MI50 build with concrete throughput, power, and network conditions. It stays in the 72–77 band because it is a niche Reddit benchmark, not a broader product or model release.

Hacker News front page

Show HN: Pu.sh – a full coding-agent harness in 400 lines of shell

Pu.sh ships a coding-agent harness in about 400 lines of shell, using only sh, curl, and awk. It supports Anthropic and OpenAI, 7 tools, REPL, auto-compaction, checkpoint/resume, pipe mode, and 90 no-API tests. It excludes TUI, streaming, images, OAuth, and Windows.

Why it matters: HKR-H/K/R all pass, but this is a small Show HN open-source tool, not a model or platform release. HN frontpage plus a reproducible 400-line implementation clears the featured bar.