Skip to content

MCP & tool use

How models connect to the outside world: the MCP ecosystem, function calling and tool integrations.

760 picksRelated topicsAgentsAI codingOpen source

Latest picks

421–440 of 760

May 2Saturday

r/LocalLLaMA

Qwen3.6-27B hits 72 tok/s on RTX 3090 with native vLLM on Windows

Reddit user One_Slip1455 released a native Windows vLLM launcher for Qwen3.6-27B, reaching 72 tok/s on an RTX 3090. It reports 64.5 tok/s at ~25k tokens, 53.4 tok/s at 127k ctx on one GPU, and 160k ctx with PP=2 on 2×3090. The key detail is no WSL or Docker, an OpenAI-compatible endpoint, and an INT4 quant path.

Why it matters: HKR-H/K/R all pass: native Windows on an RTX 3090 is the hook, the post gives tok/s and ctx figures, and it hits local-inference cost concerns. Reddit single-source limits it to the lower featured band.

QbitAI · WeChat

Apple Support App Accidentally Shipped Claude.md, Revealing Internal Claude Code Use

Apple Support v5.13 shipped a Claude.md file on May 1 and was pulled within 24 hours. The file describes Juno AI and Live Agents switching through a Protocol layer, with client, agent, and assistant messages handled in one flow. The key issue is release review; the post does not disclose how the file entered production.

Why it matters: HKR-H/K/R all pass, but this is still an app-packaging incident, not a model or platform release. Apple scale and Claude.md details clear the featured bar; the review-chain failure is not disclosed.

May 1Friday

r/LocalLLaMA

PFlash: 10x prefill speedup over llama.cpp at 128K on an RTX 3090

PFlash cuts Qwen3.6-27B Q4_K_M 128K TTFT to 24.8s on an RTX 3090, versus 248.4s cold for llama.cpp. It uses a Qwen3-0.6B drafter to score token importance, keeps 5% of spans, and runs C++/CUDA without Python, Triton, or PyTorch. The quality caveat is clear: only NIAH single-needle passes from 32K to 128K; RULER and multi-needle results are not disclosed.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit claim with quality evidence limited to single-needle NIAH 32K–128K. RULER and multi-needle results are not disclosed, so it stays at featured threshold.

The Verge · AI

Pentagon strikes classified AI deals with OpenAI, Google, and Nvidia, but not Anthropic

The Pentagon signed classified AI-use deals with 7 firms: OpenAI, Google, Microsoft, Amazon, Nvidia, xAI, and Reflection. Anthropic was excluded as a supply-chain risk; the post does not disclose contract value, model scope, or deployment terms.

Why it matters: HKR-H/K/R all pass: a classified Pentagon AI vendor list includes OpenAI, Google, Nvidia and 4 others, while Anthropic is absent. Contract value, model scope, and deployment terms are not disclosed, keeping it below 85.

r/LocalLLaMA

MiMo-V2.5-Pro: the actual best open-weights model

Reddit user cjami benchmarked Xiaomi MiMo-V2.5-Pro in autonomous Blood on the Clocktower games. It scored 88% as Good and 48% as Evil, with 183,639 output tokens per game, $0.99 cost, and a 0.4% tool-call error rate. The key comparison is Kimi K2.6: 580,000 tokens, $2.65, and 10–15 hours per game.

Why it matters: Single Reddit benchmark limits authority, so this is not a model-release story. HKR-H/K/R all pass via a named test with win rates, token counts, cost, and tool-error data, placing it in the 78–84 featured band.

The Verge · AI

Microsoft wants lawyers to trust its new AI agent in Word documents

Microsoft launched Legal Agent in Word for legal teams, focused on tasks such as contract review. It follows legal workflows, reviews clauses against a playbook, and handles tracked changes; the post does not disclose pricing or rollout scope.

Why it matters: HKR-H/K/R all pass: Word-native legal review is a sharp enterprise-agent angle, and the playbook plus tracked-changes mechanism adds substance. Price, rollout, and customer evidence are not disclosed, so it stays at the lower featured band.

Xinzhiyuan · WeChat

Claude Code's Real Story: 98.4% of What Works Is Engineering, Not AI

VILA-Lab analyzed 512,000 lines of Claude Code v2.1.88 and found 1.6% tied to AI decision logic. The other 98.4% is deterministic infrastructure: permissions, context, tool routing, and error recovery. The key shift is harness design, not longer prompts.

Why it matters: Strong HKR: the Claude Code teardown has a sharp counter-narrative and concrete 512k LOC plus 1.6%/98.4% split. It is not an official Anthropic release and lacks full reproduction details, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

OpenAI upgrades Codex to control Macs and run cross-app tasks

OpenAI upgraded Codex with Slack, Google Workspace, and Microsoft 365 integrations. Mike Russell tested Codex on a Mac across Adobe Audition, Photoshop, and Firefly, finishing in about 8 minutes with an 85–90 score. The key shift is OS-level computer control, not code completion.

Why it matters: All HKR axes pass: OpenAI Codex moves from coding into Mac-level control, with Slack, Google Workspace, and Microsoft 365 integrations. Single-source sourcing caps the score, but the 8-minute test and OS-agent angle justify P1.

Latent Space

[AINews] Agents for Everything Else: Codex for Knowledge Work, Claude for Creative Work

OpenAI expanded Codex to non-coding work, with CUA reported 42% faster. The update connects Microsoft, Google, and Salesforce, covering docs, slides, spreadsheets, research, and planning. The key signal is GUI-agent productization, not one benchmark score.

Why it matters: HKR-H/K/R all pass: Codex moves into non-code GUI work, with a 42% speed claim and named integrations. Price, rollout scope, and reproduction details are not disclosed, so it stays below P1.

r/LocalLLaMA

32x AMD MI50 32GB runs Kimi K2.6 at 9.7 t/s TG and 264 t/s PP

Reddit user ai-infos ran Kimi K2.6 int4 on 32 AMD MI50 32GB GPUs, reaching 9.7 tok/s TG on 136 output tokens. PP hit 263 tok/s on 14,564 input tokens using vllm-gfx906-mobydick across two 16-GPU nodes over 10G Ethernet. Power was about 640W idle and 4,800W peak inference; PCIe bandwidth and the vLLM distributed stack are the real bottlenecks.

Why it matters: HKR-H/K/R all pass via an unusual 32x MI50 build with concrete throughput, power, and network conditions. It stays in the 72–77 band because it is a niche Reddit benchmark, not a broader product or model release.

Hacker News front page

Show HN: Pu.sh – a full coding-agent harness in 400 lines of shell

Pu.sh ships a coding-agent harness in about 400 lines of shell, using only sh, curl, and awk. It supports Anthropic and OpenAI, 7 tools, REPL, auto-compaction, checkpoint/resume, pipe mode, and 90 no-API tests. It excludes TUI, streaming, images, OAuth, and Windows.

Why it matters: HKR-H/K/R all pass, but this is a small Show HN open-source tool, not a model or platform release. HN frontpage plus a reproducible 400-line implementation clears the featured bar.

Bloomberg Technology

AI Payoff in Focus During Tech Earnings Bonanza | Bloomberg Tech 4/30/2026

Bloomberg Tech covered AI payoff in tech earnings, saying Alphabet and Amazon show clearer returns while Meta lags. Anthropic is weighing funding at a valuation above $900B; Stripe’s John Collison discussed AI tools and a Google partnership. The post does not disclose AI spending amounts.

Why it matters: HKR-H/K/R pass: Bloomberg frames AI ROI by company and cites a >$900B Anthropic funding valuation. The score stays below 78 because the post is a roundup video and AI spending figures are not disclosed.

TechCrunch · AI

After Dissing Anthropic for Limiting Mythos, OpenAI Restricts Access to Cyber, Too

OpenAI will first roll out GPT-5.5 Cyber only to “critical cyber defenders.” The RSS snippet does not disclose eligibility rules, pricing, or launch timing. The access-tiering model is the key detail for practitioners.

Why it matters: HKR-H/K/R all pass, but the body is RSS-only: it confirms tiered access for GPT-5.5 Cyber, not criteria, pricing, or timeline. This fits a lower-featured OpenAI safety product update.

TechCrunch · AI

Stripe introduces Link, a digital wallet autonomous AI agents can use

Stripe introduced Link, a digital wallet for cards, banks, subscriptions, and AI-agent spending. The post cites approval flows, but does not disclose fees, limits, or merchant coverage. Watch the authorization boundary for agent payments.

Why it matters: HKR-H/K/R pass: agent wallet payments are clickable, the approval-control mechanism is concrete, and spend authorization is a live practitioner concern. Missing rates, limits, and merchant coverage keep it in the 72–77 band.

The Verge · AI

Meta is running get-rich-quick ads for its AI tools

The Verge says Meta-owned Manus ran quick-money ads for AI tools after a $2B acquisition. The pitch targets local firms with no or bad websites. Manus also paid creators for Instagram, YouTube, and TikTok promotion; some TikTok accounts were removed after inquiry.

Why it matters: HKR-H/K/R all pass: the story has a strong Meta-versus-grift hook, concrete funnel details, and reputational stakes. It is investigative industry reporting, not a major model or product release.

Apr 30Thursday

Ben's Bites

Building Gets Easier

Ben’s Bites lists agent tooling updates from Cloudflare, Stripe, Cursor SDK and others, with over 10 product leads. Cloudflare lets agents create accounts, buy domains, get API tokens and deploy; Stripe adds Agentic Commerce Suite, Link CLI and agent-ready Treasury accounts. The key shift is external permissions becoming agent-readable interfaces.

Why it matters: HKR-H/K/R pass, but this is a roundup rather than one major launch. Concrete Cloudflare and Stripe agent-permission details keep it in the featured-low band.

r/LocalLLaMA

Notes on what actually breaks when you run a coding agent on small local models

A Reddit user tested small local and free-tier cloud models for weeks on multi-file coding tasks. Sub-7B structured output was unreliable; failures included markdown fences, wrong-file edits, and read/write misclassification, with post-processing and validation as fixes.

Why it matters: HKR-H/K/R pass: the post names real local coding-agent failure points, a sub-7B threshold, four failure classes, and mitigations. Reddit single-post scope keeps it below release-tier news, so 75.

r/LocalLLaMA

Qwen-Scope: Official Sparse Autoencoders (SAEs) for Qwen 3.5 models

Qwen Team released Qwen-Scope, SAEs for Qwen 3.5 models from 2B to 35B MoE. It maps residual-stream features across all layers, including Feature #6159 for Chinese activation. The key point is feature-level debugging and steering; the license discourages removing safety filters.

Why it matters: HKR-H/K/R all pass: official Qwen SAEs are novel, concrete, and useful for interpretability work. This is not a new model release, so it stays in the 78–84 recommendation band.

Xinzhiyuan · WeChat

AI Raw Proofs Pile Up on GitHub as Terence Tao Says Solving Alone Is Not Enough

Terence Tao says math is shifting from proof scarcity to proof abundance, with 20-plus AI solutions pending assessment on an Erdős problems GitHub page. The post says GPT-5.4 Pro generated an Erdős #1196 approach in 80 minutes, and Tao verified the core within 24 hours. The key issue is verification and digestion workflow, not raw proof count.

Why it matters: All HKR axes pass: Tao plus GitHub proof backlog gives HKR-H, while 20+ pending AI solutions and an 80-minute GPT-5.4 Pro claim give HKR-K. This is not a model release, so it stays below 85.

TechCrunch · AI

Microsoft says it has over 20M paid Copilot users, and they really are using it

Microsoft says Copilot has over 20M paid users, with engagement growing. The post does not disclose active usage, retention, ARPU, or the counting method.

Why it matters: HKR-K is strong because Microsoft disclosed 20M+ paid Copilot users, a rare adoption metric. The score stays near the featured floor because active rate, retention, ARPU, and methodology are not disclosed.