Skip to content

MCP & tool use

How models connect to the outside world: the MCP ecosystem, function calling and tool integrations.

760 picksRelated topicsAgentsAI codingOpen source

Latest picks

601–620 of 760

Apr 14Tuesday

X · @dotey

Developer Can Vardar says disabling telemetry in Claude Code cuts prompt cache from 1 hour to 5 minutes

Can Vardar said disabling telemetry in Claude Code drops prompt cache from 1 hour to 5 minutes; Anthropic engineer Boris Cherny said the client then falls back to the 5-minute default because experiment flags stop working. The post says 1-hour cache costs more to write and less to read, so value depends on reuse; Anthropic plans env vars to force 1 hour or 5 minutes.

Why it matters: Strong HKR-H/K/R: the privacy-vs-performance tradeoff is a sharp hook, and the post adds concrete TTL and cache-cost mechanics. It scores as high featured because it affects real Claude Code usage decisions, but not P1 because this is an engineer clarification on X, not a formal,

Apr 13Monday

最佳拍档 (BestPartners)

2027 Is the Enterprise AI Singularity Year: Sundar Pichai on 10 Years as Google CEO, Transformer and Search

Sundar Pichai said in a Stripe interview that Alphabet plans $175B-$185B in 2026 capex and that 2027 will be the breakout year for enterprise AI agent workflows. He said Google cut Search latency by 30% over five years while adding AI features, manages teams with 10 ms or 30 ms latency budgets, and sees 2026-2027 constrained by wafers, memory, power, and permitting. The point to watch is not search replacement but search evolving into an agentic manager, while TPU allocation has become Google's scarcest internal resource.

Why it matters: High-signal executive commentary rather than a product launch. HKR-H/K/R all pass on the 2027 agent call, concrete capex and latency details, and the search-plus-compute nerve hit; score stays below P1 because this is a second-hand recap, not the primary interview.

Apr 12Sunday

X · @Yuchenj_UW

MiniMax M2.7 is open-source!

MiniMax open-sourced M2.7 and said its research agent now handles 30%–50% of the R&D workflow. The post says the agent covers literature review, experiment orchestration, log debugging, code fixes, and merge requests; M2.7 also rewrote its own harness for 100+ automated rounds, with a 30% gain on internal coding evals.

Why it matters: HKR-H/K/R all pass: open-sourcing plus a research agent doing 30%-50% of R&D is a strong hook, and the post includes 100+ self-rewrite loops with +30% internal coding eval. It stays at 78 because license, repo, benchmark context, and external reproduction are not disclosed.

Apr 11Saturday

X · @dotey

Anthropic launches Claude Managed Agents beta; Michael Cohen explains secure third-party key management for agents

Anthropic added Vaults to the Claude Managed Agents beta to manage each end user's third-party credentials with a per-user vault_id and automatic injection at session runtime. The post shows a three-step flow—create a Vault, bind credentials to an MCP server address, and pass vault_id when creating a session—and prices CMA at token usage plus $0.08 per session-hour. The key design is isolation: credentials never enter Claude's context window, code runs in a sandbox, auth goes through a dedicated proxy, and the harness cannot access secrets.

Why it matters: This adds the missing implementation detail for Claude Managed Agents: third-party credential isolation. HKR-H/K/R all pass via a concrete security hook, reproducible vault_id flow, pricing, and a real operator pain point; impact stays at the developer integration layer, so it is

X · @OpenAI

OpenAI says an Axios third-party library security issue prompted macOS app certificate updates

OpenAI said an Axios third-party library security issue led it to require all macOS users to update their OpenAI apps. The post says it found no evidence of user data access, system compromise, or software tampering; the change updates macOS app certificates to reduce fake app distribution risk. The post does not disclose affected versions or a timeline.

Why it matters: This is an official OpenAI desktop security incident with a concrete macOS mitigation, so HKR-H/K/R all land. It stays in the low featured band because the post does not disclose affected versions, exposure window, discovery date, or full remediation timeline.

X · @dotey

OpenAI Codex team's Nick Baumann: build dedicated CLI tools for AI instead of feeding messy data repeatedly

OpenAI Codex engineer Nick Baumann says teams should wrap repeated data access into parameterized CLI tools with JSON output instead of repeatedly dumping logs, docs, and API responses into Codex. The post lists 3 examples in daily use: codex-threads for past sessions, slack-cli for threaded Slack search, and typefully-cli for posting workflows; access still goes through the existing auth gateway. The point for practitioners is narrower interfaces: models handle focused commands more reliably than raw, noisy source data.

Why it matters: This is a practical workflow note from an OpenAI Codex team member, not a formal launch, but it offers a reusable mechanism: wrap noisy context behind parameterized JSON-returning CLIs and shows 3 live examples. HKR-H/K/R all land; no benchmark, scale, or major product release,so

X · @dotey

Anthropic launches Claude for Word beta add-in

Anthropic released a beta Claude for Word add-in for paid Claude Team and Enterprise users, with direct sidebar editing for .docx and .docm files. Edits appear in Word’s native track changes flow, the add-in can reuse conversation context from Excel and PowerPoint, and it supports reference uploads plus reusable team Skills. The key point is shared context across Office apps; the post does not disclose pricing, regions, or a wider rollout timeline.

Why it matters: This is a substantive Anthropic product update for Team and Enterprise, not a generic integration post. HKR-H/K/R all pass on novelty, concrete mechanics, and workflow resonance, but the beta scope is limited and price, regions, and GA timing are undisclosed, so it lands in mid-"

X · @dotey

Claude Code adds ultraplan: start planning in terminal, review in browser, then run in cloud or locally

Claude Code opened a preview of ultraplan to users with the web app enabled, requiring v2.1.91+, and planning starts from /ultraplan in the terminal. Claude drafts a plan in the cloud after reading the repo, users review and annotate it in the browser, then choose cloud execution with a PR or local terminal execution. The key change is splitting planning from execution: planning moves to the cloud without blocking the terminal, and the post says token use is close to local plan mode.

Why it matters: This is more than a routine feature add: Claude Code splits planning from execution, with /ultraplan in terminal, cloud-side repo reading, browser review, and cloud PR or local execution. HKR-H/K/R all pass, with a Claude-specific bump, but it is still a preview and sourced froma

X · @claudeai

Claude for Word is now in beta

Anthropic launched Claude for Word in beta, letting users draft, edit, and revise documents from the Word sidebar on Team and Enterprise plans. The post says Claude preserves formatting and shows edits as tracked changes; it does not disclose pricing, regions, or rollout timing.

Why it matters: This is a useful but mid-weight Anthropic product update. The official post confirms Word sidebar access, Team/Enterprise availability, format retention, and tracked changes; HKR-K and HKR-R pass, but missing price, region, and rollout details keep it at the low end of featured.

Apr 10Friday

X · @dotey

Anthropic launches Advisor Tool API: cheaper models execute while pricier models advise on hard decisions

Anthropic launched the advisor tool API, letting Sonnet or Haiku execute tasks and consult Opus on hard decisions; it is in beta and requires the anthropic-beta: advisor-tool-2026-03-01 header. The RSS snippet says Sonnet+Opus gains 2.7 points on multilingual SWE-bench while cutting per-task cost by 11.9%; Haiku+Opus rises from 19.7% to 41.2% on BrowseComp at 15% of Sonnet's cost. The key detail is the call path: model switching happens inside one Messages API request, advisor and executor tokens are billed separately, and max_uses caps consultations.

Why it matters: This is a substantive Anthropic API update with concrete mechanics: in-request model routing, separate token billing, max_uses, and two benchmark/cost deltas. HKR-H/K/R all pass, so it merits featured, but it is still below a model-release tier event.

X · @OpenAI

OpenAI updates ChatGPT Pro and Plus subscriptions to support growing Codex usage

OpenAI set a new ChatGPT Pro tier at $100/month and raised Codex usage to 5x ChatGPT Plus. The tier keeps all Pro features, including the exclusive Pro model and unlimited Instant and Thinking access. Through May 31, $100 Pro subscribers get up to 10x Plus usage on Codex; the real signal is separate pricing for heavy code-agent demand.

Why it matters: This is an OpenAI product-pricing update centered on Codex usage, with HKR-K from concrete pricing/quota facts and HKR-R from a clear signal on code-agent monetization. No new model or capability is disclosed, and HKR-H is weaker, so it lands as solid featured rather than must-wr

X · @claudeai

Claude Cowork is now generally available to all paid plans.

Anthropic made Claude Cowork generally available on all paid plans. For Enterprise, it added role-based access controls, group spend limits, usage analytics, and expanded OpenTelemetry; the post does not disclose pricing, quotas, or rollout dates. The key signal is stronger admin control for org-wide deployment, but finer deployment parameters are still undisclosed.

Why it matters: Official Anthropic product update. HKR-K is supported by four concrete enterprise controls, and HKR-R lands because teams care about permissions, spend, and observability. Score stays moderate because price, quotas, and rollout timing are not disclosed, and this is not a model-cp

Apr 9Thursday

QbitAI · WeChat

Claude launches managed agent service, then faces an open-source alternative from Multica

Anthropic has opened Claude Managed Agents and charges $0.08 per session-hour plus token usage. The service supports hours-long runs, sandboxing, checkpoint recovery, and multi-agent orchestration; web search costs $10 per 1,000 searches, while some memory and orchestration features remain in research preview. The title mentions a “lobster ban,” but the post does not disclose that context; the real shift is Anthropic selling agent infrastructure to enterprises.

Why it matters: This is a substantive Anthropic product update: a managed-agent service with explicit pricing, sandboxing, checkpoint restore, and search costs, so HKR-K is strong. HKR-R is real for Claude-heavy teams weighing build vs. buy, but the scope is smaller than a model launch, so it is

X · @dotey

Anthropic launches Claude Managed Agents, a managed API for building and deploying agents, now in public beta

Anthropic launched Claude Managed Agents, a managed API for building and deploying agents, in public beta. It offers a production sandbox, long-running sessions, and multi-agent coordination; Anthropic says internal tests showed up to a 10-point success-rate gain on structured file-generation tasks versus standard prompt loops. Pricing uses standard Claude token fees plus $0.08 per active session-hour; the real signal is Anthropic moving agent infrastructure into its platform layer.

Why it matters: Anthropic packaged managed agents, sandboxing, and long-running sessions into a public-beta API, which is a real workflow update for developers. HKR-H/K/R all pass: strong platform hook, concrete facts like a 10-point gain and $0.08 per hour, and clear resonance around developer-

X · @claudeai

Introducing Claude Managed Agents: everything you need to build and deploy agents at scale.

Claude has launched Claude Managed Agents in public beta on Claude Platform, claiming to compress the path from agent prototype to launch into days. The post discloses only a performance-tuned agent harness plus production infrastructure; pricing, toolchain support, model scope, and quotas are not disclosed.

Why it matters: Anthropic gets a positive bump, and HKR-H/HKR-R pass because managed agent deployment is a strong hook for Claude-heavy builders. HKR-K is limited: the post discloses a harness and prod infra, but not pricing, toolchain support, model scope, or quotas.

Apr 8Wednesday

QbitAI · WeChat

Xiaomi unveils two AI audio frameworks: Any2Speech and Midasheng-audio-generate

Xiaomi's large-model application team introduced Xiaomi Any2Speech and Midasheng-audio-generate. Any2Speech generates up to about 10 minutes per inference, while the other model turns one text prompt into mixed audio with speech, music, and ambient sound. The post names GST labeling, dual-path planning with dimension dropout, Flow Matching, and five-field structured labels; benchmark scores, training scale, and commercial terms are not disclosed.

Why it matters: Xiaomi released two audio-generation frameworks with a clear hook and concrete mechanisms, so HKR-H and HKR-K pass. HKR-R is weaker because benchmark results, training data scale, open-source status, and commercial terms are not disclosed, so this sits at the low end of featured.

QbitAI · WeChat

Free open-source 2B Chinese speech model reproduces Mangzhuang Ren with high-speed tonguetwisters

ModelBest, OpenBMB, and Tsinghua University released VoxCPM 2, a 2B open speech model that supports 9 Chinese dialects, 30 foreign languages, and 48kHz audio. The post says generation often finishes within 1 second, recommends reference audio of at least 5 seconds, and supports denoising, LoRA, and full fine-tuning; the key detail is its tokenizer-free diffusion autoregressive continuous representation design.

Why it matters: This is a substantive open-source speech release, not a thin demo: the post gives 2B, 48kHz, 9 Chinese dialects, 30 languages, ref audio ≥5s, and a tokenizer-free route. HKR-H/K/R all pass, but the event is not large enough for a must-write P1.

Latent Space

Extreme Harness Engineering for Token Billionaires: 1M LOC, 1B toks/day, 0% human code, 0% human review

OpenAI Frontier says it built an internal beta over five months with a repo above 1M LOC, over 1B tokens per day, and 0% human-written or human-reviewed code before merge. The post says the team treated failures as missing capability, context, or structure, then used Symphony orchestration, specs, tests, observability, and sub-1-minute build loops to constrain Codex. The shift to watch is from humans reviewing code to humans designing the harness; the $2k-$3k/day cost is cited secondhand in the post.

Why it matters: HKR-H/K/R all pass: the headline is clickworthy, and the piece includes concrete workflow details plus scale numbers. It stays below p1 because this is an interview-style report, not an official launch, and key claims like 1B tokens/day and cost lack independent verification.

Apr 6Monday

X · @dotey

Xiaomi MiMo lead Luo Fuli on token costs in the Agent era

Luo Fuli said Agent workloads can resend 100k+ tokens across repeated tool calls, and global compute cannot keep up with that burn. She said OpenClaw makes several times more requests than Claude Code and can push real API cost to tens of times the subscription price; the post does not disclose a pricing formula.

Why it matters: A named Xiaomi MiMo lead makes a concrete, testable critique of agent cost: 100k+ token context replay, multi-tool-call overhead, and several-times request inflation vs Claude Code. HKR-H/K/R all pass, but missing public benchmark setup and pricing keeps it at the low end of the

Apr 4Saturday

X · @dotey

Anthropic ends Claude subscription coverage for third-party tools like OpenClaw

Anthropic said that from 12:00 pm PT on April 4, Claude Pro and Max subscriptions will no longer cover usage generated through third-party tools such as OpenClaw. Existing subscribers get a one-time credit equal to one month of fees; extra usage must go through prepaid credits or usage-based API keys, and refund links will be emailed. The key point is enforcement is now complete: Anthropic added technical blocks in January and banned third-party OAuth token use in February terms.