Skip to content

MCP & tool use

How models connect to the outside world: the MCP ecosystem, function calling and tool integrations.

760 picksRelated topicsAgentsAI codingOpen source

Latest picks

621–640 of 760

Apr 4Saturday

X · @dotey

Mintlify uses ChromaFs to make AI document retrieval look like a file system

Mintlify routes its AI doc assistant’s grep, cat, and ls calls through ChromaFs into database queries, cutting session startup from 46s to 100ms and pushing marginal compute cost per chat near zero. Built on Vercel Labs’ just-bash, it maps pages to files and sections to directories; at 850,000 chats per month, replacing real sandboxes saves over $70,000 a year in compute. The real shift is retrieval design: not faster vector RAG, but model-led exploration of structured docs, and the post says this may not fit messy knowledge bases.

Why it matters: This is a substantive engineering write-up, not a routine product note. HKR-H/K/R all pass: the fake-filesystem angle is novel, the post includes hard numbers (46s→100ms, 850k chats/month, >$70k/yr), and it hits operator concerns around latency, cost, and retrieval design; strong

Apr 3Friday

X · @claudeai

Microsoft 365 connectors are now available on every Claude plan

Anthropic made Microsoft 365 connectors available on every Claude plan, covering Outlook, OneDrive, and SharePoint. The post confirms plan coverage and supported apps; it does not disclose pricing, permission boundaries, regional limits, or admin requirements. The real signal is broad rollout across all plans, not a new standalone connector.

Why it matters: This is a mid-weight Claude product update: Anthropic expanded Microsoft 365 connectors to every Claude plan, which changes real Outlook, OneDrive, and SharePoint access. HKR-H/K/R all pass, but missing price, permission, region, and admin details keeps it at low-end featured.

X · @op7418

Karpathy shared how he builds a local AI knowledge base

Karpathy uses Obsidian and local Markdown to build a personal wiki, stores source material in a RAW folder, then has an LLM generate summaries, indexes, concept pages, links, and visualizations. The setup can answer questions over the wiki and write reports or new files, but the post also says AI-generated content can pollute the corpus and should be separated from trusted sources; the post does not disclose the model, scale, or automation details.

Why it matters: HKR-H and HKR-R land because Karpathy’s local-first wiki workflow is inherently clickable and discussable for AI practitioners. HKR-K lands on the RAW→LLM→summary/index/link mechanism, but missing model, corpus size, and automation details keep it in the mid-70s.

X · @op7418

Google releases Gemma 4 for on-device use under Apache 2.0

Google released Gemma 4 in four variants—E2B, E4B, 26B MoE, and 31B Dense—targeting phones, edge devices, and up to single-H100 workstations. The RSS snippet says the 26B MoE activates 3.8B parameters and adds native function calling, JSON output, multimodal I/O, speech-to-text, and Apache 2.0 licensing; the post does not disclose benchmarks, context length, or rollout details.

Why it matters: Google releasing Gemma 4 is a substantive open-model update. HKR-H/K/R all pass on the size spread, 3.8B-active MoE detail, and deployment-cost relevance; it stays at 81 because benchmarks, context window, and test conditions are not disclosed here.

X · @claudeai

Computer use in Claude Cowork and Claude Code Desktop is now available on Windows

Claude has brought computer use in Claude Cowork and Claude Code Desktop to Windows. The post confirms the Windows rollout, but does not disclose supported versions, permission model, latency, pricing, or release timing. What matters is the reliability boundary for desktop agents on Windows, and the post gives no reproducible conditions yet.

Why it matters: HKR-H lands on the Windows rollout hook, and HKR-R lands because desktop agents on Windows map to real workflows. Score stays at 74: this is an official Claude update, but the post confirms availability only; versions, permissions, latency, and price are not disclosed.

X · @OpenAI

ChatGPT is now available in CarPlay

OpenAI is rolling out ChatGPT in CarPlay to iPhone users on iOS 26.4+ where CarPlay is supported. The post confirms voice mode is available in-car, but does not disclose regions, vehicle coverage, or feature limits. The key shift is distribution into the driving interface, not a new model launch.

Why it matters: This matters more as a distribution-surface shift than a model update. HKR-H and HKR-R pass on the CarPlay hook and assistant-entry competition; HKR-K stays limited because the post gives iOS 26.4+ rollout only, not regions, car support, or full feature bounds.

Apr 1Wednesday

MIT Technology Review · AI

The gig workers who are training humanoid robots at home

Micro1 hires thousands of contractors across 50+ countries to film chores at home with iPhones and sell that real-world data to humanoid robotics companies. The piece cites $15/hour pay for one worker, says robotics firms spend over $100 million a year on such data, and notes $6 billion+ went into humanoids in 2025. The real issue is data governance: workers know the footage trains robots, but the post shows they often do not know how it is stored, shared, or deleted.

Why it matters: This clears HKR-H/K/R: at-home chore videos are a strong hook, and the piece adds numbers on scale, pay, and spend. The sharper industry signal is the hidden data pipeline and weak governance on storage, sharing, and deletion, so it merits featured, not p1.

Mar 24Tuesday

Lex Fridman (YouTube RSS)

Jensen Huang: NVIDIA - The $4 Trillion Company & the AI Revolution | Lex Fridman Podcast #494

Jensen Huang said on the Lex Fridman podcast that NVIDIA uses “extreme co-design” for AI clusters, aiming to beat linear scaling across 10,000 computers. The interview cites Amdahl’s Law, model and data sharding, networking, power, and cooling as hard constraints; Huang also said he has 60+ direct reports. The key shift is that NVIDIA now competes at rack and data-center level, not only at single-GPU level.

Why it matters: A strong primary-source interview with clear HKR-H/K/R: a high-click hook, concrete system-scaling details, and direct relevance to the infra moat debate. It stays below 85 because this is analysis from a podcast, not a new product, personnel move, or fresh market-reported data.

Mar 20Friday

MIT Technology Review · AI

The Download: OpenAI is building a fully automated researcher, and a psychedelic trial blind spot

OpenAI says it plans to build an autonomous AI research intern by September 2026 for a small set of research problems, ahead of a multi-agent automated researcher targeted for 2028. The RSS snippet gives the timeline and staged plan, but the post does not disclose evals, compute budget, or research scope. The real question is whether the agent can produce verifiable research output.

Why it matters: HKR-H lands on the “fully automated researcher” hook, HKR-K on the two roadmap dates, and HKR-R on research-job substitution plus lab rivalry. It stays below must-write because the post does not disclose benchmarks, compute budget, or scope, so this is a strong roadmap signal, no

Mar 19Thursday

Ben's Bites

What makes a good AGENTS.md?

Ben's Bites says AGENTS.md should keep only behavior preferences, not tech-stack maps or key files; the post cites a study saying that hurts performance and raises cost by 20%. It recommends symlinking AGENTS.md to CLAUDE.md, using conditional blocks, and relying on folder-level dynamic loading; the study name and setup are not disclosed. The real point is not more context, but smaller persistent instructions.

Why it matters: This is a practitioner explainer for coding-agent users, not a product launch. HKR-K and HKR-R pass on the concrete 'keep AGENTS.md small' claim, the 20% cost figure, and usable patterns; HKR-H is weak, and the cited study name and setup are not disclosed, so it sits at the low '

TheValley101 (硅谷101)

Web3 101 Crossover: How to Prevent System-Level Risks Behind the OpenClaw Craze

Yuxian said OpenClaw has issued about 250 security advisories, and v3.2 added stricter defaults, yet broad permissions, network access, and Skill installs still expand risks like file deletion, data leaks, and loss of control. The discussion breaks risk into layers: readable local files, chat data sent upstream, logged-in browser sessions, malicious links or Skills, and automated tasks that fail repeatedly. The practical rule is isolation: separate devices or networks, local-only access or Tailscale, and strict caution with external inputs.

Mar 17Tuesday

OpenAI News

Introducing GPT-5.4 mini and nano

OpenAI released GPT-5.4 mini and nano on March 17, 2026 for coding and subagents; mini runs over 2x faster than GPT-5 mini. In the API, mini has a 400k context window and costs $0.75/$4.50 per 1M input/output tokens, while nano is API-only at $0.20/$1.25. The key signal is performance per latency: mini scores 54.4% on SWE-Bench Pro versus GPT-5.4 at 57.7%.

Why it matters: This is an official OpenAI model launch, not a routine patch. It includes concrete numbers—>2x speed, 400k context, API pricing, and 54.4% vs 57.7% on SWE-Bench Pro—so HKR-H/K/R all pass; scored at the low end of the 85–94 band.

MIT Technology Review · AI

Where OpenAI’s technology could show up in Iran

Just over two weeks after OpenAI’s classified-use deal with the Pentagon, MIT Technology Review outlined three places its tech could surface in Iran-related conflict. The post names target prioritization, Anduril counter-drone analysis, and GenAI.mil back-office use; it does not disclose when classified integration will finish or confirm deployment in Iran.

Why it matters: MIT Technology Review maps OpenAI’s classified-defense deal to 3 Iran-linked scenarios, giving it strong HKR-H and HKR-R. HKR-K is weaker because the piece does not confirm deployment, integration timing, or system limits, so it lands at the featured floor.

Mar 16Monday

MIT Technology Review · AI

Nurturing agentic AI beyond the toddler stage

The article says no-code tools and the open-source agent OpenClaw pushed agentic AI into a more autonomous stage between Dec. 2025 and Jan. 2026. It cites California AB 316 taking effect on Jan. 1, 2026, so firms cannot dodge liability by blaming AI, and an IDC survey sponsored by Data Robot reporting 96% of generative AI deployments and 92% of agentic AI deployments cost more than expected. The real issue is workflow-level governance: permission drift, orphaned agents, long-lived tokens, and sessions that can reach $100,000.

Mar 11Wednesday

MIT Technology Review · AI

Hustlers are cashing in on China’s OpenClaw AI craze

Beijing engineer Feng Qingyang turned OpenClaw installation support into a 100+ person business after starting in January, handling 7,000 orders at about RMB 248 each. Taobao and JD now show hundreds of related listings priced at RMB 100-700; the real story is setup friction and data-isolation risk turning an open-source agent into a service market.

Why it matters: Featured. HKR-H/K/R all pass: the side-gig-to-100-person-team angle is clickworthy, the piece adds hard market numbers, and the data-isolation risk gives it real industry resonance. This is not a product launch, but it is strong field reporting.

OpenAI News

From model to agent: Equipping the Responses API with a computer environment

OpenAI said on March 11, 2026 that Responses API now works with a shell tool and hosted container workspace, so models can execute commands in an isolated loop. The post says GPT-5.2 and later are trained to propose shell commands, while the API streams outputs and can run multiple commands concurrently across sessions; the container includes a filesystem, optional SQLite, and restricted network access. The key change is orchestration, not the “agent” label; pricing, quotas, and full security details are not disclosed in the visible post.

Why it matters: Substantive OpenAI developer update: the Responses API moves from tool calls to a managed computer environment with shell execution, streaming, parallel runs, and context compaction, so HKR-H/K/R all pass. The post is truncated and omits pricing, quotas, and full safety details,【

Mar 10Tuesday

NVIDIA Blog

NVIDIA and Thinking Machines Lab Announce Long-Term Gigawatt-Scale Strategic Partnership

NVIDIA and Thinking Machines Lab formed a multiyear deal to deploy at least 1 gigawatt of NVIDIA Vera Rubin systems, targeted for early next year, for frontier model training and customizable AI platforms. The partnership also covers training and serving system design for NVIDIA architectures and broader access to frontier and open models for enterprises and researchers; the post does not disclose the investment size. The key signal is the explicit 1-gigawatt compute commitment, not a routine cloud purchase.

Why it matters: The 1GW Vera Rubin commitment lifts this above routine partnership PR: HKR-H on scale, HKR-K on a named system with a dated deployment target, and HKR-R on frontier compute competition. It stays below P1 because the source is a vendor blog and key details—spend, ownership, and ph

OpenAI News

New ways to learn math and science in ChatGPT

OpenAI launched interactive math and science visualizations in ChatGPT on March 10, 2026, covering 70+ core concepts and rolling out globally across all plans. Users can adjust variables, manipulate formulas, and see graphs update in real time; OpenAI says 140 million people use ChatGPT weekly for math and science learning. The key point is productized interactivity, while the post does not disclose the underlying model, evaluation method, or outcome data.

Why it matters: HKR-H lands on the interactive-visual hook, HKR-K on 140M weekly learners plus 70+ concepts and live manipulation, and HKR-R on the product and edtech nerve. It is still a mid-weight product update; model details and learning-outcome evaluation are not disclosed, so it stays in a

Mar 9Monday

MIT Technology Review · AI

How AI Is Turning the Iran Conflict Into Theater

The author reviewed more than a dozen Iran-war dashboards in one week and argues they turn satellite data, ship tracking, AI summaries, and betting links into a real-time war spectator interface. The post cites a dashboard built by two Andreessen Horowitz staffers that pulls in Kalshi bets, while Craig Silverman has logged 20 similar dashboards. The point to watch is information quality: the piece cites Financial Times reporting on AI-generated satellite images spreading online, while these dashboards lack the human vetting and historical context used by intelligence agencies.

Why it matters: HKR-H lands on the war-dashboard-plus-betting hook; HKR-K lands on the named examples, counts, and Kalshi mechanism; HKR-R lands on reliability and ethics nerves for AI builders. Strong reported commentary, but not a product, model, or research milestone, so it ranks as featured,

OpenAI News

OpenAI to acquire Promptfoo

OpenAI said it will acquire Promptfoo and integrate its technology into OpenAI Frontier after closing. The post discloses that Promptfoo is used by over 25% of Fortune 500 companies, and the deal is still subject to customary closing conditions. The key signal is native agent security testing, red-teaming, and traceability in Frontier; the post does not disclose price or timeline.

Why it matters: This is not a routine partnership; OpenAI is absorbing a known eval and red-team vendor into Frontier. HKR-H/K/R all pass on novelty, concrete adoption data, and strong resonance with agent teams, but price, timing, and integration scope are still undisclosed, so it stays below p