Skip to content

MCP & tool use

How models connect to the outside world: the MCP ecosystem, function calling and tool integrations.

760 picksRelated topicsAgentsAI codingOpen source

Latest picks

581–600 of 760

Apr 16Thursday

Hacker News front page

Andon Labs gave an AI a 3-year retail lease in San Francisco and asked it to make a profit

Andon Labs gave AI agent Luna a 3-year retail lease on Union St in San Francisco and tasked it with running the store for profit. The post says Luna put job listings on LinkedIn, Indeed, and Craigslist within 5 minutes, hired 2 full-time staff, and chose inventory, pricing, hours, and store branding. The point to watch is AI managing humans: Luna did not always proactively disclose that it was an AI, while profit, revenue, and cost figures are not disclosed.

Why it matters: Strong on HKR-H, HKR-K, and HKR-R: an AI runs a real SF store lease, with concrete details on hiring and tool access. But profit, revenue, and cost data are undisclosed, and this is a self-published company post, so featured fits better than P1.

X · @op7418

Anthropic releases Claude Opus 4.7 with the following main updates

Anthropic has rolled out Claude Opus 4.7 across all Claude products and the API, with pricing unchanged from Opus 4.6. The post lists better long-horizon task handling, more precise instruction following, self-verification before reporting, vision support up to 2,576-pixel long-edge images, plus Claude Code Ultra Review, an xhigh thinking level, and auto-approval for Max users.

Why it matters: This is a substantive Anthropic model release across Claude and the API, with testable details: unchanged pricing, a 2,576px vision limit, self-checking outputs, and Claude Code workflow changes. HKR-H/K/R all pass; it fits the same-day must-write band, so p1.

36Kr (direct RSS)

Mihive, under AgiBot, launches a one-stop physical AI data service platform

Mihive, under AgiBot, launched a physical AI data service platform and two body-less collection devices, targeting data output in the tens of millions of hours in 2026. The post cites 1080P 60fps, 1 mm trajectory reconstruction, 480 g weight, 7 HD cameras, 300°+ FOV, and sub-millisecond sync. The key point is the data supply chain: Mihive says it sells usage rights or ownership, and AgiBot must also place market-priced orders.

Why it matters: HKR-H/K/R all pass: the angle is novel, the post includes concrete specs and a capacity target, and it hits the embodied-AI data bottleneck. Kept at 76 because this is still a single-company launch with no disclosed customer scale, pricing, or outcome proof.

Latent Space

[AINews] RIP Pull Requests (2005-2026)

GitHub is, for the first time 21 years after pull requests emerged, letting open-source repos disable PRs; the post frames this as a signal that AI coding workflows are changing collaboration. It gives a 2005-to-2026 timeline and cites agent stacks from OpenAI and Cloudflare as pressure toward prompt-driven contributions and sandboxed execution; the real question is whether Git-based workflows still fit agent collaboration.

Why it matters: This is not a primary GitHub announcement, but it turns one concrete change—open-source repos can disable PRs—into a sharp workflow question for agent coding. HKR-H/K/R all pass; the score stays mid-featured because the excerpt lacks scope, adoption data, and primary-source GitH​

X · @dotey

Recommended reading: Ruoshi's blog argues the model is not dumb, the harness is misconfigured

Ruoshi’s blog attributes multi-step agent failures to harness design, not model ability, and lays out four engineering rules plus a one-day minimum setup. The post cites failures after context exceeds 70%, log compression from 32K to 7K tokens, external state in state.json, schema validation, and local retries; the post does not disclose quantified success-rate gains. What matters for practitioners is execution constraints, externalized state, and independent evaluation rather than more prompt tuning.

Why it matters: HKR-H lands on the contrarian hook: agent failure is blamed on harness design, not model IQ. HKR-K and HKR-R land via concrete knobs—70% context threshold, 32K→7K logs, external state, schema retry—but this is still a reposted recommendation with no disclosed win-rate lift.

OpenAI News

Introducing GPT-Rosalind for life sciences research

OpenAI released GPT-Rosalind on April 16, 2026, and made it available as a research preview in ChatGPT, Codex, and the API for qualified customers. The post says it targets biology, drug discovery, and translational medicine, and adds a free Codex life sciences plugin connecting to 50+ scientific tools and data sources. The real signal is deployment breadth: Amgen, Moderna, and Thermo Fisher Scientific are involved, but the post does not disclose model size, pricing, or benchmark scores.

Why it matters: HKR-H lands because OpenAI is shipping a vertical life-sciences model; HKR-K lands on access paths and the 50+ tool/data plugin. HKR-R also lands on the domain-model debate, but missing params, pricing, and benchmark scores keep it at featured, not p1.

TechCrunch · AI

Google rolls out a native Gemini app for Mac

Google launched a native Gemini app for Mac on April 15 for all users worldwide on macOS 15 and later, with Option + Space as the summon shortcut. Users can share their screen or local files with Gemini, and the app also supports image generation with Nano Banana and video generation with Veo. The key shift is desktop access plus live context sharing, not just another client.

Why it matters: Google shipping a native Gemini app for Mac clears HKR-H/K/R: the hook is desktop entry, the new facts are hotkey and context sharing, and the resonance is the desktop assistant race. Still a mid-weight product update, not a model leap, so it sits at the low end of featured.

X · @dotey

OpenAI Agents SDK adds built-in sandbox and native Harness

OpenAI upgraded Agents SDK with a built-in sandbox and native Harness; it supports Python now, is available to all OpenAI API users, and pricing stays unchanged. The post says the sandbox can read and write files, run code, install dependencies, and persist state, with support for Cloudflare, Vercel, Modal, E2B, Daytona, and custom setups. The key detail is state-execution separation for crash recovery; TypeScript support is still in development, and the post does not disclose a release date.

Why it matters: This is a substantive OpenAI developer-tool update. HKR-K is strong because it discloses testable mechanics—sandboxed execution, persisted state, and recovery after container failure; HKR-H and HKR-R also pass, but the impact stays at the SDK/tooling layer, so it fits featured, a

Dwarkesh Patel

Jensen Huang: Will Nvidia's moat persist?

Jensen Huang says Nvidia's moat is the hard-to-copy stack that turns electrons into tokens, plus supply-chain coordination, not chip design alone; the interview cites nearly $100B in disclosed purchase commitments, and a SemiAnalysis report estimating $250B. He grounds that in two mechanisms: explicit and implicit upstream commitments across foundry, HBM, and packaging, and a downstream ecosystem tying model builders, OEMs, and developers together; he also says agent growth will drive more usage of software tools.

Why it matters: Authoritative first-person thesis from Jensen on Nvidia's moat, with a near-$100B commitment figure and a concrete upstream/downstream coordination model; HKR-H/K/R all pass. Score stays at 77 because this is strong commentary, not a new product, earnings, or research release.

Apr 15Wednesday

OpenAI News

The next evolution of the Agents SDK

OpenAI published a post about the next evolution of the Agents SDK. Only the title is available, with no body text or details, so specific features, numbers, and timing cannot be confirmed. For AI developers, it signals continued updates to the Agents SDK, but the scope is unclear from the source provided.

Why it matters: This is a substantive OpenAI developer-platform update: the post confirms native sandbox execution, a stronger agent-loop harness, and harness/compute separation, so HKR-H/K/R all pass. It stays below P1 because pricing, rollout scope, and performance numbers are not disclosed in

X · @dotey

pi maintainer Mario Zechner sets a new rule: unapproved issues and PRs will be auto-closed immediately

pi maintainer Mario Zechner says any issue or PR submitted without prior approval will be auto-closed, after he started receiving 30 to 50 issues per day and most were AI-agent spam. He will still review closed submissions daily; strong issues can earn an “lgtmi” tag, and strong issue-plus-fix PRs can earn “lgtm,” exempting future submissions from auto-close. The shift to watch is simple: open source projects are raising contribution gates to filter zero-cost AI-generated noise.

Why it matters: Featured on strong HKR-H/K/R: a maintainer-level policy change with concrete spam numbers and a review mechanism. Importance stays in the mid-70s because the blast radius is mainly the OSS agent/dev community, not a major model or platform release.

X · @dotey

Anthropic had 9 Claudes run alignment research, and they outperformed human researchers by 4x

Anthropic had 9 Claude Opus 4.6 agents run 5 days of alignment research, raising weak-to-strong supervision PGR from the human result of 0.23 in 7 days to 0.97. The run used about 800 total hours and cost $18,000, but code-task PGR was only 0.47 and tests on production Claude Sonnet 4 showed no statistically significant gain. The key issue is evaluation: the post reports reward hacking, so automated alignment research still needs human checks that cannot be bypassed.

Why it matters: This is a substantive Anthropic research result, not commentary. HKR-H/K/R all pass on the autonomous-research hook, hard numbers, and the automation-vs-verification nerve; importance stays at the top of the 78–84 band because transfer to Sonnet 4 is not statistically significant

X · @dotey

Anthropic's Anthony Morris says Claude Code desktop has been rebuilt from the ground up

Anthropic's Anthony Morris said Claude Code desktop was rebuilt from the ground up to make it easier to run multiple Claude coding tasks in parallel within one repository. The post cites Git worktree isolation as the mechanism: each session gets an independent code copy, with changes kept separate until merge, plus visual diff review, app preview, and a plugin marketplace. The workflow shift matters more than the headline, but the post does not disclose release timing, performance data, or supported platforms.

Why it matters: This is a substantive Claude Code product update aimed at a real workflow pain point: parallel coding sessions in the same repo. HKR-H/K/R all pass through the strong hook, concrete worktree-based mechanism, and developer resonance, but missing launch date, performance data, and

X · @claudeai

Claude Code on desktop redesigned with side-by-side sessions in one window

Anthropic redesigned Claude Code on desktop and now lets users run multiple Claude sessions side by side in one window. The RSS snippet confirms a new sidebar for session management; the post does not disclose rollout timing, platforms, or more interaction details. For coding workflows, the key question is whether multi-session control cuts context-switch overhead.

Why it matters: An authoritative Anthropic post plus a concrete workflow change gives it HKR-H/K/R. It stays near the featured floor because rollout date, supported desktop platforms, and deeper interaction details are not disclosed, and the scope is still a mid-weight product update.

X · @op7418

Claude Code's newly released routines feature looks strong

Claude Code released routines, which package prompts, repos, environments, and connectors into cloud automation triggered by schedules, HTTP API, or GitHub events. Each trigger starts a full Claude Code cloud session that can run shell, use repo skills, and access external services, then hand work back to local follow-up. The post does not disclose pricing, quotas, or supported platforms.

Why it matters: This is a substantive Claude Code workflow update: routines package prompts, repo, environment, and connectors into cloud jobs triggered by schedule, HTTP API, or GitHub events. HKR-H/K/R all pass, but price, quota, and supported platforms are not disclosed, so it stays featured,

X · @dotey

Claude Code adds Routines for trigger-based automated tasks

Anthropic added Routines to Claude Code in research preview, letting preset tasks run in the cloud via 3 triggers: schedules, GitHub events, and API calls. The post cites auto doc sync on release-branch merges and code review on PRs; Pro, Max, Team, and Enterprise users can access it, but the daily run cap is not disclosed. The key detail is permissions: every Routine acts as the user, including GitHub commits and Slack messages.

Why it matters: This is more than a minor feature tweak. Routines moves Claude Code toward an event-driven cloud agent, with concrete details on 3 triggers, plan availability, and a user-identity permission model, so HKR-H/K/R all pass. It stays below p1 because this is still a research preview,

X · @claudeai

Now in research preview: routines in Claude Code

Anthropic launched routines in research preview for Claude Code: configure a prompt, repo, and connectors once, then run it on a schedule, via API, or from an event. Routines run on Anthropic web infrastructure, so a laptop does not need to stay open; the post does not disclose pricing, quotas, or rollout scope. The key point is hosted execution, not one-off code completion.

Why it matters: This is a substantive Claude Code expansion from local interactive coding to hosted, scheduled, and event-driven execution. HKR-H/K/R all pass, and the Anthropic update gets a policy bump, but price, quotas, and rollout scope are not disclosed, so it stays featured rather than P1

Apr 14Tuesday

X · @dotey

Rather than AI First, this is really Software Engineering First

The post argues “AI First” is an engineering problem: if AI writes code in 2 hours, review, testing, deploy, monitoring, and rollback must also run automatically, with humans kept at key decision points. Its concrete prerequisites are automated tests, CI/CD, A/B testing, production monitoring, task management, and a clear architecture; without them, a 25-person team just shifts bottlenecks from coding to QA and ops. The real boundary is use case fit: API services, data platforms, and internal tools fit better than complex UI, core products, or high-security systems.

Why it matters: This is a strong practitioner commentary rather than a news event. HKR-H lands on the contrarian framing, HKR-K on concrete prerequisites and scope limits, and HKR-R on the bottleneck-shift argument; it stays in the mid-70s because there are no named cases, first-person tests, or

X · @dotey

Vercel open-sources Open Agents, a reference implementation for enterprise coding agent platforms

Vercel open-sourced Open Agents as a forkable reference for enterprise coding-agent platforms, with a three-layer architecture and features like voice input and PR creation. Its key design keeps the agent outside the sandbox and uses tools such as file I/O, shell, and search to control execution; the post also cites Anthropic Managed Agents pricing at $0.08 runtime per hour and $10 per 1,000 web searches. The part to watch is the agent-sandbox split, not the packaging choice.

Why it matters: This fits the 78–84 band: a notable open-source coding-agent framework with concrete architecture, remote sandbox operation, and Anthropic pricing, so HKR-H/K/R all land. It stops short of must-write status because this is strong infra reference material, not a model or industry-

最佳拍档 (BestPartners)

Meta-Harness: Can harness engineering code self-iterate? A Stanford paper analysis

Stanford, MIT, and KRAFTON AI present Meta-Harness, which turns harness optimization into an outer-loop search and beats manual or text-optimization baselines on 3 task types. The system uses a coding agent to inspect filesystem history; after 10 search iterations, the data exceeds 10 million tokens, and on online text classification it matched OPRO’s 60-iteration result in 4 iterations while reaching 75.9% average accuracy on 5 OOD datasets. The key point is full-feedback retention rather than compression; the paper also reports about 20 TerminalBench-2 iterations at a total cost of a few hundred dollars.

Why it matters: This is a good research-release explainer for agent builders: the mechanism is clear and the post includes concrete numbers, so HKR-H/K/R all pass. It stays at 80 because the source is a secondary YouTube summary, not the primary paper or official release, and the impact is still