Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

981–1000 of 1,465

May 13Wednesday

Hacker News front page

Show HN: Rotunda - A Browser Built for Agents with Simulated Typing

Pierce released Rotunda, a Firefox 150-based browser for agents that simulates mouse and keyboard timing with an RNN trained on one week of his own patterns, and exposes local control through a CLI or Playwright API for Claude, Codex, or other harnesses.

Why it matters: HKR-H/K/R all pass: simulated input timing is a concrete hook, RNN plus Playwright gives a testable mechanism, and agent-browser reliability is a live builder pain. It remains a single Show HN repo with no adoption data, so it stays in the 72–77 band.

AI HOT (Curated Pool)

Configuring Development Environments for Agents

Cursor released tools for cloud agent development environments, adding multi-repository support, Dockerfile-based configuration, audit logs, and environment-level network and secret controls; the post says cache hits improve build speed by 70%.

Why it matters: HKR-K and HKR-R pass: Cursor adds concrete cloud-agent environment controls, including Dockerfile setup, audit logs, permissions, and 70% faster cached builds. HKR-H is weaker, so this sits at the lower featured band.

OpenAI News

Building a Safe, Effective Sandbox for Codex on Windows

OpenAI built a secure sandbox for Codex on Windows. The RSS snippet discloses controlled file access and network restrictions, but the post does not disclose implementation details, performance data, or rollout conditions.

Why it matters: OpenAI details a Windows sandbox for Codex with file-access and network controls. It is not a major model release, but HKR-H/K/R all pass because the safety boundary matters for coding-agent adoption.

AI HOT (Curated Pool)

Miaoda App and Enterprise Edition launch with 90% self-generated code

Baidu launched the Miaoda app and Miaoda Enterprise Edition, saying 90% of the Miaoda app’s code was generated by Miaoda itself; Miaoda-generated apps have served over 10 million users and reached a total value of RMB 5 billion.

Why it matters: HKR-H/K/R all pass via the 90% dogfooding hook, concrete adoption/value figures, and coding-tool resonance. Company-source metrics lack independent context, so this stays in the lower featured band.

QbitAI · WeChat

An 8-Year-Old Turns Ideas into Apps as Baidu Launches Miaoda 3.0

Baidu launched Miaoda 3.0 at its 2026 Create conference, adding iOS and Android app generation, Android packaging, online hot updates, and an enterprise edition with three-level permissions, environment isolation, and SLA commitments.

Why it matters: HKR-H/K/R pass: Baidu’s Miaoda 3.0 adds mobile app generation, Android packaging, hot updates, and enterprise controls. This is a solid product update, not a flagship model release or must-write event.

Synced · WeChat

Lin Junyang Reportedly Starts New AI Lab Seeking $2 Billion Valuation

The Information says Lin Junyang is raising several hundred million dollars for a new AI Lab at a potential $2 billion post-financing valuation, while the lab’s research direction and final valuation remain undisclosed.

Why it matters: HKR-H/K/R all pass, but the article only gives The Information’s funding rumor and valuation; research focus, team, and product plan are not disclosed. This fits the 72–77 featured band.

r/LocalLLaMA

The Trillion-Parameter Dilemma: MiMo-V2.5-Pro Open-Sourced at 1.02T Parameters

Xiaomi open-sourced MiMo-V2.5-Pro with 1.02T parameters, 42B active parameters, a 1M context window, and an MIT license; the author ran 125 Claude Code sessions through the API, spending $70.12 for 387,380,436 tokens with a 96.3% cache hit rate.

Why it matters: HKR-H/K/R all pass: a Xiaomi 1.02T open model plus a concrete Claude Code API cost experiment. Reddit sourcing keeps it at the low end of the 85+ band, but the domestic flagship-model signal clears p1.

AI HOT (Curated Pool)

Google launches its first AI-first laptop Googlebook with Gemini integration

Google launched Googlebook, its first laptop designed around Gemini Intelligence, with three disclosed mechanisms: Magic Pointer as an AI interaction entry point, natural-language widget creation, and Android-based cross-device app and file access.

Why it matters: HKR-H/K/R all pass: a Google Gemini-first laptop with 3 named interaction mechanisms. Specs, pricing, launch timing, and demos are not disclosed, so it stays in the 78–84 band.

TechCrunch · AI

Medicare’s New Payment Model Is Built for AI, and Most of Tech Has No Idea

Medicare’s ACCESS creates a payment mechanism for an AI agent that monitors patients between visits, calls to check in, coordinates housing referrals, and confirms medication pickup. The RSS snippet does not disclose payment rates, rollout timing, participating states, or operational requirements.

Why it matters: HKR-H/K/R all pass: a Medicare reimbursement path is rarer than a routine product update and affects healthcare-agent commercialization. Payment rates, launch date, and state scope are not disclosed, so it stays mid-featured.

AI HOT (Curated Pool)

Claude Code adds /goal feature to keep tasks running until completion

Claude Code introduced a /goal feature that keeps Claude working until a task is completed; the post does not disclose the trigger mechanism, supported versions, pricing, or failure conditions.

Why it matters: HKR-H/K/R pass because /goal targets a real Claude Code reliability pain. It is a single-feature Anthropic update with sparse mechanics, so it lands at the lower featured band, not same-day major news.

AI HOT (Curated Pool)

90% of People Are Wasting Tokens

Andrej Karpathy says 90% of AI coding bills is wasted on unnecessary context, including repeated full-repository sends, expensive models for simple tasks, and missing prompt caching.

Why it matters: HKR-H/K/R all pass via the 90% claim, named waste mechanisms, and practitioner cost pain. It reaches featured, but stays at 72 because the post gives no billing sample or reproducible test.

AI HOT (Curated Pool)

Codex Enables Background Multitasking Across Apps

OpenAI Devs says computer use lets Codex click, type, and keep working across Mac apps in the background; the post does not disclose release timing, permission design, or availability scope.

Why it matters: HKR-H/K/R all pass: OpenAI Devs gives a concrete Codex mechanism for background Mac cross-app actions, with clear developer relevance. Release timing, permissions, and availability are missing, so it stays at the lower featured band.

AI HOT (Curated Pool)

How Anthropic's Cybersecurity Team Uses Claude Code to Build a Threat Detection Platform

Anthropic’s detection platform engineering team used Claude Code to build the CLUE threat detection and response platform, completing a proof of concept in one day and delivery in one week while reducing analyst log investigation from hours to minutes.

Why it matters: HKR-H/K/R all pass, but this is an Anthropic internal dogfooding case rather than a Claude Code capability launch. CLUE and the timing metrics keep it just above the featured threshold.

Hacker News front page

Show HN: Needle Distills Gemini Tool Calling into a 26M Model

Cactus open-sourced Needle, a 26M-parameter tool-calling model that reaches 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, with MIT-licensed weights released on Hugging Face.

Why it matters: HKR-H/K/R all pass: the tiny Gemini-style tool-calling angle is clickable, with concrete speed and license claims. Source is still Show HN/GitHub self-reporting, not an independent benchmark or major lab release, so it stays below the 78–84 band.

r/LocalLLaMA

Needle: We Distilled Gemini Tool Calling Into a 26M Model

Cactus Compute open-sourced Needle, a 26M-parameter tool-calling model that reaches 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, using an attention-and-gating architecture with no MLPs.

Why it matters: HKR-H/K/R all pass: a 26M tool-calling model has a strong hook and concrete speed/design claims. Single Reddit source and a less-known team keep it in the lower 78–84 band.

AI HOT (Curated Pool)

Claude Enters the Legal Industry

Anthropic released more than 20 MCP connectors and 12 legal plugins, letting Claude work inside Word and Outlook for contract drafting, revision, clause comparison, and routine legal workflows.

Why it matters: HKR-H/K/R all pass: a substantive Anthropic vertical product update with 20+ MCP connectors and Office workflows. It is not a model release or platform-wide capability, so it stays in the 72–77 band.

AI HOT (Curated Pool)

Google Launches New Android Smart Assistant

Google introduced Android Intelligence at Android Show 2026, with multi-step automation across Android apps, browser-use features for Gemini in Chrome, automatic form filling, Rambler voice-note transcription, and custom Gen UI widgets; the post does not disclose rollout timing, supported devices, or pricing.

Why it matters: HKR-H/K/R all pass: the hook is Android-level agent control, the new facts are concrete automation surfaces, and the resonance is the mobile AI platform fight. Thin source detail keeps it at the low end of the 85-94 band.

Hacker News front page

Show HN: Agentic Interface for Mainframes and COBOL

Hypercubic launched Hopper, an agentic development environment that combines a real TN3270 terminal, z/OS-aware panels for datasets, jobs, and spool output, and an AI agent; sensitive operations require approval, and the terminal remains visible during agent actions.

Why it matters: HKR-H/K/R all pass: the mainframe-agent angle is novel, with concrete TN3270, z/OS, and approval mechanics. Small-vendor Show HN status and missing customer/pricing/results data keep it at the featured floor.

AI HOT (Curated Pool)

Build long-running AI agents that pause, resume, and never lose context with ADK

Google Developers describes using ADK to build long-running agents for enterprise workflows lasting days or weeks, such as HR onboarding, with a persistent state machine, persistent session storage, event-driven webhooks, and multi-agent delegation to pause during idle time and resume after restarts without losing context.

Why it matters: HKR-H/K/R all pass: the ADK tutorial gives concrete persistence mechanisms for long-running agents. It is useful engineering guidance from Google Developers, not a major model or platform release, so it sits at the 72–77 featured threshold.

TechCrunch · AI

Everything Google announced at its Android Show, from Googlebooks to vibe-coded widgets

Google announced AI-first Googlebooks laptops, more agentic Gemini features, vibe-coded Android widgets, Gemini in Chrome, and refreshed Android Auto ahead of I/O; the RSS snippet does not disclose specs, pricing, availability, or rollout timelines.

Why it matters: HKR-H/K/R all pass because Google bundled several Gemini/Android AI entry points with named product hooks. Missing parameters, pricing, rollout dates, and testable performance keeps it in the mid-weight product-update band.