Skip to content

#Agent

39 today

May 13Wednesday

TechCrunch · AI

Medicare’s New Payment Model Is Built for AI, and Most of Tech Has No Idea

Medicare’s ACCESS creates a payment mechanism for an AI agent that monitors patients between visits, calls to check in, coordinates housing referrals, and confirms medication pickup. The RSS snippet does not disclose payment rates, rollout timing, participating states, or operational requirements.

Why it matters: HKR-H/K/R all pass: a Medicare reimbursement path is rarer than a routine product update and affects healthcare-agent commercialization. Payment rates, launch date, and state scope are not disclosed, so it stays mid-featured.

AI HOT (Curated Pool)

Claude Code adds /goal feature to keep tasks running until completion

Claude Code introduced a /goal feature that keeps Claude working until a task is completed; the post does not disclose the trigger mechanism, supported versions, pricing, or failure conditions.

Why it matters: HKR-H/K/R pass because /goal targets a real Claude Code reliability pain. It is a single-feature Anthropic update with sparse mechanics, so it lands at the lower featured band, not same-day major news.

AI HOT (Curated Pool)

90% of People Are Wasting Tokens

Andrej Karpathy says 90% of AI coding bills is wasted on unnecessary context, including repeated full-repository sends, expensive models for simple tasks, and missing prompt caching.

Why it matters: HKR-H/K/R all pass via the 90% claim, named waste mechanisms, and practitioner cost pain. It reaches featured, but stays at 72 because the post gives no billing sample or reproducible test.

AI HOT (Curated Pool)

Codex Enables Background Multitasking Across Apps

OpenAI Devs says computer use lets Codex click, type, and keep working across Mac apps in the background; the post does not disclose release timing, permission design, or availability scope.

Why it matters: HKR-H/K/R all pass: OpenAI Devs gives a concrete Codex mechanism for background Mac cross-app actions, with clear developer relevance. Release timing, permissions, and availability are missing, so it stays at the lower featured band.

AI HOT (Curated Pool)

How Anthropic's Cybersecurity Team Uses Claude Code to Build a Threat Detection Platform

Anthropic’s detection platform engineering team used Claude Code to build the CLUE threat detection and response platform, completing a proof of concept in one day and delivery in one week while reducing analyst log investigation from hours to minutes.

Why it matters: HKR-H/K/R all pass, but this is an Anthropic internal dogfooding case rather than a Claude Code capability launch. CLUE and the timing metrics keep it just above the featured threshold.

Hacker News front page

Show HN: Needle Distills Gemini Tool Calling into a 26M Model

Cactus open-sourced Needle, a 26M-parameter tool-calling model that reaches 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, with MIT-licensed weights released on Hugging Face.

Why it matters: HKR-H/K/R all pass: the tiny Gemini-style tool-calling angle is clickable, with concrete speed and license claims. Source is still Show HN/GitHub self-reporting, not an independent benchmark or major lab release, so it stays below the 78–84 band.

r/LocalLLaMA

Needle: We Distilled Gemini Tool Calling Into a 26M Model

Cactus Compute open-sourced Needle, a 26M-parameter tool-calling model that reaches 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, using an attention-and-gating architecture with no MLPs.

Why it matters: HKR-H/K/R all pass: a 26M tool-calling model has a strong hook and concrete speed/design claims. Single Reddit source and a less-known team keep it in the lower 78–84 band.

AI HOT (Curated Pool)

Claude Enters the Legal Industry

Anthropic released more than 20 MCP connectors and 12 legal plugins, letting Claude work inside Word and Outlook for contract drafting, revision, clause comparison, and routine legal workflows.

Why it matters: HKR-H/K/R all pass: a substantive Anthropic vertical product update with 20+ MCP connectors and Office workflows. It is not a model release or platform-wide capability, so it stays in the 72–77 band.

AI HOT (Curated Pool)

Google Launches New Android Smart Assistant

Google introduced Android Intelligence at Android Show 2026, with multi-step automation across Android apps, browser-use features for Gemini in Chrome, automatic form filling, Rambler voice-note transcription, and custom Gen UI widgets; the post does not disclose rollout timing, supported devices, or pricing.

Why it matters: HKR-H/K/R all pass: the hook is Android-level agent control, the new facts are concrete automation surfaces, and the resonance is the mobile AI platform fight. Thin source detail keeps it at the low end of the 85-94 band.

Hacker News front page

Show HN: Agentic Interface for Mainframes and COBOL

Hypercubic launched Hopper, an agentic development environment that combines a real TN3270 terminal, z/OS-aware panels for datasets, jobs, and spool output, and an AI agent; sensitive operations require approval, and the terminal remains visible during agent actions.

Why it matters: HKR-H/K/R all pass: the mainframe-agent angle is novel, with concrete TN3270, z/OS, and approval mechanics. Small-vendor Show HN status and missing customer/pricing/results data keep it at the featured floor.

AI HOT (Curated Pool)

Build long-running AI agents that pause, resume, and never lose context with ADK

Google Developers describes using ADK to build long-running agents for enterprise workflows lasting days or weeks, such as HR onboarding, with a persistent state machine, persistent session storage, event-driven webhooks, and multi-agent delegation to pause during idle time and resume after restarts without losing context.

Why it matters: HKR-H/K/R all pass: the ADK tutorial gives concrete persistence mechanisms for long-running agents. It is useful engineering guidance from Google Developers, not a major model or platform release, so it sits at the 72–77 featured threshold.

TechCrunch · AI

Everything Google announced at its Android Show, from Googlebooks to vibe-coded widgets

Google announced AI-first Googlebooks laptops, more agentic Gemini features, vibe-coded Android widgets, Gemini in Chrome, and refreshed Android Auto ahead of I/O; the RSS snippet does not disclose specs, pricing, availability, or rollout timelines.

Why it matters: HKR-H/K/R all pass because Google bundled several Gemini/Android AI entry points with named product hooks. Missing parameters, pricing, rollout dates, and testable performance keeps it in the mid-weight product-update band.

The Verge · AI

Gemini’s Latest Updates Are All About Controlling Your Phone

Google announced Gemini Intelligence at its pre-I/O Android showcase, placing Gemini in Chrome on Android, autofill suggestions, and in-app actions; the RSS snippet does not disclose supported device models, rollout timing, or pricing.

Why it matters: HKR-H/K/R all pass because Gemini is moving into Android phone control across Chrome, autofill, and app actions. Missing device list, launch timing, and pricing keep it in the 72–77 product-update band.

TechCrunch · AI

Google brings agentic AI and vibe-coded widgets to Android

Google is adding Gemini Intelligence to Android, and the RSS snippet only discloses Gboard-based dictation and form-filling capabilities; the post does not disclose launch timing, supported devices, pricing, or technical implementation details.

Why it matters: HKR-H/K/R pass because this is a Google Android platform AI update with named features. Missing rollout timing, device scope, and pricing keep it at the low featured band, not a must-write release.

AI HOT (Curated Pool)

Code w/ Claude SF 2026: Building on Exponential AI Growth

Anthropic expanded developer tooling at Code w/ Claude SF 2026: Claude Code rate limits doubled, Claude Opus API limits increased, and hosted agents on the Claude platform added four functions, including memory review, multi-agent delegation, output criteria, and webhooks.

Why it matters: Anthropic ships a substantive Claude Code update with concrete numbers and feature additions; HKR-H/K/R all pass. This is strong dev-tool news, not a flagship model release, so it fits the 78–84 band.

May 12Tuesday

AI HOT (Curated Pool)

Install the official Codex plugin in Claude Code

The author describes installing OpenAI’s official Codex plugin in Claude Code via the plugin marketplace: add the repository, install the plugin, reload, and configure it, then use it to build a Skill where Claude Code handles reasoning and Codex acts as moderator.

Why it matters: HKR-H/K/R all pass: cross-stack plugin use is clickable, the install path is concrete, and it matters to AI dev workflows. It stays low-featured because this is a tutorial-style tip, not a model or platform release.

r/LocalLLaMA

Local LLM Autocomplete and Agentic Coding on a Single 16GB GPU + 64GB RAM

Reddit user grumd runs Qwen2.5-Coder-7B Q6 for autocomplete and Qwen3.6-35B-A3B Q8 for agentic coding on one RTX 5080 with RAM offloading; the post reports about 145k context, 56GB RAM used with other apps open, and Qwen3.6-35B-A3B speed of tg128 at 35.29 tokens/s.

Why it matters: HKR-H/K/R all pass: a named first-person local coding experiment with concrete model, quantization, context, and throughput data. Source is a single Reddit post without replication or comparisons, so it stays in the low featured band.

Google DeepMind

Google DeepMind publishes Co-Scientist multi-agent research system

Google DeepMind published Co-Scientist research in Nature, introducing a Gemini-based multi-agent AI system that iteratively generates, debates and evolves new hypotheses for complex scientific problems.

Why it matters: The post discloses the system's three-stage collaboration mechanism and deployment cases at several labs, showing how AI takes part in scientific hypothesis generation.

Hacker News front page

Show HN: Statewright – Visual State Machines for More Reliable AI Agents

Statewright uses a Rust state-machine engine to constrain Claude Code tool access, iterations, transitions, and guards; the post says 13–20B models improved consistently on real SWE-bench tasks, but it does not disclose benchmark scores, sample size, or the exact evaluation protocol.

Why it matters: HKR-H/K/R all pass: the state-machine constraint is a clear agent-reliability hook with a testable SWE-bench claim. Exact scores and reproduction details are not disclosed, so it stays just above the featured threshold.

AI HOT (Curated Pool)

Exporting consumer data for personalized AI Agent services

The author surveyed export methods across 5 consumer platforms: Taobao supports exports, JD.com needs a Codex-built Chrome extension, Ele.me can provide Excel exports by request, Meituan Waimai has no method, and JD.com plus Dianping tools are open sourced.

Why it matters: HKR-H/K/R pass: the post has a concrete personal-agent hook, 5-platform data, and open-source JD/Dianping tools. It stays at the featured floor because this is a practitioner experiment, not a platform release.

TechCrunch · AI

AI Voice Startup Vapi Hits $500M Valuation After Winning Amazon Ring Over 40 Rivals

Vapi reached a $500 million valuation after Amazon Ring chose its AI voice platform over 40 rivals, and the RSS snippet says its enterprise business has grown tenfold since early 2025 as companies move support and sales calls to AI agents.

Why it matters: HKR-H/K/R all pass: Amazon Ring’s 40-rival selection and Vapi’s $500M valuation give concrete signal. Still a startup financing/customer win, not a model or platform release, so it sits at the featured threshold.

Xinzhiyuan · WeChat

OpenAI releases GPT-Realtime-2, described as a GPT-5-level reasoning audio model

OpenAI released GPT-Realtime-2 alongside Realtime-Translate and Realtime-Whisper, with a 128K context window, five reasoning-effort levels, and API pricing of $32 per million input tokens and $64 per million output tokens.

Why it matters: HKR-H/K/R all pass: realtime audio reasoning is a strong hook; 128K context, five reasoning levels, and $32/$64 per 1M tokens add substance; voice-agent cost and stack choices hit practitioners. This is a same-day OpenAI product update.

Xinzhiyuan · WeChat

The Largest Single Industrial Product in History Is Entering Mass Production in China

AgiBot says it had shipped 10,000 general-purpose embodied robots by the end of March, and its humanoid robots worked eight continuous hours on a Nanchang 3C production line, completing 2,283 tasks with zero errors under formal line-cycle requirements.

Why it matters: HKR-H/K/R all pass: AgiBot gives unit and factory-run numbers with clear robotics deployment resonance. The score stays in 78-84 because the key claims are company-sourced, with no third-party validation or cost data disclosed.

Latent Space

Thinking Machines' Native Interaction Models: TML-Interaction-Small 276B-A12B Advances Realtime Voice

Thinking Machines released TML-Interaction-Small, a 276B-parameter MoE model with 12B active parameters, and the post says it advances realtime voice through 200ms time-aligned microturns, encoder-free early fusion for audio and images under 200ms, and benchmark wins over GPT-Realtime-2 and Gemini 3.1-Flash.

Why it matters: HKR-H/K/R all pass: TML-Interaction-Small gives architecture, active parameters, 200ms interaction, and named rivals. Benchmarks still need replication, but a real-time voice SOTA claim is same-day material.

Financial Times · Technology

Amazon staff use AI tool for unnecessary tasks to inflate usage scores

Amazon’s in-house MeshClaw tool lets employees delegate work to AI agents and raise their position on the company’s AI leaderboard; the post does not disclose the number of staff involved, the scoring rules, or the specific unnecessary tasks.

Why it matters: FT gives a concrete Amazon AI-adoption gaming story, clearing HKR-H/K/R. Missing participant count, scoring rules, and task examples keep it at the featured threshold, not a must-write item.

QbitAI · WeChat

OpenClaw quietly updates with Peekaboo v3 for Mac computer use

OpenClaw-related Peekaboo v3 adds Mac agent capabilities for pixel-level screenshots, UI position reading, clicks, text input, hotkeys, scrolling, and drag-and-drop, with MCP server integration for Cursor, Claude Code, and Codex.

Why it matters: HKR-H/K/R all pass: Peekaboo v3 adds Mac GUI perception and action primitives plus MCP access for Cursor, Claude Code, and Codex. This is a useful open-source agent-tooling update, not a model-level event, so it sits in the low featured band.

Computing Life · Yage

How AI Caused and Fixed My Insomnia

The author used AI to build a HealthKit export app in about 5 minutes and run multivariate regression, finding that the last post-dinner AI usage time correlated negatively with sleep duration; after avoiding AI at night, average sleep increased by 1 hour and 40 minutes.

Why it matters: HKR-H/K/R all pass: a first-person quantified experiment links post-dinner AI use to shorter sleep, then reports +1h40m after stopping. Personal-blog scope keeps it below major industry-update territory.

AI HOT (Curated Pool)

What Parameter Golf Taught Us About AI-Assisted Research

OpenAI’s Parameter Golf brought together over 1,000 participants and more than 2,000 submissions to test AI-assisted machine learning research, coding agents, model quantization, and model design under strict parameter constraints.

Why it matters: OpenAI’s Parameter Golf recap clears HKR-H/K/R with a concrete contest, 1,000+ participants, and 2,000+ submissions. It is research/benchmark signal, not a model or product launch, so 78 fits the lower featured band.

The Verge · AI

OpenAI just released its answer to Claude Mythos

OpenAI launched Daybreak, a security initiative that uses the Codex Security AI agent released in March to model an organization’s code, validate likely vulnerabilities, and automate detection of higher-risk issues before attackers find them.

Why it matters: HKR-H/K/R all pass: Daybreak has a rivalry hook, concrete agent workflow, and code-security resonance. It is narrower than a model or ChatGPT capability release, so it stays in the 78–84 band.

Computing Life · Yage

How AI Caused and Fixed My Insomnia

The author used AI to build a HealthKit export tool and run multivariate regression, found that post-dinner AI use correlated negatively with sleep duration, and added 1 hour 40 minutes of average nightly sleep after avoiding AI for several weeks.

Why it matters: HKR-H/K/R all pass: the personal reversal is clickable, the HealthKit/regression setup adds testable detail, and sleep loss hits AI practitioners directly. Scope is anecdotal, so it stays at the featured floor.

The Verge · AI

Here’s What Mira Murati’s AI Company Is Up To

Thinking Machines announced work on “interaction models” that continuously take in audio, video, and text and respond or act in real time; the post does not disclose model size, release timing, pricing, or the final product format.

Why it matters: HKR-H/K/R all pass, but the body lacks parameters, launch timing, and product form. This is a high-interest startup direction reveal, not a usable model release, so it stays at the top of the 72–77 band.

Sinocism (Bill Bishop)

Trump China Visit; China’s Next Generation Industrial Policy; Standardizing AI Agents

China’s CAC, NDRC, and MIIT issued an implementation document on AI agent standardization, targeting privacy leakage, unauthorized actions, and loss of behavioral control from high-autonomy, high-permission agents, while tying the work to a 2027 target for new intelligent terminals and AI agent adoption above 70%.

Why it matters: HKR-H/K/R all pass: the China agent-policy hook is concrete, with a 2027 >70% target and named autonomy/permission risks. It clears featured, but it is policy guidance rather than a major model or product launch.

Bloomberg Technology

GitLab Says It Will Cut Jobs to Spend on Growth in the “Agentic Era”

GitLab said it will cut jobs to free up money for the market opportunity around AI agents; the RSS snippet does not disclose the number of roles, budget size, or execution timeline.

Why it matters: HKR-H and HKR-R pass: Bloomberg reports GitLab tying job cuts directly to agent investment, a strong devtools labor signal. HKR-K is weak because headcount, budget, and timing are missing.

AI HOT (Curated Pool)

Introducing Daybreak: Frontier AI for Cyber Defenders

OpenAI introduced Daybreak for cyber defenders, combining OpenAI models, Codex, and security partners; the post does not disclose pricing, launch timing, or concrete defense metrics.

Why it matters: OpenAI’s Daybreak announcement clears HKR-H/R as a security-focused product hook, but HKR-K fails: no defense metrics, access terms, or pricing. That keeps it in the 72–77 product-update band.

Bloomberg Technology

AI Chipmaker Cerebras Seeks $4.8 Billion in Upsized IPO | Bloomberg Tech 5/11/2026

Cerebras increased its IPO offering plan by one-third to as much as $4.8 billion; the post also mentions Circle’s first-quarter revenue and Google researchers’ first AI-built zero-day attack, but does not disclose details.

Why it matters: HKR-H/K/R all pass: Bloomberg reports Cerebras upsizing an IPO plan by one-third to as much as $4.8B, a major AI-infrastructure capital-markets signal. The video-style item lacks pricing, valuation, and timeline details, so it stays just below 85.

AI HOT (Curated Pool)

Using LLMs in Script Shebang Lines

Simon Willison demonstrates using an LLM command in a script shebang line, with fragments generating SVG, the -T option calling llm_time, and a YAML template defining Python tools to compute 2344×5252+134 and return 12,310,822.

Why it matters: HKR-H/K/R all pass: Simon Willison shows a reproducible LLM-in-shebang workflow with concrete flags. Impact stays within CLI/script automation, not a model or platform release, so it sits in the low featured band.

AI HOT (Curated Pool)

Replit launches parallel agents with support for 10 concurrent agents

Replit launched parallel agents that run up to 10 agents concurrently, with each agent holding an independent copy of the app, working on its own machine, and merging the results through an agent workflow.

Why it matters: HKR-H/K/R pass: the post gives a concrete 10-agent parallel workflow with isolated app copies and merge. This is a mid-weight dev-tool update, below a Cursor Agent-mode-scale launch, so it sits at the featured threshold.

AI HOT (Curated Pool)

Personal Intelligence Customizes Travel Itineraries

Gemini App says Personal Intelligence generates personalized travel itineraries when users connect Gmail, Google Photos, Google Search, and YouTube history, and the post says users can choose connected apps and manage personalization settings at any time.

Why it matters: HKR-H/K/R all pass: the Google data integration is the hook, mechanism, and practitioner nerve. Scope, permission controls, and evals are not disclosed, so this stays at the low featured band.

AI HOT (Curated Pool)

Anthropic Launches Claude Platform on AWS

Anthropic launched the Claude platform on AWS, letting AWS customers use existing authentication, billing, and committed-spend credits to access the full Claude API feature set, including hosted agents, code execution, and the Files API.

Why it matters: HKR-K and HKR-R pass: Anthropic brings Claude Platform into AWS procurement, billing, and committed spend. HKR-H is weak because this is distribution, not a model or capability launch.

May 11Monday

AI HOT (Curated Pool)

Anthropic open-sources full-stack financial AI templates

Anthropic open-sourced a financial services AI template library on GitHub, including 10 end-to-end agents, 7 vertical industry plugins, and MCP connectors for 11 financial data providers, with deployment paths from personal plugins to enterprise APIs and integrations for Microsoft 365 and private cloud.

Why it matters: HKR-H/K/R all pass: Anthropic shipped a reusable finance-agent template library with GitHub artifacts and concrete counts. It is not a model release, so it stays below 85, but the open-source MCP vertical stack clears featured.