Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

1001–1020 of 1,465

May 13Wednesday

The Verge · AI

Gemini’s Latest Updates Are All About Controlling Your Phone

Google announced Gemini Intelligence at its pre-I/O Android showcase, placing Gemini in Chrome on Android, autofill suggestions, and in-app actions; the RSS snippet does not disclose supported device models, rollout timing, or pricing.

Why it matters: HKR-H/K/R all pass because Gemini is moving into Android phone control across Chrome, autofill, and app actions. Missing device list, launch timing, and pricing keep it in the 72–77 product-update band.

TechCrunch · AI

Google brings agentic AI and vibe-coded widgets to Android

Google is adding Gemini Intelligence to Android, and the RSS snippet only discloses Gboard-based dictation and form-filling capabilities; the post does not disclose launch timing, supported devices, pricing, or technical implementation details.

Why it matters: HKR-H/K/R pass because this is a Google Android platform AI update with named features. Missing rollout timing, device scope, and pricing keep it at the low featured band, not a must-write release.

AI HOT (Curated Pool)

Code w/ Claude SF 2026: Building on Exponential AI Growth

Anthropic expanded developer tooling at Code w/ Claude SF 2026: Claude Code rate limits doubled, Claude Opus API limits increased, and hosted agents on the Claude platform added four functions, including memory review, multi-agent delegation, output criteria, and webhooks.

Why it matters: Anthropic ships a substantive Claude Code update with concrete numbers and feature additions; HKR-H/K/R all pass. This is strong dev-tool news, not a flagship model release, so it fits the 78–84 band.

May 12Tuesday

AI HOT (Curated Pool)

Install the official Codex plugin in Claude Code

The author describes installing OpenAI’s official Codex plugin in Claude Code via the plugin marketplace: add the repository, install the plugin, reload, and configure it, then use it to build a Skill where Claude Code handles reasoning and Codex acts as moderator.

Why it matters: HKR-H/K/R all pass: cross-stack plugin use is clickable, the install path is concrete, and it matters to AI dev workflows. It stays low-featured because this is a tutorial-style tip, not a model or platform release.

r/LocalLLaMA

Local LLM Autocomplete and Agentic Coding on a Single 16GB GPU + 64GB RAM

Reddit user grumd runs Qwen2.5-Coder-7B Q6 for autocomplete and Qwen3.6-35B-A3B Q8 for agentic coding on one RTX 5080 with RAM offloading; the post reports about 145k context, 56GB RAM used with other apps open, and Qwen3.6-35B-A3B speed of tg128 at 35.29 tokens/s.

Why it matters: HKR-H/K/R all pass: a named first-person local coding experiment with concrete model, quantization, context, and throughput data. Source is a single Reddit post without replication or comparisons, so it stays in the low featured band.

Google DeepMind

Google DeepMind publishes Co-Scientist multi-agent research system

Google DeepMind published Co-Scientist research in Nature, introducing a Gemini-based multi-agent AI system that iteratively generates, debates and evolves new hypotheses for complex scientific problems.

Why it matters: The post discloses the system's three-stage collaboration mechanism and deployment cases at several labs, showing how AI takes part in scientific hypothesis generation.

Hacker News front page

Show HN: Statewright – Visual State Machines for More Reliable AI Agents

Statewright uses a Rust state-machine engine to constrain Claude Code tool access, iterations, transitions, and guards; the post says 13–20B models improved consistently on real SWE-bench tasks, but it does not disclose benchmark scores, sample size, or the exact evaluation protocol.

Why it matters: HKR-H/K/R all pass: the state-machine constraint is a clear agent-reliability hook with a testable SWE-bench claim. Exact scores and reproduction details are not disclosed, so it stays just above the featured threshold.

AI HOT (Curated Pool)

Exporting consumer data for personalized AI Agent services

The author surveyed export methods across 5 consumer platforms: Taobao supports exports, JD.com needs a Codex-built Chrome extension, Ele.me can provide Excel exports by request, Meituan Waimai has no method, and JD.com plus Dianping tools are open sourced.

Why it matters: HKR-H/K/R pass: the post has a concrete personal-agent hook, 5-platform data, and open-source JD/Dianping tools. It stays at the featured floor because this is a practitioner experiment, not a platform release.

TechCrunch · AI

AI Voice Startup Vapi Hits $500M Valuation After Winning Amazon Ring Over 40 Rivals

Vapi reached a $500 million valuation after Amazon Ring chose its AI voice platform over 40 rivals, and the RSS snippet says its enterprise business has grown tenfold since early 2025 as companies move support and sales calls to AI agents.

Why it matters: HKR-H/K/R all pass: Amazon Ring’s 40-rival selection and Vapi’s $500M valuation give concrete signal. Still a startup financing/customer win, not a model or platform release, so it sits at the featured threshold.

Xinzhiyuan · WeChat

OpenAI releases GPT-Realtime-2, described as a GPT-5-level reasoning audio model

OpenAI released GPT-Realtime-2 alongside Realtime-Translate and Realtime-Whisper, with a 128K context window, five reasoning-effort levels, and API pricing of $32 per million input tokens and $64 per million output tokens.

Why it matters: HKR-H/K/R all pass: realtime audio reasoning is a strong hook; 128K context, five reasoning levels, and $32/$64 per 1M tokens add substance; voice-agent cost and stack choices hit practitioners. This is a same-day OpenAI product update.

Xinzhiyuan · WeChat

The Largest Single Industrial Product in History Is Entering Mass Production in China

AgiBot says it had shipped 10,000 general-purpose embodied robots by the end of March, and its humanoid robots worked eight continuous hours on a Nanchang 3C production line, completing 2,283 tasks with zero errors under formal line-cycle requirements.

Why it matters: HKR-H/K/R all pass: AgiBot gives unit and factory-run numbers with clear robotics deployment resonance. The score stays in 78-84 because the key claims are company-sourced, with no third-party validation or cost data disclosed.

Latent Space

Thinking Machines' Native Interaction Models: TML-Interaction-Small 276B-A12B Advances Realtime Voice

Thinking Machines released TML-Interaction-Small, a 276B-parameter MoE model with 12B active parameters, and the post says it advances realtime voice through 200ms time-aligned microturns, encoder-free early fusion for audio and images under 200ms, and benchmark wins over GPT-Realtime-2 and Gemini 3.1-Flash.

Why it matters: HKR-H/K/R all pass: TML-Interaction-Small gives architecture, active parameters, 200ms interaction, and named rivals. Benchmarks still need replication, but a real-time voice SOTA claim is same-day material.

Financial Times · Technology

Amazon staff use AI tool for unnecessary tasks to inflate usage scores

Amazon’s in-house MeshClaw tool lets employees delegate work to AI agents and raise their position on the company’s AI leaderboard; the post does not disclose the number of staff involved, the scoring rules, or the specific unnecessary tasks.

Why it matters: FT gives a concrete Amazon AI-adoption gaming story, clearing HKR-H/K/R. Missing participant count, scoring rules, and task examples keep it at the featured threshold, not a must-write item.

QbitAI · WeChat

OpenClaw quietly updates with Peekaboo v3 for Mac computer use

OpenClaw-related Peekaboo v3 adds Mac agent capabilities for pixel-level screenshots, UI position reading, clicks, text input, hotkeys, scrolling, and drag-and-drop, with MCP server integration for Cursor, Claude Code, and Codex.

Why it matters: HKR-H/K/R all pass: Peekaboo v3 adds Mac GUI perception and action primitives plus MCP access for Cursor, Claude Code, and Codex. This is a useful open-source agent-tooling update, not a model-level event, so it sits in the low featured band.

Computing Life · Yage

How AI Caused and Fixed My Insomnia

The author used AI to build a HealthKit export app in about 5 minutes and run multivariate regression, finding that the last post-dinner AI usage time correlated negatively with sleep duration; after avoiding AI at night, average sleep increased by 1 hour and 40 minutes.

Why it matters: HKR-H/K/R all pass: a first-person quantified experiment links post-dinner AI use to shorter sleep, then reports +1h40m after stopping. Personal-blog scope keeps it below major industry-update territory.

AI HOT (Curated Pool)

What Parameter Golf Taught Us About AI-Assisted Research

OpenAI’s Parameter Golf brought together over 1,000 participants and more than 2,000 submissions to test AI-assisted machine learning research, coding agents, model quantization, and model design under strict parameter constraints.

Why it matters: OpenAI’s Parameter Golf recap clears HKR-H/K/R with a concrete contest, 1,000+ participants, and 2,000+ submissions. It is research/benchmark signal, not a model or product launch, so 78 fits the lower featured band.

The Verge · AI

OpenAI just released its answer to Claude Mythos

OpenAI launched Daybreak, a security initiative that uses the Codex Security AI agent released in March to model an organization’s code, validate likely vulnerabilities, and automate detection of higher-risk issues before attackers find them.

Why it matters: HKR-H/K/R all pass: Daybreak has a rivalry hook, concrete agent workflow, and code-security resonance. It is narrower than a model or ChatGPT capability release, so it stays in the 78–84 band.

Computing Life · Yage

How AI Caused and Fixed My Insomnia

The author used AI to build a HealthKit export tool and run multivariate regression, found that post-dinner AI use correlated negatively with sleep duration, and added 1 hour 40 minutes of average nightly sleep after avoiding AI for several weeks.

Why it matters: HKR-H/K/R all pass: the personal reversal is clickable, the HealthKit/regression setup adds testable detail, and sleep loss hits AI practitioners directly. Scope is anecdotal, so it stays at the featured floor.

The Verge · AI

Here’s What Mira Murati’s AI Company Is Up To

Thinking Machines announced work on “interaction models” that continuously take in audio, video, and text and respond or act in real time; the post does not disclose model size, release timing, pricing, or the final product format.

Why it matters: HKR-H/K/R all pass, but the body lacks parameters, launch timing, and product form. This is a high-interest startup direction reveal, not a usable model release, so it stays at the top of the 72–77 band.

Sinocism (Bill Bishop)

Trump China Visit; China’s Next Generation Industrial Policy; Standardizing AI Agents

China’s CAC, NDRC, and MIIT issued an implementation document on AI agent standardization, targeting privacy leakage, unauthorized actions, and loss of behavioral control from high-autonomy, high-permission agents, while tying the work to a 2027 target for new intelligent terminals and AI agent adoption above 70%.

Why it matters: HKR-H/K/R all pass: the China agent-policy hook is concrete, with a 2027 >70% target and named autonomy/permission risks. It clears featured, but it is policy guidance rather than a major model or product launch.