Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

761–780 of 1,465

May 28Thursday

AI HOT (Curated Pool)

Latest Google Pay updates

Google Pay introduced a universal commerce protocol and a new MCP server for AI agents to manage integrations and analyze trends, while Android updates add dynamic callbacks for faster checkout, WebView payments in social apps, cross-device biometric authentication, and new transaction signals.

Why it matters: HKR-H/K/R pass: the MCP payments angle is concrete and relevant to agent commerce. Score stays in the 72–77 band because the post lists features but gives no adoption scale, pricing, or real agent transaction case.

Hugging Face Blog

ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks

Artificial Analysis and IBM published the ITBench-AA title, saying frontier models scored below 50% on an enterprise IT agent task benchmark; the post does not disclose tested models, sample size, or scoring method.

Why it matters: HKR-H/R pass: frontier models under 50% on enterprise IT agent tasks is clickable and deployment-relevant. HKR-K is weak because models, sample size, and scoring are not disclosed, so it stays near the featured floor.

AI HOT (Curated Pool)

I Think Anthropic and OpenAI Found Product-Market Fit

Anthropic and OpenAI changed enterprise pricing around April 2026, moving coding agents from heavily discounted seat plans to API-usage billing, with Anthropic Enterprise at $20 per seat per month plus API fees and OpenAI Codex billed by API token usage.

Why it matters: HKR-H/K/R all pass: the piece ties OpenAI and Anthropic PMF to a concrete billing shift for coding agents. It is influential commentary, not an official launch, so it fits the 78–84 band.

AI HOT (Curated Pool)

Interview with Google Search VP Robby Stein on the AI-Native Search Era

Robby Stein discussed Google Search’s move toward an AI-native mode at Google I/O, covering AI Mode, multi-turn query decomposition, TPU infrastructure costs, source-link selection, and publisher traffic tension, but the post does not disclose specific pricing, traffic numbers, or rollout conditions.

Why it matters: HKR-H/K/R all pass, but this is an interview summary rather than a fresh launch. No price, traffic, or cost numbers are disclosed, so it sits in the 72–77 quality-interview band.

May 27Wednesday

The Verge · AI

Robinhood will let your AI agent trade stocks and make (or lose) lots of money

Robinhood opened its trading platform to AI agents: traders can create a separate account, allocate a specific amount of money, and let the agent buy and sell stocks, while Robinhood warns agentic trading can cause the loss of the entire investment.

Why it matters: HKR-H/K/R all pass: real-money stock trading gives the hook, separate funded accounts add mechanism, and autonomy risk creates resonance. It stays in the 78–84 band because safeguards, rollout scope, and regulatory limits are not disclosed.

AI HOT (Curated Pool)

Runway launches Model Context Protocol server

Runway launched an MCP server that lets compatible agents such as Claude, ChatGPT, and Cursor generate images and videos inside chat interfaces, with access to Gen-4.5, Seedance 2.0, GPT Image 2, Kling 3.0, and Nano Banana Pro.

Why it matters: HKR-H/K/R all pass, but this is a Runway product integration, not an MCP protocol change or model release. It clears featured, with the score kept in the 72–77 band.

r/LocalLLaMA

I ran 8 open-weight models as agents in a persistent MMO for 10 days

Firespawn Studios ran 25 agents across 8 open-weight models for 10 days in Null Epoch Season 0 and released about 93,000 logged events, with roughly 70% of actions including the model’s reasoning or justification.

Why it matters: HKR-H/K/R all pass: a concrete 10-day MMO agent trial with 25 agents and 93k events. Reddit sourcing limits reach, so it lands in the 78–84 good-quality band, not P1.

TechCrunch · AI

Robinhood now lets your AI agents trade stocks

Robinhood lets AI agents read and analyze users’ portfolios and suggest investments, but order placement is limited to the pre-loaded balance in a dedicated wallet.

Why it matters: HKR-H/K/R all pass: the hook is agents trading real money, the concrete mechanism is portfolio access plus a prefunded wallet, and the resonance is agent safety. Robinhood is not a frontier AI lab, so this stays at the lower featured band.

Alibaba Technology · WeChat

From Language Emergence to Collaborative Emergence: How AI Can Make High-Quality Decisions

Lv Ruofan proposes the Agent Room model: multiple agents share context, a task ledger, Memory, Runtime, and Artifacts, and two software-engineering cases show the system moving from workflow automation toward collaborative judgment rather than predefined task routing.

Why it matters: HKR-H/K/R all pass, but this is a methodology piece rather than a model launch or open-source framework. Concrete Agent Room mechanisms and 2 R&D sites put it in the 72–77 featured band.

QbitAI · WeChat

7B Medical AI Agent Beats o3 and GPT-5 by Learning Where and How to Look

Shanghai Innovation Institute’s LeapQuest and three universities released Ophiuchus and MedScope, applying Think with Images and Think with Videos to medical AI; Ophiuchus-7B scored 68.0 on eight VQA benchmarks, above OpenAI-o3 at 62.2, Gemini 2.5 Pro at 61.8, and GPT-5 at 59.9.

Why it matters: HKR-H/K/R all pass: a 7B model beating o3/GPT-5 is a strong hook, 8 VQA benchmarks with 68.0 vs 62.2 add a testable claim, and medical specialist evaluation will trigger debate. Not a frontier-lab general model release, so it stays in 78–84.

QbitAI · WeChat

Language Models Need Sleep: Let AI Nap Before Continuing Inference

Carnegie Mellon University and the University of Maryland propose a “sleep” mechanism for language models: when the context window is nearly full, the model stops accepting new tokens, runs multiple offline recursive forward passes to compress accumulated context into fast weights, clears the KV cache, and then resumes inference; tests cover cellular automata, multi-hop graph retrieval, and GSM-Infinite reasoning tasks.

Why it matters: HKR-H/K/R all pass: the sleep metaphor is clickable, and the mechanism is concrete. Score stays below 78 because the provided body lacks benchmark gains, code, or deployment evidence.

New York Times Chinese

How Google Rebounded and Started Winning the AI Race

Google said regular Gemini users more than doubled in one year to 900 million, while ad revenue rose 16% to $77 billion last quarter, and its Siri partnership with Apple will place Gemini inside future iPhone assistant features.

Why it matters: HKR-H/K/R all pass: NYT ties Google’s comeback narrative to 900M Gemini users, ad growth, and a Siri distribution deal. This is strong industry analysis, not a model launch, so it fits the 78–84 band.

Synced · WeChat

From Foundation Models to Physical AI, Samsung Moves Into the Core LLM Race

Samsung disclosed three AI efforts—Meki, M2RL, and LiveClawBench—covering a memory-based edge architecture, multi-domain reinforcement learning, and Physical AI evaluation; the article also says Samsung has purchased tens of thousands of GPUs for AI infrastructure, but does not provide model size, training budget, or deployment timelines.

Why it matters: HKR-H, HKR-K, and HKR-R pass, but this is a Samsung research bundle plus strategy signal, not a flagship model or product launch. It fits the 72–77 featured band, below same-day must-write.

AI HOT (Curated Pool)

AI Builds AI: ModelBest Open-Sources ForgeTrain, a Training Framework Written by AI

ModelBest, Tsinghua University, and OpenBMB open-sourced ForgeTrain, described as the first production-grade LLM training framework written entirely by AI with zero human code, and ModelBest used it to pretrain MiniCPM5-1B on Huawei Ascend chips.

Why it matters: HKR-H/K/R all pass: an open-source training framework, AI-written code, and MiniCPM5-1B pretraining on Ascend give concrete hooks. This is a strong tooling story, not a top-model launch, so 80 fits featured rather than P1.

Latent Space

[AINews] New AI Infra Decacorns: Fireworks, Baseten, with OpenRouter on the Way

Latent Space says Fireworks is in talks for a $15 billion valuation round, Baseten is raising at an $11 billion valuation, and OpenRouter closed a $113 million Series C after volume grew 5x in six months.

Why it matters: HKR-H/K/R all pass: the decacorn hook is clickable, the post gives valuation, round, and usage figures, and the topic speaks to inference economics. Fireworks and Baseten are still reported as in talks or raising, so this stays in the 78–84 band.

AI HOT (Curated Pool)

Code w/ Claude London event: Rethinking the developer experience

Anthropic announced two Claude Managed Agents capabilities at Code w/ Claude London: self-hosted sandboxes in public beta and MCP tunnels in research preview, with Spotify, Base44, and Legora already using them.

Why it matters: Official Anthropic product update with two concrete Claude Managed Agents capabilities. HKR-H/K/R pass, but this is a developer-tooling update rather than a major model release, so it lands at 78.

TechCrunch · AI

DuckDuckGo installs are up 30% as users reject being force-fed Google’s AI Search

Google replaced Search’s blue links with AI agents at I/O 2026, and DuckDuckGo app installs rose 30% as users looked for an alternative search entry point.

Why it matters: HKR-H/K/R all pass: the 30% install jump is a concrete backlash signal tied to Google AI Search. Missing measurement window and method keep it at the lower featured threshold.

Computing Life · Yage

Using AI Better, Step Two: Write the Skill Before Execution

The author proposes writing a Skill before asking AI to execute a task; each Skill should include three elements—success criteria, observed pitfalls, and deterministic tools—and can be organized through index.md plus AGENTS.md or CLAUDE.md for reuse.

Why it matters: HKR-H/K/R pass via a concrete Skill-first workflow and reusable agent practice. No model release, product capability, or experiment numbers, so it sits at the featured threshold.

Computing Life · Yage

Step Two to Using AI Well: Write the Skill Before You Execute

Yage argues that users should externalize work before execution by writing reusable Skills for Claude Code, Codex, and Cursor. The post gives an Outlook email example: spend about 30 minutes documenting username, phone approval, and client choice, then have AI read that file on later runs.

Why it matters: HKR-H/K/R all pass, but this is a workflow tutorial rather than a product or model release. The concrete Skill mechanism and Outlook example clear the featured floor; weak source authority keeps it at 72.

AI HOT (Curated Pool)

How we contain Claude across different products

Anthropic describes three mechanisms for containing Claude agent deployment risks across products: sandboxing or VMs, network egress controls, system-prompt and training constraints, and fine-grained permissions for MCP servers and third-party plugins.

Why it matters: Anthropic discloses a concrete containment stack for Claude agents, stronger than a routine product note. HKR-H/K/R all pass, but this is not a model launch or major capability release, so it stays in the 78–84 band.