Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

901–920 of 1,465

May 19Tuesday

Hacker News front page

We Let AIs Run Radio Stations

Andon Labs gave four AI agents tools to host live radio shows and run a media business without humans. The post says revenue is terrible, but it does not disclose amounts or operating metrics.

Why it matters: HKR-H is strong from the AI-run radio premise; HKR-K has a concrete 4-agent experiment but weak revenue disclosure; HKR-R lands on agent economics and media automation. Niche lab post, so it sits at the featured threshold, not same-day news.

AI HOT (Curated Pool)

Cursor releases Composer 2.5 coding model

Cursor released Composer 2.5, claiming up to 10x higher efficiency on long coding tasks; the model is further trained on Moonshot’s Kimi K2.5 and uses text feedback for 100k-token-scale trajectories.

Why it matters: HKR-H/K/R all pass: Cursor is a core AI coding tool, and Composer 2.5 adds concrete claims around 10x long-task gains and Kimi K2.5 tuning. Limited sourcing and no independent eval keep it in the 78–84 band.

r/LocalLLaMA

Tried Every Hermes Agent Alternative So You Don't Have To: 2026 Roundup

A Reddit user compared 11 Hermes Agent alternatives across open-source and managed options; OpenClaw is listed with 347k GitHub stars, 24+ integrations, and 9 CVEs in four days, while TrustClaw uses OAuth-only sandboxed execution and Perplexity Computer requires a $200/month Max tier.

Why it matters: HKR-H/K/R all pass: this is a practical agent-tool comparison with 11 items and concrete integration/security figures. Reddit single-post sourcing limits confidence, so it stays near the featured threshold.

AI HOT (Curated Pool)

Take your local GitHub sessions anywhere

GitHub launched remote control sessions for Copilot, letting users start tasks in VS Code or the command line and continue them through github.com or GitHub Mobile.

Why it matters: GitHub Copilot session handoff from VS Code/CLI to web and mobile clears HKR-H/K/R, but the post only gives entry points and use case; permissions, pricing, and supported task scope are not disclosed.

May 18Monday

Hacker News front page

Show HN: InsForge – Open-source Heroku for coding agents

InsForge released an Apache 2.0 backend platform that lets coding agents deploy, operate, and debug backend systems through one CLI install command and Skills.

Why it matters: HKR-H/K/R all pass: the Heroku-for-agents framing, Apache 2.0 plus one-CLI install, and agent ops pain are concrete. Source is mainly Show HN/GitHub with no usage, benchmark, or production proof, so it sits at the featured threshold.

AI HOT (Curated Pool)

The Open Agent Leaderboard

IBM Research published the Open Agent Leaderboard on Hugging Face to evaluate agents across language understanding, tool use, and multi-step reasoning tasks; the post does not disclose dataset size, model scores, or the evaluation date.

Why it matters: HKR-H and HKR-R pass because an open agent leaderboard speaks to agent-eval pain. HKR-K fails: the article lacks scores, dataset size, and evaluation date, so it sits at the featured threshold.

Latent Space

The Autonomous Drone Tech Stack and Economics of Drones — Yaroslav Azhnyuk

Latent Space interviewed The Fourth Law founder Yaroslav Azhnyuk for a two-hour episode covering FPV drones, five levels of autonomy, eight dimensions of the autonomous battlefield, and China’s manufacturing advantage; the transcript claims Ukraine produced 4 million FPV drones last year and discusses a hypothetical Chinese capacity of 4 billion.

Why it matters: HKR-H/K/R all pass: the Latent Space interview offers concrete autonomy and battlefield frameworks. It is still commentary, not a model release, product update, or research artifact, so it stays just above the featured threshold.

AI HOT (Curated Pool)

OpenAI and Dell partner to bring Codex to hybrid and on-prem enterprise environments

OpenAI and Dell are partnering to bring the Codex coding agent to enterprise hybrid-cloud and on-premises deployments; the RSS snippet does not disclose launch timing, pricing, supported regions, or the specific security controls for sensitive data.

Why it matters: HKR-H/R are strong: Codex via Dell targets hybrid/on-prem enterprise code. HKR-K is limited to deployment path; launch date, price, regions, and security controls are absent, keeping it at the featured threshold.

QbitAI · WeChat

Agents Learn to Grow Skills from Failure: EvolveR Accepted by ICML 2026

EvolveR lets agents distill reusable experience from successful and failed trajectories, maintain a scored experience library, and train retrieval behavior with GRPO; the paper reports the best average performance on seven complex QA benchmarks using Qwen2.5-3B and 7B.

Why it matters: HKR-H/K/R all pass: the agent self-growing-skill angle is clickable, with mechanism and benchmark specifics. Since only a media summary is available and no repo, absolute scores, or reproduction details are disclosed, it stays in the 78–84 research band.

QbitAI · WeChat

openJiuwen open-sources JiuwenSwarm, a multi-agent swarm coordination framework

openJiuwen released and open-sourced JiuwenSwarm with four components: Agent Swarm, Swarm Skills, Skills Hub, and self-evolution, and the framework supports HOTS and HITS modes for human participation in multi-agent workflows.

Why it matters: HKR-H/K/R pass: the swarm angle is clickable, the post gives four modules plus HOTS/HITS, and agent builders care about orchestration choices. Lacking benchmarks or adoption data keeps it at the featured threshold.

Bloomberg Technology

Baidu AI Sales Eclipse Waning Legacy Ads for the First Time

Baidu reported a 1% revenue decline as growth in nascent AI businesses offset shrinking traditional internet revenue; the post does not disclose AI sales, advertising revenue, or details of the agentic AI pivot.

Why it matters: Baidu revenue fell 1% while AI sales topped legacy ads for the first time, so HKR-H/K/R pass. Missing AI/ad dollar splits and agentic-AI mechanics keep it in the low featured band, not p1.

r/LocalLLaMA

I built a coding agent that gets 87% on benchmarks with a 4B parameter model

SmallCode passes 87 of 100 benchmark tasks with Gemma 4 activating 4B parameters per token. The author attributes the result to compound tools, compile and lint feedback, task decomposition after two repeated failures, and optional escalation to Claude or OpenAI for one task.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post and the benchmark identity plus replication details are incomplete. It fits a concrete first-person experiment above the featured bar, not the 78+ band.

Synced · WeChat

openJiuwen releases JiuwenSwarm, an open-source multi-agent swarm framework

openJiuwen released and open-sourced JiuwenSwarm with four components: Agent Swarm, Swarm Skills, Swarm Skills Hub, and self-evolving Swarm Skills, and reports a 94.2% PinchBench score versus 91.6% for OpenClaw.

Why it matters: HKR-H/K/R all pass: an open-source agent-swarm framework with named components and a PinchBench 94.2% claim. It stays at 78 because openJiuwen is not a top lab and the summary lacks license, reproduction setup, and baselines.

AI HOT (Curated Pool)

Tencent AI Design Agent Ardot Enters Public Beta: Generates Editable Designs and Converts Them to Code

Tencent Cloud opened public beta for Ardot, an AI design agent that generates editable app pages, websites, and posters from one-sentence prompts, then converts designs to code.

Why it matters: HKR-H/K/R pass on a concrete Tencent product beta for editable design-to-code workflows. Missing pricing, model details, benchmarks, and field results keep it at the lower featured threshold.

AI HOT (Curated Pool)

Grok launches Skills feature

xAI launched Grok Skills on May 18, 2026, letting users set preferences, formatting rules, or workflows once and keep them active across all conversations on web, iOS, and Android.

Why it matters: HKR-H/K/R all pass: Grok Skills adds persistent preferences and workflows across web, iOS, and Android. This is a mid-weight xAI product update; rollout scope, limits, and pricing are not disclosed.

AI HOT (Curated Pool)

Composer 2.5 release and technical analysis

Cursor released Composer 2.5, built on a Moonshot open-source checkpoint, trained with synthetic data from real codebases at 25 times the previous scale, and updated with text-feedback reinforcement learning and a sharded Muon optimizer.

Why it matters: HKR-H/K/R all pass: Cursor is a core coding-agent surface, and the post gives concrete training details around Moonshot, 25x data, RL, and Muon. It lacks benchmarks, pricing, or user-facing capability limits, so it stays in the 78–84 band.

Google DeepMind

Google DeepMind adds Street View grounding to Project Genie

Google DeepMind has added Street View real-scene grounding to its experimental prototype Project Genie. Users can pick a US location, then pair it with a style and characters to generate a world.

Why it matters: With Street View imagery wired in, agents and robots can train and navigate in virtual environments that track real places.

May 17Sunday

Hacker News front page

Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep

MinishLab open-sourced Semble, a code-search tool for agents that combines Model2Vec embeddings, BM25, RRF fusion, and reranking; on a 63-repo benchmark, it used 98% fewer tokens than grep+read, reached 0.854 NDCG@10, and ran CPU queries in about 1.5 ms.

Why it matters: HKR-H/K/R all pass: the 98% token claim is clickworthy, the 63-repo benchmark adds substance, and coding-agent context cost is a real practitioner nerve. Impact is still toolchain-level, so it stays below must-write.

r/LocalLLaMA

MiroThinker-1.7 Open-Weight Deep Research Agent Based on Qwen3 MoE

MiroMindAI released the MiroThinker-1.7-deepresearch and mini APIs, with the mini version using 30B total parameters and 3B active parameters, weights on HuggingFace, and context management based on sliding window K=5 plus episode restarts.

Why it matters: HKR-H/K/R all pass, but the source is a Reddit thread and the lab is not top-tier. Open weights, MoE sizing, and context-management details clear featured, not same-day must-write.

Bloomberg Technology

Apple’s New ChatGPT-Like Siri App Will Have Auto-Deleting Chats

The title says Apple’s ChatGPT-like Siri app will support auto-deleting chats; the RSS snippet only adds that iOS 27 will include a Genmoji upgrade, and the post does not disclose retention periods, release timing, or feature details.

Why it matters: HKR-H and HKR-R pass because Bloomberg frames a specific Apple Siri privacy angle; HKR-K fails since retention and feature mechanics are missing, so this stays at the low featured threshold.