Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

1081–1100 of 1,465

May 7Thursday

Ben's Bites

Elon Doubled Limits

Ben’s Bites says Anthropic doubled Claude usage on paid plans via SpaceX’s Colossus 1. The issue also lists GPT-5.5 Instant, ChatGPT spreadsheet integration, and three Claude Managed Agents features. The title names Elon, but the post does not disclose exact limits.

Why it matters: HKR-H/K/R pass: the SpaceX Colossus 1 angle, 2x Claude usage, and quota pressure are all concrete. Missing exact caps, pricing, and rollout scope keep it in the low featured band.

AI HOT (Curated Pool)

Consistent web search and scraping for all models

OpenRouter released tools for tool-calling models to run web search and page scraping. The post says multiple search and scraping engines are supported, but does not disclose names, pricing, or limits. The key item is cross-model tool interface consistency.

Why it matters: HKR-H/K/R all pass, but engines, pricing, and limits are not disclosed. This is a mid-weight Product update: useful for model-agnostic agent stacks, not a major model or capability release.

AI HOT (Curated Pool)

Anthropic Institute Outlines Four Core Research Areas

Anthropic Institute named four research areas: economic diffusion, threats and resilience, real-world AI systems, and AI-driven R&D. The post says it will publish a more granular Anthropic Economic Index and study how AI tools speed AI research. The results will inform Anthropic’s Long-Term Benefit Trust.

Why it matters: HKR-K comes from 4 named research tracks and the Economic Index plan; HKR-R is strong on labor and governance. It is an agenda, not a model, product, or finished result, so it stays in the 72–77 band.

Latent Space

Anthropic-SpaceXAI's 300MW/$5B/yr Deal for Colossus I, ARR Growth Is 8000% Annualized

Anthropic announced a SpaceX compute partnership, doubled Claude Code’s 5-hour limits for Pro, Max, Team, and seat-based Enterprise, raised Opus API limits, and said Claude inference would ramp on Colossus within days; the post treats the 300MW and $5B-per-year figures as widely circulated but not canonized in Anthropic’s own announcement.

Why it matters: HKR-H/K/R all pass: the compute-deal numbers and Claude Code limit changes are concrete and practitioner-relevant. The 300MW/$5B/year claim is unofficial, so it stays below P1.

Xinzhiyuan · WeChat

Claude Managed Agents Add Dreaming, With Reported Task Completion Up to 6x

Anthropic added Dreaming, Outcomes, and multi-agent orchestration to Claude managed agents; Harvey reports about 6x higher task completion. Dreaming reads up to 100 sessions; one demo distilled 5.3M tokens into 98 rules, while Outcomes raised success by up to 10 points. Opus 4.7 and Sonnet 4.6 require access, with $0.08 per session-hour runtime fees.

Why it matters: HKR-H/K/R all pass: Anthropic adds Dreaming, Outcomes, and multi-agent orchestration with 100-session memory, $0.08/session-hour runtime, and Harvey’s ~6x completion claim. This is a same-day Claude agent update.

Bloomberg Technology

Kimi Chatbot Maker Moonshot AI Valued at $20 Billion in Meituan-Led Round

Moonshot AI raised about $2 billion, reaching a $20 billion valuation. The title says Meituan led the round; the post does not disclose investors, stake size, or use of funds. It signals strong demand for Chinese AI startups.

Why it matters: Bloomberg reports Moonshot AI raised about $2B at a $20B valuation, a major capital event for a Chinese model lab. HKR-H/K/R all pass; investor details and use of funds are not disclosed, so this sits in the lower 85–94 band.

AI HOT (Curated Pool)

Amp releases Neo CLI as coding agents shift toward long-horizon workflows

Amp released Neo, a CLI tool covering remote orchestration, automatic context compression, and a Plugin API. Neo lets local threads be controlled remotely, allows all operations by default, and moves safety control to plugins; the post does not disclose version, pricing, or performance gains.

Why it matters: HKR-H/K/R all pass: Neo adds remote orchestration, context compression, Plugin API, and default-allow permissions. Amp’s reach and missing price/version/perf data keep it in the 72–77 band.

Synced · WeChat

TACO Lets CLI Agents Drop Useless Context Through Self-Evolving Compression

TACO proposes a training-free terminal-observation compression framework, improving success rate and token efficiency on TerminalBench 1.0/2.0 and related benchmarks. It evolves rules within tasks, writes validated rules to a global pool, and finds 24.6%–44.1% low-value redundancy in TerminalBench 2.0 raw prompts. The key signal is stability: Top-30 rule retention exceeds 90% after multiple evolution rounds.

Why it matters: HKR-H/K/R all pass: the paper targets CLI-agent context bloat with a no-training rule-pool mechanism and concrete TerminalBench numbers. It is strong agent research, not a major model or product launch, so it sits in the 78–84 featured band.

Synced · WeChat

Claude, GPT and Gemini score 0% completion on ProgramBench

ProgramBench tested Claude Opus 4.7, GPT-5.4 and Gemini 3.1 Pro, with 0% full completion on rebuilding software projects. It gives only executables and usage docs, removes source/tests, and grades behavioral equivalence via agent-driven fuzzing. The key signal is system-level engineering, not function-level code generation.

Why it matters: HKR-H/K/R all pass: the 0% result is clickable, the setup is concrete, and the coding-agent gap matters to practitioners. Still, it is a single benchmark report, below a major model or product release.

AI HOT (Curated Pool)

Open Slide lets AI write PPT code

Open Slide builds PPTs with React, using a workflow designed for AI agents. It integrates SVGL with 1,500+ brand logos, supports manual edits, and lets AI read user comments for revisions.

Why it matters: HKR-H/K/R pass: the programmable-slide angle is clickable, with concrete React and 1500+ logo details, and deck work is a real practitioner pain. No usage metrics or hands-on test keeps it at the featured threshold.

Computing Life · Share · Yage

Agent Filesystems: From Feeding Models Memory to Letting Models Browse Files

The article frames agent filesystems as a three-stage shift from raw context to memory systems to filesystem-as-context, covering design choices from Turso, Anthropic, Vercel, and Manus, and listing four overlooked blind spots.

Why it matters: HKR-H/K/R all pass, but this is design commentary rather than a product or research release. Named comparisons across Turso, Anthropic, Vercel, and Manus justify featured, not the 78+ band.

The Verge · AI

Google shuts down Project Mariner

Google shut down Project Mariner on May 4, 2026. The experimental web-task agent once supported up to 10 concurrent tasks. Its technology moved into Google products, including Gemini Agent.

Why it matters: HKR-H/K/R all pass, but the disclosed facts are limited to shutdown timing, a 10-task limit, and migration into Gemini Agent. Strong source authority supports low featured, not a major launch.

Bloomberg Technology

Meta-Backed Scale AI Wins $500 Million Defense Department Deal

The Pentagon awarded Scale AI a $500 million contract to sift data and support decisions. The post says Meta Platforms backs Scale AI, but does not disclose term, deployment scope, or model details. The key signal is US military AI spending moving into data workflows.

Why it matters: HKR-H/K/R pass on the $500M Pentagon contract and defense-AI procurement angle. Missing term, deployment scope, and model details keep it in the 72–77 featured band.

r/LocalLLaMA

Analysis of 922 Agentic Task Traces Finds DeepSeek v4’s Cost Edge in Caching

A Reddit user analyzed 922 agentic task traces and reported $0.01 per task for DeepSeek v4 Flash versus $1.52 for Opus 4.7. Both used about 960K tokens per task, but DeepSeek showed a 97% cache hit rate versus 87%, with a 0.02 cache read/write price ratio versus 0.08. The key issue is caching, not headline pricing.

Why it matters: HKR-H/K/R all pass: 922 agent traces tie a large cost gap to cache hit rate and cache read/write pricing. Reddit single-source data and incomplete method detail keep it in the 78–84 band.

May 6Wednesday

r/LocalLLaMA

CopilotKit (MIT): Open-source building blocks for agent apps and generative UI

CopilotKit offers MIT-licensed React components and claims 30k GitHub stars. It covers chat, streaming, tool calls, HITL, and generative UI, with AG-UI support for LangGraph, CrewAI, LlamaIndex, and other backends. The key point is decoupling the UI layer from agent frameworks.

Why it matters: HKR-H/K/R all pass: MIT open source, 30k stars, and AG-UI links to major agent backends. Kept in 72–77 because the post lacks a new version, benchmark, or named production adopter.

TechCrunch · AI

Apple to Pay $250M to Settle Lawsuit Over Siri's Delayed AI Features

Apple agreed to pay $250 million to settle a class action over delayed Siri AI features. The snippet cites overpromised rollout claims; the post does not disclose user count, payout rules, or feature timing.

Why it matters: HKR-H/K/R all pass: Apple’s $250M Siri AI settlement has a strong hook, a concrete number, and product-liability resonance. Missing payout rules, affected-user count, and feature timeline keep it in the 78–84 band.

r/LocalLLaMA

An Open Benchmark for Testing RAG on Realistic Company-Internal Data

EnterpriseRAG-Bench released a 500k-document corpus for testing RAG on company-internal data. It simulates Redwood Inference across 9 sources and includes 500 questions over 10 retrieval failure modes. Baselines show BM25 beats vector search overall, while agentic/bash retrieval has the best completeness at higher cost and latency.

Why it matters: HKR-H/K/R all pass: the benchmark targets a real enterprise RAG pain point, with 500k docs and testable BM25-vs-vector results. Single Reddit-source benchmark release keeps it below same-day must-write.

Latent Space

AINews: Silicon Valley Gets Serious About Services

Anthropic and OpenAI announced enterprise services companies: Anthropic’s unnamed JV is funded with $1.5 billion, while OpenAI’s The Deployment Company has raised about $4 billion at a $10 billion pre-money valuation.

Why it matters: HKR-H/K/R all pass: the hook is labs turning into services operators, with $1.5B and ~$4B figures. The scale and OpenAI/Anthropic names put it in must-write territory.

Synced · WeChat

DeepSeek Version of Claude Code Tops Trending Chart With 8,700 Stars

DeepSeek TUI topped GitHub trending with over 8,700 stars. Hunter Bown built it in Rust for local terminal use with DeepSeek V4, supporting chat, file edits, shell commands, and task management. The key detail is RLM mode: up to 16 V4 Flash subtasks, plus a 1M-token context window and approval gates.

Why it matters: HKR-H/K/R all pass: the 8,700-star hook is strong, RLM adds concrete mechanisms, and coding-agent competition resonates. It is a third-party open-source tool, not an official DeepSeek model release, so it stays in the 78–84 band.

Synced · WeChat

Two Chinese open-source projects turn Mac into a private AI workstation

Mininglamp open-sourced Cider and Mano-P 1.0 for Apple Silicon local inference and GUI agents. Cider speeds Qwen3-VL-2B prefill by 57%–61% on M5 Pro; Mano-P 1.0-72B scores 58.2% on OSWorld. The key constraint is W8A8 memory: on 16GB devices accuracy falls from 58.0% to 54.0%, so 32GB+ is recommended.

Why it matters: HKR-H/K/R all pass: the Mac-local workstation angle is clickable, and Cider/Mano-P include testable numbers. Score stays at 80 because the source entity is not a top-tier model lab.