Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

1321–1340 of 1,465

Apr 14Tuesday

最佳拍档 (BestPartners)

Meta-Harness: Can harness engineering code self-iterate? A Stanford paper analysis

Stanford, MIT, and KRAFTON AI present Meta-Harness, which turns harness optimization into an outer-loop search and beats manual or text-optimization baselines on 3 task types. The system uses a coding agent to inspect filesystem history; after 10 search iterations, the data exceeds 10 million tokens, and on online text classification it matched OPRO’s 60-iteration result in 4 iterations while reaching 75.9% average accuracy on 5 OOD datasets. The key point is full-feedback retention rather than compression; the paper also reports about 20 TerminalBench-2 iterations at a total cost of a few hundred dollars.

Why it matters: This is a good research-release explainer for agent builders: the mechanism is clear and the post includes concrete numbers, so HKR-H/K/R all pass. It stays at 80 because the source is a secondary YouTube summary, not the primary paper or official release, and the impact is still

Apr 13Monday

最佳拍档 (BestPartners)

2027 Is the Enterprise AI Singularity Year: Sundar Pichai on 10 Years as Google CEO, Transformer and Search

Sundar Pichai said in a Stripe interview that Alphabet plans $175B-$185B in 2026 capex and that 2027 will be the breakout year for enterprise AI agent workflows. He said Google cut Search latency by 30% over five years while adding AI features, manages teams with 10 ms or 30 ms latency budgets, and sees 2026-2027 constrained by wafers, memory, power, and permitting. The point to watch is not search replacement but search evolving into an agentic manager, while TPU allocation has become Google's scarcest internal resource.

Why it matters: High-signal executive commentary rather than a product launch. HKR-H/K/R all pass on the 2027 agent call, concrete capex and latency details, and the search-plus-compute nerve hit; score stays below P1 because this is a second-hand recap, not the primary interview.

Apr 12Sunday

X · @dotey

UC Berkeley team used a cheating AI to break 8 major agent benchmarks and score near perfect without solving tasks

A UC Berkeley team used a cheating AI with no LLM calls to break 8 major agent benchmarks, scoring 73% to 100% without solving tasks. The post cites three cases: a 10-line Python hook bypassed SWE-bench tests across 500 tasks, WebArena exposed answers via file://, and FieldWorkArena gave full credit to an empty {} reply. The real issue is benchmark isolation failure; the team is turning its scanner into the open-source BenchJack project.

Why it matters: HKR-H/K/R all pass: the claim is clicky, concrete, and directly threatens trust in agent evals. I stop at 84, not 85+, because the current input is a social summary; paper status, full methods, and outside replication are not disclosed here.

X · @Yuchenj_UW

MiniMax M2.7 is open-source!

MiniMax open-sourced M2.7 and said its research agent now handles 30%–50% of the R&D workflow. The post says the agent covers literature review, experiment orchestration, log debugging, code fixes, and merge requests; M2.7 also rewrote its own harness for 100+ automated rounds, with a 30% gain on internal coding evals.

Why it matters: HKR-H/K/R all pass: open-sourcing plus a research agent doing 30%-50% of R&D is a strong hook, and the post includes 100+ self-rewrite loops with +30% internal coding eval. It stays at 78 because license, repo, benchmark context, and external reproduction are not disclosed.

Apr 11Saturday

X · @dotey

Anthropic launches Claude Managed Agents beta; Michael Cohen explains secure third-party key management for agents

Anthropic added Vaults to the Claude Managed Agents beta to manage each end user's third-party credentials with a per-user vault_id and automatic injection at session runtime. The post shows a three-step flow—create a Vault, bind credentials to an MCP server address, and pass vault_id when creating a session—and prices CMA at token usage plus $0.08 per session-hour. The key design is isolation: credentials never enter Claude's context window, code runs in a sandbox, auth goes through a dedicated proxy, and the harness cannot access secrets.

Why it matters: This adds the missing implementation detail for Claude Managed Agents: third-party credential isolation. HKR-H/K/R all pass via a concrete security hook, reproducible vault_id flow, pricing, and a real operator pain point; impact stays at the developer integration layer, so it is

X · @dotey

OpenAI Codex team's Nick Baumann: build dedicated CLI tools for AI instead of feeding messy data repeatedly

OpenAI Codex engineer Nick Baumann says teams should wrap repeated data access into parameterized CLI tools with JSON output instead of repeatedly dumping logs, docs, and API responses into Codex. The post lists 3 examples in daily use: codex-threads for past sessions, slack-cli for threaded Slack search, and typefully-cli for posting workflows; access still goes through the existing auth gateway. The point for practitioners is narrower interfaces: models handle focused commands more reliably than raw, noisy source data.

Why it matters: This is a practical workflow note from an OpenAI Codex team member, not a formal launch, but it offers a reusable mechanism: wrap noisy context behind parameterized JSON-returning CLIs and shows 3 live examples. HKR-H/K/R all land; no benchmark, scale, or major product release,so

QbitAI · WeChat

OpenClaw-style methods reach multimodal generation, with a 6B model beating Nano Banana 2 on some tasks

A team led by Shanghai AI Laboratory introduced GEMS, adding Agent Loop, Memory, and Skills to multimodal generation, and reports that 6B Z-Image-Turbo beats Nano Banana 2 on some tasks. The post reports +14.22 average gains on 5 mainstream tasks and +8.92 over the best baseline on 4 downstream tasks; the paper and code are public, but the post does not disclose Nano Banana 2's full setup.

Why it matters: Strong HKR-H/K/R: the hook is a 6B multimodal model beating Nano Banana 2, and the post includes mechanism plus testable deltas (+14.22 / +8.92) with paper and code. It stays below P1 because the article does not disclose the full Nano Banana 2 comparison setup.

X · @dotey

Anthropic launches Claude for Word beta add-in

Anthropic released a beta Claude for Word add-in for paid Claude Team and Enterprise users, with direct sidebar editing for .docx and .docm files. Edits appear in Word’s native track changes flow, the add-in can reuse conversation context from Excel and PowerPoint, and it supports reference uploads plus reusable team Skills. The key point is shared context across Office apps; the post does not disclose pricing, regions, or a wider rollout timeline.

Why it matters: This is a substantive Anthropic product update for Team and Enterprise, not a generic integration post. HKR-H/K/R all pass on novelty, concrete mechanics, and workflow resonance, but the beta scope is limited and price, regions, and GA timing are undisclosed, so it lands in mid-"

X · @dotey

Claude Code adds ultraplan: start planning in terminal, review in browser, then run in cloud or locally

Claude Code opened a preview of ultraplan to users with the web app enabled, requiring v2.1.91+, and planning starts from /ultraplan in the terminal. Claude drafts a plan in the cloud after reading the repo, users review and annotate it in the browser, then choose cloud execution with a PR or local terminal execution. The key change is splitting planning from execution: planning moves to the cloud without blocking the terminal, and the post says token use is close to local plan mode.

Why it matters: This is more than a routine feature add: Claude Code splits planning from execution, with /ultraplan in terminal, cloud-side repo reading, browser review, and cloud PR or local execution. HKR-H/K/R all pass, with a Claude-specific bump, but it is still a preview and sourced froma

Apr 10Friday

最佳拍档 (BestPartners)

LLM self-evolution: Shinka Evolve, AlphaEvolve, and sample efficiency

Sakana AI open-sourced Shinka Evolve and uses a UCB bandit to switch among GPT-5, Claude Sonnet 4.5, Gemini, and others, aiming to cut the thousands of program evaluations common in AlphaEvolve-style search. The post says it beat AlphaEvolve’s classic circle-packing result with fewer evaluations and adds full-file rewrites, crossover, editable-region guards, and a meta-notebook; the post does not disclose exact metrics, cost, or the repo link. The part to watch is surrogate-task design and hard verification: the system still needs humans to define problems.

Why it matters: Featured, not P1: HKR-H/K/R all pass. The piece has a strong hook, concrete mechanisms like UCB model routing and program crossover, and a real nerve around eval cost and hard verification. It stays at 80 because key metrics, cost, and the primary release link are not disclosed.

QbitAI · WeChat

Claude bug mixes up speaker roles, issues self-instructions, and blames the user

A developer said Claude 3.5 and Claude 4 can confuse user, assistant, and system roles under complex or malicious context, and the Hacker News post drew heavy discussion. The post cites inputs like <stop> and <end prompt> as a repro clue; Anthropic's fix status and scope are not disclosed. The real issue is control-data separation, not a single prompt failure.

Why it matters: This clears all HKR axes: the angle is clickworthy, the post includes a concrete repro clue, and the failure mode matters to anyone shipping agents. I kept it below P1 because scope, affected versions, and Anthropic’s fix status are not disclosed.

X · @dotey

Anthropic launches Advisor Tool API: cheaper models execute while pricier models advise on hard decisions

Anthropic launched the advisor tool API, letting Sonnet or Haiku execute tasks and consult Opus on hard decisions; it is in beta and requires the anthropic-beta: advisor-tool-2026-03-01 header. The RSS snippet says Sonnet+Opus gains 2.7 points on multilingual SWE-bench while cutting per-task cost by 11.9%; Haiku+Opus rises from 19.7% to 41.2% on BrowseComp at 15% of Sonnet's cost. The key detail is the call path: model switching happens inside one Messages API request, advisor and executor tokens are billed separately, and max_uses caps consultations.

Why it matters: This is a substantive Anthropic API update with concrete mechanics: in-request model routing, separate token billing, max_uses, and two benchmark/cost deltas. HKR-H/K/R all pass, so it merits featured, but it is still below a model-release tier event.

X · @claudeai

We're bringing the advisor strategy to the Claude Platform.

Claude is adding the advisor strategy to Claude Platform, with Opus as the advisor and Sonnet or Haiku as the executor. The RSS snippet says this yields near-Opus-level agent intelligence at lower cost; the post does not disclose pricing, benchmark scores, or rollout timing.

Why it matters: Anthropic ships a substantive Claude Platform update, and HKR-H/K/R all pass: the Opus-advisor plus Sonnet/Haiku-executor setup is novel, concrete, and directly relevant to agent builders. The score stays below P1 because price, benchmarks, and rollout timing are not disclosed.

Apr 9Thursday

QbitAI · WeChat

Beyond MoE, Tencent introduces MoT: a 2B embodied model ranks first in 16 of 22 evaluations

Tencent Hunyuan and Robotics X released HY-Embodied-0.5; its MoT-2B uses 4B total params with 2B active and ranks first in 16 of 22 embodied evaluations. The post says it uses 100M+ embodied data, 600B+ pretraining tokens, 30M+ mid-training samples, plus visual latent tokens, bidirectional attention, RFT, RL, and online distillation. The key point is a rebuilt edge-oriented embodied stack, not a simple VLM fine-tune.

Why it matters: Strong on HKR-H/K/R: the headline has a real hook, the body includes concrete numbers and training mechanisms, and the edge-robotics angle lands with practitioners. I keep it at 83, not 85+, because this is a high-quality embodied-model release, not a broad same-day industry-def

QbitAI · WeChat

Claude launches managed agent service, then faces an open-source alternative from Multica

Anthropic has opened Claude Managed Agents and charges $0.08 per session-hour plus token usage. The service supports hours-long runs, sandboxing, checkpoint recovery, and multi-agent orchestration; web search costs $10 per 1,000 searches, while some memory and orchestration features remain in research preview. The title mentions a “lobster ban,” but the post does not disclose that context; the real shift is Anthropic selling agent infrastructure to enterprises.

Why it matters: This is a substantive Anthropic product update: a managed-agent service with explicit pricing, sandboxing, checkpoint restore, and search costs, so HKR-K is strong. HKR-R is real for Claude-heavy teams weighing build vs. buy, but the scope is smaller than a model launch, so it is

X · @op7418

Meta releases Muse Spark model

Meta released the Muse Spark model with native multimodal reasoning, tool use, visual chain-of-thought, and multi-agent orchestration, but it is only available in the Meta AI app and is not open source for now. The snippet says its Contemplating mode coordinates multiple parallel agents for reasoning, and its Artificial Analysis score is below Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. The post does not disclose model size, pricing, or rollout timing.

Why it matters: A major-lab model launch plus the “poached team’s first output” angle lands HKR-H/K/R. The score stays near the featured floor because the post offers capability claims and relative benchmark placement only; params, pricing, rollout timing, and access scope are not disclosed.

X · @dotey

Anthropic launches Claude Managed Agents, a managed API for building and deploying agents, now in public beta

Anthropic launched Claude Managed Agents, a managed API for building and deploying agents, in public beta. It offers a production sandbox, long-running sessions, and multi-agent coordination; Anthropic says internal tests showed up to a 10-point success-rate gain on structured file-generation tasks versus standard prompt loops. Pricing uses standard Claude token fees plus $0.08 per active session-hour; the real signal is Anthropic moving agent infrastructure into its platform layer.

Why it matters: Anthropic packaged managed agents, sandboxing, and long-running sessions into a public-beta API, which is a real workflow update for developers. HKR-H/K/R all pass: strong platform hook, concrete facts like a 10-point gain and $0.08 per hour, and clear resonance around developer-

X · @claudeai

Introducing Claude Managed Agents: everything you need to build and deploy agents at scale.

Claude has launched Claude Managed Agents in public beta on Claude Platform, claiming to compress the path from agent prototype to launch into days. The post discloses only a performance-tuned agent harness plus production infrastructure; pricing, toolchain support, model scope, and quotas are not disclosed.

Why it matters: Anthropic gets a positive bump, and HKR-H/HKR-R pass because managed agent deployment is a strong hook for Claude-heavy builders. HKR-K is limited: the post discloses a harness and prod infra, but not pricing, toolchain support, model scope, or quotas.

Apr 8Wednesday

MIT Technology Review · AI

Mustafa Suleyman: AI development won’t hit a wall anytime soon—here’s why

Mustafa Suleyman argues frontier AI training compute rose from about 10^14 to over 10^26 FLOPs since 2010, a 1 trillion-fold increase, so AI development is not near a wall. He cites a 7x Nvidia chip gain in six years, 3x more HBM3 bandwidth, and Epoch AI estimates that compute needed for fixed performance halves every eight months. The piece is commentary from Microsoft AI’s CEO, not an independent study; the post does not disclose a reproducible basis for the 200GW-by-2030 claim.

Why it matters: HKR-H/K/R all pass: Suleyman takes a hard line in the scaling-wall debate and cites 10^26 flops, 7x chip gains, 3x bandwidth, and 8-month efficiency halving. Held at 82 because this is executive commentary, not independent research, and the 2030 200GW math is not disclosed.

X · @dotey

Hermes Agent is gaining traction; I installed it and the experience was decent

Nous Research open-sourced Hermes Agent in late February, and the post says it reached nearly 30,000 GitHub stars in under two months. The post describes a closed learning loop: after complex tasks with 5+ tool calls, Hermes writes Markdown skills, with one Reddit report claiming 3 skills in 2 hours and a 40% speedup on repeated research work. The key angle is its self-hosted agent engine that combines skill generation, SQLite-based memory retrieval, and five-layer safety controls.

Why it matters: HKR-H/K/R all pass: the piece combines strong OSS momentum, concrete mechanics, and a real builder nerve around self-hosted learning agents. It stays at 78 because the evidence is mostly social commentary and light user feedback, not a primary release or broad independent eval.