Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

1281–1300 of 1,465

Apr 19Sunday

Xinzhiyuan · WeChat

Amap unveiled ABot-Claw and its quadruped robot Tutu at the Yizhuang Half Marathon

Amap unveiled the ABot-Claw agent system and the quadruped robot Tutu, claiming an autonomous guide-dog demo in the 2026 Yizhuang robot half marathon. The post gives three concrete numbers: ABot-M0 reached 80.5% on Libero-Plus, nearly 30% above Pi0; ABot-N0 hit SOTA on 7 navigation benchmarks; the open UniACT dataset contains 6 million trajectories and 9,500+ hours. What matters is Map as Memory, cloud-edge control, and closed-loop self-correction; the post does not disclose race ranking, pricing, or launch timing.

Why it matters: HKR-H/K/R all pass: the open-environment half-marathon demo is a strong hook, and the post includes concrete benchmark numbers plus a 6M-trajectory release. Kept below p1 because rank, pricing, ship date, and independent replication are not disclosed, and the impact is narrower a

r/LocalLLaMA

I tested 8 LLMs as tabletop GMs: a 27B model beat the 405B on narrative quality

The author tested 8 LLMs on 6 fixed tabletop-GM scenarios, and google/gemma-3-27b-it ranked first in narrative quality with a 4.33 overall score. The probe used 8 auto metrics plus 3 LLM-judge scores, and the full run cost about $0.02; the title says a 27B beat a 405B, but the snippet does not disclose the 405B model name or full rankings.

Why it matters: A named first-person benchmark with a strong surprise hook clears HKR-H, HKR-K, and HKR-R. I kept it at featured, not higher: the source is Reddit, the post is truncated, and the 405B model name plus full ranking are not disclosed.

r/LocalLLaMA

Deep dive into LangGraph’s Pregel execution model, checkpointing internals, and DeepAgents

A technical post breaks down LangGraph as a high-level wrapper over a Pregel runtime, with PregelNodes, channels, and reducers as the core primitives. The RSS snippet cites four Postgres checkpoint tables, a Plan/Execute/Update superstep flow, and compile() preflight validation; the post does not disclose benchmark numbers in the snippet. The real takeaway is the unified runtime view of parallel execution, checkpoint write amplification, and subgraph boundaries.

Why it matters: HKR-H/K/R all pass: the post reframes LangGraph as a Pregel runtime and adds concrete internals like 4 checkpoint tables and Plan/Execute/Update supersteps. Kept at 74 because this is a Reddit deep dive, not an official release, and no benchmark or production case is disclosed.

Apr 18Saturday

QbitAI · WeChat

OpenClaw has reached the milk tea business

Guming and Intime Retail said OpenClaw tests exposed 5 deployment risks: default port 18789 exposure, at least 8% malicious Skills, privilege overreach, 20+ minutes of runaway token use, and weak legacy defenses. Reported incidents include an agent closing a normal bastion-host port and locking out ops staff, plus requests for unrelated permissions like microphone access. The real issue is not chat UX but agents touching enterprise networks, credentials, and production systems.

Why it matters: This is not generic AI-safety commentary; it documents five concrete deployment risks and one ops outage, so HKR-H/K/R all pass. It stays below P1 because the evidence is still case-level testing, with no official fix, broad rollout impact, or cross-source cluster.

Synced · WeChat

What is OpenAI prioritizing under compute limits?

Greg Brockman said OpenAI narrowed priorities under hard compute limits to two bets: a personal assistant and AI workers that solve hard user problems, and current compute cannot fully support both. The snippet says Sora resources were reduced while focus shifted to reasoning models, a unified AI layer, and the next base model Spud; it does not disclose the claimed compute budget, timeline, or model specs. The key point is not a B2B retreat but a compute-driven reprioritization.

Why it matters: HKR-H/K/R all pass: the compute-ceiling angle is strong, the piece adds concrete priority shifts, and OpenAI roadmap triage hits cost and dependency nerves. It stays at 80 because this is secondary reporting; spend, timing, and technical details are not disclosed.

Xinzhiyuan · WeChat

Claude Opus 4.7 splits users 48 hours after launch: benchmark lead, reasoning tests drop

Anthropic's Claude Opus 4.7 drew split reactions within 48 hours: Artificial Analysis scored it at 57, tied for No.1, while NYT Connections Extended fell from 94.7% on 4.6 to 41.0%. The post says a new tokenizer raises token usage to 1.0-1.35x on the same text, and old thinking parameters can return 400 errors; Anthropic also cites a 1753 Elo GDPval-AA score, 79 points above No.2. The real issue is migration cost and capability trade-offs, not a single leaderboard.

Why it matters: The signal is not the “backlash” framing but the four concrete shifts: benchmark lead, reasoning drop, higher token use, and API breakage. HKR-H/K/R all land, but this is secondary analysis 48 hours after launch, not the primary Anthropic release, so it stays below p1.

Xinzhiyuan · WeChat

Bilibili debate: Hermes responds to plagiarism claims for the first time, as MiniMax moves early on Harness

MiniMax says its M2.7 model now handles 30%-50% of daily workflows in its RL team, ran over 100 self-optimization loops, and improved evals by 30%. The post also says Hermes Agent grew from 2B to nearly 300B daily tokens, while M2.7 exceeds 25B daily tokens on OpenRouter; Hermes lead Tommy Eastman denied copying EvoMap in a livestream. The real signal is Harness: the post cites 20-40ms or 80ms sandbox startup and 15k to 600k instances per minute, showing competition is shifting from benchmark scores to agent execution infrastructure.

Why it matters: HKR-H/K/R all pass: the plagiarism-response angle pulls clicks, and the story carries concrete metrics on workflow share, self-optimization loops, sandbox latency, and concurrency. It stays at 83 because this is a dense secondary report, not a primary launch or official technical

X · @dotey

Anthropic designer Ryan Mather shares Claude Design tips while covering 7 product lines

Anthropic designer Ryan Mather shared 9 Claude Design workflow tips while covering 7 product lines. The RSS snippet says to spend 1 hour building a design system, use chat for large changes, comments for small edits, specify feedback like 8px spacing, and attach only the target component folder instead of a full monorepo. The key shift is process: from human-do/human-review to Claude-do/human-review.

Why it matters: This is a strong practitioner workflow note: an Anthropic insider shares concrete, reusable tactics, so HKR-H/K/R all pass. It stays below the 80s because this is not a formal Claude product release and the post does not disclose harder outcome data such as time saved or task win

Hacker News front page

Show HN: AI Subroutines – Run automation scripts inside your browser tab

rtrvr.ai introduced AI Subroutines, which turn a recorded browser task into a callable tool and replay it at zero token cost and zero LLM inference delay. The script runs inside the active tab, reusing auth, CSRF, TLS sessions, and signed headers; recording trims about 300 requests to about 5 and falls back to DOM-only when GraphQL operation IDs are volatile. The part to watch is batching: one LLM call can assign parameters for a 500-row sheet and launch 500 subroutines.

Why it matters: This clears HKR-H/K/R: the hook is zero-token browser automation, the post gives concrete mechanics (300→5 requests, DOM fallback, 500-row fan-out), and it hits agent reliability/cost pain. Kept to mid-featured because it is a single-company Show HN post, not a market-wide event.

Apr 17Friday

Xinzhiyuan · WeChat

Behind OpenClaw's surge, only 8.6% of users detect anomalies: a multi-university empirical study

NTU, KTH, and William & Mary ran a 303-person study and found only 8.6% noticed agent-mediated deception, while 2.7% identified the mechanism correctly. Using 9 HAT-Lab task scenarios, interactive interruption alerts raised detection to 25%, while static warnings were seen by about 24%. The key issue is human-agent cognitive failure, not just model bugs.

Why it matters: Strong HKR-H/K/R: the 8.6% detection hook is sharp, and the 303-person, 9-task study plus 25% alert lift gives testable detail. This is a solid agent-safety research release, not a market-moving product, model, or policy event, so it lands in featured, not p1.

Xinzhiyuan · WeChat

Yixin says its finance Agent harness runs single tasks for 16 hours and plans an H2 open-source release

Yixin says its finance Agent harness can run a single task for 16 hours across 12 sessions, with 65% autonomous delivery. The post adds a 50k-token cap per case, projected approval speedups above 150%, and projected unit cost at one-fifth of human work; it says an open-source release is planned for H2 2026, but does not disclose the repo, license, or reproducible evals. The key signal is governance design, not the “smarter over time” framing.

Why it matters: This clears HKR-H/K/R with a rare production claim: a finance agent runs 16 hours, spans 12 sessions, hits 65% autonomous delivery, and stays under a 50k-token cap. It stays below 85 because the evidence is self-reported and the post does not disclose a repo, license, or reproduc

Tencent Technology · WeChat

From Vibe Coding to Agentic Engineering: Rebuilding the Full Backend Development Workflow

Tencent engineers report a one-week practice that used Claude Code plus custom Skills, Commands, and MCP servers to run an 11-stage backend workflow in one terminal session. The post gives reproducible details: one requirement-exploration step used 20 tool calls, 93.8k tokens, and 56 seconds; execution was split into 4 tasks and produced 3 commits. The real point is workflow orchestration, not raw code generation; human review remains at plan, deploy, and review gates.

Why it matters: HKR-H/K/R all pass: the story turns agentic engineering into a measured backend workflow test, with tool-call, token, timing, plan-length, task, and commit data. Stronger than generic coding hype, but still a practitioner case study rather than a major product or model release.

X · @dotey

Seedance 2.0 API is now available on Volcano Engine and BytePlus

Volcano Engine has released the Seedance 2.0 API for enterprises, individual developers, and overseas users via BytePlus; China pricing is RMB 46 per million tokens, or about RMB 1 per second for pure video generation. The post says it supports text, image, audio, and video inputs, plus face verification, portrait authorization, and 10,000+ preset avatars for workflow automation; overseas pricing is not disclosed here. The part to watch is orchestration: the post cites up to 10x efficiency gains, but does not disclose a common benchmark or model specs.

Why it matters: HKR-H/K/R all pass: the overseas rollout is a real hook, and the post includes usable pricing and modality details for practitioners. It stays at 74 because this is an API availability update, not a major model launch, and the post does not disclose model params, benchmark method

X · @op7418

Seedance 2.0 API is now fully open

Volcano Engine has opened the Seedance 2.0 API to domestic users, while BytePlus serves overseas access; the API currently accepts 4 input modalities: text, image, audio, and video. The post also confirms face registration, portrait authorization, and preset virtual avatars, but does not disclose pricing, rate limits, model variants, or regional availability. The real watchpoint is whether video-agent workflows can be wired through Skills and MCP, not the ecosystem rhetoric.

Why it matters: This is a real product update from ByteDance’s stack: HKR-H on full API availability, HKR-K on 4-modal input and consent mechanics, and HKR-R on builder demand for deployable video APIs. I keep it at 75 because pricing, rate limits, regional rollout details, and quality evidence

最佳拍档 (BestPartners)

Turn your coworker into a Skill? GitHub viral project and Anthropic Skills explained

The video says the open-source “coworker.skill” project gained over 13,000 GitHub stars in days, but it produces a standardized SKILL.md prompt package, not a digital worker replacement. It gives a timeline: Anthropic launched Claude Skills on Oct 16, 2025, then published Agent Skills as an open standard on Dec 18; the mechanism keeps only a short summary in context until a task matches. The real point is scope: it fits standardized workflows like reports, docs, and code review, while the post does not disclose cross-platform compatibility rates or any settled legal standard.

Why it matters: This clears HKR-H/K/R: the coworker-to-Skill hook is sticky, the post adds dates/stars/mechanism, and the labor/IP angle resonates. I kept it at 76 because it is secondary commentary, not a primary release or first-hand test, and key compatibility/legal facts are still undiscolse

X · @dotey

Boris Cherny shares practical tips from recent heavy use of Claude Opus 4.7

Boris Cherny outlined five ways to use Claude Opus 4.7, centered on Auto mode approving safe commands and a /go skill chaining tests, code simplification, and PR creation. The post names Auto mode, Recaps, Focus mode, effort level, and computer use; pricing, launch date, and benchmark data are not disclosed. The real shift is workflow, not just the model itself.

r/LocalLLaMA

PSA: Qwen3.6 ships with preserve_thinking. Make sure you have it on.

Qwen3.6 adds a preserve_thinking flag to keep prior reasoning in context and address the KV cache invalidation issue seen with the Qwen3.5 template. The post cites the Qwen3.6-35B-A3B model page and gives a two-turn 20-digit-number test: with preserve_thinking on, the model can return the second number from its earlier reasoning. The practical point is cross-turn reasoning retention for agent and tool workflows; LM Studio does not support it yet, and an oMLX PR is open.

Why it matters: HKR-H, K, and R all pass: the story has a strong hidden-setting hook, a concrete two-turn repro, and a clear nerve for local-model and agent users. I keep it in the low 70s because this is a Reddit PSA rather than a primary release note, and the impact is concentrated in Qwen/OSS

TechCrunch · AI

OpenAI upgrades Codex with more control over your desktop

OpenAI upgraded Codex on April 16, 2026, expanding its desktop control, and the headline frames it as a move against Anthropic. The truncated post only confirms more desktop power for Codex and says Claude Code has become a preferred tool for many businesses; the post does not disclose exact features, pricing, rollout, or permission limits. The key issue is the permission boundary, not the coding-tool label.

Why it matters: TechCrunch reports an OpenAI Codex desktop-control upgrade framed as a direct move against Anthropic, so HKR-H and HKR-R land. But HKR-K is limited: the article confirms broader permissions only, with no action list, pricing, or rollout details, so it stays at the featured floor.

X · @dotey

Codex major update: from a coding tool to an assistant that can operate your computer

OpenAI upgraded Codex into a Mac desktop agent and says it serves 3M+ weekly developers. It can see screens, click, type, run parallel agents, and adds 90+ plugins. Desktop rollout starts now for ChatGPT sign-ins; computer control is macOS-first.

Why it matters: OpenAI expands Codex from coding help into a Mac-operating agent, with parallel agents, 90+ plugins, and a claimed 3M weekly developer base. HKR-H/K/R all pass; missing safety boundary and pricing details keep it in the high-80s, not the 90s.

X · @OpenAI

Codex for (almost) everything.

OpenAI said Codex can now use apps on Mac, connect to more tools, and handle ongoing and repeatable tasks. The post also claims image creation, learning from prior actions, and remembering user preferences; it does not disclose app coverage, integration method, pricing, or rollout timing.

Why it matters: This is an official OpenAI product update, and Codex moves from coding help toward desktop control, tool use, and memory, so HKR-H/K/R all pass. The post still omits supported apps, integration method, pricing, and launch timing, keeping it in the 78–84 band.