Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

1181–1200 of 1,465

Apr 28Tuesday

QbitAI · WeChat

NTU REI-Bench Tests Vague Human Instructions, With Success Rates Dropping Up to 36.9%

NTU MARS Lab released REI-Bench, a benchmark with 9 ambiguity levels for vague human instructions. Tests used 4 robot planning frameworks and 6 small LLMs; LLaMA3.1-8B+SayCan fell from 57.7% to 46.9% in standard multi-turn context. The key issue is implicit reference resolution, where baseline success dropped 7.4% to 36.9%.

Why it matters: HKR-H/K/R all pass: the 36.9% drop is a strong hook, and the setup gives 9 ambiguity levels, 4 frameworks, and 6 models. This is a solid embodied-AI benchmark, not a major model release, so it fits the 78–84 band.

QbitAI · WeChat

Xiaomi open-sources MiMo-V2.5 series; Pro builds a macOS-like desktop in 4 hours

Xiaomi open-sourced MiMo-V2.5 weights, covering Pro Agent, multimodal base, TTS, and ASR models. MiMo-V2.5-Pro built a 54-app macOS-like desktop in 4 hours without human takeover; it scored 233/233 on SysY with 672 tool calls in 4.3 hours. Key details for practitioners are the 1M context, 100T-token program, and free Agent-framework access.

Why it matters: HKR-H/K/R all pass: Xiaomi open-sourced MiMo-V2.5 weights with concrete agent and coding-task numbers. Domestic flagship model release bump puts it in the must-write same-day band.

QbitAI · WeChat

Open-source SenseNova-U1 unifies image understanding and generation

SenseTime open-sourced two SenseNova-U1 models: an 8B version and a 38B-total MoE version using NEO-unify. The architecture removes VE and VAE, processes pixels directly, and generates 2048×2048 images in about 9 seconds on one H100/H200 node. The key item is interleaved text-image reasoning; 32K context, long-text rendering, and beta interleaved creation remain limits.

Why it matters: HKR-H/K/R all pass: the architecture hook is concrete, the post gives model sizes and latency, and open multimodal work matters to builders. It stays in 78–84 because it is not a top-tier general-model launch.

Hacker News front page

Xiaomi releases MiMo-v2.5 weights with strong coding and agent benchmarks

Xiaomi released MiMo-v2.5 family weights; the title cites strong coding and agent benchmarks. The RSS body only lists URLs, 13 HN points and 2 comments; the post does not disclose size, license, or scores.

Why it matters: HKR-H/K/R pass because a Xiaomi coding/agent weights release is concrete and practitioner-relevant. Sparse sourcing holds it near the featured floor: no parameters, license, or benchmark numbers are disclosed.

The Verge · AI

Attack of the Killer Script Kiddies

The Verge discusses Claude Mythos and AI bug finding, citing DARPA AIxCC scans over 54 million code lines. Teams found most seeded flaws plus over a dozen unseeded bugs; the RSS snippet does not disclose Mythos benchmarks, pricing, or access terms.

Why it matters: HKR-H/K/R all pass: the hook is strong, DARPA AIxCC supplies concrete numbers, and the security angle resonates. No Claude Mythos benchmark, pricing, or access terms are disclosed, so it stays in the featured-threshold band.

Hacker News front page

GitHub Copilot code review will start consuming GitHub Actions minutes

GitHub will make Copilot code reviews consume GitHub Actions minutes starting June 1, 2026. Private-repo reviews use plan entitlements, with overages billed at standard Actions rates; public repos stay free. The change covers Copilot Pro, Pro+, Business, and Enterprise, including direct org billing for unlicensed users.

Why it matters: Official GitHub billing change for Copilot code review hits CI quotas and org invoices; HKR-H/K/R all pass, but it is a pricing rule, not a capability release, so it sits low in 72–77.

Xinzhiyuan · WeChat

NUS and NTU Release Pask with Streaming Intent Detection and Persistent Memory

NUS and NTU released Pask, with paper arXiv:2604.08000. Pask uses DD, MM, and PAS, with IntentFlow detecting intent in 1.5 seconds. The key bet is real-time intent detection, not longer execution chains.

Why it matters: HKR-H/K/R all pass: Pask offers a concrete real-time intent layer for proactive agents. No open-source status, benchmark table, or production deployment is disclosed, so it stays at 78 rather than P1.

Xinzhiyuan · WeChat

Claude bans hit 110-person firm; Cursor incident deletes database in 9 seconds

Anthropic allegedly suspended 110 Claude accounts at a US agtech firm, while API billing continued. The post says appeals went unanswered for 36 hours, and PocketOS says Claude Opus 4.6 via Cursor deleted production data and volume backups in 9 seconds. The key issue is access control: no RBAC, no environment isolation, and no delete confirmation.

Why it matters: HKR-H/K/R all pass: the incident has a strong hook and concrete details: 110 accounts, 36 hours, 9 seconds, and no RBAC. Kept at 82 because it is still a single-source allegation without an Anthropic postmortem.

X · @op7418

Xiaomi open-sources the MiMo-V2.5 model series

Xiaomi open-sourced the MiMo-V2.5 model series under the MIT license for commercial use, retraining, and fine-tuning. It also launched Orbit 100T Token, offering approved AI builders up to 1.6B credits worth 659 yuan. Agent framework teams can apply for free MiMo token access; the post does not disclose model size or benchmark results.

Why it matters: HKR-H/K/R all pass: Xiaomi MiMo-V2.5 open source, MIT terms, and Orbit 100T credits matter to builders. Missing params and benchmarks keep it in the 78–84 band, below P1.

r/LocalLLaMA

Local coding models have reached a threshold for real work

Antigma tested 27B–32B open-weight models; Qwen 3.6-27B scored 38.2% on Terminal-Bench 2.0. The run used 89 tasks and the default per-task timeout, while verified SOTA is about 80%. The key claim is deployment lag: offline coding is about 6–8 months behind hosted frontier models.

Why it matters: HKR-H/K/R all pass: the post gives a real-work threshold claim, a 38.2%/89-task Terminal-Bench result, and a 6–8 month offline gap. Reddit single-post sourcing keeps it in the low featured band.

Computing Life · Share · Yage

Agentic Creative Tools: From Photoshop Actions to Claude for Creative Work

Anthropic released 9 creative-tool Connectors for Claude for Creative Work. The post frames agentic creative tools around programmable APIs, connector protocols, and perceptual feedback loops. The post does not disclose the Connector list.

Why it matters: HKR-H/K/R all pass: Claude creative agents have a clear hook, 9 connectors add a fact, and creator workflow pressure adds resonance. Missing connector names and access terms keep it below must-write.

X · @dotey

GitHub Copilot switches to usage-based billing on June 1

GitHub Copilot will switch to AI Credits billing on June 1 while keeping subscription prices unchanged. Credits count input, output, and cached tokens; Pro includes $10 monthly credits and Pro+ includes $39. Watch Copilot Agent long-task costs.

Why it matters: HKR-H/K/R all pass: Copilot billing moves from subscription expectations to token/cache consumption with date and credit amounts. Single-source X context lacks enterprise details and overage rates, so it stays in the 78–84 band.

X · @dotey

Cursor 3 feedback: users want a reliable AI development workspace

Eric Zakariasson’s Cursor 3 feedback thread summarizes 431 replies, with users asking for a stable AI development workspace. Requests center on Agent Window retaining LSP, debugging, Git, terminal and diff workflows, plus multi-agent worktrees and model-cost transparency. The key issue is workflow reliability, not a flashier IDE.

Why it matters: All HKR axes pass: 431 user replies, concrete workflow requests, and strong resonance for Cursor users. Kept in the low featured band because this is feedback synthesis, not an official Cursor release or roadmap.

Apr 27Monday

Dwarkesh Patel podcast

What I've been Thinking About This Weekend: Open Questions, Intelligence vs Power, Verification in Science

Dwarkesh lists open AI questions, including that five hyperscalers own over 70% of global AI compute. He asks about coding agents, KV cache costs, merging training with inference, and online learning; the post gives questions, not experimental answers.

Why it matters: HKR-H/K/R all pass: Dwarkesh adds a concrete compute-concentration claim and practitioner-relevant questions. No experiment, release, or policy change, so it stays in the 72–77 commentary band.

TechCrunch · AI

China blocks Meta’s $2B Manus deal after months-long probe

China ordered Meta to unwind its $2B Manus acquisition. The title says the probe lasted months; the post does not disclose the legal mechanism, timeline, or Meta’s next steps. The deal risk directly hits Meta’s AI agents push.

Why it matters: HKR-H/K/R all pass: a $2B Meta-Manus AI-agent deal was blocked by China after a months-long probe. The article lacks the legal mechanism and Meta’s next step, so it lands at 88, not 90+.

Hacker News front page

Show HN: OSS Agent Dirac topped TerminalBench on Gemini-3-flash-preview

Dirac-run released Dirac and says it topped TerminalBench using Gemini-3-flash-preview. The repo claims 50-80% lower API costs via Hash Anchored edits, parallel operations, and AST manipulation; the post does not disclose full scores.

Why it matters: HKR-H/K/R all pass: an OSS coding agent claims a TerminalBench lead with cost and mechanism details. Held to 78 because the post relies on repo claims and lacks full leaderboard scores or reproduction logs.

Mistral AI

Mistral AI opens public preview of Workflows

Mistral AI has put Workflows, its enterprise AI orchestration layer, into public preview. It offers durable execution, observability and human-in-the-loop approvals. ASML, ABANCA and CMA-CGM are already using it to automate critical processes.

Why it matters: It lays out Workflows' orchestration features, deployment model and customer cases, showing the engineering bar for enterprise AI processes.

Xinzhiyuan · WeChat

First Spatio-Temporal Time-Series Reasoning Framework for LLMs | ACL'26

Emory University, Microsoft, and partners introduced STReasoner for spatio-temporal time-series reasoning, with ST-Bench covering four task types. It uses Network SDE plus Multi-Agent data generation, then Align, SFT+CoT, and S-GRPO training. The article claims inference cost is 0.004× closed models, with code on GitHub.

Why it matters: HKR-H and HKR-K pass: the story has a “first framework” hook plus ST-Bench, S-GRPO, 0.004× cost, and code release. HKR-R is weak because spatiotemporal reasoning is a narrower research lane.

QbitAI · WeChat

Stanford-led LLM-as-a-Verifier claims SOTA on Terminal-Bench 2.0

Stanford, Berkeley and Nvidia introduced LLM-as-a-Verifier, claiming SOTA on Terminal-Bench 2.0 and SWE-Bench Verified. It selects trajectories via score-token granularity, repeated checks and criteria decomposition; ForgeCode accuracy reached 86.4%.

Why it matters: HKR-H/K/R all pass: Stanford, Berkeley, and NVIDIA offer a concrete verifier mechanism and benchmark numbers. It is still a benchmark research release, not a major model or product launch, so it fits the 78–84 band.

QbitAI · WeChat

DeepSeek V4 Cuts Prices Permanently; Cached Inputs Get 90% Off, Coding Test Costs Drop 83%

DeepSeek V4 cut prices twice in two days: input/output pricing is 75% lower, with cached inputs getting another 90% off. QbitAI’s coding test fell from 31.73 yuan for 35M tokens to 5.34 yuan under new pricing, an 83% drop. The key case is high cache-hit workloads, with V4-Pro at about 95–96% cache hits.

Why it matters: HKR-H/K/R all pass: DeepSeek V4 pricing has a sharp cost hook, concrete test numbers, and strong cost resonance. It is still a pricing update, not a new model release, so it stays below the 85 P1 band.