Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

681–700 of 1,465

Jun 2Tuesday

AI HOT (Curated Pool)

Anthropic Developer Shares a Claude Code Understanding-Verification Workflow

An Anthropic developer shared a Claude Code understanding-verification workflow with 8 steps, using incremental teaching, user restatement, checklists, and quizzes to confirm the human can defend the problem, solution, and impact before moving to the next stage.

Why it matters: HKR-H/K/R all pass: a concrete Claude Code workflow with an 8-step verification loop and a strong oversight hook. It is a practical tutorial, not a product release, so it sits at the lower featured band.

Computing Life · Share · Yage

AI agents don't need to be hacked; persuasion is enough

The article says an AI agent with password-reset permission can be abused when an attacker persuades it they are a legitimate user; the snippet only discloses a three-layer architecture that separates what from who, not concrete attack steps or implementation details.

Why it matters: HKR-H/K/R all pass: the hook is strong, the post offers a what/who three-layer design, and agent permissions are a live security worry. No real incident, success rate, or product comparison keeps it at the featured threshold.

Computing Life · Share · Yage

The Next Form of AI Agents: From Chat Windows to Background Daemons

Gemini Spark is described as the first consumer-facing always-on background agent from a major platform; the post covers four product generations and a periodic versus reactive automation framework.

Why it matters: HKR-H/K/R all pass, but this is a single commentary item; the body summary does not disclose launch date, rollout scope, or hands-on results for Gemini Spark. Score stays at the lower featured band.

r/LocalLLaMA

I spent months inside verl, forked it, then stopped: internals, fork costs, and an NCCL bug

ReinforcedKnowledge analyzes ByteDance’s verl RLHF loop, covering DataProto plus rollout, reward, advantage, and update paths. The author stopped a private fork because near-daily upstream changes made sync cost exceed refactoring work, and describes an NCCL hang fixed on one node by setting NCCL_SOCKET_IFNAME=lo.

Why it matters: Niche but useful RL post-training field report, not an industry release. HKR-H comes from the fork-then-quit twist; HKR-K has verl’s five paths and NCCL_SOCKET_IFNAME=lo; HKR-R hits the cost of maintaining open-source training forks.

AI HOT (Curated Pool)

Google AI Studio adds app-building support for Gmail and other apps

Google AI Studio has added app-building support for connected Gmail, Drive, and Sheets apps, and users can add testers inside AI Studio; the post does not disclose a launch date for full public sharing.

Why it matters: HKR-H/K/R all pass, but this is a mid-weight product update: Workspace connections and tester support are confirmed, while sharing, permission details, and pricing are not disclosed.

TechCrunch · AI

Nvidia chases $200B CPU market with AI agent PCs from Microsoft, Dell, and HP

The title says Nvidia is targeting the $200B CPU market with AI agent PCs from Microsoft, Dell, and HP; the RSS snippet does not disclose specifications, pricing, launch timing, or the safety mechanism for bringing agents to consumer PCs.

Why it matters: HKR-H/K/R all pass, but specs, price, and launch timing are not disclosed. Treat it as a mid-weight product/ecosystem update, with Nvidia plus Microsoft/Dell/HP enough for low featured.

The Verge · AI

Gemini’s New AI Agent Is About as Good as Google’s Demo

The Verge tested Google Gemini Spark for one week and says the 24/7 agent can run multi-step tasks in the background, but the RSS snippet does not disclose pricing, privacy terms, or the full hands-on results.

Why it matters: HKR-H/K/R pass: a Verge hands-on stress-tests Google’s Gemini Spark demo claim and confirms background multi-step tasks. Missing price, privacy terms, and full results keep it in the 72–77 band.

AI HOT (Curated Pool)

Meta AI Exploit Used to Hijack Instagram Accounts

Meta’s AI chatbot was found vulnerable to an account-takeover exploit against Instagram accounts. Attackers could ask the AI to link a new email address, and the failure condition was the agent’s ability to execute account-management actions directly; the RSS snippet does not disclose affected account counts, patch status, or reproduction details.

Why it matters: HKR-H/K/R all pass: a Meta AI support agent allegedly enabled Instagram account takeover via add-email requests. Impact scale, fix timeline, and reproducible steps are not disclosed, so it stays in the 78–84 band.

Hacker News front page

Hackers Used Meta's AI Support Bot to Seize Instagram Accounts

The title says hackers used Meta's AI support bot to seize Instagram accounts; the RSS snippet lists 40 points and 14 comments, but the post does not disclose the attack mechanism.

Why it matters: HKR-H and HKR-R pass: a Meta AI support bot allegedly enabled Instagram account takeovers, a Krebs-sourced security angle. HKR-K fails because the feed lacks mechanism or scale, so it sits at the featured floor.

AI HOT (Curated Pool)

Perplexity Releases Search as Code Architecture

Perplexity released Search as Code, an architecture where agents write Python code to call its search stack directly instead of looping through function calls; it is now available in the Perplexity Agent API and is the default option for Computer.

Why it matters: HKR-H/K/R pass: Perplexity gives a concrete agent-search mechanism and Agent API integration. Single-source post lacks performance, pricing, and rollout scope, so this stays a low featured product update.

Jun 1Monday

Latent Space

Why Video Agent Models Are Next — Ethan He on xAI Grok Imagine

Ethan He says a small xAI team built Grok Imagine from zero to one in 3 months, and the episode discusses video agents, audio-video alignment, inference speedups, and the storage, egress, and GPU-hour costs behind large video datasets.

Why it matters: HKR-H/K/R all pass, but the body is interview-level signal: beyond the 3-month build and mechanism themes, it gives no benchmarks, cost figures, or reproducible test. Strong xAI video-agent context, not same-day must-write.

AI HOT (Curated Pool)

Open and Closed Models Are on Different Exponentials

Nathan Lambert argues that closed frontier labs will capture high-margin demand in coding-agent workflows, citing a personal willingness to pay $2,000 per month and projecting OpenAI and Anthropic valuations of $2-10 trillion over 5-10 years.

Why it matters: HKR-H/K/R all pass: the essay has a clear open-vs-closed hook, concrete price and valuation claims, and practitioner resonance. It remains single-source commentary, so it sits in the featured-threshold band.

AI HOT (Curated Pool)

Wang Xing: Meituan AI Agent Xiao Mei to Partner Deeply with Tencent Yuanbao

Meituan CEO Wang Xing said Xiao Mei will connect with Tencent Yuanbao, routing local service requests into food ordering, delivery, and related Meituan scenarios; Meituan reported Q1 2026 revenue of RMB 91.039 billion and a net loss of RMB 6.827 billion.

Why it matters: HKR-H/K/R all pass, but the deal is still “coming soon”; launch timing, UX entry point, and revenue split are not disclosed. This fits a mid-weight product partnership at the featured floor.

AI HOT (Curated Pool)

Tutorial: Turning Books into AI Skills with Claude Opus 4.8

The author used Claude Opus 4.8 to turn Nonviolent Communication into an AI Skill through a six-step workflow, taking about 45 minutes, using roughly 300,000 tokens, and costing under RMB 20.

Why it matters: HKR-H/K/R all pass: this is a numbered first-person Claude workflow with concrete cost and token details. It stays in the lower featured band because it is a personal tutorial, not an Anthropic release or model update.

AI HOT (Curated Pool)

Apache RocketMQ Releases an AI-Focused Messaging Engine

Apache RocketMQ released RocketMQ for AI, a messaging engine for long-running sessions, multi-agent workflows, and fair scheduling, with Lite-Topics, ordered messages, and traffic shaping; the post does not disclose a version number or performance figures.

Why it matters: HKR-H/K/R pass: the AI-specific RocketMQ angle has a real agent-infra hook and named mechanisms. Score stays in the 72–77 band because version, benchmarks, and production cases are not disclosed.

Xinzhiyuan · WeChat

400 tokens/s: StepFun Step 3.7 Flash cuts Agent task costs

StepFun released Step 3.7 Flash, a sparse MoE model with 196B parameters plus a 1.8B ViT, activating 11B parameters per inference and reaching up to 400 tokens per second.

Why it matters: HKR-H/K/R all pass with concrete speed and parameter numbers. The feed does not disclose pricing, benchmark setup, or open-source terms, so this stays in the 78–84 quality update band.

Synced · WeChat

OpenAI recruits for robotics team led by Sora creator Aditya Ramesh

OpenAI has listed more than a dozen San Francisco robotics roles for OpenAI Robotics, a team that evolved from Aditya Ramesh’s Worldsim work, with the actuator design engineer role offering $342,000 to $445,000 in base cash pay plus PPU incentives.

Why it matters: HKR-H/K/R all pass: OpenAI robotics hiring adds a strong hook, plus concrete roles, leader, and salary range. This is still a hiring signal, not a model or product release, so it stays in the featured-threshold band.

Synced · WeChat

World models get a “save state”: VAST releases Project Eden

VAST released Project Eden, a three-layer world-model architecture that separates persistent state evolution from visual rendering, and disclosed nearly $200 million across its A+ and A++ funding rounds.

Why it matters: HKR-H/K/R all pass: Project Eden has a product hook, architecture detail, and funding scale. VAST is not a top foundation-model lab, and benchmarks or access terms are not disclosed, so this lands in 78–84.

AI HOT (Curated Pool)

Tencent Hunyuan Releases Long-Term Memory Plugin Hy-Memory

Tencent Hunyuan released Hy-Memory for long-term collaborative agents such as OpenClaw, using a six-layer memory framework and System1/System2 dual system, with memory count reduced by over 70% and token consumption down 35% in ultra-long-context scenarios.

Why it matters: Tencent Hunyuan’s Hy-Memory clears HKR-H/K/R with a concrete memory architecture and cost-reduction figures. The score stays at the featured floor because the source is an official short post without reproducible tests, license details, or third-party benchmarks.

QbitAI · WeChat

VAST Raises Nearly $200M and Discloses Its Project Eden World Model Roadmap

VAST raised nearly $200 million in A+ and A++ rounds and disclosed Project Eden, a world model architecture that separates state evolution from visual rendering through a structured state layer, a conditional interface layer, and a generative rendering layer.

Why it matters: HKR-H/K/R all pass: the $200M A+/A++ financing is sizable, and Project Eden gives a concrete three-layer world-model mechanism. VAST is not a top-tier foundation-model lab and no metrics or release details are disclosed, so this stays in the 78–84 band.