Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

721–740 of 1,465

May 30Saturday

AI HOT (Curated Pool)

Codex now supports computer use on Windows

OpenAI added Windows computer-use support for Codex, letting users start, review, and guide tasks on a Windows PC through the ChatGPT mobile app; the post states this is an early experience and does not disclose pricing or rollout scope.

Why it matters: HKR-H/K/R all pass: OpenAI adds Windows computer use to Codex, controlled through ChatGPT mobile. The post gives the workflow and early-stage condition, but not permissions, pricing, or rollout scope, so this stays at the featured threshold.

AI HOT (Curated Pool)

What Happens When Companies Become Too AI-Pilled?

Aaron Levie says leaders replacing employees with AI often understand the work least; he calls it “AI psychosis.” ClickUp cut 22% of staff for AI agent deployment, and 2026 tech layoffs are already near the full-year 2025 total.

Why it matters: HKR-H/K/R all pass: the “AI-pilled” framing has bite, the story adds ClickUp’s 22% layoff figure and 2026 layoff context, and it hits the jobs nerve. Not a model, product, or policy event, so it stays near the featured floor.

AI HOT (Curated Pool)

xAI Releases Grok Build 0.1 Public Beta

xAI released grok-build-0.1 as a public beta through its API; the same model powers the Grok Build CLI, targets agentic coding, and is priced at $1 per million input tokens and $2 per million output tokens.

Why it matters: HKR-H/K/R all pass, but the post is thin: beta, CLI, pricing, and agent-coding positioning only; no benchmarks, context window, or hands-on results. This fits a mid-weight product update.

May 29Friday

The Verge · AI

Adobe’s Conversational AI Agent Is a Mediocre Design Intern

The Verge tested Adobe Firefly AI Assistant in beta. It can operate Adobe design apps as a conversational middleman, rather than only generating images or video. The post says it explains edit steps clearly, but the results were not impressive. The RSS snippet does not disclose pricing, release timing, or the full list of supported apps.

Why it matters: HKR-H/K/R all pass because this is a Verge hands-on of Adobe’s Firefly AI Assistant beta with a clear negative usability hook. Missing pricing, launch timing, and supported-app details keep it in the 72–77 featured-threshold band.

QbitAI · WeChat

Tencent unveils Code Craft, an AI game creation platform for beginners and developers

Tencent Games unveiled Code Craft, an AI game creation platform that turns natural-language prompts into runnable 2D or 3D games, with a planning knowledge base, Skill system, visual tuning panels, and more than 20,000 free cloud assets; the post does not disclose release timing, pricing, model details, or supported engines.

Why it matters: HKR-H/K/R pass on the Tencent game-creation hook, runnable 2D/3D output, and 20,000+ assets. Pricing, access scope, and model limits are not disclosed, so it stays in the lower featured band.

AI HOT (Curated Pool)

Google DeepMind CEO Demis Hassabis Says AGI Could Arrive Within Three Years

Demis Hassabis predicts AGI could arrive around 2029 to 2030, with mature multimodal capabilities and autonomous decision-making as key conditions, while warning that society remains underprepared and needs rules and safeguards before deployment.

Why it matters: HKR-H/K/R all pass: Hassabis gives a 2029-2030 AGI window and names multimodal plus autonomous decision-making as conditions. High-interest commentary, but thinner than a model release or major product update.

Hacker News front page

Undisclosed Addition in jqwik Instructed AI Coding Agents to Delete App Output

The title says an undisclosed jqwik addition instructed AI coding agents to delete app output; the RSS body only lists the URL, 24 points, and 16 comments, and does not disclose the code location or impact scope.

Why it matters: HKR-H/K/R all pass: the hook is sharp, the mechanism is concrete, and AI-coding safety resonates. Sparse body detail keeps it near the featured threshold: no code location, affected versions, or impact scope disclosed.

Xinzhiyuan · WeChat

Three DeepSeek Models Enter OpenRouter Monthly Top 10 With Over 17 Trillion Tokens

DeepSeek placed three models in OpenRouter’s monthly top 10 with more than 17 trillion tokens combined, including V4 Flash at 9.13T tokens; the article says Ascend’s MegaMoE operator raised Prefill throughput by 20% to 30% on DeepSeek V3.1 and Qwen3-235B tests.

Why it matters: HKR-H/K/R all pass: the story has a 17T-token hook plus concrete OpenRouter and MegaMoE Prefill numbers. It stays at 82 because the compute-sovereignty framing is strong, while reproducible test conditions are not disclosed.

Xinzhiyuan · WeChat

Claude Opus 4.8 tests split users: strong at high effort, costly under rate limits

The article says Claude Opus 4.8 scores 63 on an Extra-High senior engineering benchmark, 30 points above Opus 4.7, but drops to 42 at High effort, while $200/month Max users report hitting rate limits within hours on complex agent tasks.

Why it matters: Anthropic/Claude relevance plus concrete test numbers clears HKR-H/K/R: the hook is strength versus cost, K has benchmark and quota details, and R hits agent-budget anxiety. Source is a media test rather than an official release, so this lands at low P1.

r/LocalLLaMA

Liquid AI releases LFM2.5-8B-A1B

Liquid AI released LFM2.5-8B-A1B with a 128K context window, 38T pre-training tokens, large-scale reinforcement learning, doubled vocabulary for non-Latin tokenization, and availability on Hugging Face.

Why it matters: HKR-H/K/R pass: 8B/A1B, 128K context, and 38T tokens are concrete hooks for local inference. No benchmarks, license, or deployment limits are disclosed, so it stays in the mid featured band.

Synced · WeChat

Meta Uses 183B Tokens to Turn Math Textbooks into a Large Lean Library

Meta released ATLAS, a Lean 4 formalization library covering 26 math textbooks and 46,203 declarations, using 183.157 billion tokens to generate 630,999 lines of code, with 42,837 completed proofs and a 92.7% proof pass rate.

Why it matters: HKR-H/K/R all pass: the token scale, Lean corpus size, and verified-proof count are concrete. It stays below P1 because this is a specialized research/open-source release, not a broad model or product launch.

Latent Space

Anthropic raises $65B Series H, releases Opus 4.8 and Dynamic Workflows

Anthropic announced a $65B Series H at a $965B post-money valuation, disclosed a $47B revenue run rate, and released Claude Opus 4.8 plus Claude Code Dynamic Workflows as a research preview for parallel subagent orchestration.

Why it matters: HKR-H/K/R all pass: this combines a frontier-lab financing event with an Anthropic model and Claude Code workflow release. I score using the summary’s $65B raise and $965B post-money valuation because the title’s dollar figure conflicts with it.

AI HOT (Curated Pool)

Cursor team releases Developer Habits Report

Cursor’s report says developers’ weekly code output rose from about 3.6K to 8.6K lines, while AI agents increased tool calls per session by roughly 30%.

Why it matters: HKR-H/K/R all pass: Cursor’s own report gives concrete 3.6K→8.6K and +30% figures for AI coding work. It is not a product launch or cross-source event, so 78–84 fits better than the must-write band.

r/LocalLLaMA

StepFun 3.7 Flash

StepFun released Step 3.7 Flash with 196B total parameters, 11B active MoE, a built-in 1.8B ViT, and local execution on 128GB RAM.

Why it matters: HKR-H/K/R pass via the 196B/11B MoE specs and 128GB local-run claim. Sparse Reddit sourcing leaves license, eval method, and access conditions undisclosed, so it stays in the lower featured band.

Ruan YiFeng's Weblog

Technology Enthusiasts Weekly Issue 398: Token Costs Are Hard to Afford

Peter Steinberger posted one month of usage showing 7.6 million requests and 603 billion tokens, with CodexBar estimating a $1.3 million value under preset rates rather than his actual spend as an OpenAI employee.

Why it matters: HKR-H/K/R all pass: the CodexBar case turns token economics into concrete usage and cost. This is strong practitioner commentary, not a model or platform release, so it fits the 72–77 featured band.

AI HOT (Curated Pool)

StepFun Releases Step 3.7 Flash, Focused on Agent Efficiency

StepFun released the open-source Step 3.7 Flash model with a 198B-parameter MoE architecture, about 11B active parameters, a 256K context window, and a 67.1 score on ClawEval-1.1.

Why it matters: HKR-H/K/R all pass: the release has a clear sparse-model hook, concrete context and benchmark numbers, and practitioner resonance around open agent efficiency. Official-post sourcing and no independent eval keep it in the 78–84 band.

Computing Life · Share · Yage

Claude Code Dynamic Workflow: Where Is the Determinism Boundary Drawn?

The article analyzes Anthropic’s dynamic workflow across three boundaries: code handles control flow, agents handle execution, and multiple agents cross-check validation.

Why it matters: HKR-H/K/R all pass: the piece has a clear Claude Code reliability hook and a concrete workflow mechanism. It stays in the 72–77 band because it is commentary, not an Anthropic release, and no experiment numbers are disclosed.

AI HOT (Curated Pool)

Skill distillation

Skill distillation has Opus 4.7, GPT-5.1, and Gemini 3 Pro write standardized SKILL.md procedure files, while local Qwen 35B and Gemma 26B models execute those files step by step.

Why it matters: HKR-H/K/R pass: the agent-skill distillation pattern is concrete and practitioner-relevant. The summary lacks success rates, cost data, or task outcomes, so it sits at the featured threshold, not must-write.

Latent Space

The Age of Async Agents — Cognition's Walden Yan and OpenInspect's Cole Murray

Latent Space discusses async coding agents with Cognition’s Walden Yan and OpenInspect’s Cole Murray, citing Devin’s 7x merged PR growth and an increase from 16% to 80% of commits across Cognition repos.

Why it matters: HKR-H/K/R all pass: the Cognition repo numbers make this more than agent rhetoric. It stays in the 78 band because it is an interview/trend piece, not a major model or product release.

TechCrunch · AI

Anthropic releases Opus 4.8 with new Dynamic Workflows tool

Anthropic released Opus 4.8 with a Dynamic Workflows tool for coordinating swarms of subagents. The RSS snippet does not disclose pricing, context window size, benchmarks, or a rollout schedule.

Why it matters: HKR-H/K/R all pass: an Anthropic model release plus an agent orchestration tool fits the 85–94 same-day band. Missing price, context window, and rollout detail keep it below the top of the band.