Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

921–940 of 1,465

May 17Sunday

Google DeepMind

Google DeepMind launches Gemini for Science toolset

Google DeepMind released Gemini for Science, which includes three experimental tools on Google Labs: Hypothesis Generation, built on Co-Scientist.

Why it matters: Google is packaging research prototypes like Co-Scientist and AlphaEvolve into apply-to-use science tools, showing what agentic research looks like in practice.

AI HOT (Curated Pool)

Microsoft AI CEO predicts AI will automate all white-collar jobs within 18 months

Mustafa Suleyman predicts AI will reach human-level performance within 18 months and automate most professional tasks, including accounting, law, marketing, and project management.

Why it matters: HKR-H and HKR-R are strong, and HKR-K passes on the testable 18-month timeline. The score stays in the low 78–84 band because this is a CEO forecast, not evidence, benchmarks, or a shipped capability.

r/LocalLLaMA

DeepSeek V4's 1M Context Window: The Breaking Point

A Reddit user tested DeepSeek V4 on 45k, 180k, and 520k-token codebases and found 150k-250k tokens best for coding work. Past 300k tokens, line-number precision degraded; at 520k, outputs shifted toward architecture summaries and skipped implementation details.

Why it matters: A single Reddit post limits authority, but HKR-H/K/R all pass: it is a numbered first-person test with a concrete long-context failure pattern. The right band is featured, not 78+, because replication and model details are thin.

Synced · WeChat

AI agents may spend 1,000x more tokens without better results: the hidden bill

Researchers used OpenHands to analyze traces from 8 frontier models on 500 swe-bench-verified tasks, finding that agentic coding reached a 154:1 input-output token ratio and that human difficulty labels correlated weakly with token use at Kendall tau 0.32.

Why it matters: All HKR axes pass: strong cost-performance hook, concrete benchmark setup and correlation numbers, and direct resonance with coding-agent economics. It is not a model or platform launch, so it fits the 78–84 quality-recommendation band.

Synced · WeChat

What Are World Models? Their History and the $10 Billion Bet

Jiqizhixin translated a MoE Capital blog tracing two world-model lineages. The article says more than $10 billion entered the category over 18 months, and cites DreamDojo as using 44,711 hours of first-person video pretraining to reach r=0.995 correlation with real-world robot policy outcomes.

Why it matters: HKR-H/K/R all pass: the hook is strong and the article gives concrete figures, but it is a compiled explainer rather than a new release. It fits the featured-threshold band for a strong commentary/tutorial.

Synced · WeChat

Peter Steinberger Says His Monthly Token Bill Hit $1.3M, Covered by OpenAI

Peter Steinberger used 603 billion tokens across 7.6 million requests in 30 days, with the bill exceeding $1.3 million; he said disabling fast mode cut the price by 70%, and OpenAI does not charge him for the tokens.

Why it matters: HKR-H/K/R all pass: the story has a sharp cost hook, concrete usage numbers, and strong practitioner resonance. It is a first-person bill disclosure, not an OpenAI pricing or product launch, so it sits just above the featured threshold.

AI HOT (Curated Pool)

MagicPath Integrates with Codex to Combine Design and Development

MagicPath AI CEO @skirano demonstrated MagicPath running inside Codex as a native canvas, with users configuring it through one command, dragging UI elements, and letting Codex generate and edit code in real time.

Why it matters: HKR-H/K/R pass: MagicPath puts a draggable design canvas inside Codex with one-command setup and live code edits. Single-demo sourcing and missing framework support, permissions, and reproducible cases keep it at the lower featured band.

AI HOT (Curated Pool)

Study on the Cognition–Action Disconnect in Tool-Using Agents

An interpretability paper studies tool-using agents and finds models often recognize when to call a tool but fail to act, with a cognition-to-action mismatch rate of 26%–54%.

Why it matters: HKR-H/K/R all pass: the story has a sharp agent-failure hook, a 26%-54% mismatch rate, and clear relevance to tool-use reliability. Source detail is thin, with paper name, models, and task setup not disclosed.

AI HOT (Curated Pool)

Ring-2.6-1T Open-Sourced and Listed on OpenRouter for Agent Workflows

AntLingAGI open-sourced Ring-2.6-1T and listed it on OpenRouter with a 75% discount through the end of May; the trillion-scale reasoning model targets agent workflows, including planning, tool use, context maintenance, and complex task execution, using Async RL and IcePop training methods.

Why it matters: HKR-H/K/R all pass: a 1T open agent model is clickable, with OpenRouter access, discount, and training methods disclosed. Score stays at 74 because benchmarks, license, and context window are not given.

May 16Saturday

TechCrunch · AI

OpenAI co-founder Greg Brockman takes charge of product strategy

Greg Brockman has officially taken charge of OpenAI’s product strategy, and Wired reports that he described a plan in a staff memo to combine ChatGPT and Codex into one unified experience.

Why it matters: HKR-H/K/R all pass: OpenAI co-founder product control plus a reported ChatGPT-Codex unification matters. No launch date, feature boundary, or rollout plan is disclosed, so this stays below a major product release.

AI HOT (Curated Pool)

Anthropic Founder’s Playbook warns AI can raise startup failure rates

Anthropic published Founder’s Playbook, arguing that AI tools such as Claude Code reduce prototyping cost but increase startup failure risk across the Idea, MVP, Launch, and Scale stages through false validation, confirmation bias, agentic technical debt, and founder decision bottlenecks.

Why it matters: HKR-H/K/R pass: the Anthropic founder playbook has a sharp counterintuitive angle, a four-stage mechanism, and clear founder resonance. It stays near the featured floor because no dataset or reproducible test is disclosed.

AI HOT (Curated Pool)

Researchers use Anthropic Mythos to build a macOS kernel exploit bypassing Apple M5 MIE

Three researchers used Anthropic Mythos to develop a macOS kernel exploit in six days, moving from discovery on April 25 to completion on May 1, bypassing Apple’s MIE memory-integrity system for M5 and A19 chips and gaining root via standard unprivileged system calls; the full technical report will follow Apple’s patch.

Why it matters: HKR-H/K/R all pass: Anthropic Mythos, a 6-day macOS kernel exploit, and M5/A19 MIE bypass create real dual-use signal. Kernel-exploit depth and single X-source sourcing keep it below the 85 must-write band.

AI HOT (Curated Pool)

Codex adds multi-device remote control and shared context

Codex controls multiple devices through ChatGPT, switches by project to access each device’s context and files, and supports remote SSH setup for other VMs.

Why it matters: HKR-H/K/R all pass, but the item is a thin X-post summary with no official release note, pricing, permission model, or reproducible demo. Treat it as a mid-weight coding-agent product update at the featured threshold.

r/LocalLLaMA

Qwen3.6-35B-A3B and 9B land on the public Terminal-Bench 2.0 leaderboard

little-coder × Qwen3.6-35B-A3B scored 24.6% ±3.2 on Terminal-Bench 2.0, above Gemini 2.5 Pro on Gemini CLI at 19.6% and Qwen3-Coder-480B on Terminus 2 at 23.9%.

Why it matters: HKR-H/K/R all pass, but this is a Reddit post with leaderboard numbers only; test setup and reproducibility details are not disclosed. Strong code-agent benchmark signal, not a 78+ release story.

AI HOT (Curated Pool)

OpenAI Restructures as Brockman Takes Over Product Strategy

OpenAI merged ChatGPT, Codex, and API into one product organization, with Greg Brockman taking over product strategy; the post says Anthropic’s valuation reached $900 billion, but it does not disclose the restructuring timeline.

Why it matters: HKR-H/K/R all pass: this is an OpenAI top-level product reorg covering ChatGPT, Codex, and API. Single-source summary keeps it below the highest band, but it is same-day must-write news.

QbitAI · WeChat

Codex Integrates HeyGen for Prompt-Based Video Generation and Editing

Codex integrates the HeyGen plugin to run image generation, talking-avatar video, subtitles, and edits from natural-language prompts; the article tests roughly one-minute avatar generation, trimming content after 10 seconds, and deleting a blink at the eighth second.

Why it matters: HKR-H/K/R all pass, backed by a numbered hands-on test. The scope is still one Codex-to-HeyGen plugin workflow, not a model or platform release, so it lands in the 72-77 featured band.

Google DeepMind

Google DeepMind releases Gemini 3.5 Flash

Google DeepMind released the Gemini 3.5 model family, with the first model, Gemini 3.5 Flash, available the same day in the Gemini app, Google Search AI Mode, Google Antigravity, the Gemini API and Gemini Enterprise.

Why it matters: Google published 3.5 Flash's coding and agent benchmark scores and where it is available, enough to judge its place in long-horizon workflows.

AI HOT (Curated Pool)

Ignoring Token Costs, Using 100 AI Instances to Automate an Open Source Project

The OpenClaw team runs about 100 Codex instances to handle code review, security analysis, issue deduplication, test reproduction, task creation from meetings, spam filtering, and performance regression monitoring.

Why it matters: HKR-H/K/R all pass: 100 Codex instances running open-source maintenance is a strong operational anecdote with concrete task types. Single X post, no cost, outcome metrics, or reproducible setup, so it stays in the lower featured band.

AI HOT (Curated Pool)

Runway Agent Generates Complete Ads in One Session

Runway Agent turns product photos and ideas into fully produced ads in one session; the post does not disclose the model, pricing, generation length, or regional availability.

Why it matters: Runway’s ad-generation Agent clears HKR-H/K/R as a mid-weight product update. Missing model, pricing, duration, and region details keep it at the featured threshold, not a must-write release.

The Verge · AI

OpenAI keeps shuffling executives to win the AI agent battle

OpenAI announced a reorganization Friday that makes Greg Brockman the official lead for product, and its memo says the company will combine ChatGPT and Codex into one unified agentic experience.

Why it matters: HKR-H/K/R all pass: the power shuffle is clickable, the product merge is new, and OpenAI's agent roadmap matters. It stays at 83 because no capability shipped and timing, pricing, and technical details are absent.