Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

641–660 of 1,465

Jun 3Wednesday

AI HOT (Curated Pool)

Microsoft and OpenAI Split as Both Prepare to Compete Directly

Microsoft and OpenAI have shifted from partnership to direct competition, and Microsoft AI chief Mustafa Suleyman said Microsoft must prove from scratch that it can independently complete the required work; the post does not disclose a product roadmap or timeline.

Why it matters: HKR-H and HKR-R pass: Microsoft/OpenAI rivalry affects agent-platform strategy. HKR-K is weak because the article gives no roadmap or testable technical detail, so it sits just above the featured threshold.

AI HOT (Curated Pool)

Meta's AI Agent for WhatsApp Business Is Now Available Globally

Meta made its WhatsApp Business AI agent available to merchants globally and will charge businesses based on model token usage; the post does not disclose pricing, model names, or a market-by-market availability list.

Why it matters: HKR clears all three: a global WhatsApp Business agent rollout, token-based billing, and direct platform pressure on SMB automation. Missing price, model name, and market list keep it in the lower featured band.

MIT Technology Review · AI

The Download: Trump’s New AI Order, and Smart Glasses for Warfare

President Donald Trump signed a new AI order asking companies to voluntarily submit frontier models for government review 30 days before release, without mandatory licensing; the newsletter also says Anduril and Meta are prototyping a military AR headset that envisions drone-strike orders through eye tracking and voice commands.

Why it matters: HKR-H/K/R all pass: the article gives a concrete 30-day frontier-model review mechanism and a Meta/Anduril AR warfare prototype. A presidential AI order affecting release compliance clears the must-write band.

AI HOT (Curated Pool)

Build 2026: Microsoft tops Google in image generation while catching up on reasoning

Microsoft announced seven in-house AI models at Build 2026, including its first reasoning model, one new tuning method, and one autonomous background AI agent; the RSS snippet does not disclose model names, benchmarks, or release dates.

Why it matters: HKR-H/K/R all pass: Microsoft shipped seven in-house AI models across reasoning, tuning, and a background agent. Model names, benchmark details, and availability are not disclosed, so this stays at the top of 78–84, not P1.

Alibaba Technology · WeChat

Rethinking R&D Infrastructure When Agents Become First-Class Citizens

Xu Xiaobin argues that agent-based development compresses the intent-to-code loop from weeks or months to minutes, using a weekly-report system, a multi-role agent development setup, and image-repository provisioning as examples; the article identifies mismatches in Git, CI, code review, release flows, permissions, harness setup, and dry-run validation.

Why it matters: HKR-H/K/R all pass, but this is infrastructure commentary rather than a model or product launch. The named cases and week/month-to-minutes claim put it in the 72–77 featured band.

AI Chat-Group Daily (群聊日报)

2026-06-02 Chat Group Daily

The chat group daily says Microsoft released MAI-Thinking-1 with 35B active parameters and about 1T MoE, matching Opus 4.6 on SWE-Bench Pro and scoring 97% on AIME 2025.

Why it matters: HKR-H/K/R all pass: a Microsoft reasoning-model claim with concrete benchmark numbers. Source authority is weak, and the summary lacks official release, access terms, and full eval setup, so it stays below P1.

QbitAI · WeChat

Papers with Code returns with CVPR coverage and Hugging Face-led rebuild

Hugging Face’s open-source team launched paperswithcode.co in May 2026, using AI agents to parse papers and restore SOTA leaderboards tied to the original platform’s 9,300-plus benchmarks.

Why it matters: HKR-H/K/R all pass: a beloved research portal returns, with 9,300 restored leaderboards and agent-based paper parsing. The impact is strong for research workflows, not model-release scale.

QbitAI · WeChat

Coze 3.0 test: phone can remotely control agents on your computer

Coze 3.0 adds project-based agent collaboration across iOS, Android, Mac, Windows, and web, supports importing local agents such as Claude Code, Codex CLI, and OpenClaw, and can read a desktop PDF from a phone after user authorization.

Why it matters: HKR-H/K/R all pass: Coze 3.0 adds cross-device agent control and imports Claude Code/Codex CLI. It remains a single product update, with price, rollout scope, and security limits not disclosed, so it sits at the featured threshold.

AI HOT (Curated Pool)

Qwen3.7 Released with Upgrades to Reasoning and Agent Capabilities

Qwen released Qwen3.7, and the post says it upgrades reasoning, tool use, coding, and long-horizon agent tasks; the post does not disclose model size, pricing, benchmark scores, or release conditions.

Why it matters: HKR-H and HKR-R pass because Qwen3.7 is a flagship Alibaba model update with practitioner relevance. HKR-K fails: the post names capability areas but gives no params, pricing, benchmarks, or access terms.

r/LocalLLaMA

Microsoft Aion 1.0 Instruct and Aion 1.0 Plan models

Microsoft announced two on-device Aion 1.0 models at Build 2026. Aion 1.0 Plan is a 14B-parameter reasoning and tool-calling model with 32K context, shipping in-box with Windows on capable devices, while Aion 1.0 Instruct targets summarization, rewriting, intents, accessibility, Edge integration, and open-weight availability.

Why it matters: Microsoft announced Aion 1.0 Instruct and Plan at Build 2026, with Plan listed as a 14B, 32K-context model for eligible Windows devices. HKR-H/K/R all pass, but licensing, benchmarks, and hardware requirements are not disclosed, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

OpenAI’s Greg Brockman and a 9-Year Rift With Anthropic Co-Founder Dario Amodei

A WSJ-based profile says Dario Amodei once barred Greg Brockman from an internal OpenAI project that later led to ChatGPT, and the article says Brockman now oversees OpenAI product strategy with nearly 1,500 people under that function.

Why it matters: HKR-H/K/R all pass: the WSJ-sourced ban detail and the nearly 1,500-person scope give this more signal than gossip. It is not a model release or current executive departure, so it stays in the good-quality featured band.

AI HOT (Curated Pool)

Complete Practical Tips for Agent Engineering

@mvanhorn shared an agent engineering workflow centered on a Research→Plan→Work loop, plan.md constraints, and 22 practical tips; the snippet says it covers planning, parallel execution, input methods, and remote control, but the post does not disclose the full tool stack list.

Why it matters: HKR-H/K/R all pass, but this is a practitioner methods post, not a model or product release. The full tool stack is not disclosed, so it sits at the featured threshold.

New York Times Chinese

Tech Companies Are Cutting Jobs: Is AI the Cause or the Excuse?

Meta, Coinbase, and Block each cut at least 10% of staff in recent months, totaling about 13,000 jobs, while citing AI for part of the reductions. Layoffs.fyi says more than 150 tech companies have cut at least 115,000 workers this year, as analysts question whether AI is the cause or a cover for overhiring and weaker businesses.

Why it matters: HKR-H/K/R all pass: the NYT piece ties concrete layoff numbers to the AI-as-cause-or-excuse debate. It stays at the featured threshold because this is macro labor reporting, not a model, product, or policy update.

Computing Life · Share · Yage

After vibe coding: the industrialization of AI programming

MAI filtered 265,000 trainable tasks from 4.87 million open-source PRs and built a three-layer judging system. The key change after vibe coding is the industrialization of training infrastructure.

Why it matters: HKR-H/K/R all pass via the post-vibe-coding angle, 4.87M PR corpus, 265K tasks, and code-agent infra stakes. No model scores, open-source scope, or product access are disclosed, so it stays below P1.

AI HOT (Curated Pool)

NVIDIA launches NemoClaw platform for autonomous AI engineers in industrial software

NVIDIA released NemoClaw at COMPUTEX as an open blueprint for long-running AI agents, and more than a dozen industrial software vendors are using it to build autonomous AI engineers for CAE and EDA workflows that compress weeks-long simulation and design tasks into hours.

Why it matters: HKR-H/K/R pass: NVIDIA’s NemoClaw targets industrial agents with 10+ vendors and a weeks-to-hours claim. The NVIDIA-blog sourcing and missing technical detail keep it at the lower featured band.

AI HOT (Curated Pool)

Google DeepMind open-sources a toolkit for scientific agents

Google DeepMind released Science Skills on GitHub for scientific-discovery agent workflows; the post does not disclose the license, benchmark results, or numeric token-efficiency gains.

Why it matters: Passes HKR-H/K/R: DeepMind, open source, and science agents make it relevant. Missing license, benchmarks, and efficiency data keep it in the 78–84 band, not P1.

AI HOT (Curated Pool)

Claude Code Adds Dynamic Workflows

Claude Code added dynamic workflows that execute JavaScript files at runtime to create and coordinate multiple subagents; each subagent has its own context window, and the feature is described for research, security analysis, and code review tasks.

Why it matters: HKR-H/K/R all pass: Claude Code gets runtime JS workflows coordinating isolated-context subagents. Anthropic update earns a bump, but this is a feature release rather than a model or platform launch, so it sits in the 78–84 band.

AI HOT (Curated Pool)

Claude Code launches dynamic workflows for task-specific frameworks

Claude Code added dynamic workflows that execute JavaScript files to coordinate subagents, with configurable model choice and workspace isolation level, but the post does not disclose token overhead figures or release availability details.

Why it matters: HKR-H/K/R all pass, but the post gives mechanism-level detail only; token overhead, rollout scope, and pricing are not disclosed. Claude Code relevance lifts this to the high end of a mid-weight product update.

NVIDIA Blog

NVIDIA Partners With Microsoft on Unified Stack for Agentic AI Deployment

NVIDIA and Microsoft announced a unified agentic AI deployment stack at Build across Windows, Azure, and local environments; RTX Spark provides 1 petaflop of AI performance, while DGX Station for Windows offers 20 petaflops of FP4 performance and up to 748GB of coherent memory.

Why it matters: HKR-H/K/R pass: the NVIDIA-Microsoft stack spans Windows, Azure, and local devices, with 1 PFLOP and 20 PFLOPs FP4 specs. Vendor-source limits the score: pricing, benchmarks, and migration details are not disclosed.

AI HOT (Curated Pool)

Claude Platform Adds CLI Tool

Claude Platform added a CLI that runs every API endpoint from the terminal, calls the Messages API, launches Claude-hosted agents, and pipes results directly into the shell.

Why it matters: Claude Platform CLI clears HKR-H/K/R as a practical developer-tooling update, but the post only gives capability scope; install flow, permissions, safety limits, and pricing are not disclosed.