Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

701–720 of 1,465

Jun 1Monday

AI HOT (Curated Pool)

NVIDIA Releases FOX Factory Operations Blueprint for Autonomous Factory Management Agents

NVIDIA released the FOX factory operations blueprint at GTC Taipei, and Foxconn used it to build the MoMClaw multi-agent system with an expected 80% reduction in root-cause analysis time.

Why it matters: HKR-H/K/R pass: NVIDIA is pushing an agent blueprint into factory ops, with Foxconn’s MoMClaw and an expected 80% RCA time cut. Kept at the featured floor because the source is a vendor blog and the result is projected.

AI HOT (Curated Pool)

NVIDIA Releases RTX Spark and Local AI Agent Security and Performance Updates

NVIDIA released RTX Spark, a Windows PC for local AI agents with 1 petaflops of AI compute and 128GB of unified memory. OpenShell uses new Windows security primitives with Microsoft, while llama.cpp optimizations raise Qwen 27B throughput by up to 2x.

Why it matters: HKR-H/K/R all pass: NVIDIA frames RTX Spark for local agents and gives hard specs: 1 petaflops, 128GB, and up to 2x llama.cpp throughput. Vendor-blog framing keeps it in the low 78–84 band.

AI HOT (Curated Pool)

MiniMax M3: Frontier coding, 1M-token context, and native multimodal model

MiniMax released M3 as an open-source unified model with coding, agent, and native multimodal capabilities, supporting a 1M-token context window and using MiniMax Sparse Attention to cut per-token compute at 1M context to 1/20 of its predecessor, with over 9x faster prefill and over 15x faster decoding.

Why it matters: HKR-H/K/R all pass: MiniMax M3 has a 1M-token context hook, MSA with a claimed 20x cost cut, and open-source China-model resonance. Single official-source release keeps it in the 78–84 band, not P1.

AI HOT (Curated Pool)

Qwen3.7-Plus: Multimodal Agent Intelligence

Qwen Studio lists seven capability areas: chatbots, image and video understanding, image generation, document processing, web search integration, tool use, and artifact generation; the post does not disclose Qwen3.7-Plus parameters, pricing, or release timing.

Why it matters: HKR-H/K/R pass, but the facts are thin: 7 capability categories, no params, pricing, benchmarks, or launch terms. A Qwen flagship update clears featured, not p1.

AI HOT (Curated Pool)

MWC26 Shanghai to Host First Humanoid Robot Penalty Shootout With Unitree and 7 Other Teams

MWC26 Shanghai will host a humanoid robot penalty shootout in June 2026, with eight Chinese embodied intelligence teams competing under rules that require autonomous play without human control or preset scripts.

Why it matters: HKR-H/K/R all pass: the robot penalty shootout is clickable, with rules banning teleoperation and scripts. It stays in 72–77 because this is an event preview, not a model release or reproducible result.

May 31Sunday

AI HOT (Curated Pool)

Apple WWDC AI Upgrade: Gemini-Distilled Model Runs Locally, With Heavy External Dependencies

Apple will present Siri and on-device AI upgrades at next month’s WWDC, with iPhones running a smaller Gemini-distilled model locally while complex queries route to Google Cloud using Nvidia confidential computing.

Why it matters: HKR-H/K/R all pass: the Apple-Google-Nvidia stack is a strong WWDC AI hook with a concrete routing mechanism and clear industry tension. Capped at 82 because this is a single X-sourced claim with no model size, latency, pricing, or contract terms disclosed.

r/LocalLLaMA

PolyRange: Contamination-resistant offensive-AI benchmark for web targets

PolyRange v1.0 ships 84 WSTG-derived classes across 12 OWASP testing-guide categories. It generates fresh targets per deploy with a chosen LLM, adds two defense tiers, uses an agent-submits-flag oracle, and runs via a single-command CLI on Fly.io or Docker.

Why it matters: HKR-H/K/R all pass: PolyRange turns web-security targets into a dynamic agent benchmark with 84 WSTG classes and two defense levels. Single-source Reddit origin and security niche keep it at 78.

r/LocalLLaMA

Use any model and provider with the official OpenAI Codex Desktop App without modifying its code

Reddit user thibautrey describes a 3-step setup: edit Codex Desktop config.toml, store an API key, and use a multicodex proxy alias to map gpt-5.3-codex to MiniMax-Latest. The post lists a local base_url of 127.0.0.1:1455 and says the proxy disguises returned model names as gpt-5.3-codex.

Why it matters: This is a reproducible developer workflow trick, not an official release. HKR-H comes from the lock-in workaround, HKR-K has concrete config details, and HKR-R hits cost and model-choice pressure, placing it at the tutorial featured threshold.

Synced · WeChat

Rubrics Survey: How to Define a Good Answer in the Agent Era

Renmin University Gaoling School of Artificial Intelligence released a 40-page survey on rubrics for LLMs, organizing the topic into five parts: definitions, construction methods, training uses, evaluation scenarios, and open challenges.

Why it matters: HKR-H/K/R all pass, but this is a survey rather than a model or product launch. The 40-page rubric framework is useful for agent evaluation, placing it at the featured threshold.

Synced · WeChat

Microsoft open-sources SkillOpt for training Agent skill documents, reaching 3.3k stars in a week

Microsoft open-sourced SkillOpt, a text-space optimization framework that trains Agent skill documents without changing model weights; the paper reports best or tied-best results across 52 combinations covering 7 target models, 6 benchmarks, and 3 execution environments.

Why it matters: Microsoft’s open-source SkillOpt is a strong Agent tooling and research release. HKR-H has the 3.3k-star/trainable-skill hook, HKR-K has the text-parameter mechanism and 52 eval setups, and HKR-R hits agent engineering pain, so it lands in featured at 82.

Xinzhiyuan · WeChat

Fudan-Linked Team Releases STI-WM Spatiotemporally Integrated World Model

MouShen Intelligence released STI-WM, a spatiotemporally integrated world-action model for robotics, claiming support for RGB, point-cloud, and proprioceptive inputs, hundred-second task planning, and disclosing five funding rounds in six months plus a RMB 300 million Pre-A round.

Why it matters: HKR-H/K/R pass: STI-WM combines RGB, point clouds, and proprioception for 100-second planning, plus 5 funding rounds and a RMB300m Pre-A. Company-claim framing lacks public benchmarks or reproducible access, so it stays near the featured threshold.

QbitAI · WeChat

Fudan and Tongyi introduce ToolCUA for GUI-Tool path selection in agents

Fudan University and Tongyi Lab introduced ToolCUA-8B, which reaches 46.85% accuracy on OSWorld-MCP after training with about 4k synthetic tools and 180k interleaved GUI-Tool trajectory steps.

Why it matters: HKR-H/K/R all pass: the tool-selection failure hook is concrete, with OSWorld-MCP 46.85% and 180k steps. It stays in the 78–84 band because this is a research release, not a major model or product launch.

QbitAI · WeChat

NVIDIA’s MacBook Pro-like laptop reportedly uses an in-house CPU

NVIDIA, Microsoft, and Arm posted the same “new era of PC” teaser, and the article says the rumored N1X laptop may use a 20-core Arm CPU, a Blackwell GPU, 6,144 CUDA cores, and 128GB of LPDDR5X unified memory, while bandwidth and x86 translation remain the stated constraints.

Why it matters: HKR-H/K/R all pass, but the story rests on hints and rumored specs; launch date, price, and production plan are not confirmed. Treat it as a strong hardware rumor, not a same-day must-write release.

AI HOT (Curated Pool)

Tesla FSD completes a 6,000 km zero-intervention autonomous drive across Canada

Tesla FSD V14.3.3 completed a 6,051 km zero-intervention drive from Vancouver to Halifax in 4 days and 21 hours, with the system handling lane changes, complex road conditions, and parking without disengagements or human corrections.

Why it matters: HKR-H/K/R all pass: Tesla FSD V14.3.3 has a concrete 6,051 km zero-intervention claim. It stays below 85 because the item gives the result but lacks independent validation, route detail, and failure boundaries.

May 30Saturday

TechCrunch · AI

I put Google’s 24/7 AI assistant Gemini Spark to work, and it’s actually pretty useful

TechCrunch tested Google’s Gemini Spark as a 24/7 AI assistant for inbox summaries and local event planning; the RSS snippet does not disclose pricing, release timing, or why Google made it a separate product.

Why it matters: HKR-H/K/R pass: the hands-on angle is clickable, and inbox plus local-planning automation gives concrete substance. The score stays in the low featured band because price, launch timing, and product positioning are not disclosed.

Xinzhiyuan · WeChat

Opus 4.8 Builds a Historical Rebirth Simulator for 117 Billion Humans

Ethan Mollick used Claude Opus 4.8 to generate The Veil of History, a website that weights a random human life by 117 billion historical births and, according to the article, uses 4,000 Monte Carlo runs to estimate regional and era distributions.

Why it matters: HKR-H/K/R all pass: Mollick’s Claude Opus 4.8 demo has a strange hook, concrete numbers, and a builder-relevant prototyping angle. It is not an Anthropic release, so it stays in the lower featured band.

Financial Times · Technology

UK military looks at allowing lethal strikes without human approval

The FT headline says the UK military is examining lethal strikes without human approval, but the accessible body is a subscription page and does not disclose the weapon types, approval mechanism, legal conditions, or deployment timeline.

Why it matters: HKR-H and HKR-R are strong: the FT headline points at a lethal-autonomy policy red line. HKR-K fails because the accessible body is a subscribe page with no mechanism, timeline, or scope.

QbitAI · WeChat

RUC and Zhizhi Institute Open-Source Claw Agent Data, Training, and Evaluation Pipeline

Renmin University of China and Zhizhi Institute open-sourced ClawGym, a Claw Agent framework with 13.5K synthetic executable tasks, 200 benchmark tasks, model checkpoints, training data, and training code; ClawGym-30B-A3B scores 56.82 on ClawGym-Bench and exceeds Qwen3-235B-A23B in the reported evaluation.

Why it matters: HKR-H/K/R all pass: ClawGym bundles data, code, checkpoints, and eval tasks rather than just a leaderboard. Its impact is developer-facing, below a major lab model release or market-moving event.

Synced · WeChat

NVIDIA and Tsinghua Team's Gamma-World Tops Hugging Face Daily Chart

NVIDIA, Tsinghua, University of Toronto, and Vector Institute released Gamma-World, a multi-agent world model using simplex-based positional encoding and hub tokens to cut interaction cost from quadratic to linear, with 8-player latency dropping from 17.6 ms to 4.5 ms.

Why it matters: HKR-H/K/R all pass: Gamma-World has a concrete mechanism and latency claim from NVIDIA/Tsinghua. Scope remains multi-agent world-model research, so it sits in the 78–84 good-quality band rather than must-write.

AI HOT (Curated Pool)

Codex Can Manage Conversation Threads and Parallel Tasks

Codex can now create, search, organize, and pin conversation threads inside the Codex interface, and start worktrees for parallel tasks.

Why it matters: HKR-H/K/R pass: Codex gets concrete thread-management and parallel-worktree mechanics that matter to coding-agent users. Scope, pricing, and performance data are not disclosed, so this stays in the lower featured band.