Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

781–800 of 1,465

May 26Tuesday

AI HOT (Curated Pool)

Sundar Pichai on AI, the Future of Search, and Changes to the Web

Sundar Pichai said after Google I/O that Google is integrating Gemini into a new smart search box and the Gemini Spark agent platform; the post does not disclose model parameters, launch dates, or traffic impact numbers.

Why it matters: HKR-H and HKR-R pass: Pichai’s interview touches Google Search as an AI entry point and web traffic allocation. HKR-K is weak because the article gives Gemini-in-Search and Spark, but no rollout timing or technical detail.

r/LocalLLaMA

SkillOpt treats markdown skill files as trainable parameters with proper optimization machinery

SkillOpt uses a frontier model to propose add, delete, and replace edits to markdown skill files, then accepts only strict gains on a held-out validation set; the best skills usually converge after 1 to 4 accepted edits.

Why it matters: HKR-H/K/R all pass: the hook is trainable markdown skills, with held-out validation and 1-4 accepted edits. Single Reddit/project source and no broad adoption data keep it at 78, featured not p1.

AI HOT (Curated Pool)

Qwen3.7-Max Becomes the World’s No. 2 AI Coding Model

Qwen3.7-Max scored 1541 on Code Arena and ranked behind Claude; the post says it can run 35-hour tasks and perform more than 1,000 tool calls.

Why it matters: HKR-H/K/R all pass, but the source is a single Alibaba Cloud post and the evidence is benchmark plus vendor claims. This fits a strong product/benchmark update, not P1 without independent validation.

QbitAI · WeChat

Chinese AI-Written Pretraining Framework ForgeTrain Trains MiniCPM5-1B

ModelBest released ForgeTrain and MiniCPM5-1B, saying ForgeTrain was written by AI and trains 10% faster than NVIDIA Megatron under the same hardware conditions. MiniCPM5-1B is a 1B-parameter edge model with about 2GB FP16 weights and about 0.5GB INT4/Q4 weights.

Why it matters: HKR-H/K/R all pass: an AI-written trainer, a 10% same-hardware Megatron speed claim, and a 0.5GB 1B edge model are concrete hooks. Score stays at 80 because the first-ever claim and benchmark lack third-party reproduction.

Synced · WeChat

ACL 2026 Main: Spatial-Agent Generates Executable Geospatial Analysis Workflows for LLMs

Spatial-Agent inserts a GeoFlow Graph between natural-language questions and map tools, and Spatial-Agent with GPT-4o-mini reaches 45.15% accuracy on MapEval-API versus a 23.00% API baseline.

Why it matters: ACL Main gives a concrete mechanism and testable numbers, so HKR-H/K pass. The GIS focus limits HKR-R, placing it at the featured threshold rather than a must-write item.

Synced · WeChat

Grok keeps updating after xAI disbandment as Musk announces a new model

Elon Musk said the 1.5T-parameter Grok V9-Medium has finished training, will enter reinforcement learning in a few days, and is planned for release in two to three weeks. Grok Build supports up to 8 parallel sub-agents, a 256K-token context window, Plan Mode, Arena Mode, MCP, and ACP.

Why it matters: HKR-H/K/R all pass, but this is a Grok V9-Medium preview before RL and release, with no benchmarked capability yet. That fits a strong model-race/product update at 82, featured but not p1.

Synced · WeChat

AI-written training framework trains 1B edge model MiniCPM5-1B

ModelBest open-sourced MiniCPM5-1B and ForgeTrain; the 1B edge model scores 17.9 on AA-Index, while the AI-written ForgeTrain framework matches Megatron’s training results and runs 10% faster on Nvidia H100 under the article’s reported setup.

Why it matters: HKR-H/K/R all pass: the AI-written training framework hook is strong, with concrete AA-Index and H100 speed claims. It is not a flagship model release, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

Chinese agent SkyClaw targets Opus 4.6-level performance with free trial

Kunlun Tech released SkyClaw-v1.0 and SkyClaw-v1.0-lite with a 2-4 week free trial, claiming SkyClaw-v1.0 input costs are 1/24 of DeepSeek V4 Pro and about 1/43 of Sonnet 4.6.

Why it matters: HKR-H/K/R all pass: SkyClaw-v1.0 has a sharp cost hook, concrete trial and pricing ratios, and budget resonance. Source facts remain vendor claims, so it stays at the low featured band.

New York Times Chinese

The Shared U.S.-China AI Anxiety: Being Harvested by the Future

Yi-Ling Liu compares U.S. and Chinese AI anxiety through labor, companionship, and agency: over 70% of U.S. teenagers report using chatbots as companions, while China is projected to reach 200 million single-person households by 2030.

Why it matters: HKR-H/K/R all pass, but this is commentary rather than a model, product, or policy release. Its signal comes from two social data points and a US-China framing, so it fits the featured threshold for an insightful opinion piece.

r/LocalLLaMA

Update on a 12×32GB SXM V100 Cluster for Local Legal Drafting

A lawyer runs a local legal-drafting pipeline across 16 GPUs, with Qwen3.5-122B-A10B reaching about 50 tok/s on four V100s, while a verifier blocks ungrounded citations, dates, and Bates numbers before any final document is used.

Why it matters: HKR-H/K/R all pass: this is a first-person local-LLM experiment with concrete numbers, not a vendor post. Reddit source limits authority, so it stays at the low featured band rather than p1.

AI HOT (Curated Pool)

Apple reportedly uses a custom 1.2T-parameter Google model for next-generation Siri

Apple is reportedly using a custom 1.2T-parameter Google model to run parts of the next-generation Siri, while simpler queries are expected to run on-device; the post says response speed for everyday questions is the key constraint.

Why it matters: HKR-H/K/R all pass, but this is a single X-sourced reported claim; the post gives architecture details but not sourcing documents, rollout timing, or scope. Keep it at the featured threshold, below the 78+ band.

AI HOT (Curated Pool)

Grok Build Beta Opens to SuperGrok Users

xAI opened Grok Build Beta to all SuperGrok and X Premium+ users, with Plan Mode, Imagine-based image and video creation, and a CLI for automation or orchestrator workflows at x.ai/cli.

Why it matters: HKR-H/K/R all pass: xAI opened a paid beta with named workflow features. The score stays at the featured floor because the post lacks capability limits, pricing detail, and test results.

May 25Monday

r/LocalLLaMA

The reason small-model agent stacks aren't the default is not whether they work

A Reddit post argues small-model agent stacks are not default for business reasons, not capability limits: Gemma 4 31B reaches 86.4% on tau2-bench, and DeepSeek V4-Flash output tokens are priced about 89x below Claude Opus 4.6. The operational risk is verification, because 7–9B models produced broken reasoning for roughly half to two-thirds of correct answers in a cited audit.

Why it matters: HKR-H/K/R all pass: the angle is contrarian, with benchmark, cost, and verifier-failure numbers. Reddit-source uncertainty keeps it in the 78–84 recommendation band, not P1.

r/LocalLLaMA

Computer-use sandbox framework for Codex on headless Linux

superSmitty9999 released ai-sandbox-manager as a PoC that uses LXC templates to give Codex sudo access, browser use, Docker, and shared GPU access, with a hook that blocks git push while the agent works inside isolated copies.

Why it matters: HKR-H/K/R all pass, but this is a Reddit personal PoC with mechanisms only, not adoption, benchmarks or maturity evidence. It fits the featured floor for practical agent-sandbox work.

QbitAI · WeChat

Reasonix for DeepSeek V4 reaches 99.82% cache hit rate and cuts costs to 20%

Reasonix uses an append-only loop for DeepSeek V4 and reports a 99.82% cache hit rate in long coding sessions, cutting an example 400M-token bill from $61 to $12.

Why it matters: HKR-H/K/R all pass, but this is a third-party cost tool around DeepSeek V4, not a model launch or platform update. Concrete mechanism and billing numbers put it in the 72–77 featured band.

AI HOT (Curated Pool)

Harness, Scaffold, and AI Agent Terminology Explained

Hugging Face’s post frames an agent as three layers: Model, Scaffolding, and Harness; Scaffolding defines behavior through prompts and tool descriptions, while Harness runs model calls, tool calls, and control loops.

Why it matters: HKR-H/K/R pass: the Hugging Face post gives a concrete agent-stack taxonomy. It clears featured on practitioner relevance, but lacks a release, benchmark, or deployment case, so it stays at the threshold.

AI HOT (Curated Pool)

TrapDoor Supply Chain Attack Makes AI Assistants a New Attack Surface

TrapDoor hit npm, PyPI, and Crates.io with 34 malicious packages, using manipulated CLAUDE.md and .cursorrules files in pull requests to make Claude Code and Cursor treat attacker content as trusted instructions and run malicious commands.

Why it matters: HKR-H/K/R all pass: AI coding assistants become the execution surface, with 34 malicious packages across three registries. Single-post sourcing lacks IOCs, timeline, and victim scale, so this stays in the 78–84 band.

May 24Sunday

r/LocalLLaMA

Using llama.cpp native tools for web RAG inside llama-server WebUI

A Reddit user describes using llama.cpp native tools for web RAG inside llama-server WebUI with a 7-step setup: enable get_datetime and exec_shell_command, then run wget through firejail, a separate Linux user, and an Alpine OCI VM sandbox.

Why it matters: HKR-H/K/R all pass: the post gives a concrete local web-RAG recipe with sandboxing. It is a community tutorial, not a model or product launch, so the narrow reach and source authority keep it at the low featured band.

Synced · WeChat

Meta layoff survivors face a difficult choice

Meta is pushing some post-layoff employees into new roles: some engineering managers are returning to IC work, while some Infra and AI engineers are being reassigned to data labeling; the article cites a manager-to-report ratio shift from 1:8 to 1:50 and says Meta holds a 49% stake in Scale AI.

Why it matters: HKR-H/K/R all pass: the piece has a concrete oddity, numbers, and a job-security nerve. It is still workforce reporting rather than a model launch or executive departure, so it sits in the lower featured band.

Xinzhiyuan · WeChat

AI Agent Completes Chip Design from 219 Words to 7nm GDSII Without Engineer Input

Verkor’s Design Conductor generated an ASAP7 7nm GDSII layout for the VerCore RISC-V CPU from a 219-word English spec in 12 hours, with no engineer in the design loop; the reported result scored 3,261 CoreMark at 1.48GHz, but it has not been fabricated and lacks cache implementation.

Why it matters: HKR-H/K/R all pass, but VerCore is not taped out and lacks cache, so the claim stays at demo-and-benchmark level. Concrete numbers and test conditions put it in the 78–84 recommendation band.