Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

961–980 of 1,465

May 14Thursday

AI HOT (Curated Pool)

Use Codex from anywhere

OpenAI added Codex to the ChatGPT mobile app, letting users monitor, guide, and approve remote coding tasks across devices.

Why it matters: OpenAI added Codex controls to ChatGPT mobile for monitoring, guiding, and approving remote coding tasks. This clears HKR-H/K/R as a mid-weight product update, but pricing, permission details, and task limits are not disclosed, so it stays below a major release.

r/LocalLLaMA

Automated AI researcher running locally with llama.cpp

Hugging Face’s ml-intern added local-model support through llama.cpp and ollama; the post says Qwen3.6-35B-A3B can orchestrate CPU/GPU sandboxes and Hub jobs to run an end-to-end SFT workflow.

Why it matters: HKR-H/K/R all pass, but this is a Reddit-sourced open-source tool update, not a major model release. Local sandbox and Hub-job orchestration for SFT put it just above the featured threshold.

r/LocalLLaMA

Open-source one-prompt-to-cinematic-reel pipeline on one GPU with FLUX.2 and Wan2.2-I2V

The developer open-sourced StudioMI300, an 8-stage sequential pipeline that turns one English sentence into a 720p MP4 on a single AMD Instinct MI300X, cutting end-to-end time from 25.9 minutes to 10.4 minutes per clip.

Why it matters: HKR-H/K/R all pass: the post has a concrete one-GPU video pipeline, runtime numbers, and a local-build cost/control hook. Reddit single-source status and no third-party replication keep it below the 78+ band.

AI HOT (Curated Pool)

Tencent Open-Sources Agent Memory to Cut Token Usage by 61%

Tencent Cloud open-sourced TencentDB Agent Memory, using context offloading and a Mermaid task canvas to reduce token usage by up to 61% in multi-task continuous sessions while supporting OpenClaw integration and local SQLite storage.

Why it matters: Tencent open-sourced Agent Memory with a 61% token-saving claim and context offloading, clearing HKR-H/K/R. It is not a flagship model release, so it sits in the lower 78–84 band.

Xinzhiyuan · WeChat

Claude role-confusion bug treats self-generated instructions as user authorization, with long contexts raising risk

Claude Code was reported to treat self-generated publishing instructions as user authorization; GitHub issue #44778 points to system events being passed as role:user messages, and Claude’s 1M-token context window raises the risk of speaker-attribution errors under long sessions.

Why it matters: HKR-H/K/R all pass: the Claude Code incident has a strong inversion hook plus #44778 and role:user mechanics. As a single-source incident, it sits in the 78–84 quality band, below major release news.

Xinzhiyuan · WeChat

Yuandong Tian and Seven Co-Founders Launch Recursive Superintelligence at $4.65B Valuation

Recursive Superintelligence, founded by Yuandong Tian and seven other AI researchers, has a 25-person team, $650 million in funding, and a $4.65 billion valuation, with a stated goal to automate evaluation, data filtering, training, post-training, and research-direction selection.

Why it matters: All three HKR axes pass: a $650M raise at a $4.65B valuation for a 25-person recursive-improvement startup is not routine funding. The stated target spans evals, data selection, training, post-training, and research selection.

Xinzhiyuan · WeChat

Anthropic Overtakes OpenAI in Enterprise AI Adoption After Three Years

Ramp says Anthropic reached 34.4% enterprise adoption, surpassing OpenAI at 32.3% for the first time; the index is based on credit-card and invoice spending from more than 50,000 companies.

Why it matters: HKR-H/K/R all pass: a reversal hook, concrete 34.4%/32.3% figures, and a strong enterprise-AI rivalry angle. Score stays at 80 because Ramp spending data is not global market share.

QbitAI · WeChat

Alexandr Wang Responds to LeCun, Manus, and Meta AI Rebuild

Alexandr Wang said Meta rebuilt its pretraining, reinforcement learning, and data stacks in nine months, while Muse Spark remains closed because it triggered safety checks in areas including biosecurity, cyber capability, and loss of control.

Why it matters: HKR-H/K/R all pass: the named conflict draws clicks, the 9-month Meta stack rebuild and Muse Spark safety hold add facts, and open-source safety hits a real practitioner nerve. This is an interview, not a model launch, so it sits in the 78-84 band.

Latent Space

[AINews] Codex Rises, Claude Meters Programmatic Usage

Anthropic changed paid Claude plans to include monthly API credits equal to the subscription price, so a $200 plan includes $200 for programmatic usage outside Anthropic-owned harnesses, while OpenAI promoted Codex enterprise switching incentives in the same news cycle.

Why it matters: HKR-H/K/R all pass: the story ties Claude metering to Codex competition and gives a concrete $200 credit detail. This is a meaningful developer-cost update, not a major model or capability launch, so it sits in mid featured.

AI HOT (Curated Pool)

WeChat Group Chat Summary Skill Added, Depends on wx-cli Configuration

baoyu-skills added a WeChat group chat summary Skill that depends on wx-cli for data reading; the post provides two GitHub links and says Claude Code plus Claude Opus 4.6 gives the best results.

Why it matters: A small open-source tool update, but the workflow is highly relevant: WeChat data via wx-cli into Claude Code for group summaries. HKR-H/K/R pass; limited detail keeps it at the featured threshold.

AI HOT (Curated Pool)

OpenSquilla Open-Source Project Uses Smart Routing and Local Retrieval to Cut LLM Costs

OpenSquilla combines local model routing, vector retrieval, incremental sending, and cache hits to reduce transmitted tokens by more than 90%, while routing simple tasks to cheaper models and complex tasks to stronger models without spending tokens on the routing decision.

Why it matters: HKR-H/K/R all pass, but the source appears to be a single X project post; repo traction, test setup, and limits are not disclosed. Score lands at the featured threshold for practical open-source cost tooling.

AI HOT (Curated Pool)

xAI launches early beta of Grok Build

xAI launched an early beta of Grok Build for SuperGrok Heavy subscribers, offering a terminal-based coding agent with plan review, parallel subagents for large tasks, and a headless mode for scripting and automation.

Why it matters: HKR-H/K/R all pass: xAI enters terminal coding agents with plan mode, parallel subagents, and headless mode. Early beta access for SuperGrok Heavy keeps it below the 85 same-day must-write band.

The Verge · AI

Microsoft Edge Copilot update uses AI to pull information from across your tabs

Microsoft Edge will let Copilot gather information from all open tabs so users can ask questions, compare products, and summarize articles; the snippet says users can choose which experiences to enable, but the post does not disclose a rollout date.

Why it matters: HKR-H/K/R pass, but the post gives tab-wide reading, product comparison, and summaries without launch timing or deeper execution. This fits the lower featured band for a mid-weight product update.

TechCrunch · AI

Notion just turned its workspace into a hub for AI agents

Notion launched a developer platform that lets teams connect AI agents, external data sources, and custom code directly inside its workspace; the RSS snippet does not disclose pricing, rollout timing, supported models, or limits for the new platform.

Why it matters: HKR-H/K/R all pass, but price, launch timing, and supported models are not disclosed, keeping it in the 72–77 mid-weight product-update band. TechCrunch authority supports featured, not same-day must-write.

AI HOT (Curated Pool)

Best Practices for Computer and Browser Use with Claude

Anthropic published guidance for Claude computer and browser use, with Claude 4.6 API screenshots capped at a 1,568-pixel long edge and 1.15 million total pixels, while Opus 4.7 raises the limits to 2,576 pixels and 3.75 million total pixels.

Why it matters: Anthropic’s first-party Claude computer/browser guide has actionable screenshot limits, not just promo copy. HKR-H/K/R all pass, but this is a practice guide rather than a major model or capability launch, so it sits in the 72–77 band.

AI HOT (Curated Pool)

Claude paid plans will offer monthly coding usage credits

Claude paid plans can claim monthly coding usage credits from June 15, covering Claude Agent SDK, claude -p, Claude Code GitHub Actions, and third-party apps built on the Agent SDK.

Why it matters: HKR-H/K/R all pass: the update names a date, quota mechanism, and covered Claude coding surfaces. Importance stays in the low featured band because this is a billing/access change, not a model release.

AI HOT (Curated Pool)

Introducing Runway Agent

Runway launched Runway Agent, a video creation agent that turns one natural-language conversation into multi-scene videos with narration, dialogue, and music; new free-plan users receive 1,500 credits for their first video.

Why it matters: HKR-H/K/R pass: a notable AI-video vendor ships an agentic multi-scene workflow with a 1,500-credit free plan. Score stays in the 72–77 band because the post is still a vendor announcement without pricing, limits, or independent tests.

AI HOT (Curated Pool)

Anthropic Launches Claude for Small Business Package

Anthropic launched Claude for Small Business with connectors and 15 ready-made automation workflows for QuickBooks, PayPal, HubSpot, and related business tools; users run tasks through Claude Cowork and manually approve key steps.

Why it matters: HKR-H/K/R all pass: the Anthropic SMB bundle has 15 workflows, named connectors, and a manual approval mechanism. It is a substantive Claude product update, but pricing, rollout scope, and usage data are not disclosed, so it stays below must-write.

May 13Wednesday

AI HOT (Curated Pool)

Open-source psql_bm25s speeds up PostgreSQL retrieval for multi-agent systems by 23x

The team open-sourced psql_bm25s, a native PostgreSQL access method for exact BM25 retrieval, and says it runs about 23x faster than pg_search on standard benchmarks.

Why it matters: HKR-H/K/R pass via the 23x retrieval-speed hook, named Postgres access method, and RAG latency pressure. Single-source release details lack independent reproduction and production constraints, so it stays in the lower featured band.

TechCrunch · AI

WhatsApp Adds an Incognito Mode in Meta AI Chats

WhatsApp added an incognito mode for Meta AI chats; Meta says these conversations are not saved, and messages disappear by default once the chat is closed.

Why it matters: HKR-H/K/R all pass: the privacy hook is clear, the retention mechanism is concrete, and WhatsApp gives it scale. Still, this is a single product feature, not a model or platform shift, so it sits at the featured threshold.