Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

1381–1400 of 1,465

Feb 26Thursday

OpenAI News

Pacific Northwest National Laboratory and OpenAI partner to accelerate federal permitting

OpenAI and Pacific Northwest National Laboratory evaluated coding agents on NEPA drafting tasks from 18 federal agencies, finding 1-5 hours saved per subsection, or about 15% less drafting time. The DraftNEPABench benchmark was designed with 19 experts and covers 102 tasks, using Codex CLI with GPT-5 for long-document synthesis, cross-checking, and structured writing. The key limit is explicit: this measures well-scoped drafting work, not full real-world permitting decisions.

Why it matters: HKR-H/K/R pass: federal permitting is an unusual hook; the post gives 19 experts, 102 tasks, and 1–5 hours saved; the debate is agents entering regulated workflows. Score stays below major product news because this is a scoped benchmark, not a shipped capability.

OpenAI News

OpenAI Codex and Figma launch code-to-design roundtrip workflow

OpenAI and Figma launched a Codex integration on Feb. 26, 2026 that turns code into editable Figma designs and brings Figma Design, Figma Make, and FigJam content back into code. The workflow uses MCP via the Figma MCP Server in the Codex desktop app; OpenAI says Codex has 1M+ weekly users and usage is up 400%+ since the start of the year. The key issue is whether roundtrip context stays intact; the post does not disclose supported models, permission boundaries, or pricing.

Why it matters: This is a solid OpenAI/Figma workflow update with clear HKR-H/K/R: a bidirectional code↔design loop via MCP and Figma MCP Server. It stays below 85 because the post does not disclose model support, permission boundaries, pricing, or roundtrip reliability.

Feb 15Sunday

Computing Life · Yage

OpenClaw deep dive: why it suddenly took off, and what it means for us

OpenClaw surged in late January 2026 because it plugged local coding agents into Slack, WhatsApp, and Feishu, giving non-technical users file access, command execution, and persistent memory in a chat UI. The article also names the costs: 12% of third-party skills contained malicious code, and the $CLAWD token scam took $16 million; the chat interface remains linear, low-density, and hard to observe. The real takeaway is not to copy OpenClaw blindly, but to reuse its unified context, file-based memory, and composable skills in a controllable stack like OpenCode.

Why it matters: This is more than a recap: it breaks down OpenClaw's adoption mechanism, downside, and reusable design pattern. HKR-H/K/R all pass with two hard facts—12% malicious skills and a $16M scam—but as a personal analysis rather than an official release or industry event, it lands in `+

Computing Life · Yage

OpenClaw Deep Dive: Why It Went Viral and What It Means for You

The post says OpenClaw went viral in late January 2026, changed names 3 times in one week, and a $CLAWD scam token took $16 million. It cites two concrete risks: 12% of third-party skills had malicious code, and some users exposed consoles to the public internet without passwords. The excerpt is truncated, but the core claim is distribution: OpenClaw put agentic AI into WhatsApp, Slack, and Lark for non-technical users.

Why it matters: HKR-H/K/R all pass: the viral arc is dramatic, the post includes a 12% malicious-skills figure and a specific exposed-console risk, and the distribution angle matters to agent builders. It is still a secondary deep-dive, not a primary launch or official research, so 78 and tiered

Feb 14Saturday

Ruan YiFeng's Weblog

Using ByteDance's Seed 2.0 and TRAE with Skills for app building and deployment

Ruanyifeng used ByteDance's Seed 2.0 Code and TRAE to generate one ASCII-to-Excalidraw web app and preview it at localhost:8080. The post says Seed 2.0 includes Pro, Lite, Mini, and Code models, and shows Skills as YAML-headed Markdown files, including Anthropic's frontend-design and Vercel deploy examples.

Why it matters: HKR-H and HKR-K land because the post turns Seed 2.0 Code + TRAE into a runnable mini app and explains the Skill mechanism with concrete setup details. HKR-R also lands for coding-agent workflow reuse, but this is a strong tutorial, not a major ByteDance launch, so it sits at the

Feb 12Thursday

Lex Fridman (YouTube RSS)

OpenClaw: The Viral AI Agent Behind the Hype - Peter Steinberger | Lex Fridman Podcast #491

Lex Fridman’s episode #491 interviews Peter Steinberger about the open-source AI agent OpenClaw; the transcript says it reached 175k-180k GitHub stars. The post says it can connect to Telegram, WhatsApp, Signal, and iMessage, and use models such as Claude Opus 4.6 and GPT 5.3 Codex; it does not fully disclose the architecture, evals, or security boundaries. The real point is system-level access and self-modifying behavior: this is not chat, but an agent that can take actions.

Why it matters: This is more than a routine podcast. OpenClaw scores on HKR-H/K/R with 175k-180k GitHub stars, messaging integrations, and self-modifying behavior. It stays at featured, not p1, because the post does not disclose architecture, evaluations, or safety boundaries.

Ruan YiFeng's Weblog

Hands-on with Zhipu's flagship GLM-5: compared with Claude Opus 4.6 and GPT-5.3-Codex

Ruan Yifeng compared GLM-5, Claude Opus 4.6, and GPT-5.3-Codex on 4 coding tasks, and judged GLM-5 competitive with the two closed models overall. The post covers web redesign, a 3D sandbox, an Angry Birds clone, and Laravel-to-Next.js migration; in the migration task, GLM-5 and GPT-5.3 took about 5 minutes, while Opus 4.6 took about 20. The key point: this is a single-author hands-on comparison, not a standardized benchmark.

Why it matters: This clears HKR-H/K/R because it is a named first-person test with 4 tasks, video evidence, and a 5-minute versus ~20-minute gap. I did not score it higher because it is one author's evaluation, not a standardized benchmark or a broad multi-source release event.

MIT Technology Review · AI

Is a secure AI assistant possible?

OpenClaw was uploaded to GitHub in November 2025 and went viral in late January, extending LLMs into email, browsing, and local files with larger security risks. The post names prompt injection as the central threat, says there are likely “hundreds of thousands” of OpenClaw agents online, and notes a public warning from the Chinese government. The key point: the article says there is no silver-bullet defense yet, and the truncated body does not disclose the full mitigation details.

Why it matters: This is not a launch, but it clears HKR-H/K/R: the question is a strong hook, the piece adds concrete scale plus 'no silver-bullet' defense, and it hits the agent-builder safety nerve. Featured, not p1, because the article does not disclose reproducible mitigations.

Feb 10Tuesday

36Kr (direct RSS)

Embodied AI company Noematrix raises several hundred million yuan in Series A, with overseas funds joining

Noematrix closed a Series A worth several hundred million yuan, led by C Capital, with Sea Limited and Puhua Capital participating, and Prosperity7 Ventures increasing its stake. Founded in Nov. 2023, the company says its Noematrix Brain has been deployed on wheeled single-arm, wheeled dual-arm, and humanoid dual-arm robots in retail pharmacies and hotel laundries; the post does not disclose valuation or revenue. The sharper signal is its claimed hundreds of thousands of hours of real-robot data and its data-model-scenario loop.

Why it matters: HKR-H/K/R all pass: the funding hook is strong, and the body adds real-world data plus deployed robot forms and scenarios. It stays at the low end of featured because this is still a single-company financing scoop, and valuation, revenue, and customer counts are not disclosed.

MIT Technology Review · AI

Why the Moltbook frenzy was like Pokémon

MIT Technology Review compares the Moltbook AI-agent social experiment to 2014 Twitch Plays Pokémon: lots of spectacle, limited signal about the future. The post cites 1 million concurrent players in the Pokémon case; Moltbook also mixed in crypto scams, and some “agent” posts were actually steered by humans. The real gap is explicit: shared memory, coordination, and shared goals are still missing.

Feb 9Monday

36Kr (direct RSS)

Qwen’s 10 Million Milk Teas: How Alibaba’s Massive AI Freebie Campaign Unfolded

Alibaba’s Qwen drove over 10 million orders via a Feb. 6 free-order campaign, but the app slowed and crashed from 10 a.m. to noon as load exceeded capacity; orders had already passed 2 million before noon. 36Kr says initial server capacity was only about one-third of the expected peak, and the subsidy pool was framed as 3 billion yuan; the real signal is not a model leap but a paid test of AI commerce entry and consumer acquisition.

Why it matters: HKR-H lands on the free-milk-tea plus outage hook, while HKR-K lands on concrete scale and capacity numbers. HKR-R also lands because the story speaks to AI distribution, subsidy economics, and infra reliability, but it remains a single-company promo test rather than a market-shi

36Kr (direct RSS)

Former Baichuan co-founder Jiao Ke bets on AI audio to build AI hosts

Jiao Ke said Laifu Radio now has 15 Chinese AI hosts and 2 English ones, and raised over $10 million across two rounds by H2 2025. He said users average about 30 minutes per day, AI can prepare timely audio in under an hour, and the team treats DTU plus long-memory infra as the key moat. The real bet is not an AI podcast tool but interactive AI hosts that remember user preferences; the post also says it is working with some automakers on in-car personalized AI radio.

Why it matters: HKR-H lands because the story reframes audio AI as persistent hosts, not a podcast tool. HKR-K is strong on numbers and mechanism; HKR-R lands via memory plus in-car distribution. Early-stage company scope keeps it at featured, not p1.

Feb 7Saturday

MIT Technology Review · AI

Moltbook was peak AI theater

Moltbook went viral within hours, and the platform says it now has 1.7 million agent accounts, 250,000 posts, and 8.5 million comments, but the article argues the activity is mostly human-scripted mimicry. It says OpenClaw can connect Claude, GPT-5, or Gemini to tools like email and browsers; cited operators say the agents lack shared goals, shared memory, and self-directed autonomy, and some viral posts were written by humans posing as bots. The key takeaway is risk: agents tied to private data such as passwords or bank details were active on a site filled with spam and potentially malicious instructions.

Why it matters: This is strong anti-hype commentary, not a market-moving event. HKR-H/K/R all pass: the hook is sharp, the piece adds 1.7M/250k/8.5M plus concrete critique on memory and goals, and the security angle lands with practitioners, so it clears featured but stays mid-70s.

Feb 6Friday

TechCrunch · AI

OpenAI launches new agentic coding model minutes after Anthropic releases its own

OpenAI launched an agentic coding model minutes after Anthropic released a similar one, and the model is meant to accelerate Codex, which OpenAI launched earlier this week. The RSS snippet gives only the timing and purpose; the post does not disclose the model name, benchmarks, pricing, context length, or availability. The signal is direct competition in agentic coding, not a substantiated performance claim.

Why it matters: Major-lab product news plus a minutes-apart Anthropic clash gives this HKR-H and HKR-R. The score stays in the low featured band because HKR-K is weak: the post lacks the model name, benchmarks, price, context window, and availability.

Feb 4Wednesday

TheValley101 (硅谷101)

E224 | Why Clawdbot became the first breakout product of 2026 amid the Mac mini rush | Moltbot | MoltBook | OpenClaw

The podcast says Clawdbot passed 100k GitHub stars within days and reached 146k on Feb. 2, while being renamed to Moltbot and then OpenClaw within a week. It attributes the traction to a stack of Claude, long-term memory, IM-based messaging, and proactive heartbeat workflows; the title mentions a Mac mini rush, but the post does not disclose sales figures. The real signal is the interaction layer rather than a new model release: this is industry commentary and user anecdotes, not an official spec sheet.

Why it matters: This is a commentary-led breakdown of a hot agent phenomenon, not a primary launch. HKR-H/K/R all pass: the 146k-star surge and rename chain are novel, the post explains memory + IM + heartbeat mechanics, and it hits nerves on agent UX, dedicated hardware, and security bills; the

Feb 3Tuesday

Computing Life · Yage

Beyond Tutorial Thinking: Why AI Education Should Add Engineering Infrastructure, Not Just Content

The team says it ran 4 courses over 2 years for 2,500+ learners, yet only a minority shipped usable products; drop-off centered on setup, experimentation, deployment, and context handling friction. The post says AI Builder Space gives students a no-card unified API, one-click deployment to <name>.ai-builders.space free for 1 year, and MCP access for Cursor and Claude Code via one command. The point is productized teaching infra, not more tutorials; retention, conversion, and cost are not disclosed.

Why it matters: The piece turns a familiar complaint into operational detail: 2500+ learners, 4 failure points, and a concrete platform response with API, deployment, and MCP access. HKR-H/K/R all pass, but missing conversion, retention, and cost data keeps it at the low end of featured.

Feb 2Monday

Import AI (Jack Clark)

Import AI 443: Into the Mist: Moltbook, Agent Ecologies, and the Internet in Transition

Jack Clark writes that Moltbook has pushed AI agents into a public social network at tens-of-thousands scale, shifting conversation from humans to agents. He says it combines an agent social feed with OpenClaw-style computer access, but the post does not disclose active-agent, retention, or transaction metrics. A separate July 2025 workshop report says closed-loop AI R&D automation could raise productivity from 10x to 100x to 1000x; the key issue is measurement and outside transparency.

Why it matters: Featured: HKR-H/K/R all pass. The post has a strong hook—a public social space filled by agent ecologies—and a concrete 10x/100x/1000x closed-loop R&D claim, but it lacks Moltbook activity, retention, and transaction data, so it stays at 78.

Feb 1Sunday

Lex Fridman (YouTube RSS)

State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast #490

Lex Fridman, Sebastian Raschka, and Nathan Lambert discuss the 2026 AI race in podcast #490 and frame DeepSeek R1’s January 2025 release as a key inflection point. The episode names Claude Opus 4.5, Gemini 3, Z.ai GLM, Minimax, and Kimi Moonshot, but the post does not disclose a shared benchmark, cost table, or reproducible eval. The useful takeaway is the lens: gaps look more like compute, budget, and org culture than secret ideas.

Why it matters: High-quality commentary, not a news break. HKR-H and HKR-R pass because Lex Fridman, Sebastian Raschka, and Nathan Lambert frame China, agents, GPUs, and AGI for practitioners. HKR-K misses: the post names models and DeepSeek R1 but provides no shared benchmarks, cost table, or a

Jan 30Friday

Ruan YiFeng's Weblog

Technology Enthusiast Weekly #383: What Level of AI Programming Are You?

Steve Yegge frames AI coding into 8 levels and says he is at level 8, where an orchestrator manages parallel AI coding sessions. The post lays out a path from IDE copilots to YOLO acceptance, 3-5 windows, 10+ windows, then orchestration; it also says his AI-built tool Gas Town has 225,000 lines of Go code, which he has never read, and had 6,000 stars as of last week. The real signal is black-box programming as a workflow choice, with cost and failure risk stated plainly.

Why it matters: Strong HKR-H/K/R: the 8-level framing is sticky, and the post carries concrete workflow and project numbers. The score stays below 78 because this is secondary commentary, not a primary model, product, or research release.

Jan 29Thursday

Ruan YiFeng's Weblog

Kimi’s integrated stack vs. Manus’s layered approach

Kimi released the K2.5 model and K2.5 Agent together, with an agent mode already available on its website. The post cites 1,500-step long-horizon actions, up to 100 agents in parallel, and visual coding from design files or web videos; pricing, context window, and API terms are not disclosed. The key point is product shape: not just a model launch, but a bundled model-plus-agent release.

Why it matters: HKR-H lands on the integrated release angle; HKR-K lands on the 1,500-step, 100-agent, visual-programming details; HKR-R lands on the stack-design debate. Missing price, context window, and API terms, plus a commentary source, keep it below p1.