Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

1241–1260 of 1,465

Apr 23Thursday

Xinzhiyuan · WeChat

Zhejiang University open-sources multi-agent evolution system OpenStory: Sun Wukong turns the Grand View Garden into an empty city

Zhejiang University open-sourced OpenStory, a multi-agent narrative system, and inserted a Sun Wukong agent into a 1:1 Dream of the Red Chamber sandbox; within minutes, agents fled the scene. The memory module broadcast “Sun Wukong killed innocents,” fear overrode daily logic, and Wang Xifeng’s physical removal cascaded into an empty Grand View Garden. What matters is the fragility of memory and consensus links; the post does not disclose the base models, metrics, or reproducible setup.

Why it matters: HKR-H/K/R all pass: the stress test is vivid, and the story includes a specific memory-broadcast failure mode with clear agent-safety relevance. Missing model details, metrics, and reproducible setup keep it in the good-featured band, not 85+.

Bloomberg Technology

Alibaba Adds China Eastern Flight Booking to Flagship Qwen App

Alibaba added China Eastern flight booking to the Qwen app, letting users book flights directly; the snippet says this is the first time its agentic AI tech has opened to a major commercial partner. The RSS snippet does not disclose launch regions, fare classes, payment flow, or revenue terms. The real signal is Qwen moving from chat entry to transaction flow, not just another assistant feature.

Why it matters: Featured on HKR-H/K/R: Qwen moves from answers to booking, a concrete agent-commerce step. Kept at 76 because the brief does not disclose rollout scope, payment flow, rev-share, or fulfillment details.

X · @dotey

Microsoft makes Copilot Agent Mode the default in Word, Excel, and PowerPoint

Microsoft made Copilot Agent Mode the default experience in Word, Excel, and PowerPoint, available now to Microsoft 365 Copilot and Premium subscribers, including Personal and Family plans. Microsoft reports internal test gains: Excel engagement up 67% and approval up 65%, Word engagement up 52%, and PowerPoint new-user retention up 36%. What matters is that multi-step in-document execution is now the default, with preview and rollback controls.

Why it matters: This clears HKR-H/K/R: the default flip is the hook, the post includes rollout scope, concrete pilot metrics, and preview/keep/revert controls, and the Office distribution angle will get discussed. I keep it at 84 because this is a large product update, not a new model release or

Hugging Face Blog

How to Use Transformers.js in a Chrome Extension

Hugging Face published a guide for a Transformers.js Chrome extension using Gemma 4 E2B. It defines three MV3 entry points: background service worker, side panel, and content script. The key design keeps local inference in the background and uses messaging plus a tool loop.

Why it matters: HKR-H/K/R all pass, but this is a Hugging Face implementation tutorial, not a model or platform release. Score sits at the featured threshold for a concrete MV3 architecture walkthrough.

Financial Times · Technology

Tesla boosts spending plans to $25bn as Musk doubles down on AI bet

Tesla raised its spending plan to $25bn, with Musk directing more capital toward AI-linked projects. The RSS snippet names self-driving taxis, trucks, robots, and chip factories, and says the increase will be “very significant”; the post does not disclose the time frame, line items, or model details. The key signal is that Tesla is funding a full stack, not just model training.

Why it matters: FT reports a concrete capex jump to $25bn tied to robotaxis, trucks, robots and chip factories. HKR-H/K/R all pass on scale and strategic relevance, but missing timing, line-item spend and model specifics keep it in mid-featured, not must-write.

The Verge · AI

OpenAI now lets teams make custom bots that can do work on their own

OpenAI opened ChatGPT cloud “workspace” agents to Business, Enterprise, Edu, and Teachers plans, letting teams build custom bots for business tasks. Examples include finding product feedback on the web and sending a report to Slack, plus drafting follow-up sales emails in Gmail; the post does not disclose pricing, rollout details, or capability limits. The real shift is workflow execution inside ChatGPT, not a one-off chat reply.

Why it matters: OpenAI moves ChatGPT from chat into team workflow agents, so HKR-H/K/R all pass: the hook is strong, the post adds plan coverage and task examples, and it hits enterprise automation demand. It stops short of P1 because price, rollout detail, and capability boundaries are not disl

X · @dotey

OpenAI launches ChatGPT Workspace Agents for cross-tool enterprise workflow automation

OpenAI released ChatGPT Workspace Agents as a research preview for ChatGPT Business, Enterprise, Edu, and Teachers paid plans. The snippet says agents can connect Slack, Gmail, Google Drive, Salesforce, Notion, Linear, and Atlassian, and take actions like updating tickets, creating docs, and replying in Slack. The key point is admin controls for permissions, approvals, and monitoring; the post does not disclose pricing, quotas, or rollout regions.

Why it matters: This is a substantive OpenAI product update: ChatGPT now offers cross-tool workspace agents with named enterprise integrations, so HKR-H/K/R all pass. I keep it at 86 because it is still a research preview and pricing, quotas, and rollout regions are not disclosed.

Latent Space

Shopify’s AI Phase Transition: 2026 Usage Explosion, Unlimited Opus-4.6 Budget, Tangle

Shopify CTO Mikhail Parakhin detailed its AI stack across 3 projects: Tangle, Tangent, and SimGym. The post says Shopify is a 20-year, $200B software company, but does not disclose exact 2026 usage figures. The key shift is from code generation to review, CI/CD, and deployment stability.

Why it matters: HKR-H/K/R all pass: the CTO interview has a clear hook, names internal tools, and maps the coding-agent bottleneck to review and CI/CD. Missing usage numbers keep it in 78–84, not P1.

Hacker News front page

OpenAI: Workspace agents for business

OpenAI is offering Workspace agents in research preview for ChatGPT Business, Enterprise, Edu, and Teachers plans. The page says agents can run on schedules, use tools like Slack, Google Drive, and Microsoft apps, and support approval gates, audit logs, and role-based access control; pricing, model details, and rollout timing are not disclosed.

Hacker News front page

Website streamed live directly from a model

Flipbook generates an entire clickable website in real time with an image model, where each page is a pixel image and every click spawns a deeper image. The post says all on-screen text is drawn by the image model with no HTML or text overlays, and content comes from agentic web search plus model knowledge. The key point is the interaction model, not a standard generative UI; the live video stream remains an experimental, resource-heavy toggle.

Why it matters: HKR-H/K/R all pass: a live, clickable site rendered entirely as model-generated pixels is a strong hook, and the post explains the mechanism (no HTML/text overlay, agentic web search). Kept at 76 because latency, cost, model stack, and usage are not disclosed.

X · @OpenAI

Introducing workspace agents in ChatGPT—shared agents for complex tasks and long-running workflows

OpenAI announced workspace agents in ChatGPT, described as shared agents that work across tools and teams for complex and long-running workflows. Only the title and RSS snippet are disclosed; the post does not disclose supported tools, pricing, access tier, permission model, or rollout timing. The key issue to watch is the collaboration boundary of shared agents, not the headline claim alone.

Hacker News front page

Introducing Parallel Agents in Zed

Zed released Parallel Agents on April 22, 2026, letting multiple agents run in parallel in one window. The new Threads Sidebar sets per-thread folder and repo access, and supports stop, archive, and new-thread actions; the new default layout is opt-in for existing users. The key detail is permission scoping and thread orchestration, not just “multiple agents.”

Why it matters: First-party product update with clear HKR-H/K/R: parallel agents in one window plus thread-level repo and folder access control. It stays in the mid-70s because the post gives no performance delta, pricing impact, adoption data, or external validation; this is still a single-tool

Hacker News front page

Startups Brag They Spend More Money on AI Than Human Employees

Swan AI CEO Amos Bar-Joseph said his 4-person startup spent $113,000 on Claude in one month and treated that bill as headcount budget spent on AI instead of hires. The post says Swan targets $10M ARR with fewer than 10 people and cites Fundable AI claiming AI can replace a 15-person document team; the real signal is that token spend is being used as a growth metric, not proven ROI.

Why it matters: HKR-H lands on the payroll-vs-AI-bill inversion; HKR-K lands on the $113k/month Claude spend from a 4-person team. HKR-R is strong because it speaks to hiring, burn, and replacement anxiety, but this is still a trend piece with a thin sample, not a market-moving event.

Apr 22Wednesday

The Verge · AI

Meta will track employees’ computer activity to train its AI agents

Meta is installing its MCI tool on US employees’ computers and using mouse movements, clicks, keystrokes, and occasional screenshots from work apps and sites to train AI agents. Reuters says the data is meant to teach models to operate computers more like humans and automate tasks employees already do; Meta says it will not be used for performance reviews. The key gap is scope: the post discloses US staff and work contexts, but not retention, opt-out, or full rollout details.

Why it matters: HKR-H lands on the surveillance-for-agents hook. HKR-K lands on concrete collection details: US staff, mouse/keyboard events, occasional screenshots. HKR-R lands on privacy plus job-automation nerves. Strong reporting, but not a shipped product or model release, and key scope/ret

Hacker News front page

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

Qwen released the open-weight 27B dense model Qwen3.6-27B and made it available in Qwen Studio. It scores 77.2 on SWE-bench Verified vs. 76.2 for Qwen3.5-397B-A17B, and 59.3 on Terminal-Bench 2.0 under a 256K context and 3-hour timeout. The real takeaway is deployment: this is not a larger MoE, but a denser 27B model with stronger coding results.

Why it matters: Qwen3.6-27B is a substantive flagship-model release with open weights, concrete coding benchmarks, and a practical dense-deployment angle. HKR-H/K/R all pass, and per policy a major Chinese model launch should score on par with an equivalent US-lab release.

OpenAI News

Introducing workspace agents in ChatGPT

OpenAI introduced workspace agents in ChatGPT, describing them as Codex-powered agents that automate complex workflows in the cloud. The RSS snippet confirms secure work across tools for teams, but the post does not disclose pricing, availability, supported tools, or performance metrics.

Why it matters: This is a substantive OpenAI product update inside ChatGPT. HKR-H lands on the jump from chat to workspace agents, HKR-K on Codex-powered cloud execution across tools, and HKR-R on team workflow automation; the score stops at 86 because pricing, rollout, tool support, and metrics

OpenAI News

Speeding up agentic workflows with WebSockets in the Responses API

OpenAI says WebSockets in the Responses API speed up the Codex agent loop, using connection-scoped caching to cut API overhead and improve latency. The RSS snippet confirms the mechanism, but the post does not disclose latency deltas, throughput numbers, or workload conditions. The key point is transport-layer optimization, not a new model.

Why it matters: This is a developer-facing OpenAI product update at the systems layer: WebSockets plus connection-scoped caching target agent-loop round-trip cost. HKR-H/K/R all pass, but the post does not disclose latency gains, throughput, or workload bounds, so it stays mid-featured rather än

QbitAI · WeChat

SenseAuto's Sage with 3B active params claims to beat GPT-5.4 and Opus 4.6 in cars

SenseAuto released Sage, an in-car multimodal edge model with 32B total params and 3B active params, and says it scored 94% on PinchBench, above Claude Opus 4.6 at 93.3% and GPT-5.4 at 90.5%. The post says Sage runs on Nvidia OrinX with about 0.5s TTFT, 0.03s TPOT, and 80 tok/s throughput; its SCOUT training method cuts GPU hours by about 60%, and ERL raises complex-task completion by 20%. The key point is not the headline race but whether a 3B-active model can sustain multi-step tool use on device.

Why it matters: HKR-H/K/R all pass: the 3B-active-vs-GPT hook is strong, and the post gives concrete OrinX latency, throughput, and benchmark numbers. I keep it at 79 because the evidence is self-reported and the impact is narrower than a general model launch.

Synced · WeChat

Honor preinstalls YOYO Claw on MagicBook, calling it the world's first "agent laptop"

Honor said it preinstalls its YOYO Claw on MagicBook and claims 50% lower total token use than an OpenClaw setup. The post says it ships with 5 primary agents and 23 sub-agents, plus local processing, second-step confirmation, and kernel-level encryption. The practical angle is packaging agents as a device default, but the post does not disclose model names, hardware specs, pricing, or launch timing.

Why it matters: This clears HKR-H/K/R: the factory-installed agent angle is novel, and the post includes concrete details on 5/23 agents, 50% token reduction, local handling, confirmation gates, and kernel-level encryption. It stops at 76 because the model, hardware, price, and ship date are not

The Verge · AI

SpaceX cuts a deal to maybe buy Cursor for $60 billion

SpaceX announced an either-or deal: buy AI coding platform Cursor for $60 billion or pay a $10 billion fee. The RSS snippet says this could help xAI's coding tools chase Anthropic; the post does not disclose the structure, timing, or IPO linkage. Watch the $10 billion breakup fee, not just the tentative acquisition headline.

Why it matters: All three HKR axes land: the headline has a strong unexpected hook, and the report gives two hard facts — a $60B price and a $10B breakup fee. I keep it at featured, not P1, because only top-line terms are disclosed; structure, timing, and the exact xAI linkage are still undiscol