Skip to content

#Agent

39 today

May 22Friday

AI HOT (Curated Pool)

Codex Enables Secure Cross-Device Mac Control Around the Clock

OpenAI Devs says Codex can use apps on a Mac from a phone while the Mac remains locked and the screen is off; the post does not disclose permission boundaries, pricing, or a release timeline.

Why it matters: HKR-H/K/R all pass: OpenAI Devs disclosed a concrete Codex Mac-control condition. Missing permission boundaries, pricing, and launch timing keep it below the 85+ band.

AI HOT (Curated Pool)

Kotlin ADK and Android ADK 0.1.0 Released for Building AI Agents

Google released Kotlin ADK and Android ADK 0.1.0 for developers, with Kotlin ADK targeting backend agent workflows and Android ADK providing mobile-specific functions for building AI agents.

Why it matters: Google’s Kotlin ADK and Android ADK 0.1.0 release is a mid-weight agent tooling update. HKR-H/K/R pass, but the disclosed facts stop at platforms and version, with no performance data, examples, or ecosystem scale.

Hacker News front page

Launch HN: Runtime (YC P26) – Sandboxed coding agents for everyone on a team

Runtime launched open-source sandbox infrastructure for coding agents, supporting Claude Code, Codex, Cursor, Copilot, Gemini, and Devin, with hosted access, a free tier, and pricing based on a flat platform fee plus compute without token markup.

Why it matters: HKR-H/K/R pass: this is not a major-lab launch, but open-source sandboxes, six coding-agent types, and no token markup give teams concrete adoption signals. No usage data or marquee customers keeps it near the featured floor.

May 21Thursday

AI HOT (Curated Pool)

Anthropic Is About to Become the First Profitable AI Lab

The Wall Street Journal says Anthropic is nearing its first profitable quarter, with expected second-quarter revenue of $10.9 billion and operating profit of $559 million.

Why it matters: HKR-H/K/R all pass: the WSJ-sourced profitability claim has a clear hook, concrete revenue/profit numbers, and strong resonance around AI economics. It is a must-write business story, but still forecasted, not a finalized filing.

r/LocalLLaMA

Agent Execution Tax: New Procurement Metric for Browser Agent Benchmarks?

Fireworks ran 720 browser-agent tasks on WebVoyager and reported a 22.9% Agent Execution Tax, defined as wasted over productive inference; MiniMax M2.5 cost 2.3x less per successful task than Gemini, while GLM-5 reached 57.1% accuracy and Kimi K2.5 had 0% parse retries across 852 calls.

Why it matters: HKR-H/K/R all pass: the post adds a named procurement metric plus concrete benchmark numbers. Source scope is Reddit/Fireworks, so it stays in the 72–77 featured band rather than 78+.

MIT Technology Review · AI

Anthropic’s Code with Claude Showed Off Coding’s Future—Whether You Like It or Not

Anthropic used its two-day Code with Claude event in London to show Claude Code automation, with nearly half the room saying they shipped a pull request fully written by Claude in the past week, and many keeping their hands raised when asked whether they had shipped it without reading the code.

Why it matters: HKR-H/K/R all pass: the MIT Tech Review piece has a strong Claude Code hook, a concrete developer-behavior number, and clear resonance for programmers. It is not a model release or major product launch, so it stays in the 78–84 band.

The Verge · AI

I Can’t Believe How Fast Google Vibe Coded My First Android App

The Verge’s Sean Hollister used Google AI Studio to generate three Android apps in one afternoon; one app came from a 148-word browser prompt and installed about 10 minutes later on an Android phone prepared with USB debugging and a PC connection.

Why it matters: HKR-H/K/R all pass: the story has a personal-test hook plus concrete timing and prompt details. This is not a major Google launch, so it fits the high-quality first-person experiment band, not same-day must-write.

AI HOT (Curated Pool)

Lessons from Building Cloud Agents

Cursor summarizes lessons from building cloud agents: after migrating to Temporal, reliability rose above 99.9%, and the platform processes more than 50 million operations per day.

Why it matters: HKR-H/K/R all pass: Cursor is central to coding agents, and the post gives Temporal, 99.9%+ reliability, and 50M daily operations. Not a launch, so it stays at low-end featured.

Alibaba Technology · WeChat

Building an Agent from 0 to 1: Principles and Personal Assistant Practice

Zhan Xupeng published a roughly 50-minute article on Agent theory and a personal assistant implementation, covering memory, ReAct planning, progressive skill loading, subagents, and harness-level fault recovery.

Why it matters: HKR-K/R pass via concrete agent mechanisms and practitioner reliability pain; HKR-H is weak because the headline is a standard tutorial frame. This fits the quality-tutorial threshold, not the 78+ news band.

Xinzhiyuan · WeChat

Anthropic Acquires SDK Toolmaker Stainless, Leaving OpenAI and Others to Maintain SDKs

Anthropic has completed its acquisition of Stainless, an SDK generation company used by OpenAI, Anthropic, Meta, Cloudflare, and other infrastructure vendors; Stainless says prior SDK ownership remains with customers, but it will shut down hosted products including SDK generator and stop providing ongoing support.

Why it matters: HKR-H/K/R all pass: the deal targets API SDK generation, names OpenAI/Meta/Cloudflare as customers, and says hosted products will shut down. Anthropic bump applies, but this is not a model or core capability release, so it fits 78–84.

AI HOT (Curated Pool)

Tencent Launches OS-Level AI Assistant Mavis on Windows, Mac, and Android

Tencent launched the OS-level AI assistant Mavis on May 21 across Windows, Mac, and Android, with document parsing, image recognition, system maintenance, partial offline use, model dispatching, and desktop control of mobile apps listed as supported functions.

Why it matters: HKR-H/K/R all pass: Tencent’s OS-level assistant spans Windows, Mac, and Android with concrete tool abilities. Model, pricing, and permission design are not disclosed, so it stays at the lower featured band.

Latent Space

Railway: The Agent-Native Cloud — Jake Cooper

Railway serves 3 million users with a 35-person team, adds about 100,000 signups per week, has raised $124 million, and has moved most workloads to its own bare-metal data centers with a reported three-month payback versus rented cloud capacity.

Why it matters: HKR-H/K/R all pass: the Railway interview has concrete growth, funding, and bare-metal details tied to agent infrastructure. It is a strong practitioner story, not a core model or major AI product release, so 74 fits the featured threshold.

r/LocalLLaMA

What happened to Cohere’s Command-A series of models?

Cohere launched Command A+, describing it as its first MoE model under the Apache 2.0 license, with quantization work that lets it run well on 1 or 2 GPUs; the post says top-line performance still needs work.

Why it matters: HKR-H/K/R pass: Cohere open model news has clear local deployment facts. Reddit-level sourcing and missing parameter count, benchmarks, and context window keep it in the low featured band.

AI HOT (Curated Pool)

Google Stitch update: AI design assistant supports end-to-end building

Google updated its AI design partner Stitch with real-time streaming design builds, direct edits and feedback, codebase or Design.md imports, dynamic UI generation, shareable URL exports, and global availability.

Why it matters: HKR-H/K/R pass: Google Stitch adds streaming builds, codebase/Design.md import, and global access. It stays at the featured threshold because model details, pricing, and measured output quality are not disclosed.

The Verge · AI

Google Search’s AI Evolution Includes More Ads

Google is adding Gemini-generated product explanations to Search ads for product queries, including Sponsored Product placements and some ads with built-in chatbots; the snippet cites a compact espresso pod machine example, but the post does not disclose rollout scope or pricing.

Why it matters: HKR-H/K/R all pass: Google is inserting Sponsored Product units and ad chatbots into Gemini shopping results. Scope, pricing, and performance data are not disclosed, so this stays in low featured rather than a major product-release band.

May 20Wednesday

The Verge · AI

If Google Can’t Make AI Agents Useful, Maybe No One Can

The Verge says Google announced multiple AI agents at I/O 2026 for information gathering, event planning, and inbox or calendar summarization; the RSS snippet says the agents can run continuously in the background, but the post does not disclose launch timing, pricing, or evaluation results.

Why it matters: HKR-H and HKR-R are strong because the Verge frames Google agents as a sector test; HKR-K is limited to background-running agents. Missing launch timing, pricing, and evals keeps it in the lower featured band.

TechCrunch · AI

Figma adds an AI assistant to its collaborative canvas

Figma added an AI agent to its collaborative canvas, letting users use natural-language prompts to create new designs, edit existing ones, or automate tasks such as generating design iterations.

Why it matters: HKR-K and HKR-R pass: Figma puts an AI agent into a core collaborative design surface. Details stop at capability scope, with no model, pricing, or rollout timing, so this sits at the featured threshold.

Alibaba Technology · WeChat

Zhenwu M890 AI Chip Debuts as Agentic Compute Foundation

Alibaba released a 128-card supernode server based on T-Head’s Zhenwu M890 AI chip, with P2P latency below 150 ns and rack bandwidth at the Pb/s level; it is live on Alibaba Cloud Bailian and supports Qwen, DeepSeek, and Kimi.

Why it matters: HKR-H/K/R all pass, but the source is Alibaba’s own tech post and lacks third-party benchmarks, pricing, or production volume. Score stays in the featured-threshold band for an AI infrastructure product update.

Xinzhiyuan · WeChat

Behind Jensen Huang’s Douzhi Moment, Chinese GPUs Are Filling CUDA’s Moat

Moore Threads presented progress on its MUSA GPU ecosystem, with SDK 5.1.0 targeting CUDA 12.8 and supporting 761 driver and runtime APIs. The post says MUSA has entered SGLang’s mainline, is listed for 2026 Q2 hardware support, and supports automated library migration via MUSACODE.

Why it matters: HKR-H/K/R all pass: the headline has a meme hook, the post gives 761 APIs plus SGLang mainline support, and CUDA-lock-in anxiety is real. It remains a single-vendor ecosystem update, so it sits in mid featured rather than P1.

Xinzhiyuan · WeChat

UISEE Lists in Hong Kong as a Full-Scenario L4 Autonomous Driving Stock

UISEE listed on the Hong Kong Stock Exchange at HK$60.30 per share, with its public offering oversubscribed 6,777.29 times and a 90.5% share of the Greater China airport L4 commercial vehicle market in 2025.

Why it matters: HKR-H/K/R all pass: the IPO hook is concrete, with subscription and market-share numbers, and it ties to AV commercialization. It stays below 85 because this is not a foundation-model company IPO.

Synced · WeChat

After I/O, Google turns the search box into an agent entry point

Google announced Gemini 3.5 Flash at I/O and added AI Mode directly to Search; the company said its AI services now process over 3.2 quadrillion tokens per month, with more than 8.5 million developers using Gemini.

Why it matters: HKR-H/K/R all pass: Google I/O combines a model update, Search distribution, and concrete usage numbers. AI Mode inside the search box is heavier than a routine feature release, so it clears the same-day must-write band.

Latent Space

Google I/O 2026: Gemini 3.5 Flash, Omni, Spark, and Antigravity 2.0

Google announced Gemini 3.5 Flash at I/O 2026 with a 1M-token context window, 65k max output, four thinking levels, and pricing of $1.50 per 1M input tokens and $9.00 per 1M output tokens.

Why it matters: HKR-H/K/R all pass: this is a Google I/O model-and-product bundle with concrete context, output, thinking-tier, and pricing facts. It has same-day relevance for Claude, OpenAI, and coding-agent competition, so it clears P1.

AI HOT (Curated Pool)

Microsoft reportedly warns internally that GitHub faces existential risk as AI coding tools reduce hosting need

Microsoft internally warned that GitHub faces an existential risk from AI coding assistants such as Cursor and Claude Code, and told some teams to stop using Claude Code by the end of June 2026 and move to GitHub Copilot CLI.

Why it matters: HKR-H/K/R all pass: the angle is sharp, the summary gives a Claude Code-to-Copilot CLI deadline, and the workflow stakes are real. Single-source “reported” framing and no Microsoft response keep it below the 85 must-write band.

AI HOT (Curated Pool)

Qwen3.7: Agent Frontier

Qwen Studio released Qwen3.7 with chatbots, image and video understanding, and image generation. It also covers document processing, web search integration, tool calling, and artifact generation. The RSS snippet frames it as an agent-focused model, but the post does not disclose context length. It also omits benchmark scores, pricing, API limits, release schedule, and reproducible evaluation conditions.

Why it matters: HKR-H/K/R all pass: this is a Qwen flagship-model update with concrete capability coverage. Lack of benchmarks, pricing, and context-window details keeps it at the low end of the 85–94 band.

AI HOT (Curated Pool)

Gemini launches personal AI agent and Daily Brief

Gemini added Gemini Spark and Daily Brief: Spark acts as an always-on personal AI agent across Gmail, Google Docs, and Slides after user authorization, while Daily Brief is available to Google AI subscribers in the U.S. aged 18 or older.

Why it matters: HKR-H/K/R all pass: Google is adding Gemini Spark’s authorized actions across Gmail, Docs, and Slides, plus Daily Brief eligibility for US 18+ AI subscribers. This is a same-day Google agent product update.

The Verge · AI

Google’s AI Future Demands Trust — and Your Personal Data

Google presented Gemini Spark, Daily Brief, and expanded Gmail AI inbox access at I/O 2026; the Verge snippet says these tools depend on large amounts of personal information, but the post does not disclose detailed data-handling terms.

Why it matters: HKR-H/K/R all pass, but the body gives product names and a personal-data dependency without data-handling details. Google I/O makes it featured, not a same-day must-write.

Financial Times · Technology

Google to Release Smart Glasses and Add AI Agents to Search Engine

Google will release smart glasses and add AI agents to its search engine; CEO Sundar Pichai says features powered by a new Gemini model will narrow the gap with Anthropic and OpenAI, while the RSS snippet does not disclose specs, launch timing, or pricing.

Why it matters: HKR-H/K/R all pass: Google is moving Gemini agents into Search and smart glasses, a core entry-point product story. Missing specs, pricing, and timing keep it below the top band, but it fits the 85–94 must-write range.

AI HOT (Curated Pool)

Production Guide for Claude Operating Real User Interfaces

ClaudeDevs published a production guide for Claude computer use, and the snippet lists four mechanisms: click accuracy, thinking effort level selection, context retention in long sessions, and replayable demonstration logging.

Why it matters: HKR-H/K/R all pass: a practical Claude UI-control guide with 4 concrete mechanisms. It is not an official model or product release, so it fits the quality-tutorial band rather than same-day must-write.

AI HOT (Curated Pool)

Smarter Google AI Edge Gallery: MCP Integration, Notifications, and Session Continuity

Google AI Edge Gallery adds experimental MCP support on Android, letting Gemma 4 coordinate external data sources including Google Workspace and Google Maps; the update also adds scheduled notifications and persistent chat history for faster restoration of long-session context.

Why it matters: HKR-H/K/R all pass: Google’s developer update adds experimental MCP, notifications, and session continuity to AI Edge Gallery. It is a mid-weight product update, not a model release or major capability launch.

TechCrunch · AI

Google takes a page from Meta, announces audio-powered smart glasses at I/O 2026

Google announced “audio glasses” at I/O 2026, letting users issue voice commands across its apps and services, including Gemini; the RSS snippet does not disclose price, launch timing, or hardware specifications.

Why it matters: HKR-H/K/R pass: Google announced Gemini-linked audio glasses at I/O 2026, a credible AI-hardware platform move. Missing price, launch date, and specs keep it in the low featured band.

AI HOT (Curated Pool)

Google I/O Announces Multiple Gemini Updates

Google announced multiple Gemini updates at Google I/O, including a new experience design using neural expression technology, upcoming agent features with Daily Brief and Gemini Spark, plus Gemini Omni and 3.5 Flash; the RSS snippet does not disclose release dates, model parameters, pricing, or benchmark results.

Why it matters: HKR-H/K/R all pass: Google I/O brings concrete Gemini hooks in Omni, 3.5 Flash, and agents, with clear competitive resonance. Missing launch timing, specs, and pricing keep it in the 78–84 band.

AI HOT (Curated Pool)

Gemini 3.5 Released: A New Model Family Combining Intelligence and Action

Google AI Developers announced the Gemini 3.5 model family, saying it combines intelligence with action capabilities; the post does not disclose parameters, benchmarks, pricing, availability, or context window details.

Why it matters: HKR-H and HKR-R pass: an official Gemini 3.5 family launch has flagship-model pull and competitive resonance. HKR-K fails because the post gives no params, benchmarks, pricing, or context window, so this stays below the 85+ band.

AI HOT (Curated Pool)

Empirical Research Assistant ERA: From Nature Publication to Computational Discovery

Google Research published its Gemini-based Empirical Research Assistant in Nature and opened early access through the Google Labs trusted tester program.

Why it matters: HKR-H/K/R all pass: Google moves Gemini-based ERA from a Nature paper to a Labs trusted-tester trial. Score stays at 78 because the provided text lacks metrics, benchmark setup, or reproducible workflow details.

AI HOT (Curated Pool)

Gemini 3.5 Series Launches With Stronger Agent and Coding Performance

Google AI launched the Gemini 3.5 series with Gemini 3.5 Flash first, stating it targets agent and coding performance; the post does not disclose parameters, pricing, benchmark scores, or context window size.

Why it matters: HKR-H/K/R all pass: Google’s Gemini 3.5 series is a first-tier model update. Sparse details on price, context, and benchmarks keep it at the low end of the must-write band.

AI HOT (Curated Pool)

Google Search gets its biggest redesign in 25 years with AI-driven interaction changes

Google announced at I/O 2026 the biggest Search redesign in 25 years, powered by Gemini 3.5 Flash, with AI Mode exceeding 1 billion monthly active users and query volume doubling each quarter.

Why it matters: Google Search is an internet entry-point product; 1B+ AI Mode MAU and quarterly query doubling put this beyond a routine feature update. HKR-H/K/R all pass, so it lands in p1.

AI HOT (Curated Pool)

Google releases Gemini 3.5 Flash with 55 intelligence score

Google released Gemini 3.5 Flash, raising its intelligence score by 9 points to 55, exceeding 280 output tokens per second, and increasing operating cost by 5.5 times versus the previous generation.

Why it matters: A Google Gemini 3.5 Flash release is a top-lab model update, backed by Artificial Analysis numbers for speed, intelligence, and cost. HKR-H/K/R all pass, with the 5.5x cost jump making it more than a routine launch.

TechCrunch · AI

With Gemini 3.5 Flash, Google bets its next AI wave on agents, not chatbots

Google launched Gemini 3.5 Flash at its annual developer conference, describing it as its strongest coding and agentic AI model yet; the RSS snippet says it can autonomously execute complex tasks and build software from scratch, but the post does not disclose parameters, pricing, or context window.

Why it matters: HKR-H/K/R all pass: a Google model release with an agent-first and code-building claim is same-day material. Missing parameters, pricing, and context window keep it at the low end of the must-write band.

The Verge · AI

Google wants to compete with Anthropic’s Mythos

Google invited select experts at I/O to test the CodeMender API, an AI agent for code security that flags and fixes vulnerabilities; the RSS snippet does not disclose launch timing, pricing, benchmark results, or concrete details about Anthropic’s Claude Mythos Preview.

Why it matters: HKR-H/K/R all pass, but the post only confirms closed expert testing and the flag/fix mechanism; availability, pricing, and eval results are not disclosed, so this stays at the featured threshold.

TechCrunch · AI

Google Search as You Know It Is Over

Google is changing Search from a link list into an AI experience with conversational answers, autonomous agents, and interactive interfaces; the RSS snippet does not disclose a rollout timeline, publisher traffic impact numbers, or product parameters.

Why it matters: HKR-H/K/R all pass: Google is recasting Search around AI answers, agents, and interactive UI. The post lacks rollout timing and traffic numbers, but this is still a major Google core-product update.

AI HOT (Curated Pool)

Google releases Gemini 3.5 Flash for complex agent workflows

Google introduced Gemini 3.5 Flash at Google I/O for long-running agent workflows; it outscored 3.1 Pro on Terminal-Bench and MCP Atlas, runs up to 4x faster than other frontier models, and reaches up to 12x speed gains in Google Antigravity.

Why it matters: HKR-H/K/R all pass: Google launched Gemini 3.5 Flash for long-horizon agents with benchmark and speed claims. This is a same-day major model update, below industry-shaking tier.