Skip to content

All news

25 today

May 13Wednesday

Hacker News front page

Show HN: Rotunda - A Browser Built for Agents with Simulated Typing

Pierce released Rotunda, a Firefox 150-based browser for agents that simulates mouse and keyboard timing with an RNN trained on one week of his own patterns, and exposes local control through a CLI or Playwright API for Claude, Codex, or other harnesses.

Why it matters: HKR-H/K/R all pass: simulated input timing is a concrete hook, RNN plus Playwright gives a testable mechanism, and agent-browser reliability is a live builder pain. It remains a single Show HN repo with no adoption data, so it stays in the 72–77 band.

r/LocalLLaMA

AIDC-AI/Ovis2.6-80B-A3B on Hugging Face

AIDC-AI released Ovis2.6-80B-A3B, a multimodal MoE model with 80B total parameters and about 3B active parameters at inference, supporting a 64K-token context window and images up to 2880×2880 resolution.

Why it matters: HKR-H/K/R pass: the open multimodal MoE has concrete specs and a real efficiency hook. Score stays near the featured floor because the post gives no benchmarks, license details, or hands-on results.

AI HOT (Curated Pool)

Configuring Development Environments for Agents

Cursor released tools for cloud agent development environments, adding multi-repository support, Dockerfile-based configuration, audit logs, and environment-level network and secret controls; the post says cache hits improve build speed by 70%.

Why it matters: HKR-K and HKR-R pass: Cursor adds concrete cloud-agent environment controls, including Dockerfile setup, audit logs, permissions, and 70% faster cached builds. HKR-H is weaker, so this sits at the lower featured band.

OpenAI News

Building a Safe, Effective Sandbox for Codex on Windows

OpenAI built a secure sandbox for Codex on Windows. The RSS snippet discloses controlled file access and network restrictions, but the post does not disclose implementation details, performance data, or rollout conditions.

Why it matters: OpenAI details a Windows sandbox for Codex with file-access and network controls. It is not a major model release, but HKR-H/K/R all pass because the safety boundary matters for coding-agent adoption.

AI HOT (Curated Pool)

Miaoda App and Enterprise Edition launch with 90% self-generated code

Baidu launched the Miaoda app and Miaoda Enterprise Edition, saying 90% of the Miaoda app’s code was generated by Miaoda itself; Miaoda-generated apps have served over 10 million users and reached a total value of RMB 5 billion.

Why it matters: HKR-H/K/R all pass via the 90% dogfooding hook, concrete adoption/value figures, and coding-tool resonance. Company-source metrics lack independent context, so this stays in the lower featured band.

QbitAI · WeChat

An 8-Year-Old Turns Ideas into Apps as Baidu Launches Miaoda 3.0

Baidu launched Miaoda 3.0 at its 2026 Create conference, adding iOS and Android app generation, Android packaging, online hot updates, and an enterprise edition with three-level permissions, environment isolation, and SLA commitments.

Why it matters: HKR-H/K/R pass: Baidu’s Miaoda 3.0 adds mobile app generation, Android packaging, hot updates, and enterprise controls. This is a solid product update, not a flagship model release or must-write event.

r/LocalLLaMA

The Trillion-Parameter Dilemma: MiMo-V2.5-Pro Open-Sourced at 1.02T Parameters

Xiaomi open-sourced MiMo-V2.5-Pro with 1.02T parameters, 42B active parameters, a 1M context window, and an MIT license; the author ran 125 Claude Code sessions through the API, spending $70.12 for 387,380,436 tokens with a 96.3% cache hit rate.

Why it matters: HKR-H/K/R all pass: a Xiaomi 1.02T open model plus a concrete Claude Code API cost experiment. Reddit sourcing keeps it at the low end of the 85+ band, but the domestic flagship-model signal clears p1.

New York Times Chinese

China Seeks AI Technology Self-Reliance, Weakening Washington’s Leverage Over Beijing

DeepSeek optimized its latest model for inference on Huawei chips for the first time, while two semiconductor sources said training still relies on Nvidia chips; Huawei says it plans to release a training chip this year, but matching current Nvidia performance will take another year.

Why it matters: HKR-H/K/R all pass: NYT ties DeepSeek-Huawei chip optimization and Huawei's training-chip timeline to US export-control leverage. It is not a model launch and lacks benchmark results, so it stays in the 78–84 band.

AI HOT (Curated Pool)

Google launches its first AI-first laptop Googlebook with Gemini integration

Google launched Googlebook, its first laptop designed around Gemini Intelligence, with three disclosed mechanisms: Magic Pointer as an AI interaction entry point, natural-language widget creation, and Android-based cross-device app and file access.

Why it matters: HKR-H/K/R all pass: a Google Gemini-first laptop with 3 named interaction mechanisms. Specs, pricing, launch timing, and demos are not disclosed, so it stays in the 78–84 band.

AI HOT (Curated Pool)

Claude Code adds /goal feature to keep tasks running until completion

Claude Code introduced a /goal feature that keeps Claude working until a task is completed; the post does not disclose the trigger mechanism, supported versions, pricing, or failure conditions.

Why it matters: HKR-H/K/R pass because /goal targets a real Claude Code reliability pain. It is a single-feature Anthropic update with sparse mechanics, so it lands at the lower featured band, not same-day major news.

Bloomberg Technology

CME to Create Futures Market for AI Computing Power

CME Group and Silicon Data are partnering to create a futures market for AI computing power, with the market functionally setting a price for compute backing AI infrastructure; the post does not disclose contract specifications, launch timing, or expected trading volume.

Why it matters: HKR-H/K/R all pass: CME entering AI compute futures is a novel infrastructure-finance angle and hits compute-price risk. Missing contract specs, launch date, and scale keep it near the featured threshold.

AI HOT (Curated Pool)

Codex Enables Background Multitasking Across Apps

OpenAI Devs says computer use lets Codex click, type, and keep working across Mac apps in the background; the post does not disclose release timing, permission design, or availability scope.

Why it matters: HKR-H/K/R all pass: OpenAI Devs gives a concrete Codex mechanism for background Mac cross-app actions, with clear developer relevance. Release timing, permissions, and availability are missing, so it stays at the lower featured band.

AI HOT (Curated Pool)

Step Image Edit 2 image model released with leading performance and efficiency

StepFun released the 3.5B-parameter Step Image Edit 2 model, which ranks first in KRIS-Bench overall, factual, and conceptual categories, and is now available on the Stepfun Open Platform.

Why it matters: HKR-H/K/R all pass: the hook is a 3.5B image-editing model topping KRIS-Bench, with concrete launch details. Vendor-only sourcing and no independent test or pricing keep it at the low featured band.

AI HOT (Curated Pool)

How Anthropic's Cybersecurity Team Uses Claude Code to Build a Threat Detection Platform

Anthropic’s detection platform engineering team used Claude Code to build the CLUE threat detection and response platform, completing a proof of concept in one day and delivery in one week while reducing analyst log investigation from hours to minutes.

Why it matters: HKR-H/K/R all pass, but this is an Anthropic internal dogfooding case rather than a Claude Code capability launch. CLUE and the timing metrics keep it just above the featured threshold.

Bloomberg Technology

CME Plans Computing Power Futures Market

CME Group and Silicon Data plan to create a futures market for computing power; the RSS snippet says Bloomberg Tech discusses the rationale and mechanics, but the post does not disclose contract specifications, launch timing, or pricing methodology.

Why it matters: HKR-H and HKR-R pass: CME moving into compute futures is a fresh hook and targets AI compute-cost anxiety. HKR-K is weak because contract specs, timeline, and pricing method are not disclosed.

AI HOT (Curated Pool)

Claude Opus 4.7 Fast Mode Opens Research Preview

Claude Opus 4.7 Fast Mode is now available as a research preview in the API and Claude Code. The post does not disclose model parameters, pricing, rate limits, or a general availability date.

Why it matters: HKR-H/K/R pass because this is a Claude fast-mode preview in API and Claude Code, directly tied to developer latency and workflows. Thin disclosure on pricing, limits, parameters, and GA timing keeps it at the featured threshold, not 78+.

AI HOT (Curated Pool)

Claude Enters the Legal Industry

Anthropic released more than 20 MCP connectors and 12 legal plugins, letting Claude work inside Word and Outlook for contract drafting, revision, clause comparison, and routine legal workflows.

Why it matters: HKR-H/K/R all pass: a substantive Anthropic vertical product update with 20+ MCP connectors and Office workflows. It is not a model release or platform-wide capability, so it stays in the 72–77 band.

AI HOT (Curated Pool)

GitHub Copilot Individual Plans Add Flex Allotments and a New Max Plan

GitHub will update Copilot individual plans on June 1 by adding flex allotments to Pro and Pro+ and introducing a new Max plan; the post does not disclose pricing, quota limits, or the exact allocation rules in the provided snippet.

Why it matters: HKR-H/K/R all land lightly because Copilot plan quotas affect many developers. Missing price, caps, and allocation rules keep it at the low featured threshold, not a major capability update.

AI HOT (Curated Pool)

Google Launches New Android Smart Assistant

Google introduced Android Intelligence at Android Show 2026, with multi-step automation across Android apps, browser-use features for Gemini in Chrome, automatic form filling, Rambler voice-note transcription, and custom Gen UI widgets; the post does not disclose rollout timing, supported devices, or pricing.

Why it matters: HKR-H/K/R all pass: the hook is Android-level agent control, the new facts are concrete automation surfaces, and the resonance is the mobile AI platform fight. Thin source detail keeps it at the low end of the 85-94 band.

Hacker News front page

Show HN: Agentic Interface for Mainframes and COBOL

Hypercubic launched Hopper, an agentic development environment that combines a real TN3270 terminal, z/OS-aware panels for datasets, jobs, and spool output, and an AI agent; sensitive operations require approval, and the terminal remains visible during agent actions.

Why it matters: HKR-H/K/R all pass: the mainframe-agent angle is novel, with concrete TN3270, z/OS, and approval mechanics. Small-vendor Show HN status and missing customer/pricing/results data keep it at the featured floor.

AI HOT (Curated Pool)

Build long-running AI agents that pause, resume, and never lose context with ADK

Google Developers describes using ADK to build long-running agents for enterprise workflows lasting days or weeks, such as HR onboarding, with a persistent state machine, persistent session storage, event-driven webhooks, and multi-agent delegation to pause during idle time and resume after restarts without losing context.

Why it matters: HKR-H/K/R all pass: the ADK tutorial gives concrete persistence mechanisms for long-running agents. It is useful engineering guidance from Google Developers, not a major model or platform release, so it sits at the 72–77 featured threshold.

TechCrunch · AI

Everything Google announced at its Android Show, from Googlebooks to vibe-coded widgets

Google announced AI-first Googlebooks laptops, more agentic Gemini features, vibe-coded Android widgets, Gemini in Chrome, and refreshed Android Auto ahead of I/O; the RSS snippet does not disclose specs, pricing, availability, or rollout timelines.

Why it matters: HKR-H/K/R all pass because Google bundled several Gemini/Android AI entry points with named product hooks. Missing parameters, pricing, rollout dates, and testable performance keeps it in the mid-weight product-update band.

The Verge · AI

Gemini’s Latest Updates Are All About Controlling Your Phone

Google announced Gemini Intelligence at its pre-I/O Android showcase, placing Gemini in Chrome on Android, autofill suggestions, and in-app actions; the RSS snippet does not disclose supported device models, rollout timing, or pricing.

Why it matters: HKR-H/K/R all pass because Gemini is moving into Android phone control across Chrome, autofill, and app actions. Missing device list, launch timing, and pricing keep it in the 72–77 product-update band.

TechCrunch · AI

Google brings agentic AI and vibe-coded widgets to Android

Google is adding Gemini Intelligence to Android, and the RSS snippet only discloses Gboard-based dictation and form-filling capabilities; the post does not disclose launch timing, supported devices, pricing, or technical implementation details.

Why it matters: HKR-H/K/R pass because this is a Google Android platform AI update with named features. Missing rollout timing, device scope, and pricing keep it at the low featured band, not a must-write release.

TechCrunch · AI

The AI legal services industry is heating up — Anthropic is getting in on the action

Anthropic introduced tools for law firms that cover five clerical workflows: document search and review, case law resources, deposition preparation, document drafting, and related tasks; the RSS snippet does not disclose pricing, launch timing, or model details.

Why it matters: HKR-H/K/R pass: Anthropic is moving into a high-value legal workflow with five named use cases. No pricing, customer scale, or new model capability is disclosed, so this stays just above the featured threshold.

TechCrunch · AI

Google adds Gemini-powered dictation to Gboard, which could be bad news for dictation startups

Google adds Gemini-powered dictation to Gboard for an initial launch on Samsung Galaxy and Google Pixel phones; the post does not disclose supported languages, pricing, offline behavior, or a rollout date.

Why it matters: HKR-H/K/R all pass: Gemini dictation lands inside Gboard with Galaxy and Pixel named, and the platform-bundling angle matters to AI startups. Missing language, pricing, offline mode, and timing keep it at the low featured band.

AI HOT (Curated Pool)

Code w/ Claude SF 2026: Building on Exponential AI Growth

Anthropic expanded developer tooling at Code w/ Claude SF 2026: Claude Code rate limits doubled, Claude Opus API limits increased, and hosted agents on the Claude platform added four functions, including memory review, multi-agent delegation, output criteria, and webhooks.

Why it matters: Anthropic ships a substantive Claude Code update with concrete numbers and feature additions; HKR-H/K/R all pass. This is strong dev-tool news, not a flagship model release, so it fits the 78–84 band.

Financial Times · Technology

CME plans to launch futures market for AI computing power

CME plans to launch futures contracts tied to GPU rental prices, allowing traders and companies to bet on or hedge future costs; the RSS snippet does not disclose contract specifications, launch timing, or the reference index.

Why it matters: FT reports CME plans GPU rental-price futures, clearing HKR-H/K/R through novelty, mechanism, and compute-cost resonance. Missing contract specs, launch timing, and index details keep it at featured threshold, not P1.

May 12Tuesday

AI HOT (Curated Pool)

Install the official Codex plugin in Claude Code

The author describes installing OpenAI’s official Codex plugin in Claude Code via the plugin marketplace: add the repository, install the plugin, reload, and configure it, then use it to build a Skill where Claude Code handles reasoning and Codex acts as moderator.

Why it matters: HKR-H/K/R all pass: cross-stack plugin use is clickable, the install path is concrete, and it matters to AI dev workflows. It stays low-featured because this is a tutorial-style tip, not a model or platform release.

AI HOT (Curated Pool)

Dungeons & Desktops: Building a Procedurally Generated Roguelike with GitHub Copilot CLI

A GitHub employee used GitHub Copilot CLI to build an extension that parses any codebase into one Roguelike-style dungeon layout, with procedural level generation used as the core mechanism for a creative coding and game prototyping demo.

Why it matters: HKR-H and HKR-K pass: an official GitHub tutorial has a novel demo and a clear mechanism. It is not a major Copilot capability release, and lacks production metrics, pricing, or benchmark data, so it sits at the tutorial-featured floor.

Hacker News front page

Show HN: Statewright – Visual State Machines for More Reliable AI Agents

Statewright uses a Rust state-machine engine to constrain Claude Code tool access, iterations, transitions, and guards; the post says 13–20B models improved consistently on real SWE-bench tasks, but it does not disclose benchmark scores, sample size, or the exact evaluation protocol.

Why it matters: HKR-H/K/R all pass: the state-machine constraint is a clear agent-reliability hook with a testable SWE-bench claim. Exact scores and reproduction details are not disclosed, so it stays just above the featured threshold.

AI HOT (Curated Pool)

Exporting consumer data for personalized AI Agent services

The author surveyed export methods across 5 consumer platforms: Taobao supports exports, JD.com needs a Codex-built Chrome extension, Ele.me can provide Excel exports by request, Meituan Waimai has no method, and JD.com plus Dianping tools are open sourced.

Why it matters: HKR-H/K/R pass: the post has a concrete personal-agent hook, 5-platform data, and open-source JD/Dianping tools. It stays at the featured floor because this is a practitioner experiment, not a platform release.

Synced · WeChat

Unitree launches GD01 civilian piloted transforming mech from RMB 3.9 million

Unitree launched the GD01 piloted transforming mech with a starting price of RMB 3.9 million, biped and quadruped modes, and a stated payload of about 500 kilograms.

Why it matters: HKR-H/K/R all pass: Unitree’s GD01 has a rare manned-mech hook plus concrete price, modes, and payload. It is a notable robotics product launch, not a foundation-model-scale event, so it sits at the lower featured band.

Xinzhiyuan · WeChat

OpenAI releases GPT-Realtime-2, described as a GPT-5-level reasoning audio model

OpenAI released GPT-Realtime-2 alongside Realtime-Translate and Realtime-Whisper, with a 128K context window, five reasoning-effort levels, and API pricing of $32 per million input tokens and $64 per million output tokens.

Why it matters: HKR-H/K/R all pass: realtime audio reasoning is a strong hook; 128K context, five reasoning levels, and $32/$64 per 1M tokens add substance; voice-agent cost and stack choices hit practitioners. This is a same-day OpenAI product update.

Xinzhiyuan · WeChat

The Largest Single Industrial Product in History Is Entering Mass Production in China

AgiBot says it had shipped 10,000 general-purpose embodied robots by the end of March, and its humanoid robots worked eight continuous hours on a Nanchang 3C production line, completing 2,283 tasks with zero errors under formal line-cycle requirements.

Why it matters: HKR-H/K/R all pass: AgiBot gives unit and factory-run numbers with clear robotics deployment resonance. The score stays in 78-84 because the key claims are company-sourced, with no third-party validation or cost data disclosed.

QbitAI · WeChat

OpenClaw quietly updates with Peekaboo v3 for Mac computer use

OpenClaw-related Peekaboo v3 adds Mac agent capabilities for pixel-level screenshots, UI position reading, clicks, text input, hotkeys, scrolling, and drag-and-drop, with MCP server integration for Cursor, Claude Code, and Codex.

Why it matters: HKR-H/K/R all pass: Peekaboo v3 adds Mac GUI perception and action primitives plus MCP access for Cursor, Claude Code, and Codex. This is a useful open-source agent-tooling update, not a model-level event, so it sits in the low featured band.

The Verge · AI

OpenAI just released its answer to Claude Mythos

OpenAI launched Daybreak, a security initiative that uses the Codex Security AI agent released in March to model an organization’s code, validate likely vulnerabilities, and automate detection of higher-risk issues before attackers find them.

Why it matters: HKR-H/K/R all pass: Daybreak has a rivalry hook, concrete agent workflow, and code-security resonance. It is narrower than a model or ChatGPT capability release, so it stays in the 78–84 band.

The Verge · AI

Here’s What Mira Murati’s AI Company Is Up To

Thinking Machines announced work on “interaction models” that continuously take in audio, video, and text and respond or act in real time; the post does not disclose model size, release timing, pricing, or the final product format.

Why it matters: HKR-H/K/R all pass, but the body lacks parameters, launch timing, and product form. This is a high-interest startup direction reveal, not a usable model release, so it stays at the top of the 72–77 band.

AI HOT (Curated Pool)

Introducing Daybreak: Frontier AI for Cyber Defenders

OpenAI introduced Daybreak for cyber defenders, combining OpenAI models, Codex, and security partners; the post does not disclose pricing, launch timing, or concrete defense metrics.

Why it matters: OpenAI’s Daybreak announcement clears HKR-H/R as a security-focused product hook, but HKR-K fails: no defense metrics, access terms, or pricing. That keeps it in the 72–77 product-update band.

AI HOT (Curated Pool)

Replit launches parallel agents with support for 10 concurrent agents

Replit launched parallel agents that run up to 10 agents concurrently, with each agent holding an independent copy of the app, working on its own machine, and merging the results through an agent workflow.

Why it matters: HKR-H/K/R pass: the post gives a concrete 10-agent parallel workflow with isolated app copies and merge. This is a mid-weight dev-tool update, below a Cursor Agent-mode-scale launch, so it sits at the featured threshold.