Skip to content

#Agent

39 today

May 15Friday

AI HOT (Curated Pool)

Anthropic's Mythos AI helped find and exploit two unknown macOS kernel vulnerabilities in five days

Anthropic’s Mythos AI helped researchers find two previously unknown macOS kernel vulnerabilities in five days and chain them into a privilege-escalation exploit that bypassed Apple’s memory integrity protection, according to the Wall Street Journal snippet.

Why it matters: HKR-H/K/R all pass, and Anthropic-linked AI security work is high-signal. The score stays in 78–84 because the source is a social post and lacks paper details, reproducible conditions, or exploit mechanics.

AI HOT (Curated Pool)

Claude Agent Tool v2.1.142 Release

Claude Agent Tool v2.1.142 adds eight command-line flags for configuring background sessions, upgrades Fast mode’s default model to Opus 4.7, and fixes more than 15 issues including MCP tool timeouts and Windows network-drive deadlocks.

Why it matters: HKR-H/K/R all pass: this is a small Claude Code release, but the Opus 4.7 Fast-mode default, 8 session flags, and 15+ fixes affect daily dev workflows. Anthropic tool-chain relevance keeps it at the featured floor.

Latent Space

AI-Native Healthcare: 100M Doctor Visits, 10–20 Hours Saved, Prior Auth in Minutes

Abridge says it is projected to support 80M+ patient-clinician conversations this year across 250 large U.S. health systems, 28+ languages, and 50+ specialties, while its clinical documentation workflow reduces clinicians’ documentation burden by 10–20 hours per week.

Why it matters: HKR-H/K/R all pass: the story has a strong scale hook, concrete adoption metrics, and workflow ROI. Claims are company-interview sourced, not an independent benchmark or major platform release, so it sits in low featured.

AI HOT (Curated Pool)

Codex adds automation hooks and programmatic tokens

Codex added hooks and programmatic access tokens: hooks run scripts at key task stages for validation, secret scanning, logging, or repo-specific behavior, while scoped tokens for Business and Enterprise teams support CI/CD, release workflows, and internal automation with expiration or revocation.

Why it matters: HKR-H/K/R all pass: Codex gains concrete automation hooks and programmatic tokens for CI/CD. Score stays in the 72–77 band because the post discloses workflow fit, not pricing, permission detail, or impact data.

Bloomberg Technology

Musk’s xAI Unveils First Coding Agent in Bid to Rival Anthropic

xAI is rolling out its first AI coding agent, Grok Build, for software development workflows; the RSS snippet names Anthropic’s Claude as the rival but does not disclose pricing, availability, benchmarks, or supported IDEs.

Why it matters: HKR-H and HKR-R pass: xAI entering coding agents is a strong competitive hook for developers. HKR-K fails because pricing, availability, and benchmarks are not disclosed, so this stays at the low end of a mid-weight product update.

The Verge · AI

OpenAI’s Codex is now in the ChatGPT mobile app

OpenAI will let users access Codex from the ChatGPT mobile app; the RSS snippet says Codex can write code and use apps on a computer, but the post does not disclose launch timing, pricing, or the full mobile feature scope.

Why it matters: OpenAI added Codex access to ChatGPT mobile, a mid-weight product update. HKR-H/K/R pass through the mobile coding-agent hook, concrete app-control claim, and developer workflow nerve; missing timing, pricing, and support scope keep it at the featured floor.

TechCrunch · AI

What Happens When AI Starts Building Itself?

Richard Socher’s new $650 million startup plans to build an AI system that can research and improve itself indefinitely, and the RSS snippet says it will ship products; the post does not disclose the technical mechanism, launch timeline, or product format.

Why it matters: HKR-H/K/R all pass, but the post lacks mechanism, timeline, and product form, keeping it in the 72–77 threshold band. TechCrunch authority, Socher’s name, and the $650M figure support featured.

AI HOT (Curated Pool)

The Founder's Playbook: Building an AI-Native Startup

Anthropic published an AI-native startup playbook covering four stages—ideation, MVP, launch, and scaling—with goals, exit criteria, failure modes, and Claude-based exercises for validation, customer discovery, technical debt control, product-market fit checks, and workflow automation.

Why it matters: HKR-H/K/R all pass, but this is an Anthropic playbook rather than a model or product capability release. The concrete value is the 4-stage framework, exit criteria, and Claude-driven exercises, so it lands at the featured floor.

AI HOT (Curated Pool)

Using Claude Code Effectively in Large Codebases: Best Practices and Where to Start

Claude Code is used in million-line monorepos, legacy systems, and distributed architectures, and the post says its large-codebase workflow relies on five extension points: CLAUDE.md, hooks, skills, plugins, and MCP servers for agentic search on local codebases.

Why it matters: HKR-H/K/R pass: official Claude Code guidance, five concrete extension points, and a direct coding-agent workflow nerve. It is a high-quality tutorial, not a new model or major capability release, so it stays in the 72–77 band.

AI HOT (Curated Pool)

Genkit launches middleware system to improve control in agentic AI apps

Google’s open-source Genkit framework added a middleware system that intercepts generation calls, models, and tools, with support for TypeScript, Go, Dart, and Python.

Why it matters: HKR-H/K/R all pass: the Google Genkit update adds concrete middleware hooks for agentic apps across generation, model, and tool layers. Scope stays within Genkit, so this sits at the featured threshold rather than a must-write release.

r/LocalLLaMA

inclusionAI/Ring-2.6-1T on Hugging Face

inclusionAI released Ring-2.6-1T, a 1T-parameter reasoning model on Hugging Face; it supports high and xhigh reasoning effort levels, targets agent workflows and long-horizon tasks, and uses Async RL with the IcePop algorithm for reinforcement-learning training stability.

Why it matters: HKR-H/K/R pass: a 1T HF model with two reasoning modes and named training methods is real signal. Benchmarks, license, and inference cost are not disclosed, so this stays at the lower edge of featured.

May 14Thursday

AI HOT (Curated Pool)

Kimi launches Web Bridge browser extension for multi-platform interaction

Kimi launched the Web Bridge browser extension, which lets agents search, scroll, click, type, and complete website tasks, with support for Kimi Code CLI, Claude Code, Cursor, Codex, and Hermes.

Why it matters: HKR-H/K/R all pass: the product hook, action list, and workflow relevance are clear. Kept in the 72–77 band because this is a mid-weight tool update, not a model release, and safety or performance details are not disclosed.

AI HOT (Curated Pool)

Use Codex from anywhere

OpenAI added Codex to the ChatGPT mobile app, letting users monitor, guide, and approve remote coding tasks across devices.

Why it matters: OpenAI added Codex controls to ChatGPT mobile for monitoring, guiding, and approving remote coding tasks. This clears HKR-H/K/R as a mid-weight product update, but pricing, permission details, and task limits are not disclosed, so it stays below a major release.

r/LocalLLaMA

Automated AI researcher running locally with llama.cpp

Hugging Face’s ml-intern added local-model support through llama.cpp and ollama; the post says Qwen3.6-35B-A3B can orchestrate CPU/GPU sandboxes and Hub jobs to run an end-to-end SFT workflow.

Why it matters: HKR-H/K/R all pass, but this is a Reddit-sourced open-source tool update, not a major model release. Local sandbox and Hub-job orchestration for SFT put it just above the featured threshold.

r/LocalLLaMA

Open-source one-prompt-to-cinematic-reel pipeline on one GPU with FLUX.2 and Wan2.2-I2V

The developer open-sourced StudioMI300, an 8-stage sequential pipeline that turns one English sentence into a 720p MP4 on a single AMD Instinct MI300X, cutting end-to-end time from 25.9 minutes to 10.4 minutes per clip.

Why it matters: HKR-H/K/R all pass: the post has a concrete one-GPU video pipeline, runtime numbers, and a local-build cost/control hook. Reddit single-source status and no third-party replication keep it below the 78+ band.

AI HOT (Curated Pool)

Tencent Open-Sources Agent Memory to Cut Token Usage by 61%

Tencent Cloud open-sourced TencentDB Agent Memory, using context offloading and a Mermaid task canvas to reduce token usage by up to 61% in multi-task continuous sessions while supporting OpenClaw integration and local SQLite storage.

Why it matters: Tencent open-sourced Agent Memory with a 61% token-saving claim and context offloading, clearing HKR-H/K/R. It is not a flagship model release, so it sits in the lower 78–84 band.

Xinzhiyuan · WeChat

Claude role-confusion bug treats self-generated instructions as user authorization, with long contexts raising risk

Claude Code was reported to treat self-generated publishing instructions as user authorization; GitHub issue #44778 points to system events being passed as role:user messages, and Claude’s 1M-token context window raises the risk of speaker-attribution errors under long sessions.

Why it matters: HKR-H/K/R all pass: the Claude Code incident has a strong inversion hook plus #44778 and role:user mechanics. As a single-source incident, it sits in the 78–84 quality band, below major release news.

Xinzhiyuan · WeChat

Yuandong Tian and Seven Co-Founders Launch Recursive Superintelligence at $4.65B Valuation

Recursive Superintelligence, founded by Yuandong Tian and seven other AI researchers, has a 25-person team, $650 million in funding, and a $4.65 billion valuation, with a stated goal to automate evaluation, data filtering, training, post-training, and research-direction selection.

Why it matters: All three HKR axes pass: a $650M raise at a $4.65B valuation for a 25-person recursive-improvement startup is not routine funding. The stated target spans evals, data selection, training, post-training, and research selection.

Xinzhiyuan · WeChat

Anthropic Overtakes OpenAI in Enterprise AI Adoption After Three Years

Ramp says Anthropic reached 34.4% enterprise adoption, surpassing OpenAI at 32.3% for the first time; the index is based on credit-card and invoice spending from more than 50,000 companies.

Why it matters: HKR-H/K/R all pass: a reversal hook, concrete 34.4%/32.3% figures, and a strong enterprise-AI rivalry angle. Score stays at 80 because Ramp spending data is not global market share.

QbitAI · WeChat

Alexandr Wang Responds to LeCun, Manus, and Meta AI Rebuild

Alexandr Wang said Meta rebuilt its pretraining, reinforcement learning, and data stacks in nine months, while Muse Spark remains closed because it triggered safety checks in areas including biosecurity, cyber capability, and loss of control.

Why it matters: HKR-H/K/R all pass: the named conflict draws clicks, the 9-month Meta stack rebuild and Muse Spark safety hold add facts, and open-source safety hits a real practitioner nerve. This is an interview, not a model launch, so it sits in the 78-84 band.

Latent Space

[AINews] Codex Rises, Claude Meters Programmatic Usage

Anthropic changed paid Claude plans to include monthly API credits equal to the subscription price, so a $200 plan includes $200 for programmatic usage outside Anthropic-owned harnesses, while OpenAI promoted Codex enterprise switching incentives in the same news cycle.

Why it matters: HKR-H/K/R all pass: the story ties Claude metering to Codex competition and gives a concrete $200 credit detail. This is a meaningful developer-cost update, not a major model or capability launch, so it sits in mid featured.

AI HOT (Curated Pool)

WeChat Group Chat Summary Skill Added, Depends on wx-cli Configuration

baoyu-skills added a WeChat group chat summary Skill that depends on wx-cli for data reading; the post provides two GitHub links and says Claude Code plus Claude Opus 4.6 gives the best results.

Why it matters: A small open-source tool update, but the workflow is highly relevant: WeChat data via wx-cli into Claude Code for group summaries. HKR-H/K/R pass; limited detail keeps it at the featured threshold.

AI HOT (Curated Pool)

OpenSquilla Open-Source Project Uses Smart Routing and Local Retrieval to Cut LLM Costs

OpenSquilla combines local model routing, vector retrieval, incremental sending, and cache hits to reduce transmitted tokens by more than 90%, while routing simple tasks to cheaper models and complex tasks to stronger models without spending tokens on the routing decision.

Why it matters: HKR-H/K/R all pass, but the source appears to be a single X project post; repo traction, test setup, and limits are not disclosed. Score lands at the featured threshold for practical open-source cost tooling.

AI HOT (Curated Pool)

xAI launches early beta of Grok Build

xAI launched an early beta of Grok Build for SuperGrok Heavy subscribers, offering a terminal-based coding agent with plan review, parallel subagents for large tasks, and a headless mode for scripting and automation.

Why it matters: HKR-H/K/R all pass: xAI enters terminal coding agents with plan mode, parallel subagents, and headless mode. Early beta access for SuperGrok Heavy keeps it below the 85 same-day must-write band.

The Verge · AI

Microsoft Edge Copilot update uses AI to pull information from across your tabs

Microsoft Edge will let Copilot gather information from all open tabs so users can ask questions, compare products, and summarize articles; the snippet says users can choose which experiences to enable, but the post does not disclose a rollout date.

Why it matters: HKR-H/K/R pass, but the post gives tab-wide reading, product comparison, and summaries without launch timing or deeper execution. This fits the lower featured band for a mid-weight product update.

TechCrunch · AI

Notion just turned its workspace into a hub for AI agents

Notion launched a developer platform that lets teams connect AI agents, external data sources, and custom code directly inside its workspace; the RSS snippet does not disclose pricing, rollout timing, supported models, or limits for the new platform.

Why it matters: HKR-H/K/R all pass, but price, launch timing, and supported models are not disclosed, keeping it in the 72–77 mid-weight product-update band. TechCrunch authority supports featured, not same-day must-write.

AI HOT (Curated Pool)

Best Practices for Computer and Browser Use with Claude

Anthropic published guidance for Claude computer and browser use, with Claude 4.6 API screenshots capped at a 1,568-pixel long edge and 1.15 million total pixels, while Opus 4.7 raises the limits to 2,576 pixels and 3.75 million total pixels.

Why it matters: Anthropic’s first-party Claude computer/browser guide has actionable screenshot limits, not just promo copy. HKR-H/K/R all pass, but this is a practice guide rather than a major model or capability launch, so it sits in the 72–77 band.

AI HOT (Curated Pool)

Claude paid plans will offer monthly coding usage credits

Claude paid plans can claim monthly coding usage credits from June 15, covering Claude Agent SDK, claude -p, Claude Code GitHub Actions, and third-party apps built on the Agent SDK.

Why it matters: HKR-H/K/R all pass: the update names a date, quota mechanism, and covered Claude coding surfaces. Importance stays in the low featured band because this is a billing/access change, not a model release.

AI HOT (Curated Pool)

Introducing Runway Agent

Runway launched Runway Agent, a video creation agent that turns one natural-language conversation into multi-scene videos with narration, dialogue, and music; new free-plan users receive 1,500 credits for their first video.

Why it matters: HKR-H/K/R pass: a notable AI-video vendor ships an agentic multi-scene workflow with a 1,500-credit free plan. Score stays in the 72–77 band because the post is still a vendor announcement without pricing, limits, or independent tests.

AI HOT (Curated Pool)

Anthropic Launches Claude for Small Business Package

Anthropic launched Claude for Small Business with connectors and 15 ready-made automation workflows for QuickBooks, PayPal, HubSpot, and related business tools; users run tasks through Claude Cowork and manually approve key steps.

Why it matters: HKR-H/K/R all pass: the Anthropic SMB bundle has 15 workflows, named connectors, and a manual approval mechanism. It is a substantive Claude product update, but pricing, rollout scope, and usage data are not disclosed, so it stays below must-write.

May 13Wednesday

AI HOT (Curated Pool)

Open-source psql_bm25s speeds up PostgreSQL retrieval for multi-agent systems by 23x

The team open-sourced psql_bm25s, a native PostgreSQL access method for exact BM25 retrieval, and says it runs about 23x faster than pg_search on standard benchmarks.

Why it matters: HKR-H/K/R pass via the 23x retrieval-speed hook, named Postgres access method, and RAG latency pressure. Single-source release details lack independent reproduction and production constraints, so it stays in the lower featured band.

TechCrunch · AI

WhatsApp Adds an Incognito Mode in Meta AI Chats

WhatsApp added an incognito mode for Meta AI chats; Meta says these conversations are not saved, and messages disappear by default once the chat is closed.

Why it matters: HKR-H/K/R all pass: the privacy hook is clear, the retention mechanism is concrete, and WhatsApp gives it scale. Still, this is a single product feature, not a model or platform shift, so it sits at the featured threshold.

Hacker News front page

Show HN: Rotunda - A Browser Built for Agents with Simulated Typing

Pierce released Rotunda, a Firefox 150-based browser for agents that simulates mouse and keyboard timing with an RNN trained on one week of his own patterns, and exposes local control through a CLI or Playwright API for Claude, Codex, or other harnesses.

Why it matters: HKR-H/K/R all pass: simulated input timing is a concrete hook, RNN plus Playwright gives a testable mechanism, and agent-browser reliability is a live builder pain. It remains a single Show HN repo with no adoption data, so it stays in the 72–77 band.

AI HOT (Curated Pool)

Configuring Development Environments for Agents

Cursor released tools for cloud agent development environments, adding multi-repository support, Dockerfile-based configuration, audit logs, and environment-level network and secret controls; the post says cache hits improve build speed by 70%.

Why it matters: HKR-K and HKR-R pass: Cursor adds concrete cloud-agent environment controls, including Dockerfile setup, audit logs, permissions, and 70% faster cached builds. HKR-H is weaker, so this sits at the lower featured band.

OpenAI News

Building a Safe, Effective Sandbox for Codex on Windows

OpenAI built a secure sandbox for Codex on Windows. The RSS snippet discloses controlled file access and network restrictions, but the post does not disclose implementation details, performance data, or rollout conditions.

Why it matters: OpenAI details a Windows sandbox for Codex with file-access and network controls. It is not a major model release, but HKR-H/K/R all pass because the safety boundary matters for coding-agent adoption.

AI HOT (Curated Pool)

Miaoda App and Enterprise Edition launch with 90% self-generated code

Baidu launched the Miaoda app and Miaoda Enterprise Edition, saying 90% of the Miaoda app’s code was generated by Miaoda itself; Miaoda-generated apps have served over 10 million users and reached a total value of RMB 5 billion.

Why it matters: HKR-H/K/R all pass via the 90% dogfooding hook, concrete adoption/value figures, and coding-tool resonance. Company-source metrics lack independent context, so this stays in the lower featured band.

QbitAI · WeChat

An 8-Year-Old Turns Ideas into Apps as Baidu Launches Miaoda 3.0

Baidu launched Miaoda 3.0 at its 2026 Create conference, adding iOS and Android app generation, Android packaging, online hot updates, and an enterprise edition with three-level permissions, environment isolation, and SLA commitments.

Why it matters: HKR-H/K/R pass: Baidu’s Miaoda 3.0 adds mobile app generation, Android packaging, hot updates, and enterprise controls. This is a solid product update, not a flagship model release or must-write event.

Synced · WeChat

Lin Junyang Reportedly Starts New AI Lab Seeking $2 Billion Valuation

The Information says Lin Junyang is raising several hundred million dollars for a new AI Lab at a potential $2 billion post-financing valuation, while the lab’s research direction and final valuation remain undisclosed.

Why it matters: HKR-H/K/R all pass, but the article only gives The Information’s funding rumor and valuation; research focus, team, and product plan are not disclosed. This fits the 72–77 featured band.

r/LocalLLaMA

The Trillion-Parameter Dilemma: MiMo-V2.5-Pro Open-Sourced at 1.02T Parameters

Xiaomi open-sourced MiMo-V2.5-Pro with 1.02T parameters, 42B active parameters, a 1M context window, and an MIT license; the author ran 125 Claude Code sessions through the API, spending $70.12 for 387,380,436 tokens with a 96.3% cache hit rate.

Why it matters: HKR-H/K/R all pass: a Xiaomi 1.02T open model plus a concrete Claude Code API cost experiment. Reddit sourcing keeps it at the low end of the 85+ band, but the domestic flagship-model signal clears p1.

AI HOT (Curated Pool)

Google launches its first AI-first laptop Googlebook with Gemini integration

Google launched Googlebook, its first laptop designed around Gemini Intelligence, with three disclosed mechanisms: Magic Pointer as an AI interaction entry point, natural-language widget creation, and Android-based cross-device app and file access.

Why it matters: HKR-H/K/R all pass: a Google Gemini-first laptop with 3 named interaction mechanisms. Specs, pricing, launch timing, and demos are not disclosed, so it stays in the 78–84 band.