Skip to content

#Agent

41 today

May 28Thursday

Latent Space

Cognition Raises $1B in $26B Series D

Cognition raised a $1B Series D at a $26B valuation and projects more than $1B ARR by year-end; the post says its valuation rose 2.5× from the $10B Series C eight months earlier, while the rest of the issue summarizes agent, inference, benchmark, and multimodal AI updates from May 26–27, 2026.

Why it matters: HKR-H/K/R all pass: Cognition’s $1B Series D at a $26B valuation is large, and projected year-end ARR above $1B gives a concrete business signal. This is not a model launch, but it is must-write funding news for AI coding agents.

Xinzhiyuan · WeChat

Tsinghua Team Open-Sources PilotDeck Agent System, Claims 70% Token Cost Reduction

Tsinghua THUNLP, ModelBest, OpenBMB, and AI9stars open-sourced PilotDeck; the article says its sub-agent routing reduced cost from $12.58 to $2.83 in a Xiaohongshu content-generation test, while preserving separate WorkSpaces, editable memory, and per-session routing logs.

Why it matters: HKR-H/K/R all pass: PilotDeck has a clear agent-cost hook, a concrete routing mechanism, and $12.58 to $2.83 data. It stays in the 78–84 band because this is a tool release, not a major model or platform launch.

Synced · WeChat

ICML 2026: AutoMoT reaches SOTA on Bench2Drive and nuScenes

NTU AutoMan Lab, Harvard, and Xiaomi Auto proposed AutoMoT, a unified VLA driving model using a 4B Qwen3-VL Understanding Expert and a 1.6B Action Expert with asynchronous inference, reaching 89.42 DS and 74.09% SR on Bench2Drive with AutoMoT+.

Why it matters: HKR-H/K pass via the async VLM-driving setup and concrete Bench2Drive numbers. The autonomy focus narrows HKR-R, so this sits at the featured threshold rather than the 78+ research tier.

AI HOT (Curated Pool)

NVIDIA Releases AI Framework Polar, Raising Codex Benchmark Score by 594.74%

NVIDIA’s research team open-sourced Polar, an agent reinforcement learning framework that connects GRPO training at the model API boundary without rewriting Codex CLI, Claude Code, Qwen Code, or Pi; on Qwen3.5-4B, Polar raised Codex pass@1 on SWE-Bench Verified from 3.8% to 26.4%, while prefix_merging cut training steps from 1,185 to 218.

Why it matters: HKR-H/K/R all pass: NVIDIA open-sourced Polar with a concrete GRPO mechanism and SWE-Bench Verified numbers. This is a strong research/open-source item, not a major model or product release, so it stays in the 78–84 band.

AI HOT (Curated Pool)

Grok Build 0.1 on API

xAI released Grok Build 0.1 in public beta through the xAI API for agentic coding tasks, with throughput above 100 tokens per second and pricing at $1 per million input tokens and $2 per million output tokens.

Why it matters: HKR-H/K/R all pass, but this is a 0.1 public-beta API and pricing launch; benchmarks, context window, and task success rates are not disclosed. It fits a solid mid-weight product update at 78, featured not p1.

AI HOT (Curated Pool)

Security Changes in the AI Agent Era

Lemonade CISO Jonathan Jaffe says a single endpoint can run 200 to 10,000 agents, so security teams need to assign identity to each agent and enforce policies at the point of action, beyond current identity and access management systems.

Why it matters: HKR-H/K/R all pass, but this is an event-recap commentary rather than a product or research release. The concrete signal is the endpoint agent count and identity-control model, placing it at the 72-77 featured threshold.

AI HOT (Curated Pool)

Cognition AI raises over $1B, targets 10x software engineering productivity

Cognition AI raised over $1 billion at a $26 billion pre-money valuation, while annualized revenue grew from $37 million to about $492 million in one year, and Devin is positioned as an autonomous junior engineer that can plan, test, and deploy through multi-step workflows.

Why it matters: HKR-H/K/R all pass: the story has hard numbers on funding, valuation, and ARR, plus a direct junior-engineer automation angle. Single-post sourcing keeps it below the 95+ industry-shaking band.

AI HOT (Curated Pool)

Using LLMs to secure source code

Anthropic describes a six-step Claude Opus workflow for source-code security: threat modeling, sandboxing, vulnerability discovery, validation, triage, and remediation; in its open-source scanning work, it disclosed 1,596 vulnerabilities by May 22, 2026, with 97 already fixed.

Why it matters: HKR-H/K/R all pass: Anthropic gives a Claude Opus security-audit workflow plus 1,596/97 outcome numbers. It stays below 85 because this is not a new model or platform-level capability release.

AI HOT (Curated Pool)

Cognition becomes the world’s largest independent agent lab

Cognition announced over $1 billion in funding at a $26 billion valuation, with enterprise usage up more than 10x this year and annualized revenue reaching $492 million.

Why it matters: HKR-H/K/R all pass: Cognition’s >$1B raise, $26B valuation, and $492M ARR are hard numbers tied to coding-agent competition and developer workflows. This fits the 85–94 same-day band, with no cross-source bump or deal details.

AI HOT (Curated Pool)

OpenAI Products Support Secure Connections to Private MCP Servers

OpenAI supports ChatGPT, Codex, and the Responses API connecting to internal MCP servers through outbound-only HTTPS, while teams keep those servers inside private networks.

Why it matters: HKR-H/K/R pass: OpenAI adds private MCP server support with outbound-only HTTPS, a concrete enterprise agent integration mechanism. Missing permission model, pricing, and rollout details keep it in the lower featured band.

AI HOT (Curated Pool)

Zero-Trust Security Framework for AI Agents

Anthropic published a zero-trust framework for enterprise autonomous AI agents, saying frontier models compress vulnerability exploitation from months to hours; the post outlines a three-tier architecture, an eight-stage rollout process, and threats including prompt injection, tool poisoning, and memory poisoning.

Why it matters: Anthropic’s agent zero-trust framework clears HKR-H/K/R with a concrete exploit-cycle claim, three-layer architecture, and eight-stage process. Strong safety/agent signal, but not a model launch or major product release.

Bloomberg Technology

Meta to Sell AI Chatbot Subscriptions to Offset Spending

Meta Platforms is selling consumer subscriptions to Meta AI for the first time, aiming to offset hundreds of billions of dollars in AI investments. The RSS snippet does not disclose pricing, launch timing, markets, or feature differences versus the free chatbot.

Why it matters: HKR-H/K/R pass: Bloomberg reports Meta’s first consumer subscription plan for Meta AI, tied to AI spending payback. Missing price, launch timing, and feature split keep it below must-write range.

AI HOT (Curated Pool)

Latest Google Pay updates

Google Pay introduced a universal commerce protocol and a new MCP server for AI agents to manage integrations and analyze trends, while Android updates add dynamic callbacks for faster checkout, WebView payments in social apps, cross-device biometric authentication, and new transaction signals.

Why it matters: HKR-H/K/R pass: the MCP payments angle is concrete and relevant to agent commerce. Score stays in the 72–77 band because the post lists features but gives no adoption scale, pricing, or real agent transaction case.

Hugging Face Blog

ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks

Artificial Analysis and IBM published the ITBench-AA title, saying frontier models scored below 50% on an enterprise IT agent task benchmark; the post does not disclose tested models, sample size, or scoring method.

Why it matters: HKR-H/R pass: frontier models under 50% on enterprise IT agent tasks is clickable and deployment-relevant. HKR-K is weak because models, sample size, and scoring are not disclosed, so it stays near the featured floor.

AI HOT (Curated Pool)

I Think Anthropic and OpenAI Found Product-Market Fit

Anthropic and OpenAI changed enterprise pricing around April 2026, moving coding agents from heavily discounted seat plans to API-usage billing, with Anthropic Enterprise at $20 per seat per month plus API fees and OpenAI Codex billed by API token usage.

Why it matters: HKR-H/K/R all pass: the piece ties OpenAI and Anthropic PMF to a concrete billing shift for coding agents. It is influential commentary, not an official launch, so it fits the 78–84 band.

AI HOT (Curated Pool)

Interview with Google Search VP Robby Stein on the AI-Native Search Era

Robby Stein discussed Google Search’s move toward an AI-native mode at Google I/O, covering AI Mode, multi-turn query decomposition, TPU infrastructure costs, source-link selection, and publisher traffic tension, but the post does not disclose specific pricing, traffic numbers, or rollout conditions.

Why it matters: HKR-H/K/R all pass, but this is an interview summary rather than a fresh launch. No price, traffic, or cost numbers are disclosed, so it sits in the 72–77 quality-interview band.

May 27Wednesday

The Verge · AI

Robinhood will let your AI agent trade stocks and make (or lose) lots of money

Robinhood opened its trading platform to AI agents: traders can create a separate account, allocate a specific amount of money, and let the agent buy and sell stocks, while Robinhood warns agentic trading can cause the loss of the entire investment.

Why it matters: HKR-H/K/R all pass: real-money stock trading gives the hook, separate funded accounts add mechanism, and autonomy risk creates resonance. It stays in the 78–84 band because safeguards, rollout scope, and regulatory limits are not disclosed.

AI HOT (Curated Pool)

Runway launches Model Context Protocol server

Runway launched an MCP server that lets compatible agents such as Claude, ChatGPT, and Cursor generate images and videos inside chat interfaces, with access to Gen-4.5, Seedance 2.0, GPT Image 2, Kling 3.0, and Nano Banana Pro.

Why it matters: HKR-H/K/R all pass, but this is a Runway product integration, not an MCP protocol change or model release. It clears featured, with the score kept in the 72–77 band.

r/LocalLLaMA

I ran 8 open-weight models as agents in a persistent MMO for 10 days

Firespawn Studios ran 25 agents across 8 open-weight models for 10 days in Null Epoch Season 0 and released about 93,000 logged events, with roughly 70% of actions including the model’s reasoning or justification.

Why it matters: HKR-H/K/R all pass: a concrete 10-day MMO agent trial with 25 agents and 93k events. Reddit sourcing limits reach, so it lands in the 78–84 good-quality band, not P1.

TechCrunch · AI

Robinhood now lets your AI agents trade stocks

Robinhood lets AI agents read and analyze users’ portfolios and suggest investments, but order placement is limited to the pre-loaded balance in a dedicated wallet.

Why it matters: HKR-H/K/R all pass: the hook is agents trading real money, the concrete mechanism is portfolio access plus a prefunded wallet, and the resonance is agent safety. Robinhood is not a frontier AI lab, so this stays at the lower featured band.

Alibaba Technology · WeChat

From Language Emergence to Collaborative Emergence: How AI Can Make High-Quality Decisions

Lv Ruofan proposes the Agent Room model: multiple agents share context, a task ledger, Memory, Runtime, and Artifacts, and two software-engineering cases show the system moving from workflow automation toward collaborative judgment rather than predefined task routing.

Why it matters: HKR-H/K/R all pass, but this is a methodology piece rather than a model launch or open-source framework. Concrete Agent Room mechanisms and 2 R&D sites put it in the 72–77 featured band.

QbitAI · WeChat

7B Medical AI Agent Beats o3 and GPT-5 by Learning Where and How to Look

Shanghai Innovation Institute’s LeapQuest and three universities released Ophiuchus and MedScope, applying Think with Images and Think with Videos to medical AI; Ophiuchus-7B scored 68.0 on eight VQA benchmarks, above OpenAI-o3 at 62.2, Gemini 2.5 Pro at 61.8, and GPT-5 at 59.9.

Why it matters: HKR-H/K/R all pass: a 7B model beating o3/GPT-5 is a strong hook, 8 VQA benchmarks with 68.0 vs 62.2 add a testable claim, and medical specialist evaluation will trigger debate. Not a frontier-lab general model release, so it stays in 78–84.

QbitAI · WeChat

Language Models Need Sleep: Let AI Nap Before Continuing Inference

Carnegie Mellon University and the University of Maryland propose a “sleep” mechanism for language models: when the context window is nearly full, the model stops accepting new tokens, runs multiple offline recursive forward passes to compress accumulated context into fast weights, clears the KV cache, and then resumes inference; tests cover cellular automata, multi-hop graph retrieval, and GSM-Infinite reasoning tasks.

Why it matters: HKR-H/K/R all pass: the sleep metaphor is clickable, and the mechanism is concrete. Score stays below 78 because the provided body lacks benchmark gains, code, or deployment evidence.

New York Times Chinese

How Google Rebounded and Started Winning the AI Race

Google said regular Gemini users more than doubled in one year to 900 million, while ad revenue rose 16% to $77 billion last quarter, and its Siri partnership with Apple will place Gemini inside future iPhone assistant features.

Why it matters: HKR-H/K/R all pass: NYT ties Google’s comeback narrative to 900M Gemini users, ad growth, and a Siri distribution deal. This is strong industry analysis, not a model launch, so it fits the 78–84 band.

Synced · WeChat

From Foundation Models to Physical AI, Samsung Moves Into the Core LLM Race

Samsung disclosed three AI efforts—Meki, M2RL, and LiveClawBench—covering a memory-based edge architecture, multi-domain reinforcement learning, and Physical AI evaluation; the article also says Samsung has purchased tens of thousands of GPUs for AI infrastructure, but does not provide model size, training budget, or deployment timelines.

Why it matters: HKR-H, HKR-K, and HKR-R pass, but this is a Samsung research bundle plus strategy signal, not a flagship model or product launch. It fits the 72–77 featured band, below same-day must-write.

AI HOT (Curated Pool)

AI Builds AI: ModelBest Open-Sources ForgeTrain, a Training Framework Written by AI

ModelBest, Tsinghua University, and OpenBMB open-sourced ForgeTrain, described as the first production-grade LLM training framework written entirely by AI with zero human code, and ModelBest used it to pretrain MiniCPM5-1B on Huawei Ascend chips.

Why it matters: HKR-H/K/R all pass: an open-source training framework, AI-written code, and MiniCPM5-1B pretraining on Ascend give concrete hooks. This is a strong tooling story, not a top-model launch, so 80 fits featured rather than P1.

Latent Space

[AINews] New AI Infra Decacorns: Fireworks, Baseten, with OpenRouter on the Way

Latent Space says Fireworks is in talks for a $15 billion valuation round, Baseten is raising at an $11 billion valuation, and OpenRouter closed a $113 million Series C after volume grew 5x in six months.

Why it matters: HKR-H/K/R all pass: the decacorn hook is clickable, the post gives valuation, round, and usage figures, and the topic speaks to inference economics. Fireworks and Baseten are still reported as in talks or raising, so this stays in the 78–84 band.

AI HOT (Curated Pool)

Code w/ Claude London event: Rethinking the developer experience

Anthropic announced two Claude Managed Agents capabilities at Code w/ Claude London: self-hosted sandboxes in public beta and MCP tunnels in research preview, with Spotify, Base44, and Legora already using them.

Why it matters: Official Anthropic product update with two concrete Claude Managed Agents capabilities. HKR-H/K/R pass, but this is a developer-tooling update rather than a major model release, so it lands at 78.

TechCrunch · AI

DuckDuckGo installs are up 30% as users reject being force-fed Google’s AI Search

Google replaced Search’s blue links with AI agents at I/O 2026, and DuckDuckGo app installs rose 30% as users looked for an alternative search entry point.

Why it matters: HKR-H/K/R all pass: the 30% install jump is a concrete backlash signal tied to Google AI Search. Missing measurement window and method keep it at the lower featured threshold.

Computing Life · Yage

Using AI Better, Step Two: Write the Skill Before Execution

The author proposes writing a Skill before asking AI to execute a task; each Skill should include three elements—success criteria, observed pitfalls, and deterministic tools—and can be organized through index.md plus AGENTS.md or CLAUDE.md for reuse.

Why it matters: HKR-H/K/R pass via a concrete Skill-first workflow and reusable agent practice. No model release, product capability, or experiment numbers, so it sits at the featured threshold.

Computing Life · Yage

Step Two to Using AI Well: Write the Skill Before You Execute

Yage argues that users should externalize work before execution by writing reusable Skills for Claude Code, Codex, and Cursor. The post gives an Outlook email example: spend about 30 minutes documenting username, phone approval, and client choice, then have AI read that file on later runs.

Why it matters: HKR-H/K/R all pass, but this is a workflow tutorial rather than a product or model release. The concrete Skill mechanism and Outlook example clear the featured floor; weak source authority keeps it at 72.

AI HOT (Curated Pool)

How we contain Claude across different products

Anthropic describes three mechanisms for containing Claude agent deployment risks across products: sandboxing or VMs, network egress controls, system-prompt and training constraints, and fine-grained permissions for MCP servers and third-party plugins.

Why it matters: Anthropic discloses a concrete containment stack for Claude agents, stronger than a routine product note. HKR-H/K/R all pass, but this is not a model launch or major capability release, so it stays in the 78–84 band.

May 26Tuesday

AI HOT (Curated Pool)

Sundar Pichai on AI, the Future of Search, and Changes to the Web

Sundar Pichai said after Google I/O that Google is integrating Gemini into a new smart search box and the Gemini Spark agent platform; the post does not disclose model parameters, launch dates, or traffic impact numbers.

Why it matters: HKR-H and HKR-R pass: Pichai’s interview touches Google Search as an AI entry point and web traffic allocation. HKR-K is weak because the article gives Gemini-in-Search and Spark, but no rollout timing or technical detail.

r/LocalLLaMA

SkillOpt treats markdown skill files as trainable parameters with proper optimization machinery

SkillOpt uses a frontier model to propose add, delete, and replace edits to markdown skill files, then accepts only strict gains on a held-out validation set; the best skills usually converge after 1 to 4 accepted edits.

Why it matters: HKR-H/K/R all pass: the hook is trainable markdown skills, with held-out validation and 1-4 accepted edits. Single Reddit/project source and no broad adoption data keep it at 78, featured not p1.

AI HOT (Curated Pool)

Qwen3.7-Max Becomes the World’s No. 2 AI Coding Model

Qwen3.7-Max scored 1541 on Code Arena and ranked behind Claude; the post says it can run 35-hour tasks and perform more than 1,000 tool calls.

Why it matters: HKR-H/K/R all pass, but the source is a single Alibaba Cloud post and the evidence is benchmark plus vendor claims. This fits a strong product/benchmark update, not P1 without independent validation.

QbitAI · WeChat

Chinese AI-Written Pretraining Framework ForgeTrain Trains MiniCPM5-1B

ModelBest released ForgeTrain and MiniCPM5-1B, saying ForgeTrain was written by AI and trains 10% faster than NVIDIA Megatron under the same hardware conditions. MiniCPM5-1B is a 1B-parameter edge model with about 2GB FP16 weights and about 0.5GB INT4/Q4 weights.

Why it matters: HKR-H/K/R all pass: an AI-written trainer, a 10% same-hardware Megatron speed claim, and a 0.5GB 1B edge model are concrete hooks. Score stays at 80 because the first-ever claim and benchmark lack third-party reproduction.

Synced · WeChat

ACL 2026 Main: Spatial-Agent Generates Executable Geospatial Analysis Workflows for LLMs

Spatial-Agent inserts a GeoFlow Graph between natural-language questions and map tools, and Spatial-Agent with GPT-4o-mini reaches 45.15% accuracy on MapEval-API versus a 23.00% API baseline.

Why it matters: ACL Main gives a concrete mechanism and testable numbers, so HKR-H/K pass. The GIS focus limits HKR-R, placing it at the featured threshold rather than a must-write item.

Synced · WeChat

Grok keeps updating after xAI disbandment as Musk announces a new model

Elon Musk said the 1.5T-parameter Grok V9-Medium has finished training, will enter reinforcement learning in a few days, and is planned for release in two to three weeks. Grok Build supports up to 8 parallel sub-agents, a 256K-token context window, Plan Mode, Arena Mode, MCP, and ACP.

Why it matters: HKR-H/K/R all pass, but this is a Grok V9-Medium preview before RL and release, with no benchmarked capability yet. That fits a strong model-race/product update at 82, featured but not p1.

Synced · WeChat

AI-written training framework trains 1B edge model MiniCPM5-1B

ModelBest open-sourced MiniCPM5-1B and ForgeTrain; the 1B edge model scores 17.9 on AA-Index, while the AI-written ForgeTrain framework matches Megatron’s training results and runs 10% faster on Nvidia H100 under the article’s reported setup.

Why it matters: HKR-H/K/R all pass: the AI-written training framework hook is strong, with concrete AA-Index and H100 speed claims. It is not a flagship model release, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

Chinese agent SkyClaw targets Opus 4.6-level performance with free trial

Kunlun Tech released SkyClaw-v1.0 and SkyClaw-v1.0-lite with a 2-4 week free trial, claiming SkyClaw-v1.0 input costs are 1/24 of DeepSeek V4 Pro and about 1/43 of Sonnet 4.6.

Why it matters: HKR-H/K/R all pass: SkyClaw-v1.0 has a sharp cost hook, concrete trial and pricing ratios, and budget resonance. Source facts remain vendor claims, so it stays at the low featured band.