Skip to content

#Agent

0 today

May 16Saturday

Google DeepMind

Google DeepMind releases Gemini 3.5 Flash

Google DeepMind released the Gemini 3.5 model family, with the first model, Gemini 3.5 Flash, available the same day in the Gemini app, Google Search AI Mode, Google Antigravity, the Gemini API and Gemini Enterprise.

Why it matters: Google published 3.5 Flash's coding and agent benchmark scores and where it is available, enough to judge its place in long-horizon workflows.

May 15Friday

Alibaba Technology · WeChat

Qoder 1.0 launches as an agentic development workspace beyond AI IDE

Alibaba released Qoder 1.0 with downloads for Windows, macOS, and Linux, adding a standalone Quest workspace, cross-project parallel agent tasks, a team knowledge engine, and Experts mode with five roles for planning, research, coding, review, and testing.

Why it matters: Alibaba’s Qoder 1.0 is a mid-weight AI coding product release with concrete agent-workflow features and developer resonance. No pricing, benchmark, or task-success data is disclosed, so it stays near the featured threshold.

May 13Wednesday

OpenAI News

Building a Safe, Effective Sandbox for Codex on Windows

OpenAI built a secure sandbox for Codex on Windows. The RSS snippet discloses controlled file access and network restrictions, but the post does not disclose implementation details, performance data, or rollout conditions.

Why it matters: OpenAI details a Windows sandbox for Codex with file-access and network controls. It is not a major model release, but HKR-H/K/R all pass because the safety boundary matters for coding-agent adoption.

May 12Tuesday

Google DeepMind

Google DeepMind publishes Co-Scientist multi-agent research system

Google DeepMind published Co-Scientist research in Nature, introducing a Gemini-based multi-agent AI system that iteratively generates, debates and evolves new hypotheses for complex scientific problems.

Why it matters: The post discloses the system's three-stage collaboration mechanism and deployment cases at several labs, showing how AI takes part in scientific hypothesis generation.

May 8Friday

Alibaba Technology · WeChat

The AI-Native Era: Where R&D Organizations Go Next

Xu Xiaobin cites internal interviews showing that engineers who use AI heavily cut coding time from 30% to 5%, raised Agent conversation time from 5% to 60%, and increased end-to-end delivery efficiency by 2 to 3 times, while pure coding efficiency rose 10 times.

Why it matters: Alibaba Tech’s internal-interview numbers make HKR-H/K/R pass, but this is org-methodology commentary rather than a product or model release, so it sits just above the featured threshold.

NVIDIA Blog

Powering the Next American Century: Chris Wright and NVIDIA’s Ian Buck on Genesis Mission

The U.S. DOE and NVIDIA are building two AI supercomputers at Argonne; Equinox uses 10,000 Grace Blackwell GPUs. Solstice will use 100,000 Vera Rubin GPUs, which Buck said reach 5,000 exaflops. The key bottleneck is grid work: Wright said AI can cut interconnection studies from years to weeks or hours.

Why it matters: HKR-H/K/R all pass: the GPU counts, DOE-NVIDIA role, and grid bottleneck are concrete. NVIDIA-blog sourcing keeps it below must-write; this fits the 78–84 band.

May 6Wednesday

NVIDIA Blog

NVIDIA and ServiceNow Partner on Autonomous AI Agents for Enterprises

NVIDIA and ServiceNow expanded their partnership with Project Arc, an enterprise desktop agent. It connects via Action Fabric and uses OpenShell for sandboxed, policy-governed execution. Blackwell delivers over 50x Hopper’s token output per watt and nearly 35x lower cost per million tokens.

Why it matters: HKR-K/R pass: the post gives mechanisms and Blackwell economics. HKR-H misses because the angle is a standard vendor partnership, so this sits in the 72–77 featured-threshold band.

May 5Tuesday

OpenAI News

GPT-5.5 Instant: smarter, clearer, and more personalized

OpenAI updated ChatGPT’s default model to GPT-5.5 Instant for default chat use. The RSS snippet says answers are more accurate, hallucinations are reduced, and personalization controls improved; the post does not disclose metrics, pricing, or context window.

Why it matters: HKR-H/K/R all pass: OpenAI changed ChatGPT’s default model to GPT-5.5 Instant. The post lacks evals, pricing, and context window details, so it stays at the low end of the 85–94 band.

May 1Friday

NVIDIA Blog

Nemotron Labs: What OpenClaw Agents Mean for Every Organization

NVIDIA says OpenClaw reached 250,000 GitHub stars by March 2026, passing React within 60 days. OpenClaw is Peter Steinberger’s self-hosted persistent agent; NVIDIA introduced NemoClaw with OpenShell sandboxing and Nemotron models. The key issue is governance: the post claims reasoning AI raised token use 100x, and autonomous agents add another 1,000x.

Why it matters: HKR-H/K/R all pass: OpenClaw’s GitHub growth is a hook, and NemoClaw names concrete sandbox and access-control mechanisms. NVIDIA’s own blog keeps it in the 78–84 band.

Apr 30Thursday

Google DeepMind

Google DeepMind announces AI co-clinician medical research program

Google DeepMind announced an AI co-clinician research program, exploring how AI agents can assist patient care under a doctor's clinical supervision. In a blinded evaluation of 98 real primary care queries, the system made no critical errors in 97 cases, and doctors preferred its answers over mainstream evidence synthesis tools. On 140 consultation skills, it matched or beat primary care physicians on 68, but expert physicians were still better overall at spotting red flags and key physical exams.

Why it matters: Google DeepMind published its AI co-clinician research program and a multimodal consultation evaluation, showing where medical agents' abilities currently end.

Apr 29Wednesday

NVIDIA Blog

NVIDIA Launches Nemotron 3 Nano Omni for Vision, Audio, and Language Agents

NVIDIA launched Nemotron 3 Nano Omni, claiming up to 9x higher throughput at the same interactivity. It uses a 30B-A3B hybrid MoE with Conv3D, EVS, and 256K context, taking text, images, audio, video, documents, charts, and GUIs as input. Open weights, datasets, and training methods arrive April 28, 2026 on Hugging Face, OpenRouter, build.nvidia.com, and 25+ platforms.

Why it matters: HKR-H/K/R all pass: NVIDIA’s open multimodal model has a 9x efficiency claim, 30B-A3B MoE, and 256K context. Single-vendor sourcing keeps it in the good-quality band, below must-write.

Apr 28Tuesday

X · @claudeai

Claude Now Connects to Tools Creative Professionals Already Use

Claude added a Blender connector for scene debugging, tool building, and batch object edits from Claude. The post does not disclose versions, pricing, or rollout scope; the key issue is agent control boundaries inside DCC workflows.

Why it matters: HKR-H/K/R pass: Claude’s Blender connector is a concrete agent-tool expansion. Missing version, pricing, and rollout details keep it near the featured threshold, not a must-write.

Apr 27Monday

Mistral AI

Mistral AI opens public preview of Workflows

Mistral AI has put Workflows, its enterprise AI orchestration layer, into public preview. It offers durable execution, observability and human-in-the-loop approvals. ASML, ABANCA and CMA-CGM are already using it to automate critical processes.

Why it matters: It lays out Workflows' orchestration features, deployment model and customer cases, showing the engineering bar for enterprise AI processes.

OpenAI News

An Open-Source Spec for Orchestration: Symphony

OpenAI released Symphony, an open-source spec for Codex orchestration. The RSS snippet says it turns issue trackers into always-on agent systems; the post does not disclose spec details, license, APIs, or benchmarks.

Why it matters: HKR-H and HKR-R pass: an OpenAI open-source Codex orchestration spec is relevant to agent workflows. HKR-K is weak because license, interfaces, and reproducible mechanics are not disclosed.

Apr 25Saturday

X · @AnthropicAI

New Anthropic research: Project Deal

Anthropic announced Project Deal and had Claude buy, sell, and negotiate for employees in a San Francisco office marketplace. The setup is confirmed as an internal marketplace; the post does not disclose scale, model version, or outcome metrics.

Why it matters: This clears featured on HKR-H and HKR-R: Anthropic has attention weight, and an agent negotiating office deals is inherently discussable. It stays mid-band because HKR-K is weak; the post gives the setup, but not sample size, model version, success metrics, or controls.

Apr 24Friday

Hugging Face Blog

DeepSeek-V4: a million-token context that agents can actually use

DeepSeek released V4 with two MoE checkpoints, Pro and Flash, both supporting a 1M-token context. Pro has 1.6T total and 49B active parameters; Flash has 284B total and 13B active. The key detail is KV cost: Pro uses 27% of V3.2 single-token FLOPs and 10% of its KV cache; Flash uses 10% and 7%.

Why it matters: DeepSeek-V4 is a flagship Chinese model release with 1M-token context and KV cache at 7%–10% of V3.2. HKR-H/K/R all pass, placing it in the 85–94 same-day band.

X · @claudeai

Memory on Claude Managed Agents is now in public beta

Claude has put Memory for Managed Agents into public beta, and agents can now learn from every session. The post only says it uses an intelligence-optimized memory layer balancing performance and flexibility; it does not disclose capacity, retention, pricing, or access conditions. What matters for practitioners is when persistent memory becomes default and how it changes agent evals and state management.

Why it matters: Memory on Claude Managed Agents is a substantive Anthropic product update with clear practitioner resonance, so HKR-H and HKR-R pass. HKR-K is weak because the post omits capacity, retention, pricing, and default-on conditions, keeping it in low featured rather than p1.

X · @claudeai

Claude can now connect to more apps outside work, including Tripadvisor, Booking.com, and Resy

Claude added at least 10 consumer app connections, including Tripadvisor, Booking.com, Resy, Instacart, Spotify, Audible, AllTrails, Thumbtack, and TurboTax. The RSS snippet confirms only a product update; the post does not disclose integration method, supported actions, regions, permission scope, or rollout timing. The key question is whether Claude can act in these apps directly, not just list them.

Why it matters: Official Anthropic product update with clear HKR-H/K/R: consumer app connectors expand Claude beyond workplace tools and widen its assistant surface. The score stays at 75 because the post lists apps only; actions, permissions, regions, and rollout details are not disclosed.

X · @OpenAI

Introducing GPT-5.5

OpenAI introduced GPT-5.5, and it is now available in ChatGPT and Codex. The RSS snippet says it targets real work and agents, can understand complex goals, use tools, check its work, and carry more tasks to completion; the post does not disclose parameters, pricing, context window, or benchmark results. What matters is the execution loop, not the headline's “new class of intelligence.”

Why it matters: OpenAI launching GPT-5.5 in ChatGPT and Codex is same-day mandatory coverage. HKR-H/K/R all pass: new model release, concrete agent-workflow claims, and direct impact on daily AI work. Price, context window, params, and benchmarks are undisclosed, so it stays below 95.

Apr 23Thursday

Hugging Face Blog

How to Use Transformers.js in a Chrome Extension

Hugging Face published a guide for a Transformers.js Chrome extension using Gemma 4 E2B. It defines three MV3 entry points: background service worker, side panel, and content script. The key design keeps local inference in the background and uses messaging plus a tool loop.

Why it matters: HKR-H/K/R all pass, but this is a Hugging Face implementation tutorial, not a model or platform release. Score sits at the featured threshold for a concrete MV3 architecture walkthrough.

X · @OpenAI

Introducing workspace agents in ChatGPT—shared agents for complex tasks and long-running workflows

OpenAI announced workspace agents in ChatGPT, described as shared agents that work across tools and teams for complex and long-running workflows. Only the title and RSS snippet are disclosed; the post does not disclose supported tools, pricing, access tier, permission model, or rollout timing. The key issue to watch is the collaboration boundary of shared agents, not the headline claim alone.

Apr 22Wednesday

OpenAI News

Introducing workspace agents in ChatGPT

OpenAI introduced workspace agents in ChatGPT, describing them as Codex-powered agents that automate complex workflows in the cloud. The RSS snippet confirms secure work across tools for teams, but the post does not disclose pricing, availability, supported tools, or performance metrics.

Why it matters: This is a substantive OpenAI product update inside ChatGPT. HKR-H lands on the jump from chat to workspace agents, HKR-K on Codex-powered cloud execution across tools, and HKR-R on team workflow automation; the score stops at 86 because pricing, rollout, tool support, and metrics

OpenAI News

Speeding up agentic workflows with WebSockets in the Responses API

OpenAI says WebSockets in the Responses API speed up the Codex agent loop, using connection-scoped caching to cut API overhead and improve latency. The RSS snippet confirms the mechanism, but the post does not disclose latency deltas, throughput numbers, or workload conditions. The key point is transport-layer optimization, not a new model.

Why it matters: This is a developer-facing OpenAI product update at the systems layer: WebSockets plus connection-scoped caching target agent-loop round-trip cost. HKR-H/K/R all pass, but the post does not disclose latency gains, throughput, or workload bounds, so it stays mid-featured rather än

Apr 21Tuesday

Google DeepMind

Partnering with industry leaders to accelerate AI transformation

Google DeepMind 宣布与 Accenture、Bain & Company、BCG、Deloitte、McKinsey 合作,帮助全球企业规模化落地前沿 AI。合作方将获得包括 Gemini 系列在内的前沿模型早期访问权,并直接对接 Google DeepMind 技术团队,聚焦金融、制造、零售、媒体娱乐等行业的智能体转型。目前仅 25% 的组织成功将 AI 规模化投入生产。

Apr 17Friday

Tencent Technology · WeChat

From Vibe Coding to Agentic Engineering: Rebuilding the Full Backend Development Workflow

Tencent engineers report a one-week practice that used Claude Code plus custom Skills, Commands, and MCP servers to run an 11-stage backend workflow in one terminal session. The post gives reproducible details: one requirement-exploration step used 20 tool calls, 93.8k tokens, and 56 seconds; execution was split into 4 tasks and produced 3 commits. The real point is workflow orchestration, not raw code generation; human review remains at plan, deploy, and review gates.

Why it matters: HKR-H/K/R all pass: the story turns agentic engineering into a measured backend workflow test, with tool-call, token, timing, plan-length, task, and commit data. Stronger than generic coding hype, but still a practitioner case study rather than a major product or model release.

X · @OpenAI

Codex for (almost) everything.

OpenAI said Codex can now use apps on Mac, connect to more tools, and handle ongoing and repeatable tasks. The post also claims image creation, learning from prior actions, and remembering user preferences; it does not disclose app coverage, integration method, pricing, or rollout timing.

Why it matters: This is an official OpenAI product update, and Codex moves from coding help toward desktop control, tool use, and memory, so HKR-H/K/R all pass. The post still omits supported apps, integration method, pricing, and launch timing, keeping it in the 78–84 band.

Apr 16Thursday

X · @claudeai

Introducing Claude Opus 4.7, our most capable Opus model yet.

Claude introduced Opus 4.7 and describes it as its most capable Opus model so far. The RSS snippet gives three claims: better rigor on long-running tasks, more precise instruction following, and self-verification before replying; the post does not disclose benchmarks, context window, pricing, or rollout scope. What matters is whether those claims show up in public evals, not the tagline.

Why it matters: This is a substantive Anthropic model release and clears HKR-H/K/R: a new Opus, three testable behavior claims, and strong resonance with Claude-heavy practitioners. The score stays in the high 80s because benchmarks, pricing, context window, and rollout scope are not disclosed.

Apr 15Wednesday

OpenAI News

The next evolution of the Agents SDK

OpenAI published a post about the next evolution of the Agents SDK. Only the title is available, with no body text or details, so specific features, numbers, and timing cannot be confirmed. For AI developers, it signals continued updates to the Agents SDK, but the scope is unclear from the source provided.

Why it matters: This is a substantive OpenAI developer-platform update: the post confirms native sandbox execution, a stronger agent-loop harness, and harness/compute separation, so HKR-H/K/R all pass. It stays below P1 because pricing, rollout scope, and performance numbers are not disclosed in

X · @claudeai

Now in research preview: routines in Claude Code

Anthropic launched routines in research preview for Claude Code: configure a prompt, repo, and connectors once, then run it on a schedule, via API, or from an event. Routines run on Anthropic web infrastructure, so a laptop does not need to stay open; the post does not disclose pricing, quotas, or rollout scope. The key point is hosted execution, not one-off code completion.

Why it matters: This is a substantive Claude Code expansion from local interactive coding to hosted, scheduled, and event-driven execution. HKR-H/K/R all pass, and the Anthropic update gets a policy bump, but price, quotas, and rollout scope are not disclosed, so it stays featured rather than P1

Apr 10Friday

X · @claudeai

We're bringing the advisor strategy to the Claude Platform.

Claude is adding the advisor strategy to Claude Platform, with Opus as the advisor and Sonnet or Haiku as the executor. The RSS snippet says this yields near-Opus-level agent intelligence at lower cost; the post does not disclose pricing, benchmark scores, or rollout timing.

Why it matters: Anthropic ships a substantive Claude Platform update, and HKR-H/K/R all pass: the Opus-advisor plus Sonnet/Haiku-executor setup is novel, concrete, and directly relevant to agent builders. The score stays below P1 because price, benchmarks, and rollout timing are not disclosed.

Apr 9Thursday

X · @claudeai

Introducing Claude Managed Agents: everything you need to build and deploy agents at scale.

Claude has launched Claude Managed Agents in public beta on Claude Platform, claiming to compress the path from agent prototype to launch into days. The post discloses only a performance-tuned agent harness plus production infrastructure; pricing, toolchain support, model scope, and quotas are not disclosed.

Why it matters: Anthropic gets a positive bump, and HKR-H/HKR-R pass because managed agent deployment is a strong hook for Claude-heavy builders. HKR-K is limited: the post discloses a harness and prod infra, but not pricing, toolchain support, model scope, or quotas.

Apr 3Friday

X · @claudeai

Computer use in Claude Cowork and Claude Code Desktop is now available on Windows

Claude has brought computer use in Claude Cowork and Claude Code Desktop to Windows. The post confirms the Windows rollout, but does not disclose supported versions, permission model, latency, pricing, or release timing. What matters is the reliability boundary for desktop agents on Windows, and the post gives no reproducible conditions yet.

Why it matters: HKR-H lands on the Windows rollout hook, and HKR-R lands because desktop agents on Windows map to real workflows. Score stays at 74: this is an official Claude update, but the post confirms availability only; versions, permissions, latency, and price are not disclosed.

Google DeepMind

Google DeepMind releases the Gemma 4 open model family

Google DeepMind released Gemma 4, which it calls its most intelligent open model yet, aimed at advanced reasoning and agentic workflows under an Apache 2.0 license. The family comes in four sizes: E2B, E4B, 26B MoE and 31B Dense. The 31B ranks 3rd among open models on the Arena AI text leaderboard, and the 26B ranks 6th.

Why it matters: Gemma 4 is Apache 2.0 and spans four sizes from on-device to workstation, so you can weigh deployment and fine-tuning options for open models.

Mar 31Tuesday

Mistral AI

Spaces: A CLI Built for Humans and Agents

Mistral AI 发布 Spaces CLI,同时面向人类开发者与编码智能体。它通过 `spaces init`、`spaces dev` 等命令快速搭建多服务项目,并为每个交互式提示提供对应的 flag 与 `-y` 选项,使智能体可自主完成配置与部署。每次 init 还会生成 context.json 和 AGENTS.md,为智能体提供项目上下文与操作规则。

Mar 25Wednesday

OpenAI News

Introducing the OpenAI Safety Bug Bounty program

OpenAI launched a public Safety Bug Bounty on March 25, 2026 for AI abuse and safety issues across its products. Scope includes agentic risks, proprietary information exposure, and account or platform integrity; third-party prompt injection must reproduce at least 50% of the time. This is not a jailbreak bounty: generic policy bypasses are out of scope.

Why it matters: This clears HKR-H/K/R: the public AI-safety bounty is novel, the post gives testable scope rules, and builders care about the reporting boundary. It stays in the low featured band because this is a governance/process update, not a model or capability launch.

Mar 18Wednesday

Mistral AI

Mistral AI launches Forge, an enterprise model-training system

Mistral AI launched Forge, a system for enterprises to build frontier-class AI models on their own proprietary knowledge. It covers pre-training, post-training and reinforcement learning, supports dense and MoE architectures, and handles multimodal input. Models can be trained and governed on in-house infrastructure. Mistral AI has worked with ASML, Ericsson, the European Space Agency and Singapore's DSO National Laboratories to train models on their proprietary data.

Why it matters: It details the staged capabilities and named partners behind enterprise frontier-model training on private data.

Mar 17Tuesday

NVIDIA Blog

GTC spotlights NVIDIA RTX PCs and DGX Spark running latest open models and AI agents locally

NVIDIA used GTC to showcase RTX PCs and DGX Spark for running local AI agents, and announced Nemotron 3 Nano 4B, Nemotron 3 Super 120B, and the open source NemoClaw stack. The post says DGX Spark has 128GB unified memory for models above 120B parameters; Nemotron 3 Super scored 85.6% on PinchBench, and Qwen 3.5 supports a 262,000-token context window. The key signal is local inference for privacy and zero token cost, while the full “latest open models” lineup and pricing are not disclosed in the post.

Why it matters: HKR-H/K/R all pass: the local-agent hook is strong, and the post includes concrete specs and benchmark numbers. I keep it in featured, not higher, because the full model list and pricing are not disclosed and the source is still a vendor launch post.

Mar 12Thursday

NVIDIA Blog

NVIDIA Nemotron 3 Super delivers 5x higher throughput for agentic AI

NVIDIA launched Nemotron 3 Super, a 120B open model with 12B active parameters, and says it delivers up to 5x higher throughput for agentic AI. It has a 1M-token context window and uses hybrid MoE, latent MoE, and multi-token prediction; the post says Blackwell NVFP4 gives up to 4x faster inference than Hopper FP8, with over 10T training tokens disclosed. What matters is that NVIDIA is releasing open weights, training recipes, and RL environments for reproduction and fine-tuning.

Why it matters: This is a solid model-release story with all three HKR signals, led by strong HKR-K: parameter counts, active params, context length, training scale, and Blackwell/Hopper comparison are all concrete. It stays below 85 because the key performance claims come from NVIDIA's own blog

Mar 11Wednesday

Mistral AI

Mistral builds an agent on Vibe that writes Rails tests automatically

Mistral built an agent on its open-source coding assistant Vibe that writes Rails RSpec tests on its own. It reads source code, generates or improves tests, checks them against style rules and coverage targets, and runs unattended in CI/CD.

Why it matters: Mistral published its full method for building an auto-RSpec-test agent on Vibe, including transferable details on context engineering, skill files and custom tools.

OpenAI News

From model to agent: Equipping the Responses API with a computer environment

OpenAI said on March 11, 2026 that Responses API now works with a shell tool and hosted container workspace, so models can execute commands in an isolated loop. The post says GPT-5.2 and later are trained to propose shell commands, while the API streams outputs and can run multiple commands concurrently across sessions; the container includes a filesystem, optional SQLite, and restricted network access. The key change is orchestration, not the “agent” label; pricing, quotas, and full security details are not disclosed in the visible post.

Why it matters: Substantive OpenAI developer update: the Responses API moves from tool calls to a managed computer environment with shell execution, streaming, parallel runs, and context compaction, so HKR-H/K/R all pass. The post is truncated and omits pricing, quotas, and full safety details,【