Skip to content

#Agent

0 today

Yesterday · Sep 29Tuesday

OpenAI News

OpenAI releases GPT-6.1 Sol model

OpenAI released GPT-6.1 Sol, positioned as near-Astra-level intelligence for coding, computer use and professional work. Standard API input and output tokens cost one-fifth of Astra's price.

Why it matters: OpenAI's GPT-6.1 Sol launch shows the capability target for coding and computer use, plus the pricing shift.

OpenAI News

OpenAI releases proactive assistant dots

OpenAI released dots, a proactive assistant that keeps work moving on complex projects and everyday tasks. OpenAI says dots keeps users in control as tasks progress.

Why it matters: OpenAI's dots launch shows where the company places a proactive assistant across complex projects and daily tasks.

Sep 25Friday

GitHub Blog · AI & ML

When chat is the wrong UI

GitHub Copilot 应用推出 canvas,一种运行在应用内、无浏览器外壳的全栈小应用,可与 Copilot 智能体双向通信,并能在本地执行代码、调用第三方 API。作者认为聊天只是 AI 的通用兜底界面,用户明确任务时更该让智能体生成可复用工具,而非把智能体本身当工具、白白消耗 token。示例包括 Connect 4 游戏、Winget 包管理、SQLite 操作和开发工作流自动化。

Sep 18Friday

GitHub Blog · AI & ML

Should you read the code, is RAG dead, and did Skills kill MCP?

GitHub Podcast 最新一期拆解了五个 AI 热门观点:AI 生成的代码仍需阅读和负责,但审查力度应按风险分级;Skills 与 MCP 解决不同问题,前者是打包的团队经验,后者是连接工具与数据的标准,可组合使用;RAG 并未死亡,它为模型提供训练数据之外的相关信息,减少 token 浪费并让回答更有依据。

Sep 11Friday

GitHub Blog · AI & ML

GitHub Copilot app for Beginners: Using the diff, terminal, and browser

GitHub Copilot 应用内置 diff、终端和浏览器三个面板,让用户无需离开应用即可审查、运行和预览 AI 智能体生成的代码变更。diff 面板以绿色和红色高亮显示代码的增删改,终端面板支持直接运行项目命令并可通过 Run 按钮配置脚本,浏览器面板则提供 Pick & Polish 工具来选取页面元素并让智能体调整。

Sep 8Tuesday

OpenAI News

OpenAI CFO: GPT‑6 Astra is here, and consumer + enterprise reinforce each other

OpenAI CFO Sarah Friar published a blog framing GPT‑6 Astra as the world's most capable and aligned model. ChatGPT now has over 1B weekly active users and 2.5M business customers. Internally, the research org uses 3.1 agent-workdays per human workday. The post also claims an internal model solved the Navier–Stokes Millennium Prize Problem, but gives no technical detail. I'd treat this as a strategy narrative, not a technical report.

Why it matters: OpenAI CFO publishes a strategic framing piece for GPT-6 Astra with two concrete numbers: 1B weekly users and a 3.1x agent-workday ratio. Hits all three HKR axes. No technical details — this is narrative, not a product launch — so it stays below 85.

Sep 4Friday

GitHub Blog · AI & ML

GitHub Copilot app for Beginners: Run several agents at once

GitHub Copilot 应用支持同时运行多个智能体会话,每个会话运行在独立的 Git worktree 上,互不干扰且各自保留上下文,可随时切换并从中断处继续。用户可在会话视图中查看各任务标题与进度,例如在同一项目上并行执行 funded sort 开发、无障碍审查和测试运行。

Sep 3Thursday

Google DeepMind

Google DeepMind launches Fairwind, opening Gemini 3.8 Flash Cyber to governments and trusted partners

Google DeepMind launched the Fairwind Program, giving government agencies, critical infrastructure operators and cybersecurity partners limited access to its most advanced cyber defense capabilities. The program pairs a dedicated cyber model, Gemini 3.8 Flash Cyber, with the CodeMender harness to autonomously find, verify and fix vulnerabilities, cutting weeks of manual remediation to deployable patches generated in minutes, at lower cost than traditional frontier models.

Why it matters: The post names Fairwind's eligible users and its model-plus-tool setup, a basis for judging autonomous vulnerability patching in enterprise and government settings.

Sep 1Tuesday

Anthropic News

Anthropic launches Enterprise Frontier Safeguards with customer-held data and keys

Anthropic released Enterprise Frontier Safeguards (EFS), which pairs zero data retention (ZDR) privacy with safety monitoring for abuse detection. Data sits in the customer's own cloud infrastructure rather than at Anthropic.

Why it matters: The piece details EFS's data retention and monitoring architecture, so readers can weigh privacy against safety when deploying frontier models.

Aug 25Tuesday

Hugging Face Blog

IBM details the full pipeline behind Granite 4.2, from pre-training to agentic RL

IBM published a technical walkthrough of the Granite 4.2 model family on the Hugging Face blog. It covers architecture, pre-training, SFT data quality control, and a multi-stage RL pipeline. The RL curriculum has three phases: foundational skills, agentic RL for tool use on the 8B and 30B models, and RLHF alignment. The post also mentions FP8, FP4, and GGUF quantization. Specific benchmark scores and hardware details are not included in the provided excerpt.

Why it matters: A solid training pipeline breakdown with strong H and K, but Granite's limited community pull drags down R. The post doesn't disclose pretraining data or hardware specs, so it can't push past 78. Featured because the engineering detail is real — model trainers will bookmark this.

Aug 21Friday

Aug 20Thursday

OpenAI News

OpenAI launches Strategic Futures team and AI Futures blog on AI, power, and human agency

OpenAI announced a small Strategic Futures team and its blog AI Futures. The first post by Dean Ball frames the core problem: if states can project force and collect revenue through autonomous systems and data centers instead of human labor and consent, individual agency may erode even if formal democracy remains. It argues against radical decentralization and calls for a new balance of power, citing the Founders' Newtonian checks-and-balances model. The post is a research agenda; it does not propose specific policies.

Why it matters: OpenAI launches 'AI Futures,' a blog from its Strategic Futures team, with a debut post tackling the thorniest long-term risk: concentration of power. It has a clear analytical frame and isn't PR fluff. The cap at 78 is because this is just a blog launch — no concrete research...

Aug 19Wednesday

OpenAI News

Replit launches Free Mode powered by GPT-5.6 Luna, removing token costs for software creation

Replit introduced Free Mode running on GPT-5.6 Luna, so users can plan, ideate, and explore projects without tracking token spend. CEO Amjad Masad credits recent OpenAI price cuts for making the free tier viable at millions-of-users scale. Complex reasoning tasks get routed to GPT-5.6 Sol, then return to Luna while preserving project context. Sam Altman frames it as a step toward anyone with internet building a product or startup. The post does not disclose Free Mode quotas, concurrency limits, or the exact launch date.

Why it matters: Replit's free tier running GPT-5.6 Luna is a concrete product update with a real mechanism (dual-model handoff) and a direct CEO quote on cost economics — enough signal for featured. But it's an OpenAI customer story, not a model release, so the score stays at 72.

Aug 7Friday

OpenAI News

OpenAI says unreleased model Astra may hit its Critical cyber threshold

OpenAI disclosed on Aug 7 that internal evals of its upcoming model Astra show enough progress in agentic coding and cybersecurity that it can no longer rule out a Critical rating under its Preparedness Framework. The Critical bar means the model can autonomously find and write zero-day exploits for hardened real-world systems, or devise and execute novel end-to-end attacks given only a high-level goal. OpenAI confirmed Astra was not involved in the earlier Hugging Face incident. It has paused internal Astra work that doesn't meet tightened security controls, added isolated test environments, restricted network/tool access, encrypted model weights, deployed universal monitoring on all agentic Astra applications, and will bring in government and safety organizations for testing.

Why it matters: OpenAI voluntarily disclosed that its next-gen model Astra reached 'critical' risk level in internal testing — the first time a major lab has gone public with such an assessment before release. The post gives concrete capability definitions and touches the sensitive topic of a...

Aug 3Monday

OpenAI News

OpenAI details GPT-Live: a full-duplex voice system that drops the turn detector and streams audio continuously

OpenAI published an engineering post on Aug 3 explaining GPT-Live’s realtime voice stack. The key change: they removed the turn detector from the audio path and switched to a full-duplex model that listens and speaks simultaneously. This avoids the old problem of a tiny model guessing when the user has finished, and lets the large model stream audio directly for more natural timing. When deeper reasoning or tool use is needed, the system delegates asynchronously to frontier models like GPT-5.5 without blocking the live voice loop. The team spent six months reworking inference, context management, and media transport to keep latency low end-to-end. The post says this architecture already powers computer control and agent coordination in the ChatGPT desktop app, but it does not disclose specific latency figures or deployment scale.

Why it matters: Official OpenAI engineering post explaining the architecture shift from turn-based to full-duplex voice for GPT-Live, with concrete technical decisions. Not a product launch—it's a developer-facing deep-dive. Hits all three HKR axes. Score stays at 78 rather than 85+ because t...

Jul 31Friday

OpenAI News

OpenAI lays out its “abundant intelligence” playbook: price cuts, efficiency gains, and a full-stack flywheel

OpenAI published a strategy post on July 31 explaining its “abundant intelligence” approach. The core loop: more capable and cheaper models drive broader adoption, which generates revenue and feedback to fund the next round of R&D and infrastructure. Concrete numbers: GPT-5.6 Luna input/output prices dropped 80% to $0.20/$1.20 per million tokens; GPT-5.6 Terra dropped 20%. GPT-5.6 Sol Fast mode delivers 2.5x speed at 2x price with no intelligence change. On the engineering side, Sol helped cut end-to-end serving costs by 20% and improved speculative-decoding efficiency by over 15%. On the public ARC-AGI-3 benchmark, better retained reasoning and context management lifted Sol’s score from 13.3% to 38.3% while using 6x fewer output tokens. Product stats: ChatGPT has over 1B active users and 2M businesses; six months after signup, daily messages rise ~50% and use-case breadth roughly doubles. Agentic work via Codex now accounts for 99.8% of OpenAI’s weekly output tokens. No new model was announced—this is a strategy piece.

Why it matters: OpenAI's official blog lays out its 'abundant intelligence' strategy with concrete pricing data (GPT-5.6 Luna down 80%). Not a product launch, so it doesn't hit 85, but as a strategic signal it's worth featuring.

Jul 21Tuesday

Google DeepMind

Google DeepMind releases Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber

Google DeepMind released three new models: Gemini 3.6 Flash, 3.5 Flash-Lite, and the security-focused 3.5 Flash Cyber.

Why it matters: It gives pricing, token efficiency and benchmark comparisons for all three models, so readers can judge cost and model choice for agent workflows.

Jul 16Thursday

Hugging Face Blog

Hugging Face discloses an end-to-end autonomous AI agent intrusion into its production infrastructure

On July 16, Hugging Face disclosed that an autonomous AI agent system breached its production infrastructure through a malicious dataset. The attacker exploited remote-code loading and template injection in the dataset pipeline, escalated to node-level access, harvested cloud and cluster credentials, and moved laterally across internal clusters over a weekend. The campaign involved tens of thousands of automated actions with self-migrating C2 on public services. Hugging Face closed the initial vulnerability, rotated credentials, rebuilt compromised nodes, and tightened cluster admission controls. No tampering with public models, datasets, or Spaces was found; the software supply chain was verified clean. The post does not specify which LLM the attacker used or whether any partner/customer data was affected.

Why it matters: Hugging Face's official disclosure of a fully autonomous AI agent breaching their production environment is the first real-world case of its kind, with a complete attack chain and concrete details. All three HKR axes hit: the headline creates suspense, the body reveals specifi...

Hugging Face Blog

Ai2 shares the engineering lessons behind Shippy, a maritime AI agent

Ai2's Skylight team built Shippy, an AI assistant that helps maritime analysts query fishing activity, EEZ boundaries, and vessel tracks. The post breaks its architecture into three parts: a soul (system prompt), skills (markdown files that teach it to call APIs and interpret track data), and config (runtime settings; currently Claude Opus 4.6 with the OpenClaw framework). The core idea is wrapping a non-deterministic model in deterministic tools—every answer includes source, data cutoff, and a deep link to the Skylight map so an analyst can verify it. The post doesn't disclose error rates or latency numbers, but it stresses sandboxed hosting and evaluating the agent as a system, not just the model.

Why it matters: Ai2's three-layer agent architecture (soul/skills/config) and the Markdown-as-skill-sheet pattern are concrete engineering takeaways. But the maritime domain is too niche for broad resonance, landing right at the featured threshold.

Hugging Face Blog

Model routing is simple—until you measure real cost, not sticker price

IBM Research found that routing by model sticker price backfired in agent workloads. Across 417 AppWorld tasks, Claude Sonnet 4.6 cost $79 total vs. GPT-4.1's $155—nearly double—because Sonnet's lower cache-read pricing exploited high context reuse across steps. The post argues real cost, latency, and complexity all depend on workload-infrastructure interaction, making routing a systems optimization problem, not a classification one.

Why it matters: IBM ran 417 AppWorld tasks and found that routing by list price alone fails—Sonnet 4.6 cost $79 total while GPT-4.1 cost $155, nearly double. The core insight: when agents reuse the same context repeatedly, cache-read pricing dominates the total bill. Concrete numbers, counter...

Jun 25Thursday

OpenAI News

OpenAI publishes economic research paper on how Codex is reshaping work

OpenAI released an economic research paper on June 25, using internal and external usage data to track Codex adoption over the past year. By May 2026, 80.6% of sampled individual users had run at least one Codex task estimated to exceed 30 minutes of human work, and 25.6% had run tasks exceeding eight hours. Inside OpenAI, Codex now accounts for 99.8% of weekly output tokens; Legal and Recruiting switched their primary AI tool from ChatGPT to Codex around April 2026. Non-developer users grew fastest—137x for individuals, 189x for organizations. The paper does not disclose Codex pricing or external enterprise conversion rates.

Why it matters: OpenAI's economic research team published a paper quantifying Codex's shift from chat to long-horizon agent tasks, with 80.6% and 25.6% penetration as the core hooks. It's a self-published promotional study, not independent research, so the score stays below 85.

Jun 17Wednesday

Hugging Face Blog

Z.AI releases GLM-5.2: first open-source model with solid 1M-token context, built for long-horizon coding tasks

Z.AI open-sourced GLM-5.2, a model built for long-horizon coding tasks. It delivers a genuinely usable 1M-token context—not just accepting more tokens, but maintaining quality across long agent trajectories. IndexShare reuses one indexer across every four sparse attention layers, cutting per-token FLOPs by 2.9× at 1M context; MTP acceptance length improved by up to 20%. On FrontierSWE it beats GPT-5.5 by 1%, and on PostTrainBench it outranks both GPT-5.5 and Opus 4.7, placing second. It's the top open-source model across all three long-horizon coding benchmarks. MIT license, no regional restrictions.

Why it matters: Z.AI open-sources GLM-5.2 with a 1M-token context window and two new architectural components, explicitly targeting long-horizon agent tasks. Domestic flagship model release gets full weight per policy, but the body excerpt lacks full benchmarks, capping it below 85.

Hugging Face Blog

Hugging Face launches ARD discovery tool so agents can search for tools, skills, and other agents

Hugging Face released Discover Tool, a reference implementation of the Agentic Resource Discovery (ARD) spec. ARD is an open draft co-developed by Microsoft, Google, GoDaddy, Hugging Face, and others. It lets agents find MCP tools, A2A agents, or skills at runtime via natural-language search instead of hardcoding each one. Hugging Face's implementation wraps the Hub's existing semantic search and Agent Skills into an ARD catalog, exposed as a REST API and an MCP Tool. The post does not disclose pricing, search latency, or accuracy figures.

Why it matters: ARD tackles a real pain point—agent tool discovery—with cross-vendor backing from Microsoft, Google, and Hugging Face, plus a working reference implementation. Not scoring higher because it's still an open draft, not a ratified standard, and the post doesn't spell out adoption...

Google DeepMind

Unlocking UK house-building with AI-accelerated planning

Google DeepMind 正与英国政府、Google Cloud、Faculty 及 Barnet、Dorset、Camden 地方规划部门合作,基于 Gemini 共同开发 AI 规划原型工具,目标将住户规划申请审批时间缩短 50%。

Jun 16Tuesday

Google DeepMind

Google DeepMind publishes AI Control Roadmap for internal AI agents

Google DeepMind published an AI Control Roadmap, a framework for building and managing advanced AI deployed inside Google. It takes a defense-in-depth approach, adding system-level safety layers on top of model alignment so protections hold even when alignment is imperfect.

Why it matters: DeepMind made its internal AI Control Roadmap public, laying out a layered way to monitor and block agents as if they were insider threats.

Jun 3Wednesday

NVIDIA Blog

NVIDIA Research Presents Grasping, Autonomous Driving and Agent Training Work at CVPR

NVIDIA Research presented three physical AI papers at CVPR: GraspGen-X was trained on 2 billion simulated grasps, LCDrive cuts reasoning tokens by about half versus text-based reasoning, and NitroGen trains embodied agents across more than 1,000 games and 40,000 hours of interaction.

Why it matters: HKR-H/K/R all pass: NVIDIA’s CVPR bundle gives concrete mechanisms and scale numbers. It stays in the low 78–84 band because it is a vendor research roundup, not a major model or product launch.

Alibaba Technology · WeChat

Rethinking R&D Infrastructure When Agents Become First-Class Citizens

Xu Xiaobin argues that agent-based development compresses the intent-to-code loop from weeks or months to minutes, using a weekly-report system, a multi-role agent development setup, and image-repository provisioning as examples; the article identifies mismatches in Git, CI, code review, release flows, permissions, harness setup, and dry-run validation.

Why it matters: HKR-H/K/R all pass, but this is infrastructure commentary rather than a model or product launch. The named cases and week/month-to-minutes claim put it in the 72–77 featured band.

NVIDIA Blog

NVIDIA Partners With Microsoft on Unified Stack for Agentic AI Deployment

NVIDIA and Microsoft announced a unified agentic AI deployment stack at Build across Windows, Azure, and local environments; RTX Spark provides 1 petaflop of AI performance, while DGX Station for Windows offers 20 petaflops of FP4 performance and up to 748GB of coherent memory.

Why it matters: HKR-H/K/R pass: the NVIDIA-Microsoft stack spans Windows, Azure, and local devices, with 1 PFLOP and 20 PFLOPs FP4 specs. Vendor-source limits the score: pricing, benchmarks, and migration details are not disclosed.

May 28Thursday

NVIDIA Blog

NVIDIA Research Advances Robotics From Simulation to the Real World

NVIDIA Research presented 8 ICRA papers on sim-to-real robotics: ScheduleStream delivered a 3x speedup for multi-arm planning, COMPASS reached about 80% success across 20 real-world navigation trials, and Grasp-MPC achieved about 75% real-robot grasping success.

Why it matters: HKR-K and HKR-R are strong: the post gives concrete sim-to-real numbers from ICRA and addresses robot deployment reliability. HKR-H is moderate but passes on the real-world success-rate hook.

Mistral AI

Mistral upgrades Le Chat into unified agent Vibe, covering office work and coding

Mistral upgraded Le Chat into a unified AI agent called Vibe, with one license covering both office work and coding. Existing chats, settings and plans all carry over. Work Mode supports enterprise knowledge search, structured data analysis, document and report generation, scheduled multi-step tasks and reusable skills, and connects to Google Workspace, Outlook, SharePoint, Slack, GitHub and more.

Why it matters: It discloses Vibe's Work Mode, coding mode and CLI updates in full, so readers can judge how it plugs into existing workflows.

Alibaba Technology · WeChat

AI-Native Project Management: Two Git Repos Replace Weekly Updates, Insights, and Metrics Reports

Zhou Zhiwei describes a project-management setup that uses two Git repositories, an AI coding assistant, Shell, and Python to replace at least 80% of manual weekly-update chasing, data moving, chart generation, and engineering-metrics reporting.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, the post gives an 80% replacement claim and a two-repo mechanism, and it hits engineering-management toil. This is a strong practical workflow piece, not a model or platform launch.

Hugging Face Blog

ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks

Artificial Analysis and IBM published the ITBench-AA title, saying frontier models scored below 50% on an enterprise IT agent task benchmark; the post does not disclose tested models, sample size, or scoring method.

Why it matters: HKR-H/R pass: frontier models under 50% on enterprise IT agent tasks is clickable and deployment-relevant. HKR-K is weak because models, sample size, and scoring are not disclosed, so it stays near the featured floor.

May 27Wednesday

Alibaba Technology · WeChat

From Language Emergence to Collaborative Emergence: How AI Can Make High-Quality Decisions

Lv Ruofan proposes the Agent Room model: multiple agents share context, a task ledger, Memory, Runtime, and Artifacts, and two software-engineering cases show the system moving from workflow automation toward collaborative judgment rather than predefined task routing.

Why it matters: HKR-H/K/R all pass, but this is a methodology piece rather than a model launch or open-source framework. Concrete Agent Room mechanisms and 2 R&D sites put it in the 72–77 featured band.

May 22Friday

Mistral AI

Mistral launches Connectors in Studio with built-in and custom MCP

Mistral launched Connectors in Studio. All built-in connectors and custom MCP are now callable through the API/SDK by every model and agent. New features include direct tool calling, human-in-the-loop approval flows, and programmatic access to create, modify, list and delete connectors.

Why it matters: The original gives the API usage and code examples for Connectors, enough to judge how enterprise MCP integration gets built.

May 21Thursday

Alibaba Technology · WeChat

Building an Agent from 0 to 1: Principles and Personal Assistant Practice

Zhan Xupeng published a roughly 50-minute article on Agent theory and a personal assistant implementation, covering memory, ReAct planning, progressive skill loading, subagents, and harness-level fault recovery.

Why it matters: HKR-K/R pass via concrete agent mechanisms and practitioner reliability pain; HKR-H is weak because the headline is a standard tutorial frame. This fits the quality-tutorial threshold, not the 78+ news band.

May 20Wednesday

Alibaba Technology · WeChat

Zhenwu M890 AI Chip Debuts as Agentic Compute Foundation

Alibaba released a 128-card supernode server based on T-Head’s Zhenwu M890 AI chip, with P2P latency below 150 ns and rack bandwidth at the Pb/s level; it is live on Alibaba Cloud Bailian and supports Qwen, DeepSeek, and Kimi.

Why it matters: HKR-H/K/R all pass, but the source is Alibaba’s own tech post and lacks third-party benchmarks, pricing, or production volume. Score stays in the featured-threshold band for an AI infrastructure product update.

May 18Monday

Google DeepMind

Google DeepMind adds Street View grounding to Project Genie

Google DeepMind has added Street View real-scene grounding to its experimental prototype Project Genie. Users can pick a US location, then pair it with a style and characters to generate a world.

Why it matters: With Street View imagery wired in, agents and robots can train and navigate in virtual environments that track real places.

Google DeepMind

Introducing Google Antigravity 2.0

Google 发布智能体开发平台 Google Antigravity 2.0。该平台在 Google DeepMind 官网被列为面向开发者的 agentic development platform,与 Gemini 应用、Google AI Studio 并列。原文未披露版本功能、参数或可用性细节。

May 17Sunday

Google DeepMind

Google DeepMind launches Gemini for Science toolset

Google DeepMind released Gemini for Science, which includes three experimental tools on Google Labs: Hypothesis Generation, built on Co-Scientist.

Why it matters: Google is packaging research prototypes like Co-Scientist and AlphaEvolve into apply-to-use science tools, showing what agentic research looks like in practice.

May 16Saturday

Google DeepMind

Strengthening Singapore’s AI Future: A New National Partnership

Google DeepMind 宣布与新加坡政府达成国家 AI 合作,在新加坡推出多项计划,聚焦医疗健康、科学发现与教育。合作内容包括探索 AI 辅助临床医生、用 AlphaFold 和 Google Earth 推进东南亚传染病研究、为盲人及低视力跑者开发基于 Gemma 的跑步助手,并向中小学至初级学院教育者提供 Gemini for Education。