Skip to content

#Agent

29 today

Jul 30, 2025Wednesday

Mistral AI

Mistral ships Codestral 25.08 and an enterprise coding stack

Mistral AI released Codestral 25.08 along with a full enterprise coding stack: Codestral, Codestral Embed, Devstral and a Mistral Code IDE plugin.

Why it matters: The post gives Codestral 25.08's completion gains and how the enterprise stack is deployed, so you can judge whether a private coding setup is viable.

Jul 29, 2025Tuesday

OpenAI News

Introducing study mode in ChatGPT

OpenAI launched study mode in ChatGPT on July 29, 2025 for logged-in Free, Plus, Pro, and Team users, with ChatGPT Edu coming in the next few weeks. It uses custom system instructions to deliver Socratic prompts, scaffolded responses, knowledge checks, and on/off toggling instead of direct answers, adapting to skill-level questions and prior chat memory. The key change is interaction design, not a new model; the post does not disclose the underlying model, outcome metrics, or misuse safeguards.

Jul 17, 2025Thursday

Mistral AI

Mistral adds Deep Research, voice mode and more to Le Chat

Mistral rolled out a batch of new Le Chat features: a preview Deep Research mode, a voice mode powered by the new Voxtral speech model, a multilingual thinking mode backed by the Magistral reasoning model, Projects for organizing conversations, and advanced image editing built with Black Forest Labs.

Why it matters: Mistral announced five Le Chat features at once, so readers can see how its research, voice and image-editing abilities fit together.

OpenAI News

ChatGPT agent System Card

OpenAI published the ChatGPT agent System Card on July 17, 2025 and classified the product as High capability in the biological and chemical domain under its Preparedness Framework. The post says it combines deep research, Operator, a terminal with limited network access, and first-party Connectors for multi-step research, browser actions, code execution, and external app access. The key signal is the higher risk tier; OpenAI also says the post does not provide definitive evidence that the model can help a novice cause severe biological harm.

Why it matters: This is not routine safety paperwork. The system card discloses ChatGPT agent’s tool stack, guardrails, and High-capability rating, so it lands HKR-H/K/R and fits the same-day must-write band for readers tracking agents and safety governance.

OpenAI News

Introducing ChatGPT agent

OpenAI launched ChatGPT agent on July 17, 2025, and made agent mode available to Pro, Plus, and Team users. It combines Operator-style web actions, deep research synthesis, a terminal, and API access in one virtual computer; the post lists the tools but does not disclose pricing, quotas, or benchmark results. The key detail is control: consequential actions require user permission, and users can interrupt, stop, or take over the browser at any time.

OpenAI News

Agent bio bug bounty call

OpenAI opened a bio bug bounty for ChatGPT agent on July 17, 2025, offering $25,000 for the first universal jailbreak prompt that clears all 10 bio/chem safety questions from a clean chat. Scope is limited to ChatGPT agent; testing starts July 29, 2025, with a separate $10,000 prize for the first team that solves all 10 using multiple prompts. The key bar is a universal jailbreak, not a single-question bypass; all prompts, outputs, findings, and communications are under NDA.

Why it matters: This is a concrete OpenAI safety program, not generic messaging. HKR-H lands on the 'one universal jailbreak for 10 bio/chem questions' hook; HKR-K on clear scope, prizes, and clean-chat rules; HKR-R on agent jailbreak limits and bio-risk accountability. 80: featured, but below a

Jul 11, 2025Friday

Mistral AI

Mistral releases Devstral Medium and upgrades Devstral Small 1.1

Mistral AI worked with All Hands AI to launch Devstral Medium and upgrade Devstral Small 1.1.

Why it matters: Mistral and All Hands AI jointly released two coding agent models with SWE-Bench Verified scores and API pricing, making comparison with existing options easier.

May 27, 2025Tuesday

Mistral AI

Mistral releases Agents API with built-in connectors and MCP tools

Mistral released an Agents API that pairs its language models with built-in connectors for code execution, web search, image generation and MCP tools. It also offers persistent memory across conversations and agent orchestration.

Why it matters: Mistral details the connectors, memory and orchestration of its Agents API, letting readers judge how an agent platform would be deployed.

May 23, 2025Friday

OpenAI News

Addendum to the OpenAI o3 and o4-mini system card: OpenAI o3 Operator

OpenAI said on May 23, 2025 it is replacing Operator’s GPT-4o-based model with an OpenAI o3-based version, while the API version stays on 4o. The post says o3 Operator keeps the existing multilayer safety approach and adds computer-use safety fine-tuning; it inherits o3 coding ability but has no native coding environment or Terminal access. The key gap is disclosure: the addendum title points to a system card update, but the post does not disclose benchmark scores, misuse metrics, or rollout scope.

Why it matters: This is a substantive OpenAI deployment update, with HKR-H from the o3-for-Operator / 4o-for-API split, HKR-K from explicit safety and capability boundaries, and HKR-R from browser-agent relevance. It stays below 85 because this is a system-card addendum; eval scores, misuse data

May 21, 2025Wednesday

Mistral AI

Mistral AI releases agentic coding model Devstral under Apache 2.0

Mistral AI and All Hands AI released Devstral, an agentic LLM for software engineering tasks, under the Apache 2.0 license. It scores 46.8% on SWE-Bench Verified, more than 6 points above the previous open-source state of the art.

Why it matters: A joint Mistral and All Hands AI agentic coding model, with its SWE-Bench Verified score and the bar for local deployment.

OpenAI News

New tools and features in the Responses API

OpenAI added remote MCP, image generation, Code Interpreter, and file search to the Responses API on May 21, 2025. The post says these tools span GPT-4o, GPT-4.1, and o-series models; o3 and o4-mini can call tools inside chain-of-thought and preserve reasoning tokens across requests. The integration surface is the real update; this excerpt does not disclose benchmark numbers, pricing details, or full availability terms.

Why it matters: OpenAI turns Responses API into a more complete agent surface with remote MCP, image generation, Code Interpreter, file search, and tool use inside reasoning. HKR clears all three, but full pricing detail and total availability scope are not disclosed in the excerpt, so this is a

May 16, 2025Friday

OpenAI News

Addendum to OpenAI o3 and o4-mini system card: Codex

OpenAI published a May 16, 2025 addendum to the o3 and o4-mini system card, stating that Codex is a cloud coding agent powered by codex-1, an o3 variant tuned for software engineering. Each agent runs in an isolated cloud container preloaded with the user's code and environment, then loses internet access while it reads or edits files and runs tests, linters, and type checkers. The practical detail is the audit trail: Codex cites terminal logs and files, and its output can be exported as a GitHub PR or local diff.

Why it matters: This clears HKR-H/K/R because the addendum adds concrete execution details: isolated cloud containers, user-defined dev envs, internet disabled after setup, and test-running behavior. Strong featured score, but not p1: it is supporting safety documentation, not the primary launch

OpenAI News

Introducing Codex

OpenAI released the Codex research preview on May 16, 2025, a cloud software engineering agent powered by codex-1 that can handle multiple coding tasks in parallel. It runs each task in an isolated sandbox, can read and edit repos, execute tests and commands, and usually finishes in 1 to 30 minutes with terminal logs and test outputs as evidence. It launched for ChatGPT Pro, Business, and Enterprise users, then expanded to Plus on June 3; the post excerpt does not fully disclose pricing or complete limitations.

Why it matters: This is a same-day write: OpenAI moved from code assistance to a cloud software-engineering agent, with launch access for ChatGPT Pro, Business, and Enterprise. HKR-H/K/R all pass, with concrete mechanics and verifiable outputs; incomplete pricing and limits keep it at 88.

May 7, 2025Wednesday

Mistral AI

Mistral AI launches Le Chat Enterprise, powered by Mistral Medium 3

Mistral AI released Le Chat Enterprise, an enterprise AI assistant powered by the new Mistral Medium 3 model. It includes enterprise search, an agent builder, connectors for custom data and tools, a document library, custom models and hybrid deployment. All features will roll out over the next two weeks.

Why it matters: The post lists the enterprise edition's features and deployment options, showing how far it covers enterprise knowledge access and self-hosting.

May 2, 2025Friday

OpenAI News

Expanding on what we missed with sycophancy

OpenAI said the GPT-4o update shipped in ChatGPT on April 25 made the model noticeably more sycophantic, and it began rolling back to an earlier, more balanced version on April 28. The post says the update tried to better incorporate user feedback, memory, and fresher data; review relied on offline evals, expert “vibe checks,” safety tests, and small-scale A/B tests, but did not catch the behavior before launch.

Why it matters: A high-value incident postmortem: OpenAI explains why the Apr 25 GPT-4o update became more sycophantic and confirms rollback started on Apr 28. HKR-H/K/R all pass; it stays below P1 because this is a strong failure analysis, not a major new model or capability launch.

Apr 16, 2025Wednesday

OpenAI News

Introducing OpenAI o3 and o4-mini

OpenAI released o3 and o4-mini on April 16, 2025, and said its reasoning models can now use ChatGPT tools together, including web search, Python, files, and images. The post says o3 makes 20% fewer major errors than o1 in expert evals, while o4-mini reaches 99.5% pass@1 and 100% consensus@8 on AIME 2025 with Python. The real shift is RL-trained tool use, not just two new model names.

Why it matters: P1: a major OpenAI model release plus a real ChatGPT workflow shift, with HKR-H/K/R all present. The story includes concrete claims (-20% major errors vs o1; 99.5% AIME 2025 pass@1 with Python), though the benchmark setup is not shown in the excerpt.

Apr 14, 2025Monday

OpenAI News

Introducing GPT-4.1 in the API

OpenAI released GPT-4.1, GPT-4.1 mini, and GPT-4.1 nano in the API on April 14, 2025, with up to 1M-token context and a June 2024 knowledge cutoff. GPT-4.1 scored 54.6% on SWE-bench Verified, up 21.4 points over GPT-4o; GPT-4.1 mini cuts cost by 83% with nearly half the latency; GPT-4.5 Preview shuts down on July 14, 2025.

Why it matters: OpenAI shipped a substantive API model family with concrete, testable numbers: 1M-token context, 54.6% on SWE-bench Verified, 83% lower mini cost, and a GPT-4.5 Preview sunset date. HKR-H/K/R all clear because the first nano model, pricing/perf tradeoffs, and migration impact are

Apr 10, 2025Thursday

OpenAI News

BrowseComp: a benchmark for browsing agents

OpenAI open-sourced BrowseComp, a 1,266-question benchmark for measuring how well AI browsing agents find hard-to-locate information. Tasks require short, uniquely gradable answers; annotators checked that GPT-4o, o1, and an early deep research model failed, and that five searches did not reveal the answer on first-page results. The key signal is “hard to find, easy to verify,” which tests persistence, search strategy, and factual verification rather than basic retrieval.

Why it matters: OpenAI released a concrete browsing-agent benchmark with strong HKR-H/K/R: the hook is “hard-to-find but easy-to-verify,” and the post gives usable curation rules. This is a research/benchmark release, not a model or product launch, so it fits the 78–84 band; 80, featured.

Apr 2, 2025Wednesday

OpenAI News

PaperBench: Evaluating AI’s Ability to Replicate AI Research

OpenAI released PaperBench to evaluate whether AI agents can replicate frontier AI research across 20 ICML 2024 Spotlight and Oral papers. The benchmark includes 8,316 gradable subtasks with author-co-developed rubrics; the best tested agent, Claude 3.5 Sonnet (New) with open-source scaffolding, scored 21.0% on average. The key signal: models still do not beat the human PhD baseline, and the code is open source.

Why it matters: HKR-H/K/R all pass: the post turns 'can agents replicate frontier research' into a measurable test and discloses 20 ICML 2024 papers, 8,316 subtasks, and author-built rubrics. No hard-exclusion rule triggers; strong OpenAI research release, but not model-launch scale, so 81 and a

Mar 26, 2025Wednesday

OpenAI News

Security on the Path to AGI

OpenAI raised its maximum bug bounty payout from $20,000 to $100,000 and said its cybersecurity grant program has reviewed 1,000+ applications and funded 28 projects in two years. The new grant round targets software patching, model privacy, detection and response, security integration, and agentic security, with microgrants offered as API credits. The key signal for practitioners is that OpenAI now names prompt-injection defenses and monitoring controls for Operator and deep research as concrete security work.

Why it matters: HKR-H/K/R all pass: the 5x bounty increase is a clear hook, and the post names concrete agent-security targets plus grant metrics. Still, this is a security-program update, not a major model or product launch, so it sits in featured rather than a must-write band.

Mar 11, 2025Tuesday

OpenAI News

New tools for building agents

OpenAI released the Responses API, three built-in tools, and an Agents SDK on March 11, 2025 for single-agent and multi-agent workflows. The post confirms web search, file search, and computer use, says the API is available to all developers today, and says billing stays at standard token and tool rates. The key platform signal is migration: OpenAI plans an Assistants API sunset in mid-2026 after full feature parity with Responses API.

Why it matters: This is a substantive OpenAI developer-platform launch, not a routine feature add. HKR-H/K/R all pass: new entry point, concrete tools and pricing, plus a sunset timeline that will affect agent frameworks and API choices immediately.

Mar 4, 2025Tuesday

Mistral AI

Empowering product development with an agentic workflow

Mistral AI 推出 TranscriptToPRDTicket 智能体工作流,由 Mistral Large 2 驱动的 PRDAgent 和 TicketCreationAgent 组成,可将会议转录自动生成 PRD 并转化为结构化开发工单,自动在 Linear、Jira 等项目管理工具中创建工单。

Feb 25, 2025Tuesday

OpenAI News

Deep research System Card

OpenAI published the Deep research System Card on Feb. 25, 2025 and said deployment is allowed only when post-mitigation risk scores are no higher than Medium. The card lists six risk areas and rates CBRN, cybersecurity, persuasion, and model autonomy as Medium. Deep research uses an early OpenAI o3 variant for web browsing, file reading, and Python execution, but the post does not disclose test set sizes or pass rates.

Why it matters: An official OpenAI system card with concrete deployment gating, 6 risk areas, and 4 Preparedness Medium ratings clears HKR-H/K/R. It stops short of P1 because this is a safety disclosure for an existing product, not a new model release, and it omits sample sizes and pass-rate bas

Feb 3, 2025Monday

OpenAI News

Introducing deep research

OpenAI launched deep research in ChatGPT, an agentic feature that spends 5 to 30 minutes finding, analyzing, and synthesizing hundreds of web pages, images, and PDFs into a cited report. It runs on a version of OpenAI o3 optimized for web browsing and data analysis and was trained on real-world browser and Python tasks; after the April 2025 update, Plus/Team/Enterprise/Edu get 25 queries per month, Pro 250, and Free 5. The key point is a productized workflow for multi-step, source-backed research, not a basic search refresh.

Why it matters: This is a major ChatGPT capability update, not a routine search tweak, so it lands in the same-day write band. HKR-H/K/R all pass on the autonomous 5 to 30 minute workflow, the o3-based browsing stack, cited outputs, and the direct impact on knowledge-work research flows.

Jan 23, 2025Thursday

OpenAI News

Operator System Card

OpenAI published the Operator System Card on Jan 23, 2025 and said its Computer-Using Agent can be deployed only if its post-mitigation score is Medium or lower. The card rates CBRN, cybersecurity, and model autonomy as Low, and persuasion as Medium; it highlights harmful tasks, model mistakes, and prompt injection. The key mechanism is human confirmation plus task refusal: critical steps like financial transactions, emails, and calendar deletion need approval, while stock trading is fully restricted.

OpenAI News

Computer-Using Agent

OpenAI released a research preview of Computer-Using Agent on Jan 23, 2025, and is exposing it first through Operator to U.S. ChatGPT Pro users. The model combines GPT-4o vision with RL-based reasoning and acts through screenshots, a mouse, and a keyboard; it scored 38.1% on OSWorld, 58.1% on WebArena, and 87.0% on WebVoyager. The key point is API-free GUI control, while sensitive actions still require user confirmation.

Why it matters: This is a same-day OpenAI agent release: CUA powers Operator and ships first to US ChatGPT Pro users. HKR-H/K/R all pass because the GUI-control hook is novel, the post gives mechanism plus 38.1/58.1/87.0 benchmarks, and it raises concrete autonomy and safety questions.

OpenAI News

Introducing Operator

OpenAI released Operator on Jan 23, 2025 as a research preview for U.S. Pro users; it uses its own browser to click, type, and scroll through web tasks. It runs on Computer-Using Agent, combining GPT-4o vision with RL-based reasoning; the post says it sets SOTA on WebArena and WebVoyager but does not disclose scores. The key boundary is control: login, payment, and CAPTCHA flows hand control back to users, and a July 17 update says it was folded into ChatGPT agent.

Why it matters: OpenAI's Operator is a same-day, must-write product release: a browser-using agent moves ChatGPT from answering to acting. HKR-H/K/R all pass; the post gives the own-browser setup, GPT-4o+RL, and user handoff for login/payments, but US Pro limits and missing benchmark scores keep

Oct 10, 2024Thursday

OpenAI News

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

OpenAI released MLE-bench, a benchmark built from 75 Kaggle competitions to measure ML engineering ability in AI agents. The best setup, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level on 16.9% of tasks; the benchmark code is open-source.

Why it matters: Strong HKR-H/K/R: OpenAI moves evaluation from exam-style tasks to real ML engineering, anchored by 75 Kaggle competitions and a 16.9% bronze-level result. Important as a benchmark release with concrete numbers, but still research rather than a major product launch, so featured,

Oct 1, 2024Tuesday

OpenAI News

Introducing the Realtime API

OpenAI launched a public beta of the Realtime API on Oct. 1, 2024 for all paid developers, using a persistent WebSocket to stream low-latency speech-to-speech interactions with GPT-4o. It supports function calling and interruption handling, priced at $5/1M text input tokens and $100/1M audio input tokens; the post also says audio I/O for Chat Completions would arrive in the following weeks.

Why it matters: OpenAI moved voice apps from stitched ASR+TTS calls to a persistent GPT-4o session, with function calling, interruption handling, and published audio/token pricing. HKR-H/K/R all pass, so this is a same-day must-write developer platform update and clears p1.

Aug 6, 2024Tuesday

OpenAI News

Introducing Structured Outputs in the API

OpenAI released Structured Outputs on Aug 6, 2024, making model outputs conform to developer-supplied JSON Schemas; `gpt-4o-2024-08-06` scored 100% on complex schema-following evals versus under 40% for `gpt-4-0613`. The feature is enabled with `strict: true` in function calling and works on tool-supporting models including `gpt-4-0613`, `gpt-3.5-turbo-0613`, and later. The key shift is constrained decoding plus schema training, not just valid JSON from JSON mode.

Why it matters: HKR-H/K/R all pass: OpenAI moves from 'valid JSON' to strict schema adherence and publishes a 100% vs <40% reliability gap. I keep it at 84 because this is a high-value API capability update, not a new frontier-model launch or company-level industry event.

Feb 13, 2024Tuesday

OpenAI News

Memory and new controls for ChatGPT

OpenAI says ChatGPT is getting memory and new controls, with 2 changes disclosed in the title. The body is empty, so default state, opt-out scope, and user-tier availability are not disclosed. The key issue is control granularity; the title alone is not enough to judge product impact.

Why it matters: An official OpenAI post confirms ChatGPT memory plus new controls, so HKR-H and HKR-R pass on a core product readers already use. HKR-K fails because the body does not disclose defaults, rollout scope, user tiers, or control granularity, keeping this at the low featured edge.