Skip to content

#Agent

0 today

Mar 9Monday

OpenAI News

OpenAI to acquire Promptfoo

OpenAI said it will acquire Promptfoo and integrate its technology into OpenAI Frontier after closing. The post discloses that Promptfoo is used by over 25% of Fortune 500 companies, and the deal is still subject to customary closing conditions. The key signal is native agent security testing, red-teaming, and traceability in Frontier; the post does not disclose price or timeline.

Why it matters: This is not a routine partnership; OpenAI is absorbing a known eval and red-team vendor into Frontier. HKR-H/K/R all pass on novelty, concrete adoption data, and strong resonance with agent teams, but price, timing, and integration scope are still undisclosed, so it stays below p

Mar 6Friday

OpenAI News

Codex Security: now in research preview

OpenAI launched Codex Security in research preview on March 6, 2026 for ChatGPT Pro, Enterprise, Business, and Edu users, with free usage for the next month. Over the last 30 days, it scanned more than 1.2 million commits across external repos and reported 792 critical and 10,561 high-severity findings; noise fell by up to 84%, over-reported severity by 90%+, and false positives by 50%+. What matters is the stack: project-specific threat models, sandboxed validation, and patch proposals grounded in system context.

Why it matters: This is a substantive OpenAI product update for dev and security teams, not generic security messaging. HKR-H/K/R all pass: the angle is novel, the post includes concrete scan and false-positive metrics, and it speaks to AI coding risk plus alert fatigue; still a research preview

Feb 27Friday

OpenAI News

OpenAI and Amazon announce strategic partnership

OpenAI and Amazon announced a multi-year partnership, with Amazon investing $50 billion in OpenAI: $15 billion upfront and $35 billion tied to conditions. They will launch a Stateful Runtime Environment on Amazon Bedrock, and OpenAI will consume about 2 gigawatts of Trainium capacity on AWS. The part to watch is distribution plus compute lock-in: AWS becomes the exclusive third-party cloud distributor for OpenAI Frontier.

Why it matters: This is not a routine partnership post. The disclosed $50B staged investment, Bedrock runtime, and ~2GW Trainium commitment change OpenAI's distribution and compute posture; HKR-H/K/R all pass, so this lands in P1.

Feb 26Thursday

OpenAI News

Pacific Northwest National Laboratory and OpenAI partner to accelerate federal permitting

OpenAI and Pacific Northwest National Laboratory evaluated coding agents on NEPA drafting tasks from 18 federal agencies, finding 1-5 hours saved per subsection, or about 15% less drafting time. The DraftNEPABench benchmark was designed with 19 experts and covers 102 tasks, using Codex CLI with GPT-5 for long-document synthesis, cross-checking, and structured writing. The key limit is explicit: this measures well-scoped drafting work, not full real-world permitting decisions.

Why it matters: HKR-H/K/R pass: federal permitting is an unusual hook; the post gives 19 experts, 102 tasks, and 1–5 hours saved; the debate is agents entering regulated workflows. Score stays below major product news because this is a scoped benchmark, not a shipped capability.

OpenAI News

OpenAI Codex and Figma launch code-to-design roundtrip workflow

OpenAI and Figma launched a Codex integration on Feb. 26, 2026 that turns code into editable Figma designs and brings Figma Design, Figma Make, and FigJam content back into code. The workflow uses MCP via the Figma MCP Server in the Codex desktop app; OpenAI says Codex has 1M+ weekly users and usage is up 400%+ since the start of the year. The key issue is whether roundtrip context stays intact; the post does not disclose supported models, permission boundaries, or pricing.

Why it matters: This is a solid OpenAI/Figma workflow update with clear HKR-H/K/R: a bidirectional code↔design loop via MCP and Figma MCP Server. It stays below 85 because the post does not disclose model support, permission boundaries, pricing, or roundtrip reliability.

Jan 28Wednesday

Mistral AI

Mistral releases terminal coding agent Mistral Vibe 2.0

Mistral released Mistral Vibe 2.0, a terminal coding agent powered by the Devstral 2 model family. It adds custom subagents, multi-option clarification, slash-command skills, a unified agent mode and automatic updates.

Why it matters: The post lists Vibe 2.0's custom subagents, slash-command skills and subscription entry point, enough to judge how terminal coding agent workflows change.

Jan 21Wednesday

NVIDIA Blog

Jensen Huang on AI’s “Five-Layer Cake” at Davos: the largest infrastructure buildout in human history

Jensen Huang said at Davos that global VC investment topped $100 billion in 2025, with most capital going to AI-native startups building the AI stack’s application and infrastructure layers. He described AI as a five-layer stack: energy, chips and computing infrastructure, cloud data centers, models, and applications, and cited a US nursing shortage of about 5 million where AI can handle charting and transcription. The key point for practitioners is that the bottleneck is not just models, but the full infrastructure and labor chain.

Why it matters: This clears HKR-H/R because Jensen's Davos framing is a strong, discussable hook for practitioners. HKR-K also passes on specific facts (> $100B VC, five-layer stack, 5M nurse gap), but it is still executive commentary, not a model or product launch, so it stays in the 78-84 band

Jan 6Tuesday

NVIDIA Blog

NVIDIA DGX SuperPOD Sets the Stage for Rubin-Based Systems

NVIDIA introduced Rubin-based DGX SuperPOD systems, with DGX Vera Rubin NVL72 and DGX Rubin NVL8 slated for the second half of this year. One DGX SuperPOD can combine eight NVL72 systems for 576 Rubin GPUs, 28.8 exaflops FP4, and 600TB memory; NVIDIA says inference token cost drops by up to 10x versus the prior generation. The key detail is rack-scale design: 260TB/s NVLink per rack, which the post says removes model partitioning.

Why it matters: This is a substantive NVIDIA infra roadmap with hard numbers: 576 Rubin GPUs, 28.8 exaflops FP4, 600TB memory, 260TB/s NVLink, and up to 10x lower token cost. HKR-H/K/R all pass, but it is still a vendor roadmap post rather than a shipping model or broad product release, so it is

NVIDIA Blog

NVIDIA DRIVE AV Software Debuts in the All-New Mercedes-Benz CLA

NVIDIA said the new Mercedes-Benz CLA will be the first U.S. vehicle to ship DRIVE AV with enhanced Level 2 point-to-point driver assistance by the end of this year. The post describes a dual-stack design: end-to-end AI for core driving plus a classical safety stack built on Halos, with OTA upgrades, urban navigation, active collision avoidance, and automated parking. The launch timing is specific, but the post does not disclose pricing, sensor configuration, or the exact ODD.

Why it matters: HKR-H lands on the Mercedes CLA deployment hook. HKR-K lands on the disclosed dual-stack design and US launch timing. HKR-R lands on the shipping-autonomy debate, but missing price, sensor suite, and ODD keep it at the low end of featured.

NVIDIA Blog

NVIDIA unveils new open models, data and tools across agents, robotics, AVs and biomedicine

NVIDIA released open models, datasets and training tools spanning Nemotron, Cosmos, Alpamayo, Isaac GR00T and Clara, plus 10T language tokens, 500K robotics trajectories, 455K protein structures and 100TB of vehicle sensor data. Newly disclosed items include Nemotron Speech/RAG/Safety, Cosmos Reason 2, Transfer 2.5, Predict 2.5, GR00T N1.6 and Alpamayo 1; the key signal is that NVIDIA is opening the data stack across agents, physical AI, AVs and biomedicine.

Dec 9, 2025Tuesday

Mistral AI

Mistral releases Devstral 2 coding models and the Mistral Vibe CLI

Mistral AI released the Devstral 2 coding model family: the 123B Devstral 2 and the 24B Devstral Small 2, under a modified MIT license and Apache 2.0 respectively. Both are open source.

Why it matters: The post gives Devstral 2's SWE-bench scores, open-source licenses and deployment requirements, enough to judge the cost of running open coding models.

Oct 24, 2025Friday

Mistral AI

Mistral AI launches Mistral AI Studio production platform

Mistral AI released Mistral AI Studio, a production-grade AI platform for enterprise teams, built on three pillars: Observability, Agent Runtime and AI Registry.

Why it matters: The post lays out the three pillars of enterprise AI production and a private beta entry point, enough to judge how it differs from existing MLOps tools.

Oct 21, 2025Tuesday

OpenAI News

Introducing ChatGPT Atlas, the browser with ChatGPT built in

OpenAI launched ChatGPT Atlas on October 21, 2025, with a worldwide macOS release for Free, Plus, Pro, and Go users. Atlas embeds ChatGPT, browser memories, and page-visibility controls into the browser; agent mode preview is available for Plus, Pro, and Business. The key shift is persistent browsing context: web content is excluded from training by default unless users opt in.

Why it matters: OpenAI moving ChatGPT into its own browser is a distribution-layer product move, not a routine feature drop, so this lands at 88 and p1. HKR-H/K/R all pass: novel hook, concrete rollout/privacy details, and clear resonance around browser control, retention, and data boundaries.

Oct 6, 2025Monday

OpenAI News

Codex is now generally available

OpenAI said on October 6, 2025 that Codex is now generally available, with a Slack integration, a Codex SDK, and new admin controls. The post says daily Codex usage is up more than 10x since early August, and GPT-5-Codex served over 40 trillion tokens in three weeks; starting October 20, cloud tasks count toward usage, but the post does not disclose pricing details. The signal for practitioners is enterprise uptake: OpenAI says nearly all of its engineers use Codex, and they merge 70% more pull requests per week.

OpenAI News

Introducing apps in ChatGPT and the new Apps SDK

OpenAI launched apps inside ChatGPT on October 6, 2025 and previewed the Apps SDK for developers, for logged-in users outside the EEA, Switzerland, and the UK on Free, Go, Plus, and Pro plans. Seven partners are live and 11 more are due later this year; the SDK is open source and built on MCP, while the post does not disclose app review, listing, or revenue-share details.

Why it matters: This is a major OpenAI platform move: ChatGPT gains an app layer and developers get an SDK, so HKR-H/K/R all pass. Concrete facts include plan coverage, region limits, 7+11 partners, and an open-source MCP base; listing, review, and revenue-share terms are still undisclosed.

OpenAI News

Introducing AgentKit, new Evals, and RFT for agents

OpenAI launched AgentKit on October 6, 2025 with three agent-building components: Agent Builder, Connector Registry, and ChatKit. The post says Evals adds datasets, trace grading, automated prompt optimization, and third-party model support; Connector Registry covers Dropbox, Google Drive, SharePoint, Microsoft Teams, and third-party MCPs. The real signal is workflow versioning and safety governance; the title mentions RFT, but the provided post does not disclose its training details, pricing, or rollout scope.

Why it matters: This is a substantial OpenAI release for agent builders, with HKR-H/K/R all passing. It provides concrete mechanisms across Agent Builder, connectors, ChatKit, and Evals, but the excerpt does not disclose RFT mechanics, pricing, or rollout scope, so it stays at 84 rather than p1.

Sep 29, 2025Monday

OpenAI News

Introducing parental controls

OpenAI launched parental controls for all ChatGPT users on September 29, 2025, letting parents link with teen accounts and manage usage settings from their own account. Linked teen accounts get stronger content safeguards by default, and parents can set quiet hours, disable voice, memory, image generation, and opt out of model training. The key mechanism is the alert flow: suspected self-harm signals trigger human review, and acute distress leads to email, SMS, and push notifications to parents.

Why it matters: OpenAI rolled parental controls to all ChatGPT users and disclosed a concrete self-harm escalation flow: system detection, human review, then email/SMS/push alerts to parents. HKR-K and HKR-R are strong; this is a substantive safety product update, but not a model-level launch,so

OpenAI News

Buy it in ChatGPT: Instant Checkout and the Agentic Commerce Protocol

OpenAI launched Instant Checkout in ChatGPT on September 29, 2025, letting U.S. Plus, Pro, and Free users buy from U.S. Etsy sellers in chat; it currently supports single-item purchases. OpenAI says ChatGPT has over 700 million weekly users and open-sourced the Agentic Commerce Protocol with Stripe; Stripe merchants can enable it with as little as one line of code, while the post does not disclose the fee rate merchants pay.

Why it matters: This is a high-weight ChatGPT product expansion from discovery to completed purchases, so HKR-H/K/R all pass. The post confirms U.S. Free/Plus/Pro checkout with U.S. Etsy sellers and a Stripe-backed protocol layer; merchant fee details are not disclosed, so it stays high but sub-

Sep 25, 2025Thursday

OpenAI News

More ways to work with your team and tools in ChatGPT

OpenAI rolled out shared projects for ChatGPT Business on September 25, 2025, and made them available for Enterprise and Edu plans. Shared projects support email or link invites, two access levels, and private project memory; Enterprise and Edu have them off by default under admin control. OpenAI also added Gmail, Google Calendar, Outlook, Teams, SharePoint, GitHub, Dropbox, and Box connectors, and said ChatGPT can now choose connectors automatically per prompt.

Why it matters: HKR-H/K/R all pass: shared projects, 8 connectors, and prompt-routed connector selection are concrete workflow changes with clear admin controls. I keep it below 85 because this is a collaboration-layer product update, not a model release or a broad capability jump.

OpenAI News

Introducing ChatGPT Pulse

OpenAI previewed ChatGPT Pulse for Pro users on mobile on September 25, 2025, with one daily proactive research update. It uses memory, chat history, feedback, and optional Gmail and Google Calendar connections to generate visual cards; integrations are off by default and outputs pass safety checks. The shift to async delivery matters more than the headline, but the post does not disclose the model, pricing changes, or a Plus launch date.

Why it matters: HKR-H/K/R all pass: the novel angle is proactive outreach, and the post gives concrete scope and input sources. This is a meaningful ChatGPT product update, but model details, rollout beyond Pro mobile, update cadence, and pricing changes are not disclosed, so it stays featured,

Sep 24, 2025Wednesday

OpenAI News

SAP and OpenAI partner to launch sovereign 'OpenAI for Germany'

SAP and OpenAI announced OpenAI for Germany for the German public sector, planned for 2026 and hosted by Delos Cloud on Microsoft Azure. SAP plans to expand Delos Cloud in Germany to 4,000 GPUs for AI workloads; the post does not disclose model names, pricing, or contract size. The key point is delivery: this is a sovereign public-sector deployment focused on compliance, data residency, and AI agents inside existing workflows.

Why it matters: HKR-H/K/R all pass: the story pairs a novel sovereign-deployment angle with concrete facts like a 2026 launch, Delos Cloud on Azure, and 4,000 GPUs. It matters because sovereignty and public-sector procurement are live issues, but missing model, pricing, and deal-scope details it

Sep 16, 2025Tuesday

OpenAI News

Building towards age prediction

OpenAI is building an age-prediction system for ChatGPT so users identified as under 18 are automatically routed to a teen experience. The post says low-confidence cases default to the under-18 mode, adults can verify age to unlock adult capabilities, and parental controls will ship by the end of the month with teen account linking, memory/history toggles, and blackout hours.

Why it matters: This is not a generic safety post: OpenAI is wiring age estimation into ChatGPT routing. HKR-H/K/R all pass on the auto-teen switch, fail-closed treatment for low confidence, and the privacy/liability nerve, but it remains below a major model or platform release.

Sep 15, 2025Monday

OpenAI News

Introducing upgrades to Codex

OpenAI released GPT-5-Codex and made it the default model for Codex cloud tasks and code review; in testing, it worked independently for more than 7 hours on complex tasks. OpenAI says it used 93.7% fewer tokens than GPT-5 on the lowest 10% of employee turns, while spending 2x longer reasoning, editing, and testing on the highest 10%. The key point is one model now spans interactive coding and long-running agentic execution; pricing and full availability details are not fully disclosed in the provided body.

Why it matters: This is a substantive OpenAI developer-tool update: GPT-5-Codex becomes the default for Codex cloud tasks and code review, with concrete numbers on 7-hour autonomy and token use. HKR-H/K/R all pass; pricing and full availability are not fully disclosed in the excerpt, so it stays

OpenAI News

Addendum to GPT-5 system card: GPT-5-Codex

OpenAI published a GPT-5-Codex system card addendum on September 15, 2025, stating the model is optimized for agentic coding in Codex and is available in terminal, IDE, web, GitHub, and the ChatGPT mobile app. The post says it uses reinforcement learning on real-world coding tasks, plus safety training for harmful tasks and prompt injection, with sandboxing and configurable network access. Benchmark scores, pricing, and context window are not disclosed.

Why it matters: HKR-H/K/R all pass: this is an OpenAI coding-agent model spanning terminal, IDE, GitHub, web, and mobile, with concrete training and safety details. I kept it below 85 because benchmarks, pricing, and context window are not disclosed in the body.

Sep 12, 2025Friday

OpenAI News

Working with US CAISI and UK AISI to build more secure AI systems

OpenAI said its work with US CAISI and UK AISI found and fixed 2 novel ChatGPT Agent vulnerabilities; CAISI built a proof-of-concept exploit chain with about a 50% success rate, and OpenAI fixed it within 1 business day. The post says the bugs let attackers bypass protections under certain conditions, remotely control session-accessible systems, and impersonate logged-in users; UK AISI has red-teamed bio-misuse safeguards for ChatGPT Agent and GPT-5 since May 2025, but the truncated post does not disclose further results.

Why it matters: This is not generic safety PR. OpenAI discloses 2 new ChatGPT Agent vulns, ~50% CAISI PoC success, and a 1-business-day fix, so HKR-H/K/R all pass. Kept below 85 because the UK AISI section is truncated and the broader impact is not disclosed.

Sep 2, 2025Tuesday

Mistral AI

Make Memory work for you.

Mistral AI 为 Le Chat 上线 Memories(beta)记忆功能,可自动保存有用信息,但回忆过程可见、可溯源,并附来源链接。用户可随时关闭记忆、开启不使用记忆的隐身对话、编辑或删除单条记忆,还能导出和从外部导入记忆。同步推出 Memory Insights,基于用户自身数据提示记忆趋势与摘要。

Mistral AI

Mistral adds custom MCP connectors and Memories to Le Chat

Mistral launched a catalog of 20+ secure MCP-based connectors for Le Chat (beta), covering data, productivity, development, automation and business categories. It supports tools including Databricks, Snowflake, GitHub, Atlassian, Asana, Outlook, Box, Stripe and Zapier, and lets users add custom MCP connectors or connect to any remote MCP server.

Why it matters: The post lists the specific tool categories and deployment methods behind the 20+ connectors, showing how far Le Chat reaches into enterprise workflows.

OpenAI News

Building more helpful ChatGPT experiences for everyone

OpenAI said it will ship ChatGPT safety changes over the next 120 days and roll out Parental Controls within a month. Disclosed steps include routing conversations with signs of acute distress to reasoning models such as GPT-5-thinking, and letting parents link accounts for teens 13+, disable memory and chat history. The post does not disclose router trigger thresholds or alert false-positive rates.

Why it matters: This changes core ChatGPT behavior, so HKR-H/K/R all pass: the routing hook is novel, the post gives concrete controls, and teen safety is a live industry topic. I keep it below 85 because trigger criteria, false-positive rate, and rollout scope are not disclosed.

Aug 28, 2025Thursday

OpenAI News

Introducing gpt-realtime and Realtime API updates for production voice agents

OpenAI released the speech-to-speech model gpt-realtime and made the Realtime API generally available, adding remote MCP server support, image input, and SIP phone calling. The post reports 82.8% on Big Bench Audio versus 65.6% for the December 2024 model, and 30.5% on the audio MultiChallenge benchmark versus 20.6%. The key change is that tool access and phone connectivity now ship in the same production API.

Why it matters: This is a substantive OpenAI model + API release, not a minor refresh. HKR-H/K/R all pass: the release has a clear hook, hard benchmark deltas, and direct deployment impact for production voice agents, so it reaches p1.

Aug 7, 2025Thursday

OpenAI News

Introducing GPT-5 for developers

OpenAI released GPT-5 in its API on August 7, 2025, in three sizes: gpt-5, gpt-5-mini, and gpt-5-nano. The post reports 74.9% on SWE-bench Verified, 88% on Aider polyglot, 96.7% on τ2-bench telecom, plus new verbosity, minimal reasoning_effort, and custom tools; pricing and full availability details are not disclosed in the provided text. The real developer signal is the API surface change, not just a model rename.

Why it matters: This is an OpenAI flagship-model API launch, so it belongs in the 95–100 band. HKR-H lands on the GPT-5 debut; HKR-K lands on concrete benchmark scores and new controls; HKR-R lands on immediate developer concerns around migration, tooling, and model comparison; the excerpt omits

OpenAI News

GPT-5 and the new era of work

OpenAI launched GPT-5 on August 7, 2025, started rollout to Team users the same day, said Enterprise and Edu access would follow next week, and made it available in the API immediately. The post gives two hard numbers: 5 million paid ChatGPT business users and nearly 700 million weekly ChatGPT users; it does not disclose benchmark scores, pricing, or context length.

Why it matters: An OpenAI GPT-5 launch is a market-wide event, so HKR-H/K/R all pass. The post gives rollout timing and a 5M paid-business-user datapoint, but it omits benchmark scores, pricing, and context length, so this lands at the low end of the top band.

Aug 4, 2025Monday

OpenAI News

What OpenAI is optimizing ChatGPT for

OpenAI said on August 4, 2025 that ChatGPT is optimized to help users finish tasks and leave, not maximize time spent. Break reminders are live for long sessions, and new behavior for high-stakes personal decisions is coming soon. Evaluation now includes custom rubrics built with 90+ physicians across 30+ countries.

Why it matters: Official OpenAI guidance on ChatGPT incentives and safety, with concrete facts: rest-break reminders are live and multi-turn evals include 90+ doctors from 30+ countries. HKR-H/K/R all pass, but the high-risk decision behavior lacks shipping scope and trigger details, so this is

Jul 30, 2025Wednesday

Mistral AI

Mistral ships Codestral 25.08 and an enterprise coding stack

Mistral AI released Codestral 25.08 along with a full enterprise coding stack: Codestral, Codestral Embed, Devstral and a Mistral Code IDE plugin.

Why it matters: The post gives Codestral 25.08's completion gains and how the enterprise stack is deployed, so you can judge whether a private coding setup is viable.

Jul 29, 2025Tuesday

OpenAI News

Introducing study mode in ChatGPT

OpenAI launched study mode in ChatGPT on July 29, 2025 for logged-in Free, Plus, Pro, and Team users, with ChatGPT Edu coming in the next few weeks. It uses custom system instructions to deliver Socratic prompts, scaffolded responses, knowledge checks, and on/off toggling instead of direct answers, adapting to skill-level questions and prior chat memory. The key change is interaction design, not a new model; the post does not disclose the underlying model, outcome metrics, or misuse safeguards.

Jul 17, 2025Thursday

Mistral AI

Mistral adds Deep Research, voice mode and more to Le Chat

Mistral rolled out a batch of new Le Chat features: a preview Deep Research mode, a voice mode powered by the new Voxtral speech model, a multilingual thinking mode backed by the Magistral reasoning model, Projects for organizing conversations, and advanced image editing built with Black Forest Labs.

Why it matters: Mistral announced five Le Chat features at once, so readers can see how its research, voice and image-editing abilities fit together.

OpenAI News

ChatGPT agent System Card

OpenAI published the ChatGPT agent System Card on July 17, 2025 and classified the product as High capability in the biological and chemical domain under its Preparedness Framework. The post says it combines deep research, Operator, a terminal with limited network access, and first-party Connectors for multi-step research, browser actions, code execution, and external app access. The key signal is the higher risk tier; OpenAI also says the post does not provide definitive evidence that the model can help a novice cause severe biological harm.

Why it matters: This is not routine safety paperwork. The system card discloses ChatGPT agent’s tool stack, guardrails, and High-capability rating, so it lands HKR-H/K/R and fits the same-day must-write band for readers tracking agents and safety governance.

OpenAI News

Introducing ChatGPT agent

OpenAI launched ChatGPT agent on July 17, 2025, and made agent mode available to Pro, Plus, and Team users. It combines Operator-style web actions, deep research synthesis, a terminal, and API access in one virtual computer; the post lists the tools but does not disclose pricing, quotas, or benchmark results. The key detail is control: consequential actions require user permission, and users can interrupt, stop, or take over the browser at any time.

OpenAI News

Agent bio bug bounty call

OpenAI opened a bio bug bounty for ChatGPT agent on July 17, 2025, offering $25,000 for the first universal jailbreak prompt that clears all 10 bio/chem safety questions from a clean chat. Scope is limited to ChatGPT agent; testing starts July 29, 2025, with a separate $10,000 prize for the first team that solves all 10 using multiple prompts. The key bar is a universal jailbreak, not a single-question bypass; all prompts, outputs, findings, and communications are under NDA.

Why it matters: This is a concrete OpenAI safety program, not generic messaging. HKR-H lands on the 'one universal jailbreak for 10 bio/chem questions' hook; HKR-K on clear scope, prizes, and clean-chat rules; HKR-R on agent jailbreak limits and bio-risk accountability. 80: featured, but below a

Jul 11, 2025Friday

Mistral AI

Mistral releases Devstral Medium and upgrades Devstral Small 1.1

Mistral AI worked with All Hands AI to launch Devstral Medium and upgrade Devstral Small 1.1.

Why it matters: Mistral and All Hands AI jointly released two coding agent models with SWE-Bench Verified scores and API pricing, making comparison with existing options easier.

May 27, 2025Tuesday

Mistral AI

Mistral releases Agents API with built-in connectors and MCP tools

Mistral released an Agents API that pairs its language models with built-in connectors for code execution, web search, image generation and MCP tools. It also offers persistent memory across conversations and agent orchestration.

Why it matters: Mistral details the connectors, memory and orchestration of its Agents API, letting readers judge how an agent platform would be deployed.