Skip to content

#Agent

36 today

Feb 10Tuesday

MIT Technology Review · AI

Why the Moltbook frenzy was like Pokémon

MIT Technology Review compares the Moltbook AI-agent social experiment to 2014 Twitch Plays Pokémon: lots of spectacle, limited signal about the future. The post cites 1 million concurrent players in the Pokémon case; Moltbook also mixed in crypto scams, and some “agent” posts were actually steered by humans. The real gap is explicit: shared memory, coordination, and shared goals are still missing.

Feb 9Monday

36Kr (direct RSS)

Qwen’s 10 Million Milk Teas: How Alibaba’s Massive AI Freebie Campaign Unfolded

Alibaba’s Qwen drove over 10 million orders via a Feb. 6 free-order campaign, but the app slowed and crashed from 10 a.m. to noon as load exceeded capacity; orders had already passed 2 million before noon. 36Kr says initial server capacity was only about one-third of the expected peak, and the subsidy pool was framed as 3 billion yuan; the real signal is not a model leap but a paid test of AI commerce entry and consumer acquisition.

Why it matters: HKR-H lands on the free-milk-tea plus outage hook, while HKR-K lands on concrete scale and capacity numbers. HKR-R also lands because the story speaks to AI distribution, subsidy economics, and infra reliability, but it remains a single-company promo test rather than a market-shi

36Kr (direct RSS)

Former Baichuan co-founder Jiao Ke bets on AI audio to build AI hosts

Jiao Ke said Laifu Radio now has 15 Chinese AI hosts and 2 English ones, and raised over $10 million across two rounds by H2 2025. He said users average about 30 minutes per day, AI can prepare timely audio in under an hour, and the team treats DTU plus long-memory infra as the key moat. The real bet is not an AI podcast tool but interactive AI hosts that remember user preferences; the post also says it is working with some automakers on in-car personalized AI radio.

Why it matters: HKR-H lands because the story reframes audio AI as persistent hosts, not a podcast tool. HKR-K is strong on numbers and mechanism; HKR-R lands via memory plus in-car distribution. Early-stage company scope keeps it at featured, not p1.

Feb 7Saturday

MIT Technology Review · AI

Moltbook was peak AI theater

Moltbook went viral within hours, and the platform says it now has 1.7 million agent accounts, 250,000 posts, and 8.5 million comments, but the article argues the activity is mostly human-scripted mimicry. It says OpenClaw can connect Claude, GPT-5, or Gemini to tools like email and browsers; cited operators say the agents lack shared goals, shared memory, and self-directed autonomy, and some viral posts were written by humans posing as bots. The key takeaway is risk: agents tied to private data such as passwords or bank details were active on a site filled with spam and potentially malicious instructions.

Why it matters: This is strong anti-hype commentary, not a market-moving event. HKR-H/K/R all pass: the hook is sharp, the piece adds 1.7M/250k/8.5M plus concrete critique on memory and goals, and the security angle lands with practitioners, so it clears featured but stays mid-70s.

Feb 6Friday

TechCrunch · AI

OpenAI launches new agentic coding model minutes after Anthropic releases its own

OpenAI launched an agentic coding model minutes after Anthropic released a similar one, and the model is meant to accelerate Codex, which OpenAI launched earlier this week. The RSS snippet gives only the timing and purpose; the post does not disclose the model name, benchmarks, pricing, context length, or availability. The signal is direct competition in agentic coding, not a substantiated performance claim.

Why it matters: Major-lab product news plus a minutes-apart Anthropic clash gives this HKR-H and HKR-R. The score stays in the low featured band because HKR-K is weak: the post lacks the model name, benchmarks, price, context window, and availability.

Feb 4Wednesday

TheValley101 (硅谷101)

E224 | Why Clawdbot became the first breakout product of 2026 amid the Mac mini rush | Moltbot | MoltBook | OpenClaw

The podcast says Clawdbot passed 100k GitHub stars within days and reached 146k on Feb. 2, while being renamed to Moltbot and then OpenClaw within a week. It attributes the traction to a stack of Claude, long-term memory, IM-based messaging, and proactive heartbeat workflows; the title mentions a Mac mini rush, but the post does not disclose sales figures. The real signal is the interaction layer rather than a new model release: this is industry commentary and user anecdotes, not an official spec sheet.

Why it matters: This is a commentary-led breakdown of a hot agent phenomenon, not a primary launch. HKR-H/K/R all pass: the 146k-star surge and rename chain are novel, the post explains memory + IM + heartbeat mechanics, and it hits nerves on agent UX, dedicated hardware, and security bills; the

Feb 3Tuesday

Computing Life · Yage

Beyond Tutorial Thinking: Why AI Education Should Add Engineering Infrastructure, Not Just Content

The team says it ran 4 courses over 2 years for 2,500+ learners, yet only a minority shipped usable products; drop-off centered on setup, experimentation, deployment, and context handling friction. The post says AI Builder Space gives students a no-card unified API, one-click deployment to <name>.ai-builders.space free for 1 year, and MCP access for Cursor and Claude Code via one command. The point is productized teaching infra, not more tutorials; retention, conversion, and cost are not disclosed.

Why it matters: The piece turns a familiar complaint into operational detail: 2500+ learners, 4 failure points, and a concrete platform response with API, deployment, and MCP access. HKR-H/K/R all pass, but missing conversion, retention, and cost data keeps it at the low end of featured.

Feb 2Monday

Import AI (Jack Clark)

Import AI 443: Into the Mist: Moltbook, Agent Ecologies, and the Internet in Transition

Jack Clark writes that Moltbook has pushed AI agents into a public social network at tens-of-thousands scale, shifting conversation from humans to agents. He says it combines an agent social feed with OpenClaw-style computer access, but the post does not disclose active-agent, retention, or transaction metrics. A separate July 2025 workshop report says closed-loop AI R&D automation could raise productivity from 10x to 100x to 1000x; the key issue is measurement and outside transparency.

Why it matters: Featured: HKR-H/K/R all pass. The post has a strong hook—a public social space filled by agent ecologies—and a concrete 10x/100x/1000x closed-loop R&D claim, but it lacks Moltbook activity, retention, and transaction data, so it stays at 78.

Feb 1Sunday

Lex Fridman (YouTube RSS)

State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast #490

Lex Fridman, Sebastian Raschka, and Nathan Lambert discuss the 2026 AI race in podcast #490 and frame DeepSeek R1’s January 2025 release as a key inflection point. The episode names Claude Opus 4.5, Gemini 3, Z.ai GLM, Minimax, and Kimi Moonshot, but the post does not disclose a shared benchmark, cost table, or reproducible eval. The useful takeaway is the lens: gaps look more like compute, budget, and org culture than secret ideas.

Why it matters: High-quality commentary, not a news break. HKR-H and HKR-R pass because Lex Fridman, Sebastian Raschka, and Nathan Lambert frame China, agents, GPUs, and AGI for practitioners. HKR-K misses: the post names models and DeepSeek R1 but provides no shared benchmarks, cost table, or a

Jan 30Friday

Ruan YiFeng's Weblog

Technology Enthusiast Weekly #383: What Level of AI Programming Are You?

Steve Yegge frames AI coding into 8 levels and says he is at level 8, where an orchestrator manages parallel AI coding sessions. The post lays out a path from IDE copilots to YOLO acceptance, 3-5 windows, 10+ windows, then orchestration; it also says his AI-built tool Gas Town has 225,000 lines of Go code, which he has never read, and had 6,000 stars as of last week. The real signal is black-box programming as a workflow choice, with cost and failure risk stated plainly.

Why it matters: Strong HKR-H/K/R: the 8-level framing is sticky, and the post carries concrete workflow and project numbers. The score stays below 78 because this is secondary commentary, not a primary model, product, or research release.

Jan 29Thursday

Ruan YiFeng's Weblog

Kimi’s integrated stack vs. Manus’s layered approach

Kimi released the K2.5 model and K2.5 Agent together, with an agent mode already available on its website. The post cites 1,500-step long-horizon actions, up to 100 agents in parallel, and visual coding from design files or web videos; pricing, context window, and API terms are not disclosed. The key point is product shape: not just a model launch, but a bundled model-plus-agent release.

Why it matters: HKR-H lands on the integrated release angle; HKR-K lands on the 1,500-step, 100-agent, visual-programming details; HKR-R lands on the stack-design debate. Missing price, context window, and API terms, plus a commentary source, keep it below p1.

Jan 28Wednesday

MIT Technology Review · AI

What AI “remembers” about you is privacy’s next frontier

Google launched Personal Intelligence this month, letting Gemini use Gmail, Photos, Search, and YouTube history for personalization. The piece says OpenAI, Anthropic, and Meta are adding memory too, but current designs often pool cross-context data into one repository, increasing privacy and misuse risks. The key issue is memory architecture: segmentation, provenance tracking, user edit/delete controls, and privacy-preserving evaluation.

Mistral AI

Mistral releases terminal coding agent Mistral Vibe 2.0

Mistral released Mistral Vibe 2.0, a terminal coding agent powered by the Devstral 2 model family. It adds custom subagents, multi-option clarification, slash-command skills, a unified agent mode and automatic updates.

Why it matters: The post lists Vibe 2.0's custom subagents, slash-command skills and subscription entry point, enough to judge how terminal coding agent workflows change.

Jan 21Wednesday

NVIDIA Blog

Jensen Huang on AI’s “Five-Layer Cake” at Davos: the largest infrastructure buildout in human history

Jensen Huang said at Davos that global VC investment topped $100 billion in 2025, with most capital going to AI-native startups building the AI stack’s application and infrastructure layers. He described AI as a five-layer stack: energy, chips and computing infrastructure, cloud data centers, models, and applications, and cited a US nursing shortage of about 5 million where AI can handle charting and transcription. The key point for practitioners is that the bottleneck is not just models, but the full infrastructure and labor chain.

Why it matters: This clears HKR-H/R because Jensen's Davos framing is a strong, discussable hook for practitioners. HKR-K also passes on specific facts (> $100B VC, five-layer stack, 5M nurse gap), but it is still executive commentary, not a model or product launch, so it stays in the 78-84 band

Jan 20Tuesday

MIT Technology Review · AI

The era of agentic chaos and how data will save us

The piece says a mid-sized enterprise can run 4,000 agents, and misaligned data can directly hit revenue, compliance, and customer experience. It cites BCG saying 60% of companies see minimal gains despite heavy AI spend, while leaders report 5x revenue growth and 3x cost reduction; the article frames reliability through four quadrants—models, tools, context, and governance—and argues data debt is the main blocker, not model quality.

Why it matters: This is a sourced enterprise-AI commentary, not empty thought leadership. HKR-H comes from the '4,000 agents' chaos hook; HKR-K comes from the 60% / 5x / 3x BCG data and four-part reliability frame; HKR-R lands because data debt, compliance, and customer-experience risk are live,

MIT Technology Review · AI

The UK government is backing AI scientists that can run their own lab experiments

UK agency ARIA selected 12 AI scientist projects from 245 proposals, doubled its planned funding, and will give each team about £500,000 for nine months. ARIA defines an AI scientist as a system that hypothesizes, runs experiments, analyzes results, and iterates; the funded projects still rely on existing tools. The key signal is reproducible lab-loop execution, not press-release heat: one cited external study reports LLM agents failed to complete a scientific workflow 3 out of 4 times.

Why it matters: HKR-H/K/R all pass: 'AI runs its own lab experiments' is a strong hook, and the piece includes 12 teams, 245 proposals, ~£500k each, a 9-month term, and a cited 75% failure rate. Important for agentic science, but this is funding for early systems, not a proven breakthrough.

Jan 19Monday

Import AI (Jack Clark)

Import AI 441: My agents are working. Are yours?

Jack Clark says his research agents processed thousands of papers while he hiked or slept, and Claude finished site scraping, embeddings, local vector search, and a GUI in under one hour. The post confirms multi-agent retrieval, cross-checking, and report generation; it does not disclose model versions, cost, failure rate, or benchmark data. The point to watch is workflow friction dropping enough for AI to shift from single prompts to ongoing delegated work.

Why it matters: HKR-H lands with the challenge in the headline; HKR-K lands because Clark describes a <1 hour workflow with retrieval, cross-checking, and report generation. Missing model version, cost, failure rate, and evaluation keep it in featured, not p1.

Jan 16Friday

Ruan YiFeng's Weblog

Technology Enthusiast Weekly (Issue 381): What China's AI Foundation Model Leaders Are Thinking

Ruan Yifeng’s Issue 381 excerpts talks from Beijing’s AGI-Next summit on Jan 10, covering views from Zhipu, Alibaba Qwen, and Tencent AI leaders on China’s model roadmap. The post cites Lin Junyang saying US compute is 1-2 orders of magnitude larger, Yao Shunyu calling the odds of a China-led top AI company in 3-5 years high, while Lin puts it at 20%. The key split is strategic: Tang Jie points to RLVR in 2025, Lin bets on multimodal foundation agents, and Yao says B2B buyers pay a $200/month premium for stronger models.

Why it matters: It clears all three HKR axes: public strategic disagreement gives it a strong hook, and the post includes concrete numbers and testable claims. The score stops short of the high bands because this is a secondary synthesis of summit remarks, not a primary release or original scoop

Jan 13Tuesday

MIT Technology Review · AI

CES showed me why Chinese tech companies feel so optimistic

CES 2026 drew 148,000+ attendees and 4,100+ exhibitors, with Chinese companies making up nearly a quarter and standing out in AI hardware and robotics. The post ties their optimism to manufacturing-led iteration speed, not one breakthrough; Lenovo Qira, Nvidia Vera Rubin, and AMD Helios show the race is shifting to cloud and hybrid AI.

Why it matters: This is on-the-ground CES reporting with a competition thesis: Chinese optimism comes from manufacturing and supply-chain iteration, supported by 148k attendees, 4,100 exhibitors, and roughly one-quarter from China. HKR-H/K/R pass, but shipment, revenue, and order data are not in

Jan 12Monday

Import AI (Jack Clark)

Import AI 440: Red Queen AI, AI regulating AI, and o-ring automation

Import AI 440 highlights two threads: Sakana used GPT-4 mini to evolve Core War programs, and specialized warriors beat 89.1% of human-designed warriors. The post says DRQ uses MAP-Elites plus matches against prior champions; a separate policy proposal ties AI rules to automatability triggers, with example thresholds of <=1% false positives, <=1% false negatives, and <=$10,000 per model evaluation.

Why it matters: This is a high-signal roundup, not the primary release, so it stays below the 78+ band. HKR-H lands on the unusual 'AI regulating AI' framing; HKR-K lands on the 89.1% result and ≤1% / <$10k thresholds; HKR-R lands on automation and governance nerves.

Jan 6Tuesday

NVIDIA Blog

NVIDIA DGX SuperPOD Sets the Stage for Rubin-Based Systems

NVIDIA introduced Rubin-based DGX SuperPOD systems, with DGX Vera Rubin NVL72 and DGX Rubin NVL8 slated for the second half of this year. One DGX SuperPOD can combine eight NVL72 systems for 576 Rubin GPUs, 28.8 exaflops FP4, and 600TB memory; NVIDIA says inference token cost drops by up to 10x versus the prior generation. The key detail is rack-scale design: 260TB/s NVLink per rack, which the post says removes model partitioning.

Why it matters: This is a substantive NVIDIA infra roadmap with hard numbers: 576 Rubin GPUs, 28.8 exaflops FP4, 600TB memory, 260TB/s NVLink, and up to 10x lower token cost. HKR-H/K/R all pass, but it is still a vendor roadmap post rather than a shipping model or broad product release, so it is

NVIDIA Blog

NVIDIA DRIVE AV Software Debuts in the All-New Mercedes-Benz CLA

NVIDIA said the new Mercedes-Benz CLA will be the first U.S. vehicle to ship DRIVE AV with enhanced Level 2 point-to-point driver assistance by the end of this year. The post describes a dual-stack design: end-to-end AI for core driving plus a classical safety stack built on Halos, with OTA upgrades, urban navigation, active collision avoidance, and automated parking. The launch timing is specific, but the post does not disclose pricing, sensor configuration, or the exact ODD.

Why it matters: HKR-H lands on the Mercedes CLA deployment hook. HKR-K lands on the disclosed dual-stack design and US launch timing. HKR-R lands on the shipping-autonomy debate, but missing price, sensor suite, and ODD keep it at the low end of featured.

NVIDIA Blog

NVIDIA unveils new open models, data and tools across agents, robotics, AVs and biomedicine

NVIDIA released open models, datasets and training tools spanning Nemotron, Cosmos, Alpamayo, Isaac GR00T and Clara, plus 10T language tokens, 500K robotics trajectories, 455K protein structures and 100TB of vehicle sensor data. Newly disclosed items include Nemotron Speech/RAG/Safety, Cosmos Reason 2, Transfer 2.5, Predict 2.5, GR00T N1.6 and Alpamayo 1; the key signal is that NVIDIA is opening the data stack across agents, physical AI, AVs and biomedicine.

Jan 5Monday

Import AI (Jack Clark)

Import AI 439: AI kernels; decentralized training; and universal representations

Meta says KernelEvolve cut kernel development from weeks to hours and delivered up to 17x over PyTorch baselines in production tests. The system uses Llama, GPT, and Claude to generate kernels, validates them, and feeds results into a knowledge base across NVIDIA, AMD, and MTIA; the post also says decentralized training is growing 20x per year but still uses about 1000x less compute than frontier runs. The real signal is continuous self-optimizing infra in production, while decentralized training matters if that 1000x gap keeps shrinking.

Why it matters: HKR-H/K/R all pass: the kernel-writing angle is novel, the post includes concrete numbers and mechanism, and the decentralization thread hits cost and power-concentration nerves. I stop at 80 because this is a newsletter synthesis of technical work, not a single industry-defining

Jan 1Thursday

36Kr (direct RSS)

Escaping the user-acquisition nightmare: Moonshot AI's 10 billion yuan cash reserve and Yang Zhilin's confidence

Moonshot AI raised $500 million at a $4.3 billion post-money valuation; Yang Zhilin said the company holds over 10 billion yuan in cash and is not rushing to IPO. Named backers include IDG, with Alibaba, Tencent, Gaorong Ventures, and Capital Today reportedly taking super pro rata; the memo also says paid users grew over 170% MoM on average and overseas API revenue rose 4x from September to November. The signal that matters is the shift from paid traffic to open source, model capability, and agents: the post says K2 reached No. 2 on OpenRouter's global trending list within a week of open-sourcing.

Why it matters: Moonshot is a Chinese frontier-model company, so a fresh $500M round plus operating metrics matters. HKR-H/K/R all pass on the strategic pivot and hard numbers, but this is still funding and business reporting, not a major model or product launch, so it stays featured rather than

Dec 9, 2025Tuesday

Mistral AI

Mistral releases Devstral 2 coding models and the Mistral Vibe CLI

Mistral AI released the Devstral 2 coding model family: the 123B Devstral 2 and the 24B Devstral Small 2, under a modified MIT license and Apache 2.0 respectively. Both are open source.

Why it matters: The post gives Devstral 2's SWE-bench scores, open-source licenses and deployment requirements, enough to judge the cost of running open coding models.

Oct 24, 2025Friday

Mistral AI

Mistral AI launches Mistral AI Studio production platform

Mistral AI released Mistral AI Studio, a production-grade AI platform for enterprise teams, built on three pillars: Observability, Agent Runtime and AI Registry.

Why it matters: The post lays out the three pillars of enterprise AI production and a private beta entry point, enough to judge how it differs from existing MLOps tools.

Oct 21, 2025Tuesday

OpenAI News

Introducing ChatGPT Atlas, the browser with ChatGPT built in

OpenAI launched ChatGPT Atlas on October 21, 2025, with a worldwide macOS release for Free, Plus, Pro, and Go users. Atlas embeds ChatGPT, browser memories, and page-visibility controls into the browser; agent mode preview is available for Plus, Pro, and Business. The key shift is persistent browsing context: web content is excluded from training by default unless users opt in.

Why it matters: OpenAI moving ChatGPT into its own browser is a distribution-layer product move, not a routine feature drop, so this lands at 88 and p1. HKR-H/K/R all pass: novel hook, concrete rollout/privacy details, and clear resonance around browser control, retention, and data boundaries.

Oct 6, 2025Monday

OpenAI News

Codex is now generally available

OpenAI said on October 6, 2025 that Codex is now generally available, with a Slack integration, a Codex SDK, and new admin controls. The post says daily Codex usage is up more than 10x since early August, and GPT-5-Codex served over 40 trillion tokens in three weeks; starting October 20, cloud tasks count toward usage, but the post does not disclose pricing details. The signal for practitioners is enterprise uptake: OpenAI says nearly all of its engineers use Codex, and they merge 70% more pull requests per week.

OpenAI News

Introducing apps in ChatGPT and the new Apps SDK

OpenAI launched apps inside ChatGPT on October 6, 2025 and previewed the Apps SDK for developers, for logged-in users outside the EEA, Switzerland, and the UK on Free, Go, Plus, and Pro plans. Seven partners are live and 11 more are due later this year; the SDK is open source and built on MCP, while the post does not disclose app review, listing, or revenue-share details.

Why it matters: This is a major OpenAI platform move: ChatGPT gains an app layer and developers get an SDK, so HKR-H/K/R all pass. Concrete facts include plan coverage, region limits, 7+11 partners, and an open-source MCP base; listing, review, and revenue-share terms are still undisclosed.

OpenAI News

Introducing AgentKit, new Evals, and RFT for agents

OpenAI launched AgentKit on October 6, 2025 with three agent-building components: Agent Builder, Connector Registry, and ChatKit. The post says Evals adds datasets, trace grading, automated prompt optimization, and third-party model support; Connector Registry covers Dropbox, Google Drive, SharePoint, Microsoft Teams, and third-party MCPs. The real signal is workflow versioning and safety governance; the title mentions RFT, but the provided post does not disclose its training details, pricing, or rollout scope.

Why it matters: This is a substantial OpenAI release for agent builders, with HKR-H/K/R all passing. It provides concrete mechanisms across Agent Builder, connectors, ChatKit, and Evals, but the excerpt does not disclose RFT mechanics, pricing, or rollout scope, so it stays at 84 rather than p1.

Sep 29, 2025Monday

OpenAI News

Introducing parental controls

OpenAI launched parental controls for all ChatGPT users on September 29, 2025, letting parents link with teen accounts and manage usage settings from their own account. Linked teen accounts get stronger content safeguards by default, and parents can set quiet hours, disable voice, memory, image generation, and opt out of model training. The key mechanism is the alert flow: suspected self-harm signals trigger human review, and acute distress leads to email, SMS, and push notifications to parents.

Why it matters: OpenAI rolled parental controls to all ChatGPT users and disclosed a concrete self-harm escalation flow: system detection, human review, then email/SMS/push alerts to parents. HKR-K and HKR-R are strong; this is a substantive safety product update, but not a model-level launch,so

OpenAI News

Buy it in ChatGPT: Instant Checkout and the Agentic Commerce Protocol

OpenAI launched Instant Checkout in ChatGPT on September 29, 2025, letting U.S. Plus, Pro, and Free users buy from U.S. Etsy sellers in chat; it currently supports single-item purchases. OpenAI says ChatGPT has over 700 million weekly users and open-sourced the Agentic Commerce Protocol with Stripe; Stripe merchants can enable it with as little as one line of code, while the post does not disclose the fee rate merchants pay.

Why it matters: This is a high-weight ChatGPT product expansion from discovery to completed purchases, so HKR-H/K/R all pass. The post confirms U.S. Free/Plus/Pro checkout with U.S. Etsy sellers and a Stripe-backed protocol layer; merchant fee details are not disclosed, so it stays high but sub-

Sep 25, 2025Thursday

OpenAI News

More ways to work with your team and tools in ChatGPT

OpenAI rolled out shared projects for ChatGPT Business on September 25, 2025, and made them available for Enterprise and Edu plans. Shared projects support email or link invites, two access levels, and private project memory; Enterprise and Edu have them off by default under admin control. OpenAI also added Gmail, Google Calendar, Outlook, Teams, SharePoint, GitHub, Dropbox, and Box connectors, and said ChatGPT can now choose connectors automatically per prompt.

Why it matters: HKR-H/K/R all pass: shared projects, 8 connectors, and prompt-routed connector selection are concrete workflow changes with clear admin controls. I keep it below 85 because this is a collaboration-layer product update, not a model release or a broad capability jump.

OpenAI News

Introducing ChatGPT Pulse

OpenAI previewed ChatGPT Pulse for Pro users on mobile on September 25, 2025, with one daily proactive research update. It uses memory, chat history, feedback, and optional Gmail and Google Calendar connections to generate visual cards; integrations are off by default and outputs pass safety checks. The shift to async delivery matters more than the headline, but the post does not disclose the model, pricing changes, or a Plus launch date.

Why it matters: HKR-H/K/R all pass: the novel angle is proactive outreach, and the post gives concrete scope and input sources. This is a meaningful ChatGPT product update, but model details, rollout beyond Pro mobile, update cadence, and pricing changes are not disclosed, so it stays featured,

Sep 24, 2025Wednesday

OpenAI News

SAP and OpenAI partner to launch sovereign 'OpenAI for Germany'

SAP and OpenAI announced OpenAI for Germany for the German public sector, planned for 2026 and hosted by Delos Cloud on Microsoft Azure. SAP plans to expand Delos Cloud in Germany to 4,000 GPUs for AI workloads; the post does not disclose model names, pricing, or contract size. The key point is delivery: this is a sovereign public-sector deployment focused on compliance, data residency, and AI agents inside existing workflows.

Why it matters: HKR-H/K/R all pass: the story pairs a novel sovereign-deployment angle with concrete facts like a 2026 launch, Delos Cloud on Azure, and 4,000 GPUs. It matters because sovereignty and public-sector procurement are live issues, but missing model, pricing, and deal-scope details it

Sep 16, 2025Tuesday

OpenAI News

Building towards age prediction

OpenAI is building an age-prediction system for ChatGPT so users identified as under 18 are automatically routed to a teen experience. The post says low-confidence cases default to the under-18 mode, adults can verify age to unlock adult capabilities, and parental controls will ship by the end of the month with teen account linking, memory/history toggles, and blackout hours.

Why it matters: This is not a generic safety post: OpenAI is wiring age estimation into ChatGPT routing. HKR-H/K/R all pass on the auto-teen switch, fail-closed treatment for low confidence, and the privacy/liability nerve, but it remains below a major model or platform release.

Sep 15, 2025Monday

OpenAI News

Introducing upgrades to Codex

OpenAI released GPT-5-Codex and made it the default model for Codex cloud tasks and code review; in testing, it worked independently for more than 7 hours on complex tasks. OpenAI says it used 93.7% fewer tokens than GPT-5 on the lowest 10% of employee turns, while spending 2x longer reasoning, editing, and testing on the highest 10%. The key point is one model now spans interactive coding and long-running agentic execution; pricing and full availability details are not fully disclosed in the provided body.

Why it matters: This is a substantive OpenAI developer-tool update: GPT-5-Codex becomes the default for Codex cloud tasks and code review, with concrete numbers on 7-hour autonomy and token use. HKR-H/K/R all pass; pricing and full availability are not fully disclosed in the excerpt, so it stays

OpenAI News

Addendum to GPT-5 system card: GPT-5-Codex

OpenAI published a GPT-5-Codex system card addendum on September 15, 2025, stating the model is optimized for agentic coding in Codex and is available in terminal, IDE, web, GitHub, and the ChatGPT mobile app. The post says it uses reinforcement learning on real-world coding tasks, plus safety training for harmful tasks and prompt injection, with sandboxing and configurable network access. Benchmark scores, pricing, and context window are not disclosed.

Why it matters: HKR-H/K/R all pass: this is an OpenAI coding-agent model spanning terminal, IDE, GitHub, web, and mobile, with concrete training and safety details. I kept it below 85 because benchmarks, pricing, and context window are not disclosed in the body.

Sep 12, 2025Friday

OpenAI News

Working with US CAISI and UK AISI to build more secure AI systems

OpenAI said its work with US CAISI and UK AISI found and fixed 2 novel ChatGPT Agent vulnerabilities; CAISI built a proof-of-concept exploit chain with about a 50% success rate, and OpenAI fixed it within 1 business day. The post says the bugs let attackers bypass protections under certain conditions, remotely control session-accessible systems, and impersonate logged-in users; UK AISI has red-teamed bio-misuse safeguards for ChatGPT Agent and GPT-5 since May 2025, but the truncated post does not disclose further results.

Why it matters: This is not generic safety PR. OpenAI discloses 2 new ChatGPT Agent vulns, ~50% CAISI PoC success, and a 1-business-day fix, so HKR-H/K/R all pass. Kept below 85 because the UK AISI section is truncated and the broader impact is not disclosed.