Skip to content

#OpenAI

0 today

Jan 23, 2025Thursday

OpenAI News

Operator System Card

OpenAI published the Operator System Card on Jan 23, 2025 and said its Computer-Using Agent can be deployed only if its post-mitigation score is Medium or lower. The card rates CBRN, cybersecurity, and model autonomy as Low, and persuasion as Medium; it highlights harmful tasks, model mistakes, and prompt injection. The key mechanism is human confirmation plus task refusal: critical steps like financial transactions, emails, and calendar deletion need approval, while stock trading is fully restricted.

OpenAI News

Computer-Using Agent

OpenAI released a research preview of Computer-Using Agent on Jan 23, 2025, and is exposing it first through Operator to U.S. ChatGPT Pro users. The model combines GPT-4o vision with RL-based reasoning and acts through screenshots, a mouse, and a keyboard; it scored 38.1% on OSWorld, 58.1% on WebArena, and 87.0% on WebVoyager. The key point is API-free GUI control, while sensitive actions still require user confirmation.

Why it matters: This is a same-day OpenAI agent release: CUA powers Operator and ships first to US ChatGPT Pro users. HKR-H/K/R all pass because the GUI-control hook is novel, the post gives mechanism plus 38.1/58.1/87.0 benchmarks, and it raises concrete autonomy and safety questions.

OpenAI News

Introducing Operator

OpenAI released Operator on Jan 23, 2025 as a research preview for U.S. Pro users; it uses its own browser to click, type, and scroll through web tasks. It runs on Computer-Using Agent, combining GPT-4o vision with RL-based reasoning; the post says it sets SOTA on WebArena and WebVoyager but does not disclose scores. The key boundary is control: login, payment, and CAPTCHA flows hand control back to users, and a July 17 update says it was folded into ChatGPT agent.

Why it matters: OpenAI's Operator is a same-day, must-write product release: a browser-using agent moves ChatGPT from answering to acting. HKR-H/K/R all pass; the post gives the own-browser setup, GPT-4o+RL, and user handoff for login/payments, but US Pro limits and missing benchmark scores keep

Jan 22, 2025Wednesday

OpenAI News

Trading Inference-Time Compute for Adversarial Robustness

OpenAI reports that o1-preview and o1-mini often drive adversarial attack success rates close to zero as inference-time compute increases. The paper tests math tasks, SimpleQA prompt injection, Attack Bard images, and StrongREJECT misuse prompts; it labels the result as preliminary, and the truncated post does not fully disclose all failure cases. The key point is that this gain comes from longer reasoning at inference, not adversarial training.

Why it matters: Strong HKR-H/K/R: the hook is counterintuitive, the paper proposes a concrete mechanism, and it lands on a real safety/deployment nerve. I kept it at 82, not p1, because the post frames this as initial evidence and the excerpt does not fully disclose failure modes, cost tradeoffs

Jan 21, 2025Tuesday

OpenAI News

Announcing The Stargate Project

OpenAI, SoftBank, Oracle, and MGX launched Stargate, a new company planning to invest $500 billion over four years in US AI infrastructure for OpenAI, with $100 billion deployed immediately. SoftBank handles financing, OpenAI handles operations, and Masayoshi Son is chairman; buildout has started in Texas with Arm, Microsoft, NVIDIA, and Oracle as initial technology partners. The key signal is compute supply and control structure, not the headline rhetoric.

Why it matters: This is far above a routine partnership story: OpenAI is tying itself to a $500B, four-year infrastructure buildout with $100B to deploy immediately. HKR-H/K/R all pass because the scale is surprising, the post gives concrete capital and governance details, and the story lands on

Jan 15, 2025Wednesday

OpenAI News

Partnering with Axios expands OpenAI’s work with the news industry

OpenAI announced a content partnership with Axios and funding to expand Axios Local into 4 U.S. cities. OpenAI says it now works with nearly 20 media organizations, covering 160+ outlets, hundreds of brands, and 20+ languages. ChatGPT Search shows select summaries, excerpts, citations, and source links from partners; the post does not disclose deal value or Axios-specific technical terms.

Why it matters: This passes HKR-K and HKR-R: OpenAI gives concrete scope numbers and a specific Search distribution mechanism. It stays near the featured floor because the post is still partnership PR, and the grant size plus technical terms are not disclosed.

Jan 14, 2025Tuesday

OpenAI News

Adebayo Ogunlesi joins OpenAI's Board of Directors

OpenAI said on January 14, 2025 that Adebayo Ogunlesi joined its Board of Directors. The post identifies him as GIP's founding partner, chairman and CEO, and a senior managing director at BlackRock; OpenAI says the appointment adds infrastructure, finance, and market strategy experience to board oversight.

Why it matters: This is a high-attention personnel move: OpenAI added Adebayo Ogunlesi to its board, which carries real governance interest. HKR-H and HKR-R pass, but HKR-K is weaker because the company post gives bio details only, not board remit or structural changes, so it fits the 72–77 band

Jan 13, 2025Monday

OpenAI News

OpenAI’s Economic Blueprint

OpenAI published its Economic Blueprint on January 13, 2025, arguing the US should use nationwide AI rules and invest in chips, data, energy, and talent. The post cites $175 billion in global funds waiting for AI projects and says OpenAI will launch its Innovating for America effort at a January 30 event in Washington, DC. The real signal is policy, not product: it argues against state-by-state regulation, and the February 20 update only says it added federal AI workforce proposals.

Why it matters: This is an official OpenAI policy memo, not a product launch. HKR-K comes from the $175B investment figure and a clear federal-over-state regulatory stance; HKR-R comes from chips, energy, talent, and regulatory pressure points. HKR-H is weaker, so it lands near the featured floo

Dec 27, 2024Friday

OpenAI News

Why OpenAI’s structure must evolve to advance our mission

OpenAI says its board is evaluating changes to its nonprofit/for-profit structure, after estimating in 2019 that AGI would require about $10B. The post cites ChatGPT’s 300M+ weekly users and $137M in 2015 donations, but the specific final structure under consideration is not fully disclosed in the provided text. The key signal is financing pressure: OpenAI says investors at this scale want more conventional equity.

Dec 20, 2024Friday

OpenAI News

Deliberative alignment: reasoning enables safer language models

OpenAI published deliberative alignment on Dec 20, 2024, training o-series models to reason over written safety specs before answering. The post says o1 uses this method and needs no human-labeled CoT or answers; it says o1 beats GPT-4o on internal and external safety benchmarks, but the post does not disclose exact scores.

Why it matters: HKR-H/K/R all land: the angle is novel, the mechanism is concrete, and the topic hits a live industry debate on reasoning-model safety. I keep it at 83 because the post excerpt does not disclose key benchmark scores, so it stays in the high-quality research band, not must-write.

Dec 17, 2024Tuesday

OpenAI News

OpenAI o1 and new tools for developers

OpenAI released o1 in the API, updated the Realtime API, added Preference Fine-Tuning, and shipped beta Go/Java SDKs; o1 is rolling out first to usage tier 5 developers. Disclosed details include 60% fewer reasoning tokens than o1-preview on average, and a 60% GPT-4o audio price cut in Realtime API to $40/1M input and $80/1M output tokens. The key shift is production support for function calling, Structured Outputs, developer messages, vision, and a reasoning_effort parameter; the post is truncated, so some GPT-4o mini realtime pricing details are not disclosed here.

Why it matters: This is a substantive OpenAI developer release: o1 reaches the API with function calling, Structured Outputs, vision, and developer messages, which materially improves production readiness. HKR-H/K/R all pass; the excerpt includes concrete token and pricing data, but later GPT-4o

Dec 13, 2024Friday

OpenAI News

Elon Musk wanted an OpenAI for-profit

OpenAI said on December 13, 2024 that Elon Musk pushed in 2017 to convert OpenAI into a for-profit and sought majority equity, absolute control, and the CEO role. The post includes a timeline and email excerpts, saying Musk formed “Open Artificial Intelligence Technologies, Inc.” on September 15, 2017, and that OpenAI rejected those terms. The real signal is the capital logic: the post says the team concluded in 2017 that AGI would need billions in compute, with Ilya Sutskever referencing hardware spend below $10B.

Why it matters: HKR-H/K/R all pass: the headline has a real reversal, and the post adds specific 2017 control demands plus concrete compute-cost claims. It stays at 80 because this is a one-sided OpenAI legal narrative, not an independently verified product or research release.

Dec 9, 2024Monday

OpenAI News

Sora is here

OpenAI moved Sora out of research preview on December 9, 2024 and rolled it out to ChatGPT Plus and Pro users. Sora Turbo supports up to 1080p and 20-second videos; Plus includes up to 50 monthly 480p videos or fewer 720p generations. The key detail for practitioners is deployment scope: the UK, Switzerland, and the EEA are excluded, person uploads are limited, and OpenAI says physics and long complex actions remain weak.

Why it matters: OpenAI moved Sora from preview to paid availability, so HKR-H/K/R all pass: high-curiosity launch, concrete specs and limits, and clear impact on creator workflows. I stop below 95 because the post itself notes region blocks, restrictions on uploads with people, and instabilityon

Dec 5, 2024Thursday

OpenAI News

Introducing ChatGPT Pro

OpenAI launched ChatGPT Pro at $200 per month, with unlimited access to OpenAI o1, o1-mini, GPT-4o, Advanced Voice, and a higher-compute o1 pro mode. The post specifies a stricter 4/4 reliability metric, where a question counts only if the model answers correctly in all four attempts, but it does not disclose concrete quotas or latency figures. The key signal is compute tiering: longer reasoning time is now a paid product feature.

OpenAI News

OpenAI o1 System Card

OpenAI published the system card for o1 and o1-mini, with a deployment gate that requires post-mitigation risk scores of medium or lower. The listed Preparedness results are low for cybersecurity, medium for CBRN and persuasion, and low for model autonomy; testing covered o1-near-final-checkpoint and o1-dec5-release. The key point for practitioners is that OpenAI confirms large-scale RL for chain-of-thought reasoning, while the post does not disclose dataset mix or full benchmark scores.

Why it matters: This is a high-signal safety disclosure for a frontier OpenAI reasoning model, not routine collateral. HKR-K is strong because it publishes the deployment threshold, four Preparedness ratings, and test scope; HKR-R lands because practitioners track CoT safety, transparency, and 3

Nov 21, 2024Thursday

OpenAI News

Advancing red teaming with people and AI

OpenAI published 2 papers on Nov 21, 2024, outlining its external human red teaming process and a new automated red teaming method. The post discloses 3 concrete design choices for external testing—threat-model-based team selection, versioned model access, and structured feedback via API or ChatGPT interfaces—but this excerpt does not fully disclose the automated method's metrics or results.

Why it matters: HKR-K carries this story: OpenAI describes 2 papers and at least 3 reusable human red-team design choices. HKR-R also passes because safety and eval teams can apply the workflow; HKR-H is weaker, and the excerpt does not fully disclose automated-red-team results, so this sits at

Nov 4, 2024Monday

OpenAI News

OpenAI’s comments to the NTIA on data center growth, resilience, and security

OpenAI said on Nov. 4, 2024 that it submitted comments to the US NTIA, arguing a single 5GW data center can create or support about 40,000 jobs. The post cites $17B-$20B in state GDP impact per 5GW site and says $175B in global infrastructure funds are waiting to be deployed. The real signal is policy and AI infrastructure, not a model launch; the post does not disclose new model specs or product timelines.

Why it matters: Authoritative OpenAI policy filing. HKR-K is supported by the 5GW, jobs, GDP, and capital figures; HKR-R comes from compute supply and energy constraints. HKR-H is weak because there is no model or product update, so this sits at the low end of featured.

Oct 31, 2024Thursday

OpenAI News

Introducing ChatGPT search

OpenAI launched ChatGPT search on Oct. 31, 2024 for Plus, Team, and SearchGPT waitlist users, adding web answers with source links inside ChatGPT. It can trigger web search automatically or manually, shows a Sources sidebar, and uses a fine-tuned GPT-4o post-trained with distilled outputs from o1-preview. The shift to watch is distribution: search is folded into chat, not a separate search engine hop.

Why it matters: This is a same-day OpenAI product launch, not a minor feature tweak; search is merged into the chat UI, so HKR-H/K/R all pass. The post confirms source-linked web answers and launch conditions, and the move hits search distribution directly, which pushes it to P1.

Oct 30, 2024Wednesday

OpenAI News

Introducing SimpleQA

OpenAI open-sourced SimpleQA, a 4,326-question benchmark for factual short-answer QA and model calibration. Two independent AI trainers verified each item; a 1,000-question audit showed 94.4% agreement and an estimated inherent error rate near 3%. The key signal: it is built to challenge frontier models, and the post says GPT-4o scores below 40%.

Why it matters: This is not a routine paper post. HKR-H comes from the inversion that a 'simple' benchmark stumps frontier models; HKR-K comes from the dataset size, agreement rate, and irreducible-error estimate; HKR-R comes from the ongoing industry fixation on hallucination and calibration,so

Oct 24, 2024Thursday

OpenAI News

OpenAI’s approach to AI and national security

After the White House issued an AI National Security Memorandum on October 24, 2024, OpenAI published a framework for national security partnerships and said each use case goes through formal review by its Product Policy and National Security teams. The post names 3 existing examples: DARPA cyber defense work, USAID using ChatGPT to cut administrative burden, and bioscience collaboration with Los Alamos National Laboratory; it does not disclose pricing, model versions, or contract size. The key signal is the boundary: OpenAI says its policies ban uses that harm people, destroy property, or develop weapons, while it explores research, logistics, translation, summarization, and civilian-harm mitigation use cases with the U.S. and allies.

Why it matters: This is not a product launch, so HKR-H is weak. HKR-K and HKR-R pass on the concrete review process, 3 existing projects, and explicit weapons bans, but missing contract scale, model versions, and outcome data keep it at the low end of featured.

Oct 23, 2024Wednesday

OpenAI News

Simplifying, stabilizing, and scaling continuous-time consistency models

OpenAI introduced sCM and scaled continuous-time consistency models to 1.5B parameters on ImageNet at 512×512. The post says sCM reaches sample quality comparable to leading diffusion models in 2 sampling steps, with about 50x wall-clock speedup. Its largest model generates one sample in 0.11s on a single A100 at batch size 1 without inference optimization.

Why it matters: This clears HKR-H/K/R: the hook is 2-step sampling with diffusion-like quality, and the paper gives concrete numbers—1.5B params, ImageNet 512x512, ~50x wall-clock speed, and 0.11s per sample on one A100. Strong research release, but not a shipped product, so featured fits better

Oct 22, 2024Tuesday

OpenAI News

Dr. Ronnie Chatterji named OpenAI’s first Chief Economist

OpenAI appointed Ronnie Chatterji as its first Chief Economist on October 22, 2024, to study AI’s effects on growth, job creation, and labor markets. The post states he previously helped execute the $52 billion CHIPS and Science Act and served as Chief Economist at the US Department of Commerce. This is not a model launch; it is a new formal role at OpenAI’s economics-policy interface.

Why it matters: This is a substantive personnel and policy signal: OpenAI created a formal Chief Economist role for the first time. HKR-H and HKR-R pass, but HKR-K is limited because the post mainly gives the hire plus a $52B CHIPS credential, not a research agenda.

Oct 15, 2024Tuesday

OpenAI News

Evaluating fairness in ChatGPT

OpenAI analyzed millions of ChatGPT requests to test whether user names trigger harmful stereotypes, finding an overall rate of about 0.1%. The study used GPT-4o as a privacy-preserving evaluator; its gender-related judgments matched human raters over 90% of the time, while race and ethnicity agreement was lower. The key signal is model drift across versions: GPT-3.5 Turbo showed the highest task-level bias.

Why it matters: OpenAI provides a rare production-scale fairness audit with concrete rates, evaluator agreement, and a model-comparison result, so HKR-K is strong and HKR-R clears on trust and safety. This is a substantive research release, not a model launch or major product shift, so it lands

Oct 10, 2024Thursday

OpenAI News

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

OpenAI released MLE-bench, a benchmark built from 75 Kaggle competitions to measure ML engineering ability in AI agents. The best setup, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level on 16.9% of tasks; the benchmark code is open-source.

Why it matters: Strong HKR-H/K/R: OpenAI moves evaluation from exam-style tasks to real ML engineering, anchored by 75 Kaggle competitions and a 16.9% bronze-level result. Important as a benchmark release with concrete numbers, but still research rather than a major product launch, so featured,

Oct 9, 2024Wednesday

OpenAI News

An update on disrupting deceptive uses of AI

OpenAI says it has disrupted more than 20 operations and deceptive networks that tried to abuse its models since the start of 2024. The post ties this to election-related influence campaigns, social-media manipulation, and state-linked actors, and links an October 2024 threat report; the post does not disclose model-level breakdowns or exact enforcement mechanics.

Why it matters: OpenAI clears HKR-H/K/R here: the 20+ takedown count is a real hook, the Oct. 2024 threat-intel update adds a concrete fact, and election-linked deception is highly resonant. It stays in featured, not higher, because operation-level samples, model names, and enforcement mechanics

Oct 8, 2024Tuesday

OpenAI News

OpenAI and Hearst Content Partnership

OpenAI said on October 8, 2024 it partnered with Hearst to bring content from 20+ magazine brands and 40+ newspapers into ChatGPT and other products. The post says ChatGPT has 200 million weekly users and Hearst content will include citations and direct links; Hearst businesses outside magazines and newspapers are excluded. The key point is licensed publisher content entering the product retrieval layer, not a generic brand announcement.

Oct 3, 2024Thursday

OpenAI News

Introducing canvas, a new way to write and code with ChatGPT

OpenAI launched the canvas beta on October 3, 2024 for ChatGPT Plus and Team users, adding a GPT-4o-based workspace for writing and coding beyond chat. The post says canvas can auto-trigger or open via “use canvas,” supports targeted edits, version restore, and shortcuts like code review and bug fixing. The key signal is model training: across 20+ internal evals, trigger accuracy reached 83% for writing and 94% for coding, targeted edits beat baseline by 18%, and comment accuracy and quality improved by 30% and 16%.

OpenAI News

New Credit Facility Enhances Financial Flexibility

OpenAI said on October 3, 2024 it established a $4 billion revolving credit facility with nine banks, including JPMorgan Chase, Citi, Goldman Sachs, and HSBC. The facility was undrawn at closing; combined with its earlier $6.6 billion funding round, OpenAI said it now has over $10 billion in liquidity. The key signal is financing capacity, not a product launch; the post does not disclose interest rate, tenor, or collateral terms.

Why it matters: HKR-H/K/R all pass: a $4B undrawn revolver plus $6.6B equity is a strong, concrete financing signal from OpenAI. Important for the capex race, but missing rate, tenor, and collateral details, so it lands in featured, not P1.

Oct 2, 2024Wednesday

OpenAI News

New funding to scale the benefits of AI

OpenAI said it raised $6.6B at a $157B post-money valuation. The post says the money will fund frontier AI research, more compute, and product tools; ChatGPT has over 250M weekly users. The investor list, ownership terms, and added compute scale are not disclosed.

Why it matters: Official OpenAI funding post with hard numbers: $6.6B, $157B post-money, and 250M weekly ChatGPT users. HKR-H/K/R all pass because the scale is newsy, the facts are concrete, and the story speaks directly to the capital-compute race; missing investor and structure details keep it

Oct 1, 2024Tuesday

OpenAI News

Introducing the Realtime API

OpenAI launched a public beta of the Realtime API on Oct. 1, 2024 for all paid developers, using a persistent WebSocket to stream low-latency speech-to-speech interactions with GPT-4o. It supports function calling and interruption handling, priced at $5/1M text input tokens and $100/1M audio input tokens; the post also says audio I/O for Chat Completions would arrive in the following weeks.

Why it matters: OpenAI moved voice apps from stitched ASR+TTS calls to a persistent GPT-4o session, with function calling, interruption handling, and published audio/token pricing. HKR-H/K/R all pass, so this is a same-day must-write developer platform update and clears p1.

OpenAI News

Introducing vision to the fine-tuning API

OpenAI launched GPT-4o vision fine-tuning on Oct 1, 2024, letting paid-tier developers train with images plus text, starting from as few as 100 images. The post cites Grab improving lane-count accuracy by 20% and speed-limit sign localization by 13%, while Automat raised RPA success from 16.60% to 61.67%. The notable shift is multimodal customization in the main API; the pricing section is truncated, so full price details are not disclosed.

Why it matters: OpenAI shipped a substantive API update: GPT-4o vision fine-tuning with a 100-image floor and named gains from Grab and Automat, so HKR-H/K/R all pass. Scope is strong for builders, but the blast radius is narrower than a flagship model launch, and pricing is incomplete in the ex

OpenAI News

Prompt Caching in the API

OpenAI added automatic prompt caching to GPT-4o, GPT-4o mini, o1-preview, and o1-mini API models, giving a 50% discount on recently reused input prefixes. Caching starts at 1,024 tokens and grows in 128-token increments; caches are often cleared after 5-10 minutes of inactivity and always within 1 hour of last use. The field to watch is cached_tokens in the API usage response.

Why it matters: A substantive OpenAI API update: not a new model, but it ships a 50% input discount, a 1,024-token threshold, 128-token cache steps, and cached_tokens telemetry, so HKR-H/K/R all pass. It is highly relevant to builder cost and latency, strong enough for featured, but not a same‑y

OpenAI News

Model Distillation in the API

OpenAI launched an API distillation workflow on October 1, 2024, letting developers use outputs from GPT-4o and o1-preview to fine-tune cheaper models such as GPT-4o mini. The suite includes Stored Completions, Evals in beta, and fine-tuning; setting store:true auto-saves input-output pairs with no added latency, per the post. Pricing includes 2M free GPT-4o mini training tokens per day and 1M for GPT-4o through October 31; Evals are free up to 7 runs per week through year-end if shared with OpenAI.

Sep 26, 2024Thursday

OpenAI News

Upgrading the Moderation API with OpenAI's new multimodal moderation model

OpenAI released omni-moderation-latest on September 26, 2024, a GPT-4o-based Moderation API model for text and image inputs that is free for all developers. It adds illicit and illicit/violent text categories, supports image moderation in 6 subcategories, and improves 42% on an internal 40-language eval, with gains in 98% of languages tested.

Why it matters: Official OpenAI developer product update with strong HKR-K: new moderation classes, image coverage, and a concrete +42% result across 40 languages. HKR-R also lands because moderation and compliance affect shipping teams directly; HKR-H is weak, so this sits at the low end of the

Sep 16, 2024Monday

OpenAI News

An update on OpenAI's safety and security practices

OpenAI said on September 16, 2024 that its Safety and Security Committee will become an independent board oversight committee, chaired by Zico Kolter, for critical safeguards in model development and deployment. The committee can review major model safety evaluations and delay launches until concerns are addressed; the post also cites a 90-day review, evaluation of an AI-sector ISAC, and work with Los Alamos National Laboratory.

Why it matters: HKR-H/K/R all pass. OpenAI says an independent board committee can review major safety evaluations and delay release, which is more concrete than a generic safety post. It stays below P1 because there is no new model, external audit result, or reproducible benchmark data.

Sep 12, 2024Thursday

OpenAI News

Introducing OpenAI o1

OpenAI released o1-preview and o1-mini on Sept. 12, 2024, with access for ChatGPT Plus, Team, and tier-5 API developers. The post cites 83% vs 13% on an IMO qualifier, 84 vs 22 on a jailbreak test, and says o1-mini is 80% cheaper than o1-preview. The tradeoff is clear: the API lacks function calling, streaming, and system messages, and the models do not yet support browsing or file and image uploads.

Why it matters: A major OpenAI reasoning-model launch with all three HKR signals: HKR-H from the new “think before answering” hook, HKR-K from concrete benchmark, safety, and pricing numbers, and HKR-R from the tradeoff practitioners must manage between stronger reasoning and missing API basics.

OpenAI News

Learning to reason with LLMs

OpenAI released o1-preview and reported 74% single-sample accuracy on AIME 2024, versus 12% for GPT-4o. The post says o1 reached the 89th percentile on Codeforces and exceeded human PhD experts on GPQA Diamond; it attributes this to large-scale RL and gains from both train-time and test-time compute. The key signal is scaling reasoning with compute, not just pretraining a larger base model.

Why it matters: This is a substantive OpenAI research release with product implications. HKR-H lands on the new reasoning line, HKR-K on the disclosed benchmark jumps and compute-scaling mechanism, and HKR-R on the direct impact to model strategy and inference economics; strong 90s, not 95+.

OpenAI News

OpenAI o1-mini

OpenAI released o1-mini on Sept. 12, 2024 for Tier 5 API users at 80% lower cost than o1-preview. The post reports 70.0% on AIME and 1650 Codeforces Elo, close to o1 at 74.4% and 1673, with about 3-5x faster answers than o1-preview in one word-reasoning example. The key tradeoff is explicit: it targets STEM reasoning, while non-STEM factual knowledge is only comparable to small models like GPT-4o mini.

Why it matters: OpenAI shipped a substantive model release, so this lands in the must-write band. HKR-H comes from the 80%-cheaper/nearly-o1 tradeoff; HKR-K from AIME 70.0 and Codeforces 1650; HKR-R from immediate developer cost/performance implications.

Aug 20, 2024Tuesday

OpenAI News

OpenAI partners with Condé Nast

OpenAI said on August 20, 2024 that it partnered with Condé Nast to surface content from brands such as Vogue, The New Yorker, and Wired in ChatGPT and the SearchGPT prototype. The post names at least nine Condé Nast brands and says SearchGPT links directly to source stories; it does not disclose deal value, licensing scope, revenue terms, or launch regions.

Why it matters: This is a meaningful OpenAI licensing/distribution move: ChatGPT and SearchGPT will show Condé Nast content with direct links. HKR-K and HKR-R pass, but HKR-H is limited because deal terms, rollout scope, and revenue share are undisclosed, so it lands at low-featured.

OpenAI News

Fine-tuning now available for GPT-4o

OpenAI has opened GPT-4o fine-tuning to developers on all paid tiers, with 1M free training tokens per org per day through September 23. Training costs $25 per 1M tokens, and inference costs $3.75 per 1M input tokens and $15 per 1M output tokens on gpt-4o-2024-08-06. The signal for practitioners: partners reported 43.8% on SWE-bench Verified and 71.83% on BIRD-SQL with fine-tuned GPT-4o.

Why it matters: This is a substantive OpenAI developer release with concrete details: temporary free training quota, train/inference prices, base model version, and two benchmark datapoints. HKR-H/K/R all pass, but this is an API capability expansion, not a new frontier-model launch or platform-