Skip to content

Deployment & engineering

Running models in practice: inference optimization, memory and cost, serving architecture and infrastructure choices.

Latest picks

461–463 of 463

Oct 2, 2024Wednesday

OpenAI News

New funding to scale the benefits of AI

OpenAI said it raised $6.6B at a $157B post-money valuation. The post says the money will fund frontier AI research, more compute, and product tools; ChatGPT has over 250M weekly users. The investor list, ownership terms, and added compute scale are not disclosed.

Why it matters: Sets the scale of OpenAI's war chest and the user base behind it.

Oct 1, 2024Tuesday

OpenAI News

Prompt Caching in the API

OpenAI added automatic prompt caching to GPT-4o, GPT-4o mini, o1-preview, and o1-mini API models, giving a 50% discount on recently reused input prefixes. Caching starts at 1,024 tokens and grows in 128-token increments; caches are often cleared after 5-10 minutes of inactivity and always within 1 hour of last use. The field to watch is cached_tokens in the API usage response.

Why it matters: Halves input cost for repeated long prefixes, which changes how long-context apps are priced.

Aug 6, 2024Tuesday

OpenAI News

Introducing Structured Outputs in the API

OpenAI released Structured Outputs on Aug 6, 2024, making model outputs conform to developer-supplied JSON Schemas; `gpt-4o-2024-08-06` scored 100% on complex schema-following evals versus under 40% for `gpt-4-0613`. The feature is enabled with `strict: true` in function calling and works on tool-supporting models including `gpt-4-0613`, `gpt-3.5-turbo-0613`, and later. The key shift is constrained decoding plus schema training, not just valid JSON from JSON mode.

Why it matters: Structured Outputs makes models follow a developer's JSON Schema exactly, and the accuracy gap between new and old models shows why schema adherence is a model capability, not a prompt trick.