Prompt Caching in the API
OpenAI added automatic prompt caching to GPT-4o, GPT-4o mini, o1-preview, and o1-mini API models, giving a 50% discount on recently reused input prefixes. Caching starts at 1,024 tokens and grows in 128-token increments; caches are often cleared after 5-10 minutes of inactivity and always within 1 hour of last use. The field to watch is cached_tokens in the API usage response.
Why it matters: A substantive OpenAI API update: not a new model, but it ships a 50% input discount, a 1,024-token threshold, 128-token cache steps, and cached_tokens telemetry, so HKR-H/K/R all pass. It is highly relevant to builder cost and latency, strong enough for featured, but not a same‑y