OpenAI ships better prompt caching for GPT-6, plus a dashboard and diagnostics
OpenAI 为 GPT-6 推出改进的提示词缓存系统与诊断工具
GPT-6 prompt caching now hits more often by default, with discounts for shared prefixes reused within 30 minutes. A new dashboard tracks hit rates and a diagnostics tool pinpoints misses—e.g., a tools_changed reason costing 5,629 tokens. Developers can set explicit cache breakpoints, adjust reasoning effort without breaking cache, and prewarm context to cut latency. GitHub Copilot reports a >50% drop in tokens needing fresh processing; Manus raised cache hit rates from ~85% to >90% in under a week.
Why it matters: Official OpenAI post on GPT-6 prompt caching improvements with a diagnostic dashboard and manual breakpoints — a real cost win for agent developers. Score stays below 85 because it's infrastructure, not a new model, but the concrete numbers and tooling details make it a solid ...