25 September 2026 — OpenAI detailed changes to prompt caching for its GPT-6 family on 22 September, adding a dashboard, a tool for diagnosing cache misses and controls over which parts of a prompt can be reused. The company says the system achieves higher cache hit rates by default and offers discounts of up to 90% on cached input tokens.
Key points
- Eligible shared prefixes reused within 30 minutes receive cache discounts.
- The dashboard tracks hit rates; diagnostics identify changes that prevented reuse.
- Explicit breakpoints give developers control over which prefixes to cache.
- GPT-6 can vary reasoning effort between responses while retaining cached context.
GPT-6 discounts prefixes reused within 30 minutes
A persistent agent may make a succession of requests that carry forward instructions, descriptions of available tools and material from earlier exchanges. OpenAI caches that shared material so later requests can reuse the computation. The reusable part is a prefix: the opening stretch of a prompt, much like the pages at the front of a file that remain the same while new pages are added at the back. Under the GPT-6 update, OpenAI says eligible prefixes reused within a 30-minute window receive cache discounts.
Revising a document over several exchanges could carry the same instructions and background material into each request. If that opening material remained eligible for reuse, later responses could require less fresh processing and begin sooner. Changes to the shared material could instead require it to be processed again.
The discount concerns cached input tokens, rather than the entire cost of a request. OpenAI describes the limit as up to 90%, so the amount an application could save would depend on how much of its input qualifies for reuse. The company says its default caching changes increase hit rates, while the new controls let developers adapt caching to the requests their applications make.
The caching changes accompanied the release of GPT-6 Sol and Luna. Their API prices are 50% lower than the promotional prices of GPT-5.6 Sol and Luna, Analytics India Magazine reported on 23 September. That model-price comparison and the discount on eligible cached input describe different parts of the cost of running an application.
Prompt Caching Dashboard tracks missed reuse
OpenAI’s Prompt Caching Dashboard shows the share of an application’s input served from cache, with a view of hit rates over time and a chart separating cached from uncached tokens. That gives developers a way to watch for a drop after they change an application, rather than relying on the expected behaviour of the caching system.
For an individual miss, the diagnostics tool compares a request with a recent response. OpenAI says it identifies changes to the model, tools, settings or input that prevented reuse and estimates how many tokens were affected. The distinction matters because the dashboard can reveal a falling hit rate, while a comparison of requests can point to the change associated with a particular miss.
GitHub Chief Product Officer Mario Rodriguez said the company had reduced the share of prompt tokens needing fresh processing by more than 50% against its previous baseline. He described that result as spanning billions of requests to OpenAI models over several months. His figure measures GitHub’s share of tokens processed afresh, rather than a reduction in the price of every request.
GPT-6 breakpoints protect reusable context
Explicit cache breakpoints let developers choose which prompt prefixes to reuse. OpenAI advises keeping tool definitions, schemas and their ordering stable as an agent’s needs change. Its guidance says applications can use allowed_tools to limit which tools are callable, or set tool_choice to none, instead of removing definitions from the prompt. New developer instructions can be appended towards the end of the context.
GPT-6 also allows reasoning effort to change between responses without breaking the cache, OpenAI says. The method is to append a configuration_update while leaving request-level reasoning effort unchanged. An application can therefore request more effort for a harder task, or less for a routine follow-up, while preserving context that qualifies for reuse.
OpenAI also offers prewarming: preparing known instructions, tool definitions or reference material before a request arrives so that processing takes place outside the wait for a response. The company gives application startup, before the first question, as an example of when shared context could be prepared.
Strawberry Browser Chief Technology Officer Arian Hanifi said the dashboard and diagnostics helped his company raise cache hit rates by a few percentage points and reduce costs by 20%. He said its use of explicit breakpoints keeps stable context cached while placing content that changes frequently at the end of the prompt.