OpenAI’s September 22 GPT-6 prompt-caching update adds explicit cache boundaries, diagnostics and a dashboard for API developers. The practical setup is straightforward: keep reusable instructions and tool definitions stable at the start of the request, use implicit breakpoints or mark an explicit boundary after stable content you expect to reuse, then read actual cached-token counts from the response. A matching prefix can reduce input processing time and price, but a breakpoint does not guarantee a cache hit.
For GPT-6, the useful questions are where the reusable prefix ends, which request changes break a match, how to tell a real cache read from a diagnostic comparison, and how the first write and later reads affect total cost.
What prompt caching does
Prompt caching preserves the model’s key-value (KV) state for an eligible prompt prefix. On a later request, OpenAI can reuse that state when the rendered prefix and relevant settings match and a cache entry is available. The new request still processes any suffix that was not reused and generates a fresh response; caching does not save a finished answer.
The current guide sets the minimum cacheable prefix for GPT-5.6 and later—including GPT-6—at 1,024 visible input tokens. Tokens in OpenAI’s hidden system content do not count toward that minimum. Meeting the threshold makes a prefix eligible; it does not promise that the request will reach a machine with that entry or that the entry will still be available. See OpenAI’s prompt-caching guide for the current rules.
Place the breakpoint after stable content
For a multi-turn agent that appends messages to its history, the default implicit mode is a sensible starting point. OpenAI selects eligible breakpoints automatically. Choose explicit-only mode when you want to stop cache writes before a changing suffix that is unlikely to be reused. The content after the last explicit breakpoint is still processed at the ordinary input rate; it simply is not written to the cache on that request.
The Responses API request below marks the end of a stable developer prefix. Keep that block and the relevant request settings unchanged on later calls, then put each new question after it. The example is a source-checked template and was not executed.
{
"model": "gpt-6-sol",
"input": [
{
"role": "developer",
"content": [
{
"type": "input_text",
"text": "Shared instructions and reference material that remain the same...",
"prompt_cache_breakpoint": { "mode": "explicit" }
}
]
},
{
"role": "user",
"content": [
{ "type": "input_text", "text": "This request's changing question..." }
]
}
],
"prompt_cache_options": {
"mode": "explicit",
"ttl": "30m"
}
}
For GPT-6, prompt_cache_options.mode selects implicit or explicit behavior, and a supported content block can carry prompt_cache_breakpoint: { "mode": "explicit" }. In explicit mode, only marked boundaries participate; without a marker the request creates no cache write. You can place up to four writes in a request, but place them only where later calls can reuse the resulting prefix. The top-level instructions field cannot contain an explicit breakpoint; put reusable developer text in an input_text block inside the input instead.
Do not switch between implicit and explicit modes expecting the old entry to survive. OpenAI still needs a matching prefix and available entry. Compare response usage after changing the boundary or mode.
Keep the prefix stable across requests
Caching works on the full rendered prefix, not a label such as “system prompt.” Keep stable developer instructions, reference material and tool definitions early; put timestamps, per-user values, request-specific context and other changing content later. Append conversation history instead of rewriting earlier messages when the application can preserve it.
- Keep the same model, service tier, tools, tool order, names and schemas when you expect the same prefix to match.
- Changes to
parallel_tool_calls,text.format,reasoning.effort,text.verbosityor context compaction can change the rendered prefix or cache instructions. - To make tools unavailable for a turn, prefer
tool_choice: "none"or a narrowerallowed_toolsselection over deleting and rebuilding tool definitions. - On GPT-6, append a
configuration_updateitem to change reasoning effort mid-conversation while leaving the top-levelreasoning.effortunchanged. For example:{"type":"configuration_update","reasoning":{"effort":"high"}}.
When tools must be added mid-thread, OpenAI documents append-only updates. A developer-role additional_tools item can add tools, but it does not currently accept an explicit cache breakpoint. Check the current prompt-caching guide before redesigning a tool loop around that exception. General Responses and tool-loop setup is covered separately in our GPT-6 Astra API coding guide.
Measure cache reads with usage, not a “hit” label
Start with the response’s usage.input_tokens and usage.input_tokens_details fields. For each request, separate total input into ordinary input, cache reads and cache writes:
ordinary_input_tokens = input_tokens - cached_tokens - cache_write_tokens
In the Responses API, the fields are usage.input_tokens_details.cached_tokens and usage.input_tokens_details.cache_write_tokens. A write is part of the input-token total and receives the write rate; it is not an extra fee added on top of also charging those same tokens at the ordinary rate.
To investigate a miss, save a baseline response ID and set prompt_cache_options.comparison_response_id on the request you want to compare. Read response.prompt_cache_diagnostics for the comparison type, reason and estimated missed tokens. The comparison flag requests diagnostics; it does not load the old conversation or change caching behavior. A diagnostic cache_hit means no miss was detected against that baseline, not that the whole request was cached. Use usage.input_tokens_details.cached_tokens to measure actual reuse and calculate billing. Diagnostic estimates can differ from usage counts. OpenAI’s diagnostics guide lists the miss reasons and comparison workflow.
Use the Prompt Caching Dashboard to watch application-level cache reads and input composition over time. For one request, compare cached tokens, cache-write tokens, ordinary input, output tokens and latency before changing the prompt. That makes a break in reuse visible without treating a diagnostic estimate as a bill.
Calculate one implicit-mode write and later read
This is the implicit-mode lifecycle shown in OpenAI’s guide: a 12,000-token first request reports 12,000 cache-write tokens; a 15,000-token second request reports 12,000 cached tokens and 3,000 cache-write tokens. It is separate from the explicit-only request above. With an explicit breakpoint after the reusable 12,000-token prefix, a changing 3,000-token suffix would be ordinary input, not a cache write. Applying GPT-6 Sol’s current Standard short-context rates gives the table below. The documented token counts are paired with an illustrative 1,000 output tokens per request; this is not an observed API run.
| Request | Ordinary input | Cache read | Cache write | Output | Illustrative cost |
|---|---|---|---|---|---|
| First: 12,000 input, 1,000 output | 0 | 0 | 12,000 × $2.50/M = $0.0300 | 1,000 × $10/M = $0.0100 | $0.0400 |
| Second: 15,000 input, 1,000 output | 0 | 12,000 × $0.20/M = $0.0024 | 3,000 × $2.50/M = $0.0075 | 1,000 × $10/M = $0.0100 | $0.0199 |
| Two-request total | $0.0599 |
Without cache pricing, those same 27,000 input tokens would cost $0.054 at $2 per million, and the 2,000 output tokens would add $0.020, for $0.074 total. Under these assumptions, the two-request lifecycle costs about 19.1% less. That result depends on these specific token counts and one full read; it is not a general savings forecast. The first cache write costs 1.25 times the ordinary input rate, so a prefix that is written once and rarely reused can cost more than it saves.
This arithmetic uses GPT-6 Sol at Standard rates for requests below the long-context threshold. GPT-6 Sol’s listed short-context text rates are $2.00 per million ordinary input tokens, $0.20 for cached input, $2.50 for cache writes and $10.00 for output. If a request has more than 272,000 input tokens, OpenAI doubles input and cache rates and applies 1.5× output pricing to the full request. Regional processing, Batch, Flex and Fast mode have their own price adjustments, so recalculate with the rate for the actual processing mode. Check the GPT-6 Sol model page before using rates in a budget. This example uses Sol for arithmetic, not as a model recommendation; if you are choosing a model for your cache test, see our GPT-6 Sol vs Luna guide. For the separate Astra model rate table, see our Astra API rate table.
Account for cache lifetime and routing
For GPT-6, prompt_cache_options.ttl supports only "30m", which is also the default. A prefix remains eligible for reuse for at least 30 minutes after its latest write or reuse; OpenAI may retain it longer. Reusing the prefix refreshes that eligibility without another cache-write charge. Treat 30 minutes as a minimum eligibility window, not a timer that guarantees either a hit or deletion at exactly minute 30.
Cached state is held on individual machines. A request can miss even with an identical prefix if it reaches a machine without the matching entry, or the entry is unavailable. Caches are not shared between organizations or across regional processing boundaries. GPT-6 does not require prompt_cache_key for routing optimization; use a stable key only when you want separate cache accounting for users, customers or workspaces.
Use prewarming only when the prefix is predictable
For a known shared prefix, a Responses API request can set prompt_cache_options.prewarm to true to prepare the cache without generating a user-facing answer. Send the real request afterward with the same prefix and prewarm omitted or false. Prewarm writes are billed at the standard cache-write rate. It is most useful when the first interactive request would otherwise pay the prefix-processing delay; measure whether that extra write is worthwhile for the expected reuse.
Implementation checklist
- Choose a stable shared prefix that has a real reason to be at least 1,024 visible input tokens.
- Keep variable user content after the reusable prefix; use implicit mode for append-only conversations or mark deliberate boundaries in explicit mode.
- Preserve the model, tools and compatible settings when you expect a match. Use append-only tool updates and GPT-6 configuration updates where appropriate.
- For a representative request, record input, cached, cache-write and output token counts. Calculate each input category at its own rate.
- Use a recent response ID with diagnostics when reuse drops, then read actual cached-token counts and latency before changing the prompt.
- Recheck TTL, routing, long-context and processing-mode prices against the official docs before rolling a measured result into a budget.
Disclosure: The request shape and cost lifecycle above are derived from OpenAI’s documentation and current published prices. They are not an executed API test; your cache hits and costs depend on actual prompts, settings, routing, reuse and output volume.