{"article":{"slug":"better-prompt-caching-for-gpt-6","title":"Better prompt caching for GPT-6","subtitle":null,"summary":"OpenAI explains GPT-6 prompt-caching improvements: higher cache hit rates, new diagnostics, explicit breakpoints, and controls aimed at cutting latency and inference cost.","content_type":"changelog","language":"en","canonical_url":"https://openai.com/index/better-prompt-caching-for-gpt-6/","author":{"name":"OpenAI","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"OpenAI","url":"https://openai.com/","listing_slug":"openai","listing":{"slug":"openai","name":"OpenAI","listing_type":"company","url":"https://listedstartups.com/companies/openai"}},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Engineering","slug":"engineering","url":"https://listedarticles.com/topics/engineering"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"}],"about_listings":[{"slug":"chatgpt","name":"ChatGPT","listing_type":"product","url":"https://listedstartups.com/products/chatgpt"}],"cover_image_url":null,"license":"all-rights-reserved","word_count":623,"reading_minutes":3,"published_at":"2026-09-22T21:21:23.141Z","added_at":"2026-09-22T21:21:23.141Z","updated_at":"2026-09-22T21:21:23.141Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/better-prompt-caching-for-gpt-6","markdown_url":"https://listedarticles.com/articles/better-prompt-caching-for-gpt-6.md","example":false,"citation":"OpenAI, OpenAI. \"Better prompt caching for GPT-6.\" 22 Sept 2026. https://openai.com/index/better-prompt-caching-for-gpt-6/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://openai.com/index/better-prompt-caching-for-gpt-6/"},"body_markdown":"# Better prompt caching for GPT‑6\n\nHigher cache hit rates and new tools to help persistent agents run faster and cost less.\n\nGPT‑6 enables persistent agents to work for hours on complex tasks, from refactoring codebases to producing well-researched documents and presentations. The applications behind these agents make a series of API requests that build on one another, often carrying forward the same instructions, tool definitions, and context from earlier turns. OpenAI caches that shared context to reuse computation across requests, reducing response times and giving developers discounts of up to 90% on cached input tokens.\n\nWith the GPT‑6 family, we launched an improved prompt caching system that delivers higher cache hit rates by default. We now give cache discounts for eligible shared prefixes reused within a 30-minute window. We’re also introducing new tools to help developers monitor cache performance, diagnose misses, and choose how much of a prompt to cache.\n\nThe new [__Prompt Caching Dashboard__(opens in a new window)](https://platform.openai.com/usage?usage_section=prompt-caching) shows how much of your application’s input is served from cache. Track hit rates over time and use the input composition chart to compare cached and uncached tokens. These views help you spot drops in cache hits and evaluate how changes to your application impact caching performance.\n\nWhen you see an unexpected cache miss, use the [__prompt caching diagnostics tool__(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching/diagnostics) to understand what happened. Compare a request with a recent response to identify changes to the model, tools, settings, or input that prevented reuse. The estimated number of affected tokens helps you assess the size of the impact and decide how you can optimize your integration to maximize cache hit rates. \n\n```\n{\n  \"prompt_cache_diagnostics\": {\n    \"type\": \"cache_miss\",\n    \"reason\": \"tools_changed\",\n    \"comparison_reusable_tokens\": 5629,\n    \"cache_missed_tokens\": 5629\n  }\n}\n```\n**Choose what to cache.** Explicit cache breakpoints let you choose which prompt prefixes to reuse. The refreshed [__prompt caching guide__(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching) explains how to use them, how long cached prefixes remain eligible, and how changes to tools and inputs affect reuse.\n\n**Adjust reasoning effort without breaking cache.** On GPT‑6 models, you can now [__change reasoning effort__(opens in a new window)](https://developers.openai.com/api/docs/guides/reasoning?api-mode=responses#change-reasoning-mid-conversation) between responses without breaking cache. Raise effort for a harder task or lower it for a routine follow-up by appending a `configuration_update` while leaving request-level reasoning effort unchanged. This lets you adjust how much reasoning a task needs while preserving reusable context.\n\n**Preserve cache as tools and instructions change.** As your agent’s tool use needs change, keep tool definitions, schemas, and ordering stable so earlier context stays reusable. Use `allowed_tools` to make only the relevant tools callable, or set tool_choice to none when no tools are needed, instead of removing definitions. Use new developer messages to append new instructions towards the end of the context to override older ones. See our [__guidance on managing tool changes__(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching#manage-tools-with-append-only-updates).\n\n**Prewarm the cache to reduce latency.** [__Prewarming__(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching#prewarm-the-cache) prepares known context ahead of time so the model can start responding sooner when a request arrives. For example, an application can prewarm shared instructions, tool definitions, or reference material during startup, before the user asks their first question. This moves processing out of the user’s wait time.\n\nThese optional controls build on the engine’s default performance, helping you tailor caching to your workload.\n\n- Monitor cache hit rates in the [__Prompt Caching Dashboard__(opens in a new window)](https://platform.openai.com/usage?usage_section=prompt-caching) .\n- Investigate unexpected misses with the [__diagnostics tool__(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching/diagnostics) .\n- Follow the [__prompt caching guide__(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching) to improve your setup, or[__use Codex__(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching#how-to-optimize-prompt-caching) to review your code, apply improvements, and measure results.\n\n## Keep reading\n\n[View all](https://openai.com/news/)","body_html":"<h1 id=\"better-prompt-caching-for-gpt-6\">Better prompt caching for GPT‑6</h1>\n<p>Higher cache hit rates and new tools to help persistent agents run faster and cost less.</p>\n<p>GPT‑6 enables persistent agents to work for hours on complex tasks, from refactoring codebases to producing well-researched documents and presentations. The applications behind these agents make a series of API requests that build on one another, often carrying forward the same instructions, tool definitions, and context from earlier turns. OpenAI caches that shared context to reuse computation across requests, reducing response times and giving developers discounts of up to 90% on cached input tokens.</p>\n<p>With the GPT‑6 family, we launched an improved prompt caching system that delivers higher cache hit rates by default. We now give cache discounts for eligible shared prefixes reused within a 30-minute window. We’re also introducing new tools to help developers monitor cache performance, diagnose misses, and choose how much of a prompt to cache.</p>\n<p>The new <a href=\"https://platform.openai.com/usage?usage_section=prompt-caching\" rel=\"nofollow ugc noopener\"><strong>Prompt Caching Dashboard</strong>(opens in a new window)</a> shows how much of your application’s input is served from cache. Track hit rates over time and use the input composition chart to compare cached and uncached tokens. These views help you spot drops in cache hits and evaluate how changes to your application impact caching performance.</p>\n<p>When you see an unexpected cache miss, use the <a href=\"https://developers.openai.com/api/docs/guides/prompt-caching/diagnostics\" rel=\"nofollow ugc noopener\"><strong>prompt caching diagnostics tool</strong>(opens in a new window)</a> to understand what happened. Compare a request with a recent response to identify changes to the model, tools, settings, or input that prevented reuse. The estimated number of affected tokens helps you assess the size of the impact and decide how you can optimize your integration to maximize cache hit rates. </p>\n<pre><code>{\n  &quot;prompt_cache_diagnostics&quot;: {\n    &quot;type&quot;: &quot;cache_miss&quot;,\n    &quot;reason&quot;: &quot;tools_changed&quot;,\n    &quot;comparison_reusable_tokens&quot;: 5629,\n    &quot;cache_missed_tokens&quot;: 5629\n  }\n}</code></pre>\n<p><strong>Choose what to cache.</strong> Explicit cache breakpoints let you choose which prompt prefixes to reuse. The refreshed <a href=\"https://developers.openai.com/api/docs/guides/prompt-caching\" rel=\"nofollow ugc noopener\"><strong>prompt caching guide</strong>(opens in a new window)</a> explains how to use them, how long cached prefixes remain eligible, and how changes to tools and inputs affect reuse.</p>\n<p><strong>Adjust reasoning effort without breaking cache.</strong> On GPT‑6 models, you can now <a href=\"https://developers.openai.com/api/docs/guides/reasoning?api-mode=responses#change-reasoning-mid-conversation\" rel=\"nofollow ugc noopener\"><strong>change reasoning effort</strong>(opens in a new window)</a> between responses without breaking cache. Raise effort for a harder task or lower it for a routine follow-up by appending a <code>configuration_update</code> while leaving request-level reasoning effort unchanged. This lets you adjust how much reasoning a task needs while preserving reusable context.</p>\n<p><strong>Preserve cache as tools and instructions change.</strong> As your agent’s tool use needs change, keep tool definitions, schemas, and ordering stable so earlier context stays reusable. Use <code>allowed_tools</code> to make only the relevant tools callable, or set tool_choice to none when no tools are needed, instead of removing definitions. Use new developer messages to append new instructions towards the end of the context to override older ones. See our <a href=\"https://developers.openai.com/api/docs/guides/prompt-caching#manage-tools-with-append-only-updates\" rel=\"nofollow ugc noopener\"><strong>guidance on managing tool changes</strong>(opens in a new window)</a>.</p>\n<p><strong>Prewarm the cache to reduce latency.</strong> <a href=\"https://developers.openai.com/api/docs/guides/prompt-caching#prewarm-the-cache\" rel=\"nofollow ugc noopener\"><strong>Prewarming</strong>(opens in a new window)</a> prepares known context ahead of time so the model can start responding sooner when a request arrives. For example, an application can prewarm shared instructions, tool definitions, or reference material during startup, before the user asks their first question. This moves processing out of the user’s wait time.</p>\n<p>These optional controls build on the engine’s default performance, helping you tailor caching to your workload.</p>\n<ul><li>Monitor cache hit rates in the <a href=\"https://platform.openai.com/usage?usage_section=prompt-caching\" rel=\"nofollow ugc noopener\"><strong>Prompt Caching Dashboard</strong>(opens in a new window)</a> .</li><li>Investigate unexpected misses with the <a href=\"https://developers.openai.com/api/docs/guides/prompt-caching/diagnostics\" rel=\"nofollow ugc noopener\"><strong>diagnostics tool</strong>(opens in a new window)</a> .</li><li>Follow the <a href=\"https://developers.openai.com/api/docs/guides/prompt-caching\" rel=\"nofollow ugc noopener\"><strong>prompt caching guide</strong>(opens in a new window)</a> to improve your setup, or<a href=\"https://developers.openai.com/api/docs/guides/prompt-caching#how-to-optimize-prompt-caching\" rel=\"nofollow ugc noopener\"><strong>use Codex</strong>(opens in a new window)</a> to review your code, apply improvements, and measure results.</li></ul>\n<h2 id=\"keep-reading\">Keep reading</h2>\n<p><a href=\"https://openai.com/news/\" rel=\"nofollow ugc noopener\">View all</a></p>","headings":[{"level":1,"text":"Better prompt caching for GPT‑6","id":"better-prompt-caching-for-gpt-6"},{"level":2,"text":"Keep reading","id":"keep-reading"}]}}