▸ Free tool
Prompt Caching Calculator.
Does prompt caching save you money, and how much? Enter your prompt shape, traffic and cache prices, and compare the bill with and without caching.
▸ Presets checked October 4, 2026
Presets come from official pricing pages and are only a starting point. Every price field is editable.
▸ Per month (30 days)
- Without caching
- $0
- With caching
- $0
- You save
- $0
- Per request, no cache
- $0
- Per request, cached
- $0
- Per day, no cache / cached
- $0
- Break even hit rate
- 0%
Monthly breakdown, cached
- Cache reads (hits)
- $0
- Cache writes (misses)
- $0
- Uncached input
- $0
- Output
- $0
Live as you type · runs in your browser
Price presets.
Checked against official pricing pages on October 4, 2026, standard tier. Pick one to fill the price fields, then edit anything you like.
-
Claude Sonnet 5.5, 5 min cache
$2 in · $10 out · +25% write · 90% off reads
Anthropic pricing page
-
Claude Sonnet 5.5, 1 hour cache
$2 in · $10 out · +100% write · 90% off reads
Anthropic pricing page
-
Claude Haiku 4.5, 5 min cache
$1 in · $5 out · +25% write · 90% off reads
Anthropic pricing page
-
Claude Opus 5.5, 5 min cache
$4 in · $20 out · +25% write · 95% off reads
Anthropic pricing page
-
GPT-6.1 Sol, short context
$2 in · $10 out · +25% write · 95% off reads
OpenAI pricing page
-
Generic: no write fee, half price reads
$1 in · $4 out · +0% write · 50% off reads
Example, not a vendor
How the maths works
Call the cached prefix P, the changing input U, the output O, the input price in, the output price out, the hit rate h, the write premium w and the read discount d.
- Without caching:
(P + U) × in + O × out - With caching:
h × P × in × (1 - d) + (1 - h) × P × in × (1 + w) + U × in + O × out - Break even: caching is cheaper when
h > w / (w + d).
Worked example with the default inputs: 10,000 prefix tokens, 1,000 uncached, 500 output, $2 in and $10 out per million, +25% write, 90% off reads, 80% hits, 2,000 requests a day. Without caching a request costs 11,000 × $0.000002 + 500 × $0.00001 = $0.027, or $1,620 a month. With caching it costs $0.0016 for hits + $0.005 for misses + $0.002 + $0.005 = $0.0136, or $816 a month. That saves $804, about 50%. The break even hit rate is 0.25 / 1.15, about 22%.
Getting a high hit rate
- Fixed parts first. System prompt, tool definitions and long reference documents go at the start. User input, retrieved chunks and timestamps go at the end.
- Mind the lifetime. A cache that lives a few minutes only helps if the same prefix comes back within that time. Quiet endpoints miss a lot. A longer cache lifetime often costs a bigger write premium, so check the break even number for both.
- Do not reorder tools. Building the tool list from an unordered map, or adding one tool per request, changes the prefix and turns every request into a miss.
- Measure it. APIs report cached and written tokens in the usage data of each response. Use those numbers for the hit rate here instead of a guess.
This calculator leaves out a few things on purpose: per hour storage fees for explicit caches on some providers, higher prices above a long context threshold, batch discounts and minimum cacheable prompt lengths. If they apply to you, adjust the price fields to match.
Questions, answered
How does prompt caching pricing work?
The provider stores the start of your prompt (the prefix) after the first request. When a later request starts with the exact same prefix, those tokens are read from the cache at a discount. Writing the prefix into the cache can cost more than normal input. So each request is either a hit, which pays the cheap read price for the prefix, or a miss, which pays the write price. Tokens after the prefix and all output tokens cost the same as without caching.
What hit rate do I need for prompt caching to save money?
Divide the write premium by the write premium plus the read discount. With a 25% write premium and a 90% read discount, that is 0.25 / 1.15, about 22%. With a 100% premium (a longer cache lifetime on some providers) it is 1 / 1.9, about 53%. If there is no write premium, any hit saves money. The calculator shows this break even number for your inputs.
What counts as a cache hit?
A request that starts with the same tokens as a recent request, within the cache lifetime. Matching is on an exact prefix, so a timestamp, a user name or a changing retrieval block placed early in the prompt breaks the match for everything after it. Put fixed parts first (system prompt, tool definitions, long documents) and the changing parts last.
Why is my real hit rate lower than I expected?
Usually traffic gaps or prompt changes. If requests for the same prefix arrive further apart than the cache lifetime, the cache expires and the next request is a miss. Each deploy that edits the system prompt or tool list also starts a new cache. Providers also have a minimum prefix length before they cache anything. Check the cached token counts your API returns in its usage data and type the real hit rate in here.
Are the preset prices current?
The vendor presets were checked against the official Anthropic and OpenAI pricing pages on October 4, 2026, standard tier, short context. Prices change, and some providers charge more above a context length or for regional processing. Every field is editable, so treat the presets as a starting point and type in the numbers from your own bill or contract.
Does this calculator send my numbers anywhere?
No. The maths runs in your browser. There is no request, no account and no logging. The share button only writes your inputs into the page address so you can copy the link.
▸ Last verified:
Need this in production?
Shipping AI features? I build RAG pipelines, MCP servers and agents that hold up in production.
AI & MCP Integration
Keep reading
-
▸ Tool
LLM Cost Calculator
The full monthly bill per model, with caching as one line in it.
-
▸ Tool
Token Counter
Count your system prompt and tool definitions to get the prefix size.
-
▸ Post
/blog/redis-semantic-caching-rag/
Prompt caching saves on input. Semantic caching skips the call entirely.
-
▸ Post
/blog/prompt-engineering-vs-context-engineering/
Ordering the context so the fixed part comes first is context engineering.