Prompt Caching: Cost & Performance Analysis Across Providers

Prompt caching can cut input token costs by up to 90% and makes long context workloads viable. Compare cache pricing, discounts and API specifications across all major AI providers below.

Caching requires exact prompt matches and varies by provider - some like OpenAI and DeepSeek offer automatic caching, while others including Google, Anthropic, and Amazon require manual setup. Learn more about how it works in our introduction to prompt caching below.

Pricing

Pricing: Cache Hit, Cache Write, Input, and Output

Price (USD per M Tokens)

Cache Discount

Pricing: Cache Discount

1 - (cache hit price / input price) · Higher is better

Prompt Caching API Specifications

Provider
Model
Input (standard)
Cache write
Cache hit
Cache storage
Output (standard)
Auto-Enabled
Min tokens
Cache TTL
Notes
AnthropicAnthropic
  • Cache read tokens are 90% cheaper than base input tokens
  • Cache write tokens are 25% more expensive than base input tokens
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
$10.00
$12.50
$0.25
$20.00
$50.00
-
-
Claude Opus 5 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
$10.00
$25.00
-
-

1h cache write: $10

Amazon BedrockAmazon Bedrock
  • Amazon supports caching for Nova models and Anthropic's Claude models.
  • Pricing and usage differs between model families.
GPT-6 Astra (max)
$10.00
$12.50
$1.00
-
$50.00
-
-

30min cache write

Claude Opus 5 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
-
$25.00
-
-
OpenAIOpenAI
  • Cache read tokens are 50% cheaper than base input tokens
  • Cache persists up to one hour during off-peak periods
GPT-6 Astra (max)
$10.00
$12.50
$1.00
-
$50.00
-
-
-
GPT-5.6 Sol (max)
$4.00
$5.00
$0.40
-
$20.00
-
-
GoogleGoogle
  • Google supports caching for Gemini models and Anthropic's Claude models.
  • Pricing and usage differs between model families.
Claude Opus 5 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
-
$25.00
-
-
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
$10.00
$12.50
$1.00
-
$50.00
-
-
MetaMeta
Muse Spark 1.3 (max)
$1.25
-
$0.15
-
$4.25
-
-
Microsoft AzureMicrosoft Azure

OpenAI models:

  • Cache read tokens are 50% cheaper than base input tokens (Standard deployments)
  • Cache persists up to one hour during off-peak periods
GPT-5.6 Sol (max)
$4.00
-
$0.40
-
$20.00
-
-
Alibaba CloudAlibaba Cloud
  • Alibaba offers two cache types: implicit and explicit.
  • Implicit cache is automatically enabled. It is billed at 20% of the standard input token price.
  • Explicit cache must be activated. It creates a cache for specific content to ensure a deterministic hit within its 5-minute validity period. Tokens used to create the cache are billed at 125% of the standard input token price, while subsequent cache hits are billed at 10% of that price.
Qwen3.8 Max (0902)
$2.00
-
$0.25
-
$6.00
-
-
Qwen3.8 27B (xhigh)
$0.50
-
$0.05
-
$3.00
-
-
GLM-5.3 (max)
$1.40
-
$0.14
-
$4.40
-
-
GLM-5.3 (max)
$2.10
-
$0.21
-
$6.60
-
-
Bitdeer AIBitdeer AI
GLM-5.3 (max)
$1.40
-
$0.14
-
$4.40
-
-
Kimi K3 (max)
$2.66
-
$0.28
-
$13.30
-
-
CrusoeCrusoe
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$1.74
-
$0.15
-
$3.48
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
DeepInfraDeepInfra
  • Prompt caching is automatic — no extra parameters required.
GLM-5.3 (max)
$1.20
-
$0.12
-
$4.00
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
DigitalOceanDigitalOcean
GLM-5.3 (max)
$0.95
-
$0.20
-
$3.40
-
-
Kimi K3 (max)
$2.55
-
$0.28
-
$12.95
-
-
FireworksFireworks
  • Prompt caching is enabled by default.
  • The default discount is 50%, but the exact discount varies by model.
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
GMIGMI
GLM-5.3 (max)
$1.40
-
$0.21
-
$4.40
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
IncoInco
GLM-5.3 (max)
$2.80
-
$0.52
-
$8.80
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
MakoraMakora
GLM-5.3 (max)
$1.35
-
$0.23
-
$4.40
-
-
Kimi K3 (max)
$2.55
-
$0.26
-
$12.75
-
-
ModularModular
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
MiniMax-M3
$0.30
-
$0.06
-
$1.20
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
WaferWafer
GLM-5.3 (max)
$1.19
-
$0.26
-
$4.40
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
ZaiZai
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
SpaceXAISpaceXAI
  • Prompt caching is not 100% guaranteed.
Grok 4.6 (high)
$2.00
-
$0.50
-
$6.00
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
LithosAILithosAI
Kimi K3 (max)
$2.40
-
$0.24
-
$12.00
-
-
Kimi K3 (max)
$4.00
-
$0.40
-
$20.00
-
-
ModalModal
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
GLM 5.3 Flash
$0.15
-
$0.05
-
$0.50
-
-
Qwen3.8 27B (xhigh)
$0.40
-
$0.15
-
$3.00
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
DeepSeek V4.1 Flash (Reasoning, Max Effort)
$0.30
-
$0.01
-
$1.20
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$1.32
-
$0.04
-
$3.96
-
-
DeepSeekDeepSeek
  • Cache read tokens are 50% cheaper on average (up to 90% with cache optimization)
  • Implements Context Caching on Disk technology
  • No guarantee of 100% cache hits
DeepSeek V4.1 Flash (Reasoning, Max Effort)
$0.30
-
$0.01
-
$1.20
-
-
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$1.32
-
$0.04
-
$3.96
-
-
MistralMistral
Mistral Medium 3.5
$1.50
-
$0.15
-
$7.50
-
-
MiniMaxMiniMax
  • MiniMax supports Anthropic API compatible caching that is managed through explicit cache_control settings.
MiniMax-M3
$0.30
$0.38
$0.06
-
$1.20
-
-
Thinking MachinesThinking Machines
Inkling (xhigh)
$1.00
-
$0.17
-
$4.05
-
-
-
GroqGroq
  • The minimum cacheable prompt length varies by model, ranging from 128 to 1024 tokens depending on the specific model used.
  • All cached data automatically expires after 2 hours without use.
gpt-oss-120b (high)
$0.15
-
$0.07
-
$0.60
-
2 hrs

Introduction to Prompt Caching

What is Prompt Caching?

Prompt caching lets language model inference reuse previously processed input tokens, cutting their cost by up to 90% and making long context workloads viable. Getting your approach to caching right can deliver huge cost savings on input tokens and meaningful performance benefits.

When you send a prompt, the system first checks if that exact prompt has been processed before. If found (cache hit), it returns the stored response instead of generating a new one. If not found (cache miss), the prompt is processed normally, and the response is stored for future use.

Key Metrics to Watch

  • Input Price: The standard price you pay for input tokens
  • Cache Write Price: What you pay to save prompt tokens into the cache; sometimes higher than standard input pricing
  • Cache Hit Price: Discounted rate for prompt tokens that hit the cache
  • Cache Storage Price: Hourly cost per million cached tokens (currently unique to Google)
  • Cache TTL: The time cached tokens remain available, ranging from hours to days
  • Cache Minimum Tokens: Minimum matching token count required before a cache hit is served

How Does Prompt Caching Work?

When you send a prompt to a transformer-based language model, the attention layers process each input token into key (K) and value (V) vectors that are stored in the KV cache. By keeping these values in memory, processing on input tokens can be avoided when identical input tokens are sent into the model again.

Until recently, leveraging the speed and cost benefits of caching was only available for dedicated deployments. Now, serverless API providers—including the frontier labs—have begun passing on some of the cost benefits of caching to developers.

Optimal Use Cases

  • System instructions: Large system prompts that must be included across many interactions
  • Chat history: Conversation context that accompanies each new user turn
  • Per-user personalized context: Extensive user memories or profiles that enable deep personalization

Implementation Considerations

  • Activation method varies by provider - some require manual setup while others offer automatic caching
  • Cache hit discounts range from 50-90% off standard input token pricing - this really is worth the time to get right
  • Caching improves performance for very long prompts (50k+ tokens)