提示缓存:跨供应商的成本与性能分析

提示缓存可将输入 token 成本降低多达 90%,并使长上下文工作负载变得可行。在下方比较各大 AI 供应商的缓存价格、折扣和 API 规格。

缓存要求提示词完全匹配,且因供应商而异:OpenAI 和 DeepSeek 等提供自动缓存,而 Google、Anthropic 和 Amazon 等则需要手动设置。了解其工作原理,请参阅下方的提示缓存介绍

价格

价格:缓存命中、缓存写入、输入和输出

Price (USD per M Tokens)

缓存折扣

Pricing: Cache Discount

1 - (cache hit price / input price) · Higher is better

提示缓存 API 规格

供应商
模型
输入(标准)
缓存写入
缓存命中
缓存存储
输出(标准)
自动启用
最小 token 数
缓存 TTL
备注
AnthropicAnthropic
  • Cache read tokens are 90% cheaper than base input tokens
  • Cache write tokens are 25% more expensive than base input tokens
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
$10.00
$12.50
$0.25
$20.00
$50.00
-
-
Claude Opus 5 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
$10.00
$25.00
-
-

1h cache write: $10

Amazon BedrockAmazon Bedrock
  • Amazon supports caching for Nova models and Anthropic's Claude models.
  • Pricing and usage differs between model families.
GPT-6 Astra (max)
$10.00
$12.50
$1.00
-
$50.00
-
-

30min cache write

Claude Opus 5 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
-
$25.00
-
-
OpenAIOpenAI
  • Cache read tokens are 50% cheaper than base input tokens
  • Cache persists up to one hour during off-peak periods
GPT-6 Astra (max)
$10.00
$12.50
$1.00
-
$50.00
-
-
-
GPT-5.6 Sol (max)
$4.00
$5.00
$0.40
-
$20.00
-
-
GoogleGoogle
  • Google supports caching for Gemini models and Anthropic's Claude models.
  • Pricing and usage differs between model families.
Claude Opus 5 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
-
$25.00
-
-
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
$10.00
$12.50
$1.00
-
$50.00
-
-
MetaMeta
Muse Spark 1.3 (max)
$1.25
-
$0.15
-
$4.25
-
-
Microsoft AzureMicrosoft Azure

OpenAI models:

  • Cache read tokens are 50% cheaper than base input tokens (Standard deployments)
  • Cache persists up to one hour during off-peak periods
GPT-5.6 Sol (max)
$4.00
-
$0.40
-
$20.00
-
-
Alibaba CloudAlibaba Cloud
  • Alibaba offers two cache types: implicit and explicit.
  • Implicit cache is automatically enabled. It is billed at 20% of the standard input token price.
  • Explicit cache must be activated. It creates a cache for specific content to ensure a deterministic hit within its 5-minute validity period. Tokens used to create the cache are billed at 125% of the standard input token price, while subsequent cache hits are billed at 10% of that price.
Qwen3.8 2.4T A95B
$2.00
-
$0.25
-
$6.00
-
-
Qwen3.8 27B (xhigh)
$0.50
-
$0.05
-
$3.00
-
-
GLM-5.3 (max)
$1.40
-
$0.14
-
$4.40
-
-
GLM-5.3 (max)
$2.10
-
$0.21
-
$6.60
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
DeepInfraDeepInfra
  • Prompt caching is automatic — no extra parameters required.
GLM-5.3 (max)
$1.20
-
$0.12
-
$4.00
-
-
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
DigitalOceanDigitalOcean
GLM-5.3 (max)
$0.95
-
$0.20
-
$3.40
-
-
Kimi K3 (max)
$2.55
-
$0.28
-
$12.95
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
IncoInco
GLM-5.3 (max)
$2.80
-
$0.52
-
$8.80
-
-
Kimi K3 (max)
$6.00
-
$0.60
-
$30.00
-
-
MakoraMakora
GLM-5.3 (max)
$1.35
-
$0.23
-
$4.40
-
-
Kimi K3 (max)
$2.55
-
$0.26
-
$12.75
-
-
ModularModular
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
MiniMax-M3
$0.30
-
$0.06
-
$1.20
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
WaferWafer
GLM-5.3 (max)
$1.19
-
$0.26
-
$4.40
-
-
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
ZaiZai
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
SpaceXAISpaceXAI
  • Prompt caching is not 100% guaranteed.
Grok 4.6 (high)
$2.00
-
$0.50
-
$6.00
-
-
Bitdeer AIBitdeer AI
Kimi K3 (max)
$2.66
-
$0.28
-
$13.30
-
-
GLM-5.3-Flash
$0.07
-
$0.01
-
$0.25
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
LithosAILithosAI
Kimi K3 (max)
$2.40
-
$0.24
-
$12.00
-
-
Kimi K3 (max)
$4.00
-
$0.40
-
$20.00
-
-
ModalModal
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
GLM-5.3-Flash
$0.15
-
$0.05
-
$0.50
-
-
Qwen3.8 27B (xhigh)
$0.40
-
$0.15
-
$3.00
-
-
FireworksFireworks
  • Prompt caching is enabled by default.
  • The default discount is 50%, but the exact discount varies by model.
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
Qwen3.8 2.4T A95B
$2.00
-
$0.25
-
$6.00
-
-
GMIGMI
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$1.32
-
$0.04
-
$3.96
-
-
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$1.32
-
$0.04
-
$3.96
-
-
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$1.32
-
$0.04
-
$3.96
-
-
DeepSeekDeepSeek
  • Cache read tokens are 50% cheaper on average (up to 90% with cache optimization)
  • Implements Context Caching on Disk technology
  • No guarantee of 100% cache hits
DeepSeek V4.1 Flash (Reasoning, Max Effort)
$0.30
-
$0.01
-
$1.20
-
-
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$1.32
-
$0.04
-
$3.96
-
-
CrusoeCrusoe
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$1.74
-
$0.15
-
$3.48
-
-
gpt-oss-120b (high)
$0.05
-
$0.05
-
$0.25
-
-
MistralMistral
Mistral Medium 3.5
$1.50
-
$0.15
-
$7.50
-
-
MiniMaxMiniMax
  • MiniMax supports Anthropic API compatible caching that is managed through explicit cache_control settings.
MiniMax-M3
$0.30
$0.38
$0.06
-
$1.20
-
-
Thinking MachinesThinking Machines
Inkling (xhigh)
$1.00
-
$0.17
-
$4.05
-
-
-
GroqGroq
  • The minimum cacheable prompt length varies by model, ranging from 128 to 1024 tokens depending on the specific model used.
  • All cached data automatically expires after 2 hours without use.
gpt-oss-120b (high)
$0.15
-
$0.07
-
$0.60
-
2 hrs

提示缓存介绍

什么是提示缓存?

提示缓存让语言模型推理可以复用已经处理过的输入 token,将其成本降低多达 90%,并使长上下文工作负载变得可行。把缓存策略做对,可以在输入 token 上带来巨大节省,并带来有意义的性能提升。

当你发送提示词时,系统会先检查该精确提示词是否已经处理过。如果找到(缓存命中),就会返回已存储的响应,而不是重新生成。如果未找到(缓存未命中),提示词会正常处理,并将响应存起来供以后使用。

需要关注的关键指标

  • 输入价格:你为输入 token 支付的标准价格
  • 缓存写入价格:将提示 token 写入缓存所需支付的费用;有时高于标准输入价格
  • 缓存命中价格:命中缓存的提示 token 的折扣费率
  • 缓存存储价格:每百万已缓存 token 的每小时成本(目前为 Google 独有)
  • 缓存 TTL:已缓存 token 保持可用的时间,从数小时到数天不等
  • 缓存最小 token 数:在提供缓存命中之前所需的最小匹配 token 数量

提示缓存如何工作?

当你向基于 transformer 的语言模型发送提示词时,注意力层会将每个输入 token 处理成键(K)和值(V)向量,并存储在 KV 缓存中。将这些值保留在内存中后,当再次向模型发送相同的输入 token 时,就可以避免重复处理。

直到最近,利用缓存带来的速度和成本优势还仅限于专用部署。现在,serverless API 供应商——包括前沿实验室——已经开始把部分缓存成本优势转给开发者。

最佳使用场景

  • 系统指令:必须在多次交互中包含的大型系统提示词
  • 聊天历史:伴随用户每一轮新输入的对话上下文
  • 按用户个性化的上下文:用于深度个性化的大量用户记忆或资料

实现注意事项

  • 激活方式因供应商而异:有的需要手动设置,有的提供自动缓存
  • 缓存命中折扣相当于标准输入 token 价格的 50–90% 优惠——值得花时间把它做对
  • 缓存可以提升超长提示词(5 万+ token)的性能