提示缓存:跨供应商的成本与性能分析

提示缓存可将输入 token 成本降低多达 90%,并使长上下文工作负载变得可行。在下方比较各大 AI 供应商的缓存价格、折扣和 API 规格。

缓存要求提示词完全匹配,且因供应商而异:OpenAI 和 DeepSeek 等提供自动缓存,而 Google、Anthropic 和 Amazon 等则需要手动设置。了解其工作原理,请参阅下方的提示缓存介绍

价格

价格:缓存命中、缓存写入、输入和输出

Price (USD per M Tokens)
Reasoning models are indicated by a lightbulb icon

Price per token for cached prompts (previously processed), typically offering a significant discount compared to regular input price, represented as USD per million tokens. The values shown here are the cache hit price; cache write and cache storage are billed separately and vary by provider — see "Cache pricing by provider" for detail.

缓存折扣

Pricing: Cache Discount

1 - (cache hit price / input price) · Higher is better
Reasoning models are indicated by a lightbulb icon

Reduction in input token cost due to cache hit relative to input price. Formula: 1 - (Cache Hit Price per Token / Input Token Price), where cache hit price is the first-party cache hit price or the median provider cache hit price. Note that this discount figure does not account for all costs associated with cache hits, such as cache write and storage costs.

提示缓存 API 规格

供应商
模型
输入(标准)
缓存写入
缓存命中
缓存存储
输出(标准)
自动启用
最小 token 数
缓存 TTL
备注
Amazon BedrockAmazon Bedrock
  • Amazon supports caching for Nova models and Anthropic's Claude models.
  • Pricing and usage differs between model families.
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
-
$25.00
-
-
GPT-5.5 (xhigh)
$5.50
-
$0.55
-
$33.00
-
-
AnthropicAnthropic
  • Cache read tokens are 90% cheaper than base input tokens
  • Cache write tokens are 25% more expensive than base input tokens
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
$10.00
$12.50
$1.00
$20.00
$50.00
-
-
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
$10.00
$25.00
-
-

1h cache write: $10

GoogleGoogle
  • Google supports caching for Gemini models and Anthropic's Claude models.
  • Pricing and usage differs between model families.
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
$10.00
$12.50
$1.00
-
$50.00
-
-
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
-
$25.00
-
-
OpenAIOpenAI
  • Cache read tokens are 50% cheaper than base input tokens
  • Cache persists up to one hour during off-peak periods
GPT-5.5 (xhigh)
$5.00
-
$0.50
-
$30.00
1024
5-10 minutes
GPT-5.4 mini (xhigh)
$0.75
-
$0.07
-
$4.50
-
-
SpaceXAISpaceXAI
  • Prompt caching is not 100% guaranteed.
Grok 4.3 (high)
$1.25
$1.25
$0.20
-
$2.50
-
-

For requests greater than 200k tokens, pricing is $2.50 per 1M input tokens, $0.40 per 1M cached input tokens, and $5.00 per 1M output tokens

Kimi K2.6
$0.95
-
$0.16
-
$4.00
-
-
FireworksFireworks
  • Prompt caching is enabled by default.
  • The default discount is 50%, but the exact discount varies by model.
Kimi K2.6
$0.95
-
$0.16
-
$4.00
-
-
Kimi K2.6
$0.95
-
$0.16
-
$4.00
-
-
Kimi K2.6
$0.75
-
$0.16
-
$3.50
-
-
Microsoft AzureMicrosoft Azure

OpenAI models:

  • Cache read tokens are 50% cheaper than base input tokens (Standard deployments)
  • Cache persists up to one hour during off-peak periods
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)
$2.00
$2.50
$0.20
-
$10.00
-
-
GPT-5.4 mini (xhigh)
$0.75
-
$0.07
-
$4.50
-
-
DeepSeekDeepSeek
  • Cache read tokens are 50% cheaper on average (up to 90% with cache optimization)
  • Implements Context Caching on Disk technology
  • No guarantee of 100% cache hits
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$0.43
-
$0.00
-
$0.87
-
-
Kimi K2.6
$0.65
-
$0.15
-
$3.41
-
-
CrusoeCrusoe
Kimi K2.6
$0.70
-
$0.35
-
$3.50
-
-
DeepInfraDeepInfra
  • Prompt caching is automatic — no extra parameters required.
Kimi K2.6
$0.75
-
$0.15
-
$3.50
-
-
GMIGMI
Kimi K2.6
$0.85
-
$0.14
-
$3.60
-
-
Kimi K2.6
$0.80
-
$0.16
-
$3.40
-
-
Kimi K2.6
$0.77
-
$0.14
-
$3.40
-
-
CloudflareCloudflare
  • Prefix caching is enabled by default. To maximize cache hit rates, a header must be sent.
Kimi K2.6
$0.95
-
$0.16
-
$4.00
-
-

提示缓存介绍

什么是提示缓存?

提示缓存让语言模型推理可以复用已经处理过的输入 token,将其成本降低多达 90%,并使长上下文工作负载变得可行。把缓存策略做对,可以在输入 token 上带来巨大节省,并带来有意义的性能提升。

当你发送提示词时,系统会先检查该精确提示词是否已经处理过。如果找到(缓存命中),就会返回已存储的响应,而不是重新生成。如果未找到(缓存未命中),提示词会正常处理,并将响应存起来供以后使用。

需要关注的关键指标

  • 输入价格:你为输入 token 支付的标准价格
  • 缓存写入价格:将提示 token 写入缓存所需支付的费用;有时高于标准输入价格
  • 缓存命中价格:命中缓存的提示 token 的折扣费率
  • 缓存存储价格:每百万已缓存 token 的每小时成本(目前为 Google 独有)
  • 缓存 TTL:已缓存 token 保持可用的时间,从数小时到数天不等
  • 缓存最小 token 数:在提供缓存命中之前所需的最小匹配 token 数量

提示缓存如何工作?

当你向基于 transformer 的语言模型发送提示词时,注意力层会将每个输入 token 处理成键(K)和值(V)向量,并存储在 KV 缓存中。将这些值保留在内存中后,当再次向模型发送相同的输入 token 时,就可以避免重复处理。

直到最近,利用缓存带来的速度和成本优势还仅限于专用部署。现在,serverless API 供应商——包括前沿实验室——已经开始把部分缓存成本优势转给开发者。

最佳使用场景

  • 系统指令:必须在多次交互中包含的大型系统提示词
  • 聊天历史:伴随用户每一轮新输入的对话上下文
  • 按用户个性化的上下文:用于深度个性化的大量用户记忆或资料

实现注意事项

  • 激活方式因供应商而异:有的需要手动设置,有的提供自动缓存
  • 缓存命中折扣相当于标准输入 token 价格的 50–90% 优惠——值得花时间把它做对
  • 缓存可以提升超长提示词(5 万+ token)的性能