プロンプトキャッシュ:プロバイダー横断のコストと性能分析

プロンプトキャッシュは入力トークンのコストを最大90%削減でき、長いコンテキストのワークロードを現実的にします。主要AIプロバイダーのキャッシュ料金、割引、API仕様を以下で比較できます。

キャッシュにはプロンプトの完全一致が必要で、プロバイダーによって異なります。OpenAIやDeepSeekのように自動キャッシュを提供する一方、Google、Anthropic、Amazonなどは手動設定が必要です。仕組みの詳細は、以下のプロンプトキャッシュの紹介をご覧ください。

料金

料金:キャッシュヒット、キャッシュ書き込み、入力、出力

Price (USD per M Tokens)

キャッシュ割引

Pricing: Cache Discount

1 - (cache hit price / input price) · Higher is better

プロンプトキャッシュ API仕様

プロバイダー
モデル
入力(標準)
キャッシュ書き込み
キャッシュヒット
キャッシュ保存
出力(標準)
自動有効化
最小トークン数
キャッシュTTL
注記
AnthropicAnthropic
  • Cache read tokens are 90% cheaper than base input tokens
  • Cache write tokens are 25% more expensive than base input tokens
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)
$4.00
$5.00
$0.20
-
$20.00
-
-
-
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
$10.00
$12.50
$0.25
$20.00
$50.00
-
-
Amazon BedrockAmazon Bedrock
  • Amazon supports caching for Nova models and Anthropic's Claude models.
  • Pricing and usage differs between model families.
GPT-6 Astra (max)
$10.00
$12.50
$1.00
-
$50.00
-
-

30min cache write

Claude Opus 5 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
-
$25.00
-
-
OpenAIOpenAI
  • Cache read tokens are 50% cheaper than base input tokens
  • Cache persists up to one hour during off-peak periods
GPT-6 Astra (max)
$10.00
$12.50
$1.00
-
$50.00
-
-
-
GPT-6 Sol (max)
$2.00
$2.50
$0.20
-
$10.00
-
-
-
GoogleGoogle
  • Google supports caching for Gemini models and Anthropic's Claude models.
  • Pricing and usage differs between model families.
Claude Opus 5 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
-
$25.00
-
-
Gemini 3.8 Flash (high)
$0.75
$0.75
$0.07
-
$3.75
-
-
MetaMeta
Muse Spark 1.3 (max)
$1.25
-
$0.15
-
$4.25
-
-
Microsoft AzureMicrosoft Azure

OpenAI models:

  • Cache read tokens are 50% cheaper than base input tokens (Standard deployments)
  • Cache persists up to one hour during off-peak periods
GPT-5.6 Sol (max)
$4.00
-
$0.40
-
$20.00
-
-
SpaceXAISpaceXAI
  • Prompt caching is not 100% guaranteed.
Grok 4.7 (xhigh)
$2.00
-
$0.50
-
$6.00
-
-
Grok 4.6 (high)
$2.00
-
$0.50
-
$6.00
-
-
MiMo-V2.6-Pro
$0.43
-
$0.00
-
$0.87
-
-
Alibaba CloudAlibaba Cloud
  • Alibaba offers two cache types: implicit and explicit.
  • Implicit cache is automatically enabled. It is billed at 20% of the standard input token price.
  • Explicit cache must be activated. It creates a cache for specific content to ensure a deterministic hit within its 5-minute validity period. Tokens used to create the cache are billed at 125% of the standard input token price, while subsequent cache hits are billed at 10% of that price.
Qwen3.8 Max (0902)
$2.00
-
$0.25
-
$6.00
-
-
Qwen3.8 27B (xhigh)
$0.50
-
$0.05
-
$3.00
-
-
GLM-5.3 (max)
$1.40
-
$0.14
-
$4.40
-
-
GLM-5.3 (max)
$2.10
-
$0.21
-
$6.60
-
-
Bitdeer AIBitdeer AI
GLM-5.3 (max)
$1.40
-
$0.14
-
$4.40
-
-
Kimi K3 (max)
$2.66
-
$0.28
-
$13.30
-
-
CrusoeCrusoe
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
DeepInfraDeepInfra
  • Prompt caching is automatic — no extra parameters required.
GLM-5.3 (max)
$1.20
-
$0.12
-
$4.00
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
DigitalOceanDigitalOcean
GLM-5.3 (max)
$0.95
-
$0.20
-
$3.40
-
-
Kimi K3 (max)
$2.55
-
$0.28
-
$12.95
-
-
FireworksFireworks
  • Prompt caching is enabled by default.
  • The default discount is 50%, but the exact discount varies by model.
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
GMIGMI
GLM-5.3 (max)
$1.40
-
$0.21
-
$4.40
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
IncoInco
GLM-5.3 (max)
$2.80
-
$0.52
-
$8.80
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
MakoraMakora
GLM-5.3 (max)
$1.35
-
$0.23
-
$4.40
-
-
Kimi K3 (max)
$2.55
-
$0.26
-
$12.75
-
-
ModularModular
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
MiniMax-M3
$0.30
-
$0.06
-
$1.20
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
WaferWafer
GLM-5.3 (max)
$1.19
-
$0.26
-
$4.40
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
ZaiZai
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
Step 5 Preview
$1.00
-
$0.05
-
$2.70
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
LithosAILithosAI
Kimi K3 (max)
$2.40
-
$0.24
-
$12.00
-
-
Kimi K3 (max)
$4.00
-
$0.40
-
$20.00
-
-
ModalModal
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
GLM 5.3 Flash
$0.15
-
$0.05
-
$0.50
-
-
Qwen3.8 27B (xhigh)
$0.40
-
$0.15
-
$3.00
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
DeepSeek V4.1 Flash (Reasoning, Max Effort)
$0.30
-
$0.01
-
$1.20
-
-
GLM 5.3 Flash
$0.15
-
$0.03
-
$0.50
-
-
DeepSeek V4.1 Flash (Reasoning, Max Effort)
$0.30
-
$0.01
-
$1.20
-
-
DeepSeekDeepSeek
  • Cache read tokens are 50% cheaper on average (up to 90% with cache optimization)
  • Implements Context Caching on Disk technology
  • No guarantee of 100% cache hits
DeepSeek V4.1 Flash (Reasoning, Max Effort)
$0.30
-
$0.01
-
$1.20
-
-
MistralMistral
Mistral Medium 3.5
$1.50
-
$0.15
-
$7.50
-
-
MiniMaxMiniMax
  • MiniMax supports Anthropic API compatible caching that is managed through explicit cache_control settings.
MiniMax-M3
$0.30
$0.38
$0.06
-
$1.20
-
-
Thinking MachinesThinking Machines
Inkling (xhigh)
$1.00
-
$0.17
-
$4.05
-
-
-

プロンプトキャッシュの紹介

プロンプトキャッシュとは?

プロンプトキャッシュは、言語モデルの推論がすでに処理した入力トークンを再利用できるようにし、そのコストを最大90%削減して長いコンテキストのワークロードを現実的にします。キャッシュの方針を正しく設計すれば、入力トークンで大きなコスト削減と、意味のある性能向上が得られます。

プロンプトを送ると、システムはまずその完全に同じプロンプトが以前に処理されたかを確認します。見つかった場合(キャッシュヒット)は、新たに生成せず保存済みの応答を返します。見つからない場合(キャッシュミス)は、プロンプトを通常どおり処理し、応答を今後の利用のために保存します。

注目すべき主要指標

  • 入力料金:入力トークンに対して支払う標準料金
  • キャッシュ書き込み料金:プロンプトトークンをキャッシュに保存するために支払う料金。標準の入力料金より高い場合があります
  • キャッシュヒット料金:キャッシュにヒットしたプロンプトトークンの割引料金
  • キャッシュ保存料金:キャッシュされたトークン100万件あたりの時間単価(現在はGoogleのみ)
  • キャッシュTTL:キャッシュされたトークンが利用可能な時間。数時間から数日まで
  • キャッシュ最小トークン数:キャッシュヒットを返す前に必要な、一致するトークンの最小数

プロンプトキャッシュはどのように動くのか?

Transformerベースの言語モデルにプロンプトを送ると、アテンション層が各入力トークンをキー(K)とバリュー(V)のベクトルに処理し、KVキャッシュに保存します。これらの値をメモリに保持することで、同じ入力トークンが再びモデルに送られたときに入力トークンの処理を避けられます。

つい最近まで、キャッシュによる速度とコストの利点を活かせるのは専用デプロイに限られていました。現在は、サーバーレスAPIプロバイダー(フロンティアラボを含む)が、キャッシュのコストメリットの一部を開発者に還元し始めています。

最適なユースケース

  • システム指示:多くのやり取りで必ず含める必要がある大きなシステムプロンプト
  • チャット履歴:ユーザーの新しいターンごとに付く会話コンテキスト
  • ユーザーごとのパーソナライズコンテキスト:深いパーソナライズを可能にする、大規模なユーザーメモリやプロファイル

実装上の考慮点

  • 有効化の方法はプロバイダーによって異なります。手動設定が必要な場合もあれば、自動キャッシュを提供する場合もあります
  • キャッシュヒットの割引は、標準の入力トークン料金から50~90%です。正しく設計する時間をかける価値があります
  • キャッシュは非常に長いプロンプト(5万トークン以上)の性能を向上させます