プロンプトキャッシュ:プロバイダー横断のコストと性能分析

プロンプトキャッシュは入力トークンのコストを最大90%削減でき、長いコンテキストのワークロードを現実的にします。主要AIプロバイダーのキャッシュ料金、割引、API仕様を以下で比較できます。

キャッシュにはプロンプトの完全一致が必要で、プロバイダーによって異なります。OpenAIやDeepSeekのように自動キャッシュを提供する一方、Google、Anthropic、Amazonなどは手動設定が必要です。仕組みの詳細は、以下のプロンプトキャッシュの紹介をご覧ください。

料金

料金:キャッシュヒット、キャッシュ書き込み、入力、出力

Price (USD per M Tokens)
Reasoning models are indicated by a lightbulb icon

Price per token for cached prompts (previously processed), typically offering a significant discount compared to regular input price, represented as USD per million tokens. The values shown here are the cache hit price; cache write and cache storage are billed separately and vary by provider — see "Cache pricing by provider" for detail.

キャッシュ割引

Pricing: Cache Discount

1 - (cache hit price / input price) · Higher is better
Reasoning models are indicated by a lightbulb icon

Reduction in input token cost due to cache hit relative to input price. Formula: 1 - (Cache Hit Price per Token / Input Token Price), where cache hit price is the first-party cache hit price or the median provider cache hit price. Note that this discount figure does not account for all costs associated with cache hits, such as cache write and storage costs.

プロンプトキャッシュ API仕様

プロバイダー
モデル
入力(標準)
キャッシュ書き込み
キャッシュヒット
キャッシュ保存
出力(標準)
自動有効化
最小トークン数
キャッシュTTL
注記
Amazon BedrockAmazon Bedrock
  • Amazon supports caching for Nova models and Anthropic's Claude models.
  • Pricing and usage differs between model families.
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
-
$25.00
-
-
GPT-5.5 (xhigh)
$5.50
-
$0.55
-
$33.00
-
-
AnthropicAnthropic
  • Cache read tokens are 90% cheaper than base input tokens
  • Cache write tokens are 25% more expensive than base input tokens
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
$10.00
$12.50
$1.00
$20.00
$50.00
-
-
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
$10.00
$25.00
-
-

1h cache write: $10

GoogleGoogle
  • Google supports caching for Gemini models and Anthropic's Claude models.
  • Pricing and usage differs between model families.
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
$10.00
$12.50
$1.00
-
$50.00
-
-
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
-
$25.00
-
-
OpenAIOpenAI
  • Cache read tokens are 50% cheaper than base input tokens
  • Cache persists up to one hour during off-peak periods
GPT-5.5 (xhigh)
$5.00
-
$0.50
-
$30.00
1024
5-10 minutes
GPT-5.4 mini (xhigh)
$0.75
-
$0.07
-
$4.50
-
-
SpaceXAISpaceXAI
  • Prompt caching is not 100% guaranteed.
Grok 4.3 (high)
$1.25
$1.25
$0.20
-
$2.50
-
-

For requests greater than 200k tokens, pricing is $2.50 per 1M input tokens, $0.40 per 1M cached input tokens, and $5.00 per 1M output tokens

Kimi K2.6
$0.95
-
$0.16
-
$4.00
-
-
FireworksFireworks
  • Prompt caching is enabled by default.
  • The default discount is 50%, but the exact discount varies by model.
Kimi K2.6
$0.95
-
$0.16
-
$4.00
-
-
Kimi K2.6
$0.95
-
$0.16
-
$4.00
-
-
Kimi K2.6
$0.75
-
$0.16
-
$3.50
-
-
Microsoft AzureMicrosoft Azure

OpenAI models:

  • Cache read tokens are 50% cheaper than base input tokens (Standard deployments)
  • Cache persists up to one hour during off-peak periods
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)
$2.00
$2.50
$0.20
-
$10.00
-
-
GPT-5.4 mini (xhigh)
$0.75
-
$0.07
-
$4.50
-
-
DeepSeekDeepSeek
  • Cache read tokens are 50% cheaper on average (up to 90% with cache optimization)
  • Implements Context Caching on Disk technology
  • No guarantee of 100% cache hits
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$0.43
-
$0.00
-
$0.87
-
-
Kimi K2.6
$0.65
-
$0.15
-
$3.41
-
-
CrusoeCrusoe
Kimi K2.6
$0.70
-
$0.35
-
$3.50
-
-
DeepInfraDeepInfra
  • Prompt caching is automatic — no extra parameters required.
Kimi K2.6
$0.75
-
$0.15
-
$3.50
-
-
GMIGMI
Kimi K2.6
$0.85
-
$0.14
-
$3.60
-
-
Kimi K2.6
$0.80
-
$0.16
-
$3.40
-
-
Kimi K2.6
$0.77
-
$0.14
-
$3.40
-
-
CloudflareCloudflare
  • Prefix caching is enabled by default. To maximize cache hit rates, a header must be sent.
Kimi K2.6
$0.95
-
$0.16
-
$4.00
-
-

プロンプトキャッシュの紹介

プロンプトキャッシュとは?

プロンプトキャッシュは、言語モデルの推論がすでに処理した入力トークンを再利用できるようにし、そのコストを最大90%削減して長いコンテキストのワークロードを現実的にします。キャッシュの方針を正しく設計すれば、入力トークンで大きなコスト削減と、意味のある性能向上が得られます。

プロンプトを送ると、システムはまずその完全に同じプロンプトが以前に処理されたかを確認します。見つかった場合(キャッシュヒット)は、新たに生成せず保存済みの応答を返します。見つからない場合(キャッシュミス)は、プロンプトを通常どおり処理し、応答を今後の利用のために保存します。

注目すべき主要指標

  • 入力料金:入力トークンに対して支払う標準料金
  • キャッシュ書き込み料金:プロンプトトークンをキャッシュに保存するために支払う料金。標準の入力料金より高い場合があります
  • キャッシュヒット料金:キャッシュにヒットしたプロンプトトークンの割引料金
  • キャッシュ保存料金:キャッシュされたトークン100万件あたりの時間単価(現在はGoogleのみ)
  • キャッシュTTL:キャッシュされたトークンが利用可能な時間。数時間から数日まで
  • キャッシュ最小トークン数:キャッシュヒットを返す前に必要な、一致するトークンの最小数

プロンプトキャッシュはどのように動くのか?

Transformerベースの言語モデルにプロンプトを送ると、アテンション層が各入力トークンをキー(K)とバリュー(V)のベクトルに処理し、KVキャッシュに保存します。これらの値をメモリに保持することで、同じ入力トークンが再びモデルに送られたときに入力トークンの処理を避けられます。

つい最近まで、キャッシュによる速度とコストの利点を活かせるのは専用デプロイに限られていました。現在は、サーバーレスAPIプロバイダー(フロンティアラボを含む)が、キャッシュのコストメリットの一部を開発者に還元し始めています。

最適なユースケース

  • システム指示:多くのやり取りで必ず含める必要がある大きなシステムプロンプト
  • チャット履歴:ユーザーの新しいターンごとに付く会話コンテキスト
  • ユーザーごとのパーソナライズコンテキスト:深いパーソナライズを可能にする、大規模なユーザーメモリやプロファイル

実装上の考慮点

  • 有効化の方法はプロバイダーによって異なります。手動設定が必要な場合もあれば、自動キャッシュを提供する場合もあります
  • キャッシュヒットの割引は、標準の入力トークン料金から50~90%です。正しく設計する時間をかける価値があります
  • キャッシュは非常に長いプロンプト(5万トークン以上)の性能を向上させます