Prompt-Caching: Kosten- und Leistungsanalyse über Anbieter hinweg

Prompt-Caching kann die Kosten für Eingabetokens um bis zu 90 % senken und macht Workloads mit langem Kontext wirtschaftlich. Vergleichen Sie Cache-Preise, Rabatte und API-Spezifikationen der großen KI-Anbieter.

Caching erfordert exakte Prompt-Übereinstimmungen und unterscheidet sich je nach Anbieter – einige wie OpenAI und DeepSeek bieten automatisches Caching, andere einschließlich Google, Anthropic und Amazon erfordern eine manuelle Einrichtung. Mehr dazu in unserer Einführung in Prompt-Caching weiter unten.

Preise

Preise: Cache-Treffer, Cache-Schreiben, Eingabe und Ausgabe

Price (USD per M Tokens)

Price per token for cached prompts (previously processed), typically offering a significant discount compared to regular input price, represented as USD per million tokens. The values shown here are the cache hit price; cache write and cache storage are billed separately and vary by provider — see "Cache pricing by provider" for detail.

Cache-Rabatt

Pricing: Cache Discount

1 - (cache hit price / input price) · Higher is better

Reduction in input token cost due to cache hit relative to input price. Formula: 1 - (Cache Hit Price per Token / Input Token Price), where cache hit price is the first-party cache hit price or the median provider cache hit price. Note that this discount figure does not account for all costs associated with cache hits, such as cache write and storage costs.

API-Spezifikationen für Prompt-Caching

Anbieter
Modell
Eingabe (Standard)
Cache-Schreiben
Cache-Treffer
Cache-Speicher
Ausgabe (Standard)
Automatisch aktiviert
Min. Tokens
Cache-TTL
Hinweise
AnthropicAnthropic
  • Cache read tokens are 90% cheaper than base input tokens
  • Cache write tokens are 25% more expensive than base input tokens
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
$10.00
$12.50
$0.25
$20.00
$50.00
-
-
Claude Opus 5 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
$10.00
$25.00
-
-

1h cache write: $10

Amazon BedrockAmazon Bedrock
  • Amazon supports caching for Nova models and Anthropic's Claude models.
  • Pricing and usage differs between model families.
Claude Opus 5 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
-
$25.00
-
-
GPT-5.6 Sol (max)
$5.50
$6.88
$0.55
-
$33.00
-
-
GoogleGoogle
  • Google supports caching for Gemini models and Anthropic's Claude models.
  • Pricing and usage differs between model families.
Claude Opus 5 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
-
$25.00
-
-
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
$10.00
$12.50
$1.00
-
$50.00
-
-
OpenAIOpenAI
  • Cache read tokens are 50% cheaper than base input tokens
  • Cache persists up to one hour during off-peak periods
GPT-5.6 Sol (max)
$4.00
$5.00
$0.40
-
$20.00
-
-
GPT-5.6 Terra (max)
$2.00
$2.50
$0.20
-
$12.00
-
-
SpaceXAISpaceXAI
  • Prompt caching is not 100% guaranteed.
Grok 4.6 (high)
$2.00
-
$0.50
-
$6.00
-
-
MetaMeta
Muse Spark 1.3 (xhigh)
$1.25
-
$0.15
-
$4.25
-
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
Bitdeer AIBitdeer AI
Kimi K3 (max)
$2.66
-
$0.28
-
$13.30
-
-
Qwen3.8 2.4T A95B
$1.90
-
$0.19
-
$5.70
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
DigitalOceanDigitalOcean
Kimi K3 (max)
$2.85
-
$0.28
-
$14.25
-
-
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
FireworksFireworks
  • Prompt caching is enabled by default.
  • The default discount is 50%, but the exact discount varies by model.
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
Kimi K3 (max)
$4.50
-
$0.45
-
$22.50
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
MakoraMakora
Kimi K3 (max)
$2.55
-
$0.26
-
$12.75
-
-
GLM-5.3 (max)
$1.35
-
$0.23
-
$4.40
-
-
ModalModal
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
-
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
Kimi K3 (max)
$3.00
-
$0.30
-
$15.00
-
-
Qwen3.8 2.4T A95B
$2.50
-
$0.50
-
$6.25
-
-
DeepInfraDeepInfra
  • Prompt caching is automatic — no extra parameters required.
GLM-5.3 (max)
$1.20
-
$0.12
-
$4.00
-
-
Qwen3.8 2.4T A95B
$2.00
-
$0.20
-
$6.00
-
-
ModularModular
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
MiniMax-M3
$0.30
-
$0.06
-
$1.20
-
-
ZaiZai
GLM-5.3 (max)
$1.40
-
$0.26
-
$4.40
-
-
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
Alibaba CloudAlibaba Cloud
  • Alibaba offers two cache types: implicit and explicit.
  • Implicit cache is automatically enabled. It is billed at 20% of the standard input token price.
  • Explicit cache must be activated. It creates a cache for specific content to ensure a deterministic hit within its 5-minute validity period. Tokens used to create the cache are billed at 125% of the standard input token price, while subsequent cache hits are billed at 10% of that price.
Qwen3.8 2.4T A95B
$2.00
-
$0.25
-
$6.00
-
-
Qwen3.8 27B (xhigh)
$0.50
-
$0.05
-
$3.00
-
-
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
GMIGMI
GLM-5.3-Flash
$0.07
-
$0.01
-
$0.25
-
-
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$1.12
-
$0.04
-
$3.37
-
-
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$1.32
-
$0.13
-
$3.96
-
-
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$1.50
-
$0.14
-
$3.13
-
-
WaferWafer
GLM-5.3-Flash
$0.15
-
$0.03
-
$0.50
-
-
DeepSeekDeepSeek
  • Cache read tokens are 50% cheaper on average (up to 90% with cache optimization)
  • Implements Context Caching on Disk technology
  • No guarantee of 100% cache hits
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$1.32
-
$0.04
-
$3.96
-
-
Qwen3.8 27B (xhigh)
$0.40
-
$0.15
-
$3.00
-
-
MiniMax-M3
$0.23
-
$0.05
-
$0.96
-
-
MistralMistral
Mistral Medium 3.5
$1.50
-
$0.15
-
$7.50
-
-
MiniMaxMiniMax
  • MiniMax supports Anthropic API compatible caching that is managed through explicit cache_control settings.
MiniMax-M3
$0.30
$0.38
$0.06
-
$1.20
-
-
Thinking MachinesThinking Machines
Inkling (xhigh)
$1.00
-
$0.17
-
$4.05
-
-
-
GroqGroq
  • The minimum cacheable prompt length varies by model, ranging from 128 to 1024 tokens depending on the specific model used.
  • All cached data automatically expires after 2 hours without use.
gpt-oss-120b (high)
$0.15
-
$0.07
-
$0.60
-
2 hrs

Einführung in Prompt-Caching

Was ist Prompt-Caching?

Prompt-Caching ermöglicht es der Inferenz von Sprachmodellen, bereits verarbeitete Eingabetokens wiederzuverwenden, ihre Kosten um bis zu 90 % zu senken und Workloads mit langem Kontext wirtschaftlich zu machen. Wer Caching richtig angeht, kann bei Eingabetokens stark sparen und spürbare Leistungsvorteile erzielen.

Wenn Sie einen Prompt senden, prüft das System zuerst, ob genau dieser Prompt bereits verarbeitet wurde. Bei einem Treffer (Cache-Hit) wird die gespeicherte Antwort zurückgegeben, statt eine neue zu erzeugen. Wird nichts gefunden (Cache-Miss), wird der Prompt normal verarbeitet und die Antwort für die spätere Nutzung gespeichert.

Wichtige Kennzahlen

  • Eingabepreis: Der Standardpreis, den Sie für Eingabetokens zahlen
  • Cache-Schreibpreis: Was Sie zahlen, um Prompt-Tokens im Cache zu speichern; manchmal höher als der Standard-Eingabepreis
  • Preis bei Cache-Treffer: Reduzierter Satz für Prompt-Tokens, die den Cache treffen
  • Cache-Speicherpreis: Stündliche Kosten pro Million gecachter Tokens (derzeit nur bei Google)
  • Cache-TTL: Die Zeit, für die gecachte Tokens verfügbar bleiben, von Stunden bis Tagen
  • Mindestanzahl Cache-Tokens: Mindestzahl übereinstimmender Tokens, bevor ein Cache-Treffer ausgeliefert wird

Wie funktioniert Prompt-Caching?

Wenn Sie einen Prompt an ein transformer-basiertes Sprachmodell senden, verarbeiten die Attention-Schichten jedes Eingabetoken zu Key- (K) und Value-Vektoren (V), die im KV-Cache gespeichert werden. Bleiben diese Werte im Speicher, kann die Verarbeitung der Eingabetokens entfallen, wenn erneut identische Eingabetokens an das Modell gesendet werden.

Bis vor kurzem waren die Geschwindigkeits- und Kostenvorteile des Cachings nur in dedizierten Deployments nutzbar. Inzwischen geben serverless-API-Anbieter – einschließlich der Frontier-Labs – einen Teil der Kostenvorteile des Cachings an Entwickler weiter.

Optimale Anwendungsfälle

  • Systemanweisungen: Große Systemprompts, die in vielen Interaktionen enthalten sein müssen
  • Chatverlauf: Gesprächskontext, der jede neue Nutzeräußerung begleitet
  • Personalisierter Kontext pro Nutzer: Umfangreiche Nutzermemories oder Profile für tiefe Personalisierung

Hinweise zur Umsetzung

  • Die Aktivierung unterscheidet sich je nach Anbieter – manche erfordern eine manuelle Einrichtung, andere bieten automatisches Caching
  • Rabatte bei Cache-Treffern liegen bei 50–90 % unter dem Standardpreis für Eingabetokens – es lohnt sich, das richtig umzusetzen
  • Caching verbessert die Leistung bei sehr langen Prompts (50k+ Tokens)