프롬프트 캐싱: 제공업체별 비용 및 성능 분석

프롬프트 캐싱은 입력 토큰 비용을 최대 90%까지 줄일 수 있으며 긴 컨텍스트 워크로드를 현실적으로 만듭니다. 주요 AI 제공업체의 캐시 가격, 할인, API 사양을 아래에서 비교하세요.

캐싱은 프롬프트가 정확히 일치해야 하며 제공업체마다 다릅니다. OpenAI와 DeepSeek처럼 자동 캐싱을 제공하는 곳도 있고, Google, Anthropic, Amazon처럼 수동 설정이 필요한 곳도 있습니다. 작동 방식은 아래 프롬프트 캐싱 소개에서 자세히 알아보세요.

가격

가격: 캐시 적중, 캐시 쓰기, 입력 및 출력

Price (USD per M Tokens)
Reasoning models are indicated by a lightbulb icon

Price per token for cached prompts (previously processed), typically offering a significant discount compared to regular input price, represented as USD per million tokens. The values shown here are the cache hit price; cache write and cache storage are billed separately and vary by provider — see "Cache pricing by provider" for detail.

캐시 할인

Pricing: Cache Discount

1 - (cache hit price / input price) · Higher is better
Reasoning models are indicated by a lightbulb icon

Reduction in input token cost due to cache hit relative to input price. Formula: 1 - (Cache Hit Price per Token / Input Token Price), where cache hit price is the first-party cache hit price or the median provider cache hit price. Note that this discount figure does not account for all costs associated with cache hits, such as cache write and storage costs.

프롬프트 캐싱 API 사양

제공업체
모델
입력(표준)
캐시 쓰기
캐시 적중
캐시 저장
출력(표준)
자동 활성화
최소 토큰
캐시 TTL
참고
Amazon BedrockAmazon Bedrock
  • Amazon supports caching for Nova models and Anthropic's Claude models.
  • Pricing and usage differs between model families.
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
-
$25.00
-
-
GPT-5.5 (xhigh)
$5.50
-
$0.55
-
$33.00
-
-
AnthropicAnthropic
  • Cache read tokens are 90% cheaper than base input tokens
  • Cache write tokens are 25% more expensive than base input tokens
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
$10.00
$12.50
$1.00
$20.00
$50.00
-
-
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
$10.00
$25.00
-
-

1h cache write: $10

GoogleGoogle
  • Google supports caching for Gemini models and Anthropic's Claude models.
  • Pricing and usage differs between model families.
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
$10.00
$12.50
$1.00
-
$50.00
-
-
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
$5.00
$6.25
$0.50
-
$25.00
-
-
OpenAIOpenAI
  • Cache read tokens are 50% cheaper than base input tokens
  • Cache persists up to one hour during off-peak periods
GPT-5.5 (xhigh)
$5.00
-
$0.50
-
$30.00
1024
5-10 minutes
GPT-5.4 mini (xhigh)
$0.75
-
$0.07
-
$4.50
-
-
SpaceXAISpaceXAI
  • Prompt caching is not 100% guaranteed.
Grok 4.3 (high)
$1.25
$1.25
$0.20
-
$2.50
-
-

For requests greater than 200k tokens, pricing is $2.50 per 1M input tokens, $0.40 per 1M cached input tokens, and $5.00 per 1M output tokens

Kimi K2.6
$0.95
-
$0.16
-
$4.00
-
-
FireworksFireworks
  • Prompt caching is enabled by default.
  • The default discount is 50%, but the exact discount varies by model.
Kimi K2.6
$0.95
-
$0.16
-
$4.00
-
-
Kimi K2.6
$0.95
-
$0.16
-
$4.00
-
-
Kimi K2.6
$0.75
-
$0.16
-
$3.50
-
-
Microsoft AzureMicrosoft Azure

OpenAI models:

  • Cache read tokens are 50% cheaper than base input tokens (Standard deployments)
  • Cache persists up to one hour during off-peak periods
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)
$2.00
$2.50
$0.20
-
$10.00
-
-
GPT-5.4 mini (xhigh)
$0.75
-
$0.07
-
$4.50
-
-
DeepSeekDeepSeek
  • Cache read tokens are 50% cheaper on average (up to 90% with cache optimization)
  • Implements Context Caching on Disk technology
  • No guarantee of 100% cache hits
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
$0.43
-
$0.00
-
$0.87
-
-
Kimi K2.6
$0.65
-
$0.15
-
$3.41
-
-
CrusoeCrusoe
Kimi K2.6
$0.70
-
$0.35
-
$3.50
-
-
DeepInfraDeepInfra
  • Prompt caching is automatic — no extra parameters required.
Kimi K2.6
$0.75
-
$0.15
-
$3.50
-
-
GMIGMI
Kimi K2.6
$0.85
-
$0.14
-
$3.60
-
-
Kimi K2.6
$0.80
-
$0.16
-
$3.40
-
-
Kimi K2.6
$0.77
-
$0.14
-
$3.40
-
-
CloudflareCloudflare
  • Prefix caching is enabled by default. To maximize cache hit rates, a header must be sent.
Kimi K2.6
$0.95
-
$0.16
-
$4.00
-
-

프롬프트 캐싱 소개

프롬프트 캐싱이란?

프롬프트 캐싱은 언어 모델 추론이 이미 처리한 입력 토큰을 재사용하게 하여 비용을 최대 90%까지 낮추고 긴 컨텍스트 워크로드를 가능하게 합니다. 캐싱 방식을 잘 잡으면 입력 토큰에서 큰 비용 절감과 의미 있는 성능 이점을 얻을 수 있습니다.

프롬프트를 보내면 시스템은 먼저 그 정확한 프롬프트가 이전에 처리되었는지 확인합니다. 찾으면(캐시 적중) 새로 생성하지 않고 저장된 응답을 반환합니다. 찾지 못하면(캐시 미스) 프롬프트를 정상적으로 처리하고 응답을 이후 사용을 위해 저장합니다.

살펴볼 핵심 지표

  • 입력 가격: 입력 토큰에 대해 지불하는 표준 가격
  • 캐시 쓰기 가격: 프롬프트 토큰을 캐시에 저장하는 데 지불하는 비용. 표준 입력 가격보다 높을 수 있음
  • 캐시 적중 가격: 캐시에 적중한 프롬프트 토큰의 할인 요금
  • 캐시 저장 가격: 캐시된 토큰 100만 개당 시간당 비용(현재 Google만 해당)
  • 캐시 TTL: 캐시된 토큰이 사용 가능한 시간. 수 시간에서 수일까지
  • 캐시 최소 토큰 수: 캐시 적중을 제공하기 전에 필요한 최소 일치 토큰 수

프롬프트 캐싱은 어떻게 작동하나요?

트랜스포머 기반 언어 모델에 프롬프트를 보내면 어텐션 레이어가 각 입력 토큰을 키(K)와 값(V) 벡터로 처리해 KV 캐시에 저장합니다. 이 값을 메모리에 유지하면 동일한 입력 토큰이 다시 들어올 때 입력 토큰 처리를 건너뛸 수 있습니다.

얼마 전까지만 해도 캐싱의 속도와 비용 이점을 활용하려면 전용 배포가 필요했습니다. 이제 서버리스 API 제공업체(프론티어 랩 포함)가 캐싱의 비용 이점 일부를 개발자에게 넘기기 시작했습니다.

최적의 사용 사례

  • 시스템 지시: 여러 상호작용에 반드시 포함해야 하는 긴 시스템 프롬프트
  • 채팅 기록: 사용자의 새 턴마다 함께 가는 대화 컨텍스트
  • 사용자별 개인화 컨텍스트: 깊은 개인화를 가능하게 하는 광범위한 사용자 메모리나 프로필

구현 시 고려 사항

  • 활성화 방식은 제공업체마다 다릅니다. 수동 설정이 필요한 곳도 있고 자동 캐싱을 제공하는 곳도 있습니다
  • 캐시 적중 할인은 표준 입력 토큰 가격 대비 50–90%입니다. 제대로 맞추는 데 시간을 들일 가치가 있습니다
  • 캐싱은 매우 긴 프롬프트(5만+ 토큰)의 성능을 개선합니다