모델 비교: 지능, 성능 및 가격 분석
Microevals Playground지능
출력 속도(토큰/초)
지연 시간(초)
가격(토큰 100만 개당 $)
컨텍스트 창
주요 내용
지능
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index by Open Weights / Proprietary
Intelligence Evaluations
Agentic real-world work tasks, (Elo-500)/2000
Agentic tool use
Agentic coding & terminal use
Coding
Reasoning & knowledge
Scientific reasoning
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Agentic knowledge work, Elo
Agentic SaaS workflows
Legal agentic work, criterion pass rate
Agentic business operations
Instruction following
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
AA-Briefcase
AA-Briefcase Elo
AA-Omniscience
AA-Omniscience Index
Openness Index
Artificial Analysis Openness Index: Score
Intelligence Index 비교
Intelligence Index vs. Cost per Intelligence Index Task
토큰 사용량
Output Tokens per Intelligence Index Task
비용
Cost per Intelligence Index Task
Cost to Run Artificial Analysis Intelligence Index
Pricing: Cache Hit, Input, and Output
컨텍스트 창
Context Window
속도
출력 속도(초당 토큰 수)로 측정
Output Speed
Time per Intelligence Index Task
지연 시간
첫 토큰까지 걸린 시간(초)으로 측정
Latency: Time To First Answer Token
종단 간 응답 시간
Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed
End-to-End Response Time
모델 크기(오픈 웨이트 모델만 해당)
Model Size: Total and Active Parameters
자주 묻는 질문
Claude Opus 5 (Adaptive Reasoning, Max Effort)은 Artificial Analysis Intelligence Index에서 61점으로, 평가된 모델 175개 중 선두입니다.
Intelligence Index 기준 상위 AI 모델: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort)(61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)(60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)(60), 4. GPT-5.6 Sol (max)(59) 및 5. Claude Opus 5 (Adaptive Reasoning, High Effort)(59).
Celeris-1이 초당 2,033.7토큰으로 가장 빠르며, Mercury 2(741.8 t/s)과 LFM2.5-VL-1.6B(474.5 t/s)이 뒤를 잇습니다.
Nova Micro이 혼합 가격 기준 토큰 100만 개당 $0.03로 가장 저렴하며, Sarvam 30B (high)($0.03)과 Gemma 4 E4B (Non-reasoning)($0.03)이 뒤를 잇습니다.
Gemini 2.5 Flash-Lite (Non-reasoning)의 첫 토큰까지 걸린 시간이 0.33초로 가장 짧으며, Command A+(0.44초)과 Gemini 2.5 Flash (Non-reasoning)(0.47초)이 뒤를 잇습니다.
Kimi K3 (max)이 Intelligence Index 점수 57점으로 가장 높은 순위의 오픈 웨이트 모델입니다. 평가한 전체 모델 175개 중 오픈 웨이트 모델은 99개입니다.
Intelligence Index 기준 상위 오픈 웨이트 AI 모델: 1. Kimi K3 (max)(57), 2. GLM-5.2 (max)(51) 및 3. DeepSeek V4 Flash 0731 (Reasoning, Max Effort)(50).
Claude Opus 5 (Adaptive Reasoning, Max Effort)이 Intelligence Index 점수 61점으로 추론 모델 130개 중 선두입니다. 추론 모델은 답변하기 전에 확장 사고를 사용해 복잡한 문제를 해결합니다.
지능(품질), 가격, 출력 속도(초당 토큰 수), 지연 시간(첫 토큰까지 걸린 시간), 종단 간 응답 시간, 컨텍스트 창 크기 등 여러 차원에서 모델을 비교합니다. 표준화된 프롬프트를 사용해 모델 591개의 성능 지표를 직접 측정합니다.
차트에서 모델 이름이나 행을 클릭하면 상세 지표와 비슷한 모델과의 직접 비교가 있는 전용 페이지를 볼 수 있습니다. 모델 선택기를 사용해 각 차트에 표시할 모델을 직접 정할 수도 있습니다. 리더보드 보기