주요 내용

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
Weighted average cost (USD) per Intelligence Index task · Lower is better
새 언어 모델 평가 · 9월 13일
K2 Horizon 0.9BK2 Horizon 0.9B
새 언어 모델 평가 · 9월 13일
K2 Horizon 3.7BK2 Horizon 3.7B
새 언어 모델 평가 · 9월 13일
K2 Horizon 7BK2 Horizon 7B
새 언어 모델 평가 · 9월 13일
K2 Horizon MoVA 36B A4BK2 Horizon MoVA 36B A4B
새 언어 모델 평가 · 9월 10일
Agnes 3.0 FlashAgnes 3.0 Flash
새 언어 모델 평가 · 9월 10일
DeepSeek V4.1 Flash (Reasoning, Max Effort)DeepSeek V4.1 Flash (Reasoning, Max Effort)
새 언어 모델 평가 · 9월 10일
Ling-3.0-flash-VLLing-3.0-flash-VL
새 아티클 게시 · 9월 9일
Benchmarking GPT-6 Astra
새 아티클 게시 · 9월 7일
OpenBMB releases MiniCPM5-2B
새 아티클 게시 · 9월 7일
Announcing the Artificial Analysis Intelligence Index v4.3
방법론 업데이트 · 9월 7일
Artificial AnalysisArtificial Analysis Intelligence Index v4.3
새 언어 모델 평가 · 9월 7일
MiniCPM5-2BMiniCPM5-2B
새 아티클 게시 · 9월 4일
Announcing Artificial Analysis Intelligence Index v4.2
새 언어 모델 평가 · 9월 3일
GPT-6 Astra (low)GPT-6 Astra (low)
새 언어 모델 평가 · 9월 3일
GPT-6 Astra (medium)GPT-6 Astra (medium)
새 언어 모델 평가 · 9월 3일
GPT-6 Astra (high)GPT-6 Astra (high)
새 언어 모델 평가 · 9월 3일
GPT-6 Astra (xhigh)GPT-6 Astra (xhigh)
새 언어 모델 평가 · 9월 3일
GPT-6 Astra (max)GPT-6 Astra (max)
새 언어 모델 평가 · 9월 3일
K2 Horizon 375B A23BK2 Horizon 375B A23B
새 아티클 게시 · 9월 2일
Muse Spark 1.3: Meta reaches the frontier더 보기

지능Updated

독립적인 평가를 바탕으로 한 주요 AI 모델의 지능

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better

Intelligence Index vs. Cost per Intelligence Index Task

Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task
Most attractive quadrant
Pareto line

Intelligence Index vs. Cost per Intelligence Index Task, by Model Release

All reasoning and effort variants of each selected release · Weighted average cost (USD) per Artificial Analysis Intelligence Index task
Most attractive quadrant
Pareto line

시간에 따른 최첨단 언어 모델 지능

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

종단 간 소프트웨어 엔지니어링 작업에서 주요 코딩 에이전트의 성능, 비용, 실행 시간

Artificial Analysis Coding Agent Index

Artificial Analysis Coding Agent Index v1.5 incorporates 3 benchmarks: DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA · Higher is better

Artificial Analysis Coding Agent Index vs. 작업당 비용

Artificial Analysis Coding Agent Index vs. 작업당 평균 토큰당 종량제 API 비용 (USD)
Most attractive quadrant
Pareto line

이미지 및 동영상

95% 신뢰 구간과 함께 제공되는 Image Arena 및 Video Arena 리더보드의 상위 모델

텍스트-이미지 생성 리더보드

Image Arena의 블라인드 선호도 투표에서 산출한 Elo 점수입니다. 전체 리더보드를 확인하세요.

음성

Text to Speech Arena, 음성 텍스트 변환, 음성 대 음성 평가의 상위 모델

Provider Voice Arena Quality Elo

Arena Elo: average Elo rating of the model · Higher is better

특정 역량 및 산업에서 모델의 성능을 측정합니다.

Artificial Analysis Finance & Accounting Index

Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2, AA-Briefcase, Humanity's Last Exam, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Quantitative analysis on spreadsheets & documents

Instruction following

Agentic tool use

Long-horizon agentic tasks

Kubernetes incident root-cause analysis

Visual reasoning

AA-Briefcase

AA-Briefcase는 스프레드시트, 프레젠테이션, 메모 같은 결과물을 요구하는 실제 비즈니스 워크플로에서 에이전트를 테스트하는 장기 지식 업무용 최첨단 에이전트 평가입니다.

AA-Briefcase Elo

AA-Briefcase is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better

AA-AnalystAgent

AA-AnalystAgent는 실제 스프레드시트와 문서를 대상으로 엔드투엔드 정량 분석을 평가하는 벤치마크로, 비즈니스 분석가와 데이터 분석가가 매일 수행하는 업무를 다룹니다.

AA-AnalystAgent pass^5

Share of end-to-end quantitative analysis tasks solved on all five attempts · Higher is better

AA-Omniscience

AA-Omniscience는 정확한 답변에 보상하고 잘못된 추측에는 감점을 부여하여 여러 분야에서 사실에 근거한 신뢰할 수 있는 출력을 생성하는 모델을 종합적으로 보여 주는 지식 및 환각 벤치마크입니다.

AA-Omniscience Index

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.

GDPval-AA v2

GDPval-AA v2은 다양한 직종의 실제 경제적 가치가 있는 작업에서 AI 모델을 평가합니다.

GDPval-AA v2 Leaderboard

Elo rating for performance on real-world work tasks · Anchored to a human baseline of 1,000 · Higher is better
Human Baseline (1,000)

Artificial Analysis Openness Index는 여러 구성 요소의 가용성과 투명성을 바탕으로 모델이 얼마나 '개방적'인지 평가합니다.

Artificial Analysis Openness Index: Components

Openness Index underlying score contribution by components, up to a maximum of 18 (higher is more open)

Artificial Analysis Openness Index vs. Artificial Analysis Intelligence Index

Most attractive quadrant
Pareto line

출력 토큰

독립적인 평가를 바탕으로 한 주요 AI 모델의 출력 토큰 수

Output Tokens per Intelligence Index Task

Weighted average number of output tokens used to run one task in the Artificial Analysis Intelligence Index

비용

독립적인 평가를 바탕으로 한 주요 AI 모델의 가격 및 실제 비용

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better

Cost to Run Artificial Analysis Intelligence Index

Cost (USD) to run all evaluations in the Artificial Analysis Intelligence Index

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens)

속도 및 지연 시간

자체 API 성능 비교

Output Speed

Output tokens per second · Higher is better

Time per Intelligence Index Task

Weighted average decode time (minutes) per task; excludes TTFT and overhead time · Lower is better

제공업체

Endpoint Accuracy Index: gpt-oss-120b (high)

v1.0 · Composite of BFCL v4-500, HLE-250 and AA-LCR-25 run against each provider endpoint · Percentage of the reference endpoint, with 95% confidence interval · Higher is better

Output Speed vs. Price: gpt-oss-120b (high)

Output tokens per second · USD per 1M tokens (blended) · 10,000 input tokens
Most attractive quadrant
Pareto line

가격(캐시 적중, 입력, 출력): gpt-oss-120b (high)

Price (USD per M Tokens) · Lower is better · 10,000 input tokens

Output Speed: gpt-oss-120b (high)

Output speed: output tokens per second · 10,000 input tokens