독립적인 AI 분석
AI 환경을 이해하고 사용 사례에 가장 적합한 모델과 제공업체를 선택하세요.
주요 내용
지능Updated
독립적인 평가를 바탕으로 한 주요 AI 모델의 지능
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index by Open Weights / Proprietary
Cost per Intelligence Index Task
Intelligence Index vs. Cost per Intelligence Index Task
Intelligence Index vs. Cost per Intelligence Index Task, by Model Release
시간에 따른 최첨단 언어 모델 지능
Coding Agent IndexUpdated
종단 간 소프트웨어 엔지니어링 작업에서 주요 코딩 에이전트의 성능, 비용, 실행 시간
Artificial Analysis Coding Agent Index
Artificial Analysis Coding Agent Index vs. 작업당 비용
이미지 및 동영상
95% 신뢰 구간과 함께 제공되는 Image Arena 및 Video Arena 리더보드의 상위 모델
텍스트-이미지 생성 리더보드
음성
Text to Speech Arena, 음성 텍스트 변환, 음성 대 음성 평가의 상위 모델
Provider Voice Arena Quality Elo
Artificial Analysis Finance & Accounting Index
Intelligence Evaluations
Agentic knowledge work, (Elo-500)/2000
Agentic real-world work tasks, (Elo-500)/2000
Agentic SaaS workflows
Agentic coding & terminal use
Coding
Reasoning & knowledge
Professional document reasoning, All-pass
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Quantitative analysis on spreadsheets & documents
Instruction following
Agentic tool use
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
AA-Briefcase
AA-Briefcase는 스프레드시트, 프레젠테이션, 메모 같은 결과물을 요구하는 실제 비즈니스 워크플로에서 에이전트를 테스트하는 장기 지식 업무용 최첨단 에이전트 평가입니다.
AA-Briefcase Elo
AA-AnalystAgent
AA-AnalystAgent는 실제 스프레드시트와 문서를 대상으로 엔드투엔드 정량 분석을 평가하는 벤치마크로, 비즈니스 분석가와 데이터 분석가가 매일 수행하는 업무를 다룹니다.
AA-AnalystAgent pass^5
AA-Omniscience
AA-Omniscience는 정확한 답변에 보상하고 잘못된 추측에는 감점을 부여하여 여러 분야에서 사실에 근거한 신뢰할 수 있는 출력을 생성하는 모델을 종합적으로 보여 주는 지식 및 환각 벤치마크입니다.
AA-Omniscience Index
GDPval-AA v2
GDPval-AA v2은 다양한 직종의 실제 경제적 가치가 있는 작업에서 AI 모델을 평가합니다.
GDPval-AA v2 Leaderboard
Artificial Analysis Openness Index는 여러 구성 요소의 가용성과 투명성을 바탕으로 모델이 얼마나 '개방적'인지 평가합니다.
Artificial Analysis Openness Index: Components
Artificial Analysis Openness Index vs. Artificial Analysis Intelligence Index
출력 토큰
독립적인 평가를 바탕으로 한 주요 AI 모델의 출력 토큰 수
Output Tokens per Intelligence Index Task
비용
독립적인 평가를 바탕으로 한 주요 AI 모델의 가격 및 실제 비용
Cost per Intelligence Index Task
Cost to Run Artificial Analysis Intelligence Index
Pricing: Cache Hit, Input, and Output
속도 및 지연 시간
자체 API 성능 비교