독립적인 AI 분석
AI 환경을 이해하고 사용 사례에 가장 적합한 모델과 제공업체를 선택하세요.
주요 내용
지능
독립적인 평가를 바탕으로 한 주요 AI 모델의 지능
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index by Open Weights / Proprietary
Cost per Intelligence Index Task
Intelligence Index vs. Cost per Intelligence Index Task
시간에 따른 최첨단 언어 모델 지능
종단 간 소프트웨어 엔지니어링 작업에서 주요 코딩 에이전트의 성능, 비용, 실행 시간
Artificial Analysis Coding Agent Index
이미지 및 동영상
95% 신뢰 구간과 함께 제공되는 Image Arena 및 Video Arena 리더보드의 상위 모델
텍스트-이미지 생성 리더보드
음성
Text to Speech Arena, 음성 텍스트 변환, 음성 대 음성 평가의 상위 모델
Text to Speech Arena Leaderboard
Artificial Analysis Agentic Index
Intelligence Evaluations
Agentic real-world work tasks, (Elo-500)/2000
Agentic tool use
Agentic coding & terminal use
Coding
Reasoning & knowledge
Scientific reasoning
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Agentic knowledge work, Elo
Agentic SaaS workflows
Legal agentic work, criterion pass rate
Agentic business operations
Instruction following
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
AA-Briefcase
AA-Briefcase는 스프레드시트, 프레젠테이션, 메모 같은 결과물을 요구하는 실제 비즈니스 워크플로에서 에이전트를 테스트하는 장기 지식 업무용 최첨단 에이전트 평가입니다.
AA-Briefcase Elo
AA-Omniscience
AA-Omniscience는 정확한 답변에 보상하고 잘못된 추측에는 감점을 부여하여 여러 분야에서 사실에 근거한 신뢰할 수 있는 출력을 생성하는 모델을 종합적으로 보여 주는 지식 및 환각 벤치마크입니다.
AA-Omniscience Index
GDPval-AA v2
GDPval-AA v2은 다양한 직종의 실제 경제적 가치가 있는 작업에서 AI 모델을 평가합니다.
GDPval-AA v2 Leaderboard
Artificial Analysis Openness Index는 여러 구성 요소의 가용성과 투명성을 바탕으로 모델이 얼마나 '개방적'인지 평가합니다.
Artificial Analysis Openness Index: Components
Artificial Analysis Openness Index vs. Artificial Analysis Intelligence Index
출력 토큰
독립적인 평가를 바탕으로 한 주요 AI 모델의 출력 토큰 수
Output Tokens per Intelligence Index Task
비용
독립적인 평가를 바탕으로 한 주요 AI 모델의 가격 및 실제 비용
Cost per Intelligence Index Task
Cost to Run Artificial Analysis Intelligence Index
Pricing: Cache Hit, Input, and Output
속도 및 지연 시간
자체 API 성능 비교