Artificial Analysis 역량 지수 방법론
개요
Artificial Analysis 역량 지수는 코딩 같은 광범위한 역량부터 법률 업무 같은 전문 분야까지, 특정 사용 사례에서 모델이 얼마나 뛰어난 성능을 보이는지 측정합니다.
지수를 구성하는 모든 벤치마크는 Artificial Analysis에서 독립적으로 실행한 후 그 점수를 결합합니다. 모든 평가 지표와 마찬가지로 역량 지수에는 한계가 있으며 모든 사용 사례에 직접 적용되지는 않을 수 있지만, 특정 분야에서 중요한 업무를 기준으로 모델을 비교하는 데 유용한 종합 정보를 제공합니다.
각 기반 벤치마크의 실행 및 채점 방법은 지능 벤치마킹 방법론에 설명되어 있습니다.
지수 구성 분석
지수는 스킬 기반 또는 산업 기반으로 나뉩니다. 스킬 지수는 코딩이나 에이전트형 업무처럼 여러 분야에 적용되는 역량을 측정하며, 구성 벤치마크의 동일 가중 평균으로 계산합니다. 산업 지수는 법률, 의료 및 보건, 금융 및 회계 같은 하나의 직업 또는 분야를 대상으로 하며, 해당 분야의 실제 업무에서 각 역량이 나타나는 빈도에 따라 일련의 역량에 가중치를 부여합니다.
업무 가중치는 O*NET 방식의 업무 활동 분류 체계를 바탕으로 합니다. 모든 평가 지표와 마찬가지로 이러한 지수에는 한계가 있으며 모든 사용 사례에 직접 대응하지 않을 수 있습니다. 아래 표는 각 지수의 구성 요소와 가중치를 보여 줍니다.
| 지수 | 유형 | 가중치 | 역량 | 평가 | 설명 |
|---|---|---|---|---|---|
| Agentic Index | Skill | 50% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step execution of real-world knowledge work |
| 50% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using customer-support workflows over a knowledge base | ||
| Finance & Accounting Index | Industry | 30% | Business Knowledge | AA-Omniscience | Domain recall in accounting, corporate finance, economics, and investments |
| 30% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step task execution, such as running spreadsheets, querying ERPs, or coordinating a workflow | ||
| 20% | Reasoning | HLE | Multi-step quantitative and analytic reasoning, used for sensitivity analysis, valuation, and structured problem solving | ||
| 10% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using customer workflows over an unstructured knowledge base | ||
| 5% | Long-Context | LCR | Reading and reasoning across long financial filings, deal documents, and research notes | ||
| 5% | Non-Hallucination | AA-Omniscience | Avoiding fabricated figures or citations | ||
| Strategy & Ops Index | Industry | 30% | Business Knowledge | AA-Omniscience | Working knowledge of business processes, accounting basics, and operational concepts |
| 35% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and orchestrating multi-step office workflows end-to-end | ||
| 30% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn dialogue and tool use to resolve customer and stakeholder requests | ||
| 5% | Long-Context | LCR | Holding context across long threads, policies, and records | ||
| Legal Index | Industry | 35% | Legal Knowledge | AA-Omniscience | Recall of statutes, doctrines, and procedure across jurisdictions |
| 25% | Agentic Knowledge Work | GDPval-AA v2 | Running matter-management workflows, drafting pipelines, and tool-augmented research | ||
| 10% | Long-Context | LCR | Reading and synthesizing across contracts, discovery productions, and case-law packets | ||
| 10% | Non-Hallucination | AA-Omniscience | Avoiding fabricated case cites or invented statutes | ||
| 5% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using client intake and support workflows | ||
| 15% | Reasoning | HLE | Multi-step argumentation, statutory interpretation, and weighing conflicting authorities | ||
| Healthcare & Medical Index | Industry | 35% | Medical & Health Knowledge | AA-Omniscience | Clinical knowledge across diagnosis, pharmacology, and care pathways |
| 25% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and orchestrating EHR/pharmacy workflows end-to-end | ||
| 15% | Non-Hallucination | AA-Omniscience | Avoiding fabricated drug interactions, doses, or guidelines | ||
| 15% | Reasoning | HLE | Multi-step clinical reasoning across biology and medicine | ||
| 10% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using patient and member support workflows | ||
| Engineering Index | Industry | 35% | Engineering Knowledge | AA-Omniscience | Domain recall across civil, electrical, mechanical, and other engineering disciplines |
| 35% | Reasoning | HLE, GPQA Diamond, Crit-Pt | Multi-step quantitative reasoning for derivations, sizing calculations, and design trade-offs | ||
| 25% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step execution of engineering deliverables | ||
| 5% | Agentic Terminal Use | Terminal-Bench v2.1 | Operating real terminal environments for builds, scripts, system administration, and debugging | ||
| Economics Index | Industry | 35% | Economics Knowledge | AA-Omniscience | Recall across micro and macroeconomics, public finance, and markets |
| 35% | Reasoning | HLE | Multi-step quantitative and analytic reasoning for modeling, estimation, and inference | ||
| 15% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step execution of analytical deliverables | ||
| 15% | Long-Context | LCR | Reading and reasoning across long reports, datasets, and research notes |