Artificial Analysis 역량 지수 방법론
개요
Artificial Analysis 역량 지수는 법률, 의료, 금융 업무처럼 특정 전문 사용 사례에서 모델이 얼마나 뛰어난 성능을 보이는지 측정합니다.
지수를 구성하는 모든 벤치마크를 독립적으로 실행한 후 그 점수를 지수로 결합합니다. 모든 평가 지표와 마찬가지로 역량 지수에는 한계가 있으며 모든 사용 사례에 직접 적용되지는 않을 수 있지만, 특정 분야에서 중요한 업무를 기준으로 모델을 비교하는 데 유용한 종합 정보를 제공합니다.
각 기반 벤치마크의 실행 및 채점 방법은 지능 벤치마킹 방법론에 설명되어 있습니다.
지수 구성 분석
각 지수는 법률, 의료 및 보건, 금융 및 회계 같은 하나의 직업 또는 분야를 대상으로 하며, 해당 분야의 실제 업무에서 각 역량이 나타나는 빈도에 따라 일련의 역량에 가중치를 부여합니다.
업무 가중치는 O*NET 방식의 업무 활동 분류 체계를 바탕으로 합니다. 아래 표는 각 지수의 구성 요소와 가중치를 보여 줍니다.
| 지수 | 가중치 | 역량 | 평가 | 설명 |
|---|---|---|---|---|
| Finance & Accounting Index | 30% | Business Knowledge | AA-Omniscience | Domain recall in accounting, corporate finance, economics, and investments |
| 30% | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | Tool use, planning, and multi-step task execution, such as building spreadsheets, running an analysis, or coordinating a workflow | |
| 20% | Reasoning | HLE | Multi-step quantitative and analytic reasoning, used for sensitivity analysis, valuation, and structured problem solving | |
| 10% | Agentic Tool Use | AutomationBench-AA | Completing finance workflows across business apps such as spreadsheets, email, and accounting tools without breaking guardrails | |
| 5% | Long-Context | LCR, GDP.pdf | Reading and reasoning across long financial filings, deal documents, and PDF reports | |
| 5% | Non-Hallucination | AA-Omniscience | Avoiding fabricated figures or citations | |
| Strategy & Ops Index | 30% | Business Knowledge | AA-Omniscience | Working knowledge of business processes, accounting basics, and operational concepts |
| 35% | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | Tool use, planning, and orchestrating multi-step office workflows end-to-end | |
| 30% | Agentic Tool Use | AutomationBench-AA | Completing operations, HR, marketing, sales, and support workflows across business apps without breaking guardrails | |
| 5% | Long-Context | LCR, GDP.pdf | Holding context across long threads, policies, records, and PDF documents | |
| Legal Index | 35% | Legal Knowledge | AA-Omniscience | Recall of statutes, doctrines, and procedure across jurisdictions |
| 25% | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | Running matter-management workflows, drafting pipelines, and tool-augmented research | |
| 15% | Reasoning | HLE | Multi-step argumentation, statutory interpretation, and weighing conflicting authorities | |
| 10% | Long-Context | LCR, GDP.pdf | Reading and synthesizing across contracts, discovery productions, case-law packets, and PDF filings | |
| 10% | Non-Hallucination | AA-Omniscience | Avoiding fabricated case cites or invented statutes | |
| 5% | Agentic Tool Use | AutomationBench-AA | Completing client support and operations workflows across business apps without breaking guardrails | |
| Healthcare & Medical Index | 30% | Medical & Health Knowledge | AA-Omniscience | Clinical knowledge across diagnosis, pharmacology, and care pathways |
| 25% | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | Tool use, planning, and orchestrating EHR/pharmacy workflows end-to-end | |
| 15% | Long-Context Reasoning | MLCR-AA | Synthesising findings across long, fragmented patient records and claims files | |
| 10% | Non-Hallucination | AA-Omniscience | Avoiding fabricated drug interactions, doses, or guidelines | |
| 10% | Reasoning | HLE | Multi-step clinical reasoning across biology and medicine | |
| 10% | Agentic Tool Use | AutomationBench-AA | Completing patient support and operations workflows across business apps without breaking guardrails | |
| Engineering Index | 35% | Engineering Knowledge | AA-Omniscience | Domain recall across civil, electrical, mechanical, and other engineering disciplines |
| 30% | Reasoning | HLE, CritPt | Multi-step quantitative reasoning for derivations, sizing calculations, and design trade-offs | |
| 20% | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | Tool use, planning, and multi-step execution of engineering deliverables | |
| 15% | Agentic Terminal Use | Terminal-Bench v4.0 | Operating real terminal environments for builds, scripts, system administration, and debugging | |
| Economics Index | 35% | Economics Knowledge | AA-Omniscience | Recall across micro and macroeconomics, public finance, and markets |
| 35% | Reasoning | HLE | Multi-step quantitative and analytic reasoning for modeling, estimation, and inference | |
| 25% | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | Tool use, planning, and multi-step execution of analytical deliverables | |
| 5% | Long-Context Reasoning | LCR | Reading and reasoning across long reports, datasets, and research notes |