Artificial Analysis 역량 지수 방법론

개요

Artificial Analysis 역량 지수는 코딩 같은 광범위한 역량부터 법률 업무 같은 전문 분야까지, 특정 사용 사례에서 모델이 얼마나 뛰어난 성능을 보이는지 측정합니다.

지수를 구성하는 모든 벤치마크는 Artificial Analysis에서 독립적으로 실행한 후 그 점수를 결합합니다. 모든 평가 지표와 마찬가지로 역량 지수에는 한계가 있으며 모든 사용 사례에 직접 적용되지는 않을 수 있지만, 특정 분야에서 중요한 업무를 기준으로 모델을 비교하는 데 유용한 종합 정보를 제공합니다.

각 기반 벤치마크의 실행 및 채점 방법은 지능 벤치마킹 방법론에 설명되어 있습니다.

지수 구성 분석

지수는 스킬 기반 또는 산업 기반으로 나뉩니다. 스킬 지수는 코딩이나 에이전트형 업무처럼 여러 분야에 적용되는 역량을 측정하며, 구성 벤치마크의 동일 가중 평균으로 계산합니다. 산업 지수는 법률, 의료 및 보건, 금융 및 회계 같은 하나의 직업 또는 분야를 대상으로 하며, 해당 분야의 실제 업무에서 각 역량이 나타나는 빈도에 따라 일련의 역량에 가중치를 부여합니다.

업무 가중치는 O*NET 방식의 업무 활동 분류 체계를 바탕으로 합니다. 모든 평가 지표와 마찬가지로 이러한 지수에는 한계가 있으며 모든 사용 사례에 직접 대응하지 않을 수 있습니다. 아래 표는 각 지수의 구성 요소와 가중치를 보여 줍니다.

지수유형가중치역량평가설명
Agentic IndexSkill50%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step execution of real-world knowledge work
50%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using customer-support workflows over a knowledge base
Finance & Accounting IndexIndustry30%Business KnowledgeAA-OmniscienceDomain recall in accounting, corporate finance, economics, and investments
30%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step task execution, such as running spreadsheets, querying ERPs, or coordinating a workflow
20%ReasoningHLEMulti-step quantitative and analytic reasoning, used for sensitivity analysis, valuation, and structured problem solving
10%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using customer workflows over an unstructured knowledge base
5%Long-ContextLCRReading and reasoning across long financial filings, deal documents, and research notes
5%Non-HallucinationAA-OmniscienceAvoiding fabricated figures or citations
Strategy & Ops IndexIndustry30%Business KnowledgeAA-OmniscienceWorking knowledge of business processes, accounting basics, and operational concepts
35%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and orchestrating multi-step office workflows end-to-end
30%Agentic Customer Interaction𝜏³-BankingMulti-turn dialogue and tool use to resolve customer and stakeholder requests
5%Long-ContextLCRHolding context across long threads, policies, and records
Legal IndexIndustry35%Legal KnowledgeAA-OmniscienceRecall of statutes, doctrines, and procedure across jurisdictions
25%Agentic Knowledge WorkGDPval-AA v2Running matter-management workflows, drafting pipelines, and tool-augmented research
10%Long-ContextLCRReading and synthesizing across contracts, discovery productions, and case-law packets
10%Non-HallucinationAA-OmniscienceAvoiding fabricated case cites or invented statutes
5%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using client intake and support workflows
15%ReasoningHLEMulti-step argumentation, statutory interpretation, and weighing conflicting authorities
Healthcare & Medical IndexIndustry35%Medical & Health KnowledgeAA-OmniscienceClinical knowledge across diagnosis, pharmacology, and care pathways
25%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and orchestrating EHR/pharmacy workflows end-to-end
15%Non-HallucinationAA-OmniscienceAvoiding fabricated drug interactions, doses, or guidelines
15%ReasoningHLEMulti-step clinical reasoning across biology and medicine
10%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using patient and member support workflows
Engineering IndexIndustry35%Engineering KnowledgeAA-OmniscienceDomain recall across civil, electrical, mechanical, and other engineering disciplines
35%ReasoningHLE, GPQA Diamond, Crit-PtMulti-step quantitative reasoning for derivations, sizing calculations, and design trade-offs
25%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step execution of engineering deliverables
5%Agentic Terminal UseTerminal-Bench v2.1Operating real terminal environments for builds, scripts, system administration, and debugging
Economics IndexIndustry35%Economics KnowledgeAA-OmniscienceRecall across micro and macroeconomics, public finance, and markets
35%ReasoningHLEMulti-step quantitative and analytic reasoning for modeling, estimation, and inference
15%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step execution of analytical deliverables
15%Long-ContextLCRReading and reasoning across long reports, datasets, and research notes