Finance & Accounting Index
Assesses model performance across the finance and accounting domain. Capabilities evaluated include domain-specific knowledge (accounting, investments, corporate and markets), financial analysis and reporting, compliance and audit, market research, and more.
대표 워크플로 보기The Artificial Analysis Finance & Accounting Index combines performance across Intelligence benchmarks sliced for finance and accounting tasks. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across finance and accounting tasks. 모든 기반 벤치마크는 Artificial Analysis가 독립적으로 실행합니다. 평가 수행 방식은 지능 벤치마킹 방법론을 참조하세요.
| 역량 | 가중치 | 평가 |
|---|---|---|
| Business Knowledge | 30% | AA-Omniscience Business Accuracy |
| Agentic Knowledge Work | 30% | GDPval-AA v2 및 AA-Briefcase |
| Reasoning | 20% | HLE |
| Agentic Tool Use | 10% | AutomationBench-AA Finance |
| Long-Context | 5% | LCR 및 GDP.pdf |
| Non-Hallucination | 5% | AA-Omniscience Business Non-Hallucination |
점수
Artificial Analysis Finance & Accounting Index
Artificial Analysis Finance & Accounting Index: 역량 세부 분석
역량 세부 분석
Artificial Analysis Finance & Accounting Index: Business Knowledge
대표 워크플로
Finance & Accounting Index에서 가장 높은 가중치를 둔 역량을 평가하는 실제 워크플로입니다.
예시: Build a same-day EBITDA bridge for an acquisition target from five years of audited 10-Ks and an unaudited Q3 pack to reconcile GAAP and IFRS treatment, surface customer-concentration risk in the margin walk, and assemble a trading-comps table.
예시: Reconstruct twelve months of undocumented expense reimbursements from the general ledger and journal entries ahead of a SOX audit to tie each line to source evidence, produce an auditable trail, and flag unsupported items.
예시: Recover a twice-re-scoped SAP S/4HANA rollout where three departments are cross-blocked to map the critical-path dependencies, re-baseline milestones with explicit trade-offs, and draft stakeholder updates that state what slips if scope stays fixed.
예시: Reconcile two vendor studies reporting opposite consumer preferences for the same launch to compare their sampling and conjoint methodologies, explain plausible reasons for the split, and recommend the lower-risk go-to-market path.
예시: Diagnose a fulfilment centre whose pick error rate doubled after a floor reorganisation to rank likely root causes from shift logs and layout data, and propose interventions by expected lift.
출시일
Artificial Analysis Finance & Accounting Index vs. 출시일
비용
Artificial Analysis Finance & Accounting Index: 작업당 비용
Artificial Analysis Finance & Accounting Index: 총비용
속도
Artificial Analysis Finance & Accounting Index: 작업당 시간
출력 토큰
Artificial Analysis Finance & Accounting Index: 작업당 출력 토큰
자주 묻는 질문
Artificial Analysis Finance & Accounting Index에 따르면 현재 금융 및 회계 업무에서 가장 뛰어난 AI 모델은 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (57), Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) (56) 및 GPT-6 Astra (max) (55)입니다. 순위는 새 모델 출시에 맞춰 업데이트됩니다.
네. Artificial Analysis의 Finance & Accounting Index는 금융 및 회계 업무에서 AI 모델의 성능을 측정하는 독립적인 벤치마크입니다. 비즈니스 지식, 에이전트 지식 업무, 추론, 에이전트 도구 사용, 긴 컨텍스트 분석, 환각 방지를 평가합니다.
Finance & Accounting Index는 금융 및 회계 분야에서 모델의 성능을 평가하는 Artificial Analysis의 종합 벤치마크입니다. 회계, 투자, 기업, 시장 관련 전문 지식과 금융 분석 및 보고, 규제 준수 및 감사, 시장 조사 등을 평가합니다.
Finance & Accounting Index는 각 역량 하위 점수의 가중 평균으로 계산합니다. 하위 점수와 가중치는 Business Knowledge (30%), Agentic Knowledge Work (30%), Reasoning (20%), Agentic Tool Use (10%), Long-Context (5%) 및 Non-Hallucination (5%)입니다.
Finance & Accounting Index에는 AA-Omniscience Business Accuracy, GDPval-AA v2, AA-Briefcase, HLE, AutomationBench-AA Finance, LCR, GDP.pdf 및 AA-Omniscience Business Non-Hallucination이 포함됩니다.
결과가 공개된 모델 중 현재 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)의 Finance & Accounting Index 점수가 57로 가장 높습니다. 모델 보기
Finance & Accounting Index 점수가 높을수록 지수를 구성하는 벤치마크 전반에서 더 뛰어난 성능을 보인다는 뜻입니다. 특정 사용 사례에서는 종합 점수보다 개별 벤치마크 결과가 더 유용할 수 있습니다.