Finance & Accounting Index
Assesses model performance across the finance and accounting domain. Capabilities evaluated include domain-specific knowledge (accounting, investments, corporate and markets), financial analysis and reporting, compliance and audit, market research, and more.
查看代表性工作流The Artificial Analysis Finance & Accounting Index combines performance across Intelligence benchmarks sliced for finance and accounting tasks. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across finance and accounting tasks. 所有底层基准测试均由 Artificial Analysis 独立运行。有关评测的实施方式,请参阅我们的智能基准测试方法论。
| 能力 | 权重 | 评测 |
|---|---|---|
| Business Knowledge | 30% | AA-Omniscience Business Accuracy |
| Agentic Knowledge Work | 30% | GDPval-AA v2和AA-Briefcase |
| Reasoning | 20% | HLE |
| Agentic Tool Use | 10% | AutomationBench-AA Finance |
| Long-Context | 5% | LCR和GDP.pdf |
| Non-Hallucination | 5% | AA-Omniscience Business Non-Hallucination |
得分
Artificial Analysis Finance & Accounting Index
Artificial Analysis Finance & Accounting Index:能力明细
能力明细
Artificial Analysis Finance & Accounting Index:Business Knowledge
代表性工作流
这些真实工作流重点检验 Finance & Accounting Index 中权重最高的能力。
示例:Build a same-day EBITDA bridge for an acquisition target from five years of audited 10-Ks and an unaudited Q3 pack to reconcile GAAP and IFRS treatment, surface customer-concentration risk in the margin walk, and assemble a trading-comps table.
示例:Reconstruct twelve months of undocumented expense reimbursements from the general ledger and journal entries ahead of a SOX audit to tie each line to source evidence, produce an auditable trail, and flag unsupported items.
示例:Recover a twice-re-scoped SAP S/4HANA rollout where three departments are cross-blocked to map the critical-path dependencies, re-baseline milestones with explicit trade-offs, and draft stakeholder updates that state what slips if scope stays fixed.
示例:Reconcile two vendor studies reporting opposite consumer preferences for the same launch to compare their sampling and conjoint methodologies, explain plausible reasons for the split, and recommend the lower-risk go-to-market path.
示例:Diagnose a fulfilment centre whose pick error rate doubled after a floor reorganisation to rank likely root causes from shift logs and layout data, and propose interventions by expected lift.
成本
Artificial Analysis Finance & Accounting Index:单任务成本
Artificial Analysis Finance & Accounting Index 与单任务成本
速度
Artificial Analysis Finance & Accounting Index:单任务耗时
输出 token
Artificial Analysis Finance & Accounting Index:单任务输出 token
发布日期
Artificial Analysis Finance & Accounting Index 与发布日期
常见问题
根据 Artificial Analysis Finance & Accounting Index,目前在金融与会计工作上表现最佳的 AI 模型是 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (57)、Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) (56)和GPT-6 Astra (max) (55)。新模型发布后,排行榜会随之更新。
有。Artificial Analysis Finance & Accounting Index 是一项独立基准测试,用于衡量 AI 模型在金融与会计工作上的表现。它评估商业知识、智能体知识工作、推理、智能体工具使用、长上下文分析和避免幻觉等能力。
Finance & Accounting Index 是 Artificial Analysis 推出的综合基准测试,用于评估模型在金融与会计领域的表现。评估能力包括会计、投资、企业与市场等专业知识,以及金融分析与报告、合规与审计、市场研究等。
Finance & Accounting Index 按各项能力子分数的加权平均值计算。各项子分数及其权重为:Business Knowledge (30%)、Agentic Knowledge Work (30%)、Reasoning (20%)、Agentic Tool Use (10%)、Long-Context (5%)和Non-Hallucination (5%)。
Finance & Accounting Index 包含 AA-Omniscience Business Accuracy、GDPval-AA v2、AA-Briefcase、HLE、AutomationBench-AA Finance、LCR、GDP.pdf和AA-Omniscience Business Non-Hallucination。
在已公布结果的模型中,Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) 目前以 57 分位居 Finance & Accounting Index 榜首。 查看模型
Finance & Accounting Index 得分越高,表示模型在构成该指数的各项基准测试中整体表现越强。对于特定用例,单项基准测试结果可能比综合得分更具参考价值。