財務・会計指数
財務・会計分野におけるモデルの性能を評価します。会計、投資、企業・市場に関する専門知識、財務分析・報告、コンプライアンス・監査、市場調査などを対象とします。
代表的なワークフローを見るThe Artificial Analysis Finance & Accounting Index combines performance across Intelligence benchmarks sliced for finance and accounting tasks. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across finance and accounting tasks. 基礎となるすべてのベンチマークは、Artificial Analysisが独立して実施しています。評価方法の詳細は、知能ベンチマークの方法論をご覧ください。
| 能力 | 重み | 評価 |
|---|---|---|
| ビジネス知識 | 30% | AA-Omniscience ビジネスの正確性 |
| エージェント型知識作業 | 30% | GDPval-AA v2.1、AA-Briefcase v1.1 |
| 推論 | 20% | HLE |
| エージェント型ツール利用 | 10% | AutomationBench-AA 財務 |
| 長文脈 | 5% | LCR、GDP.pdf |
| 非ハルシネーション | 5% | AA-Omniscience ビジネスの非ハルシネーション |
スコア
Artificial Analysis 財務・会計指数
Artificial Analysis 財務・会計指数:能力の内訳
能力の内訳
Artificial Analysis 財務・会計指数:ビジネス知識
代表的なワークフロー
財務・会計指数で特に重視される能力を測る、実際の業務を想定したワークフローです。
例:Build a same-day EBITDA bridge for an acquisition target from five years of audited 10-Ks and an unaudited Q3 pack to reconcile GAAP and IFRS treatment, surface customer-concentration risk in the margin walk, and assemble a trading-comps table.
例:Reconstruct twelve months of undocumented expense reimbursements from the general ledger and journal entries ahead of a SOX audit to tie each line to source evidence, produce an auditable trail, and flag unsupported items.
例:Recover a twice-re-scoped SAP S/4HANA rollout where three departments are cross-blocked to map the critical-path dependencies, re-baseline milestones with explicit trade-offs, and draft stakeholder updates that state what slips if scope stays fixed.
例:Reconcile two vendor studies reporting opposite consumer preferences for the same launch to compare their sampling and conjoint methodologies, explain plausible reasons for the split, and recommend the lower-risk go-to-market path.
例:Diagnose a fulfilment centre whose pick error rate doubled after a floor reorganisation to rank likely root causes from shift logs and layout data, and propose interventions by expected lift.
コスト
Artificial Analysis 財務・会計指数:タスク当たりのコスト
Artificial Analysis 財務・会計指数 vs. タスク当たりのコスト
速度
Artificial Analysis 財務・会計指数:タスク当たりの時間
出力トークン
Artificial Analysis 財務・会計指数:タスク当たりの出力トークン
リリース日
Artificial Analysis 財務・会計指数 vs. リリース日
よくある質問
Artificial Analysisの財務・会計指数によると、現在金融・会計業務で最も高い性能を示すAIモデルはClaude Opus 5.5 (Max, Default Fallback) (61)、Claude Opus 5.5 (Xhigh, Default Fallback) (58)、Claude Sonnet 5.5 (Max, Default Fallback) (57)です。ランキングは新しいモデルのリリースに合わせて更新されます。
はい。Artificial Analysisの財務・会計指数は、金融・会計業務におけるAIモデルの性能を測る独立したベンチマークです。ビジネス知識、エージェント型ナレッジワーク、推論、エージェント型ツール使用、長文コンテキスト分析、ハルシネーション抑制を評価します。
財務・会計指数は、金融・会計分野におけるモデルの性能を評価するArtificial Analysisの複合ベンチマークです。会計、投資、企業、市場などの専門知識に加え、財務分析とレポート作成、コンプライアンスと監査、市場調査などを評価します。
財務・会計指数は、各能力のサブスコアを加重平均して算出します。サブスコアとその重みは、ビジネス知識 (30%)、エージェント型知識作業 (30%)、推論 (20%)、エージェント型ツール利用 (10%)、長文脈 (5%)、非ハルシネーション (5%)です。
財務・会計指数にはAA-Omniscience ビジネスの正確性、GDPval-AA v2.1、AA-Briefcase v1.1、HLE、AutomationBench-AA 財務、LCR、GDP.pdf、AA-Omniscience ビジネスの非ハルシネーションが含まれます。
結果が公開されているモデルの中では、現在Claude Opus 5.5 (Max, Default Fallback)が財務・会計指数で最高の61を記録しています。 モデルを見る
財務・会計指数のスコアが高いほど、指数を構成するベンチマーク全体での性能が優れていることを示します。特定の用途では、複合スコアよりも個別のベンチマーク結果のほうが参考になる場合があります。