法務指数
法務分野におけるモデルの性能を評価します。契約法、不法行為法、憲法に関する専門知識、法的調査・文書作成、訴訟支援、コンプライアンス審査などを対象とします。
代表的なワークフローを見るThe Artificial Analysis Legal Index combines performance across benchmarks chosen for legal practice. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across legal tasks. 基礎となるすべてのベンチマークは、Artificial Analysisが独立して実施しています。評価方法の詳細は、知能ベンチマークの方法論をご覧ください。
| 能力 | 重み | 評価 |
|---|---|---|
| 法的知識 | 35% | AA-Omniscience 法務の正確性 |
| エージェント型知識作業 | 25% | GDPval-AA v2.1、AA-Briefcase v1.1 |
| 推論 | 15% | HLE |
| 長文脈 | 10% | LCR、GDP.pdf |
| 非ハルシネーション | 10% | AA-Omniscience 法務の非ハルシネーション |
| エージェント型ツール利用 | 5% | AutomationBench-AA サポート、オペレーション |
スコア
Artificial Analysis 法務指数
Artificial Analysis 法務指数:能力の内訳
能力の内訳
Artificial Analysis 法務指数:法的知識
代表的なワークフロー
法務指数で特に重視される能力を測る、実際の業務を想定したワークフローです。
例:Determine whether a non-compete is enforceable under controlling state precedent by gathering the relevant case law, applying each holding to the employee's facts, distinguishing unfavorable rulings, and synthesizing the analysis into a report.
例:Reconcile US and EU indemnity language in a cross-border M&A share purchase agreement 48 hours before signing to flag irreconcilable conflicts, propose drafting that satisfies both regimes where possible, and deliver a partner-ready redline.
例:Advise a startup shipping a feature that may trigger unsettled state privacy rules such as the CCPA to ask the clarifying questions, lay out the trade-offs by jurisdiction, surface open legal risks, and recommend a defensible launch posture.
例:Work a 200,000-document e-discovery production delivered ten days before trial to prioritise responsive material, flag likely privilege issues for attorney review, and draft a deposition outline tied to the strongest exhibits.
例:Rewrite internal policy for a new financial regulation taking effect in 90 days that clashes with procedures in three business units to produce a unified replacement policy, an implementation plan with named owners, and a training brief grounded in the statute and existing policy library.
例:Consolidate 40 active litigation matters tracked across three incompatible case-management systems to produce one unified docket, surface conflicting court deadlines, and propose a single workflow going forward.
コスト
Artificial Analysis 法務指数:タスク当たりのコスト
Artificial Analysis 法務指数 vs. タスク当たりのコスト
速度
Artificial Analysis 法務指数:タスク当たりの時間
出力トークン
Artificial Analysis 法務指数:タスク当たりの出力トークン
リリース日
Artificial Analysis 法務指数 vs. リリース日
よくある質問
Artificial Analysisの法務指数によると、現在法務業務で最も高い性能を示すAIモデルはClaude Opus 5.5 (Max, Default Fallback) (63)、Claude Fable 5.1 (Xhigh, Default Fallback) (61)、Claude Opus 5.5 (Xhigh, Default Fallback) (61)です。ランキングは新しいモデルのリリースに合わせて更新されます。
はい。Artificial Analysisの法務指数は、法務業務におけるAIモデルの性能を測る独立したベンチマークです。法律知識、エージェント型ナレッジワーク、推論、長文文書の分析、ハルシネーション抑制、エージェント型ツール使用を評価します。
法務指数は、法務分野におけるモデルの性能を評価するArtificial Analysisの複合ベンチマークです。契約法、不法行為法、憲法などの専門知識に加え、法務調査と文書作成、訴訟支援、コンプライアンスレビューなどを評価します。
法務指数は、各能力のサブスコアを加重平均して算出します。サブスコアとその重みは、法的知識 (35%)、エージェント型知識作業 (25%)、推論 (15%)、長文脈 (10%)、非ハルシネーション (10%)、エージェント型ツール利用 (5%)です。
法務指数にはAA-Omniscience 法務の正確性、GDPval-AA v2.1、AA-Briefcase v1.1、HLE、LCR、GDP.pdf、AA-Omniscience 法務の非ハルシネーション、AutomationBench-AA サポート、オペレーションが含まれます。
結果が公開されているモデルの中では、現在Claude Opus 5.5 (Max, Default Fallback)が法務指数で最高の63を記録しています。 モデルを見る
法務指数のスコアが高いほど、指数を構成するベンチマーク全体での性能が優れていることを示します。特定の用途では、複合スコアよりも個別のベンチマーク結果のほうが参考になる場合があります。