Legal Index
Assesses model performance across the legal domain. Capabilities evaluated include domain-specific knowledge (contract law, tort law, constitutional law), legal research and drafting, litigation support, compliance review, and more.
代表的なワークフローを見るThe Artificial Analysis Legal Index combines performance across benchmarks chosen for legal practice. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across legal tasks. 基礎となるすべてのベンチマークは、Artificial Analysisが独立して実施しています。評価方法の詳細は、知能ベンチマークの方法論をご覧ください。
| 能力 | 重み | 評価 |
|---|---|---|
| Legal Knowledge | 35% | AA-Omniscience Law Accuracy |
| Agentic Knowledge Work | 25% | GDPval-AA v2、AA-Briefcase |
| Reasoning | 15% | HLE |
| Long-Context | 10% | LCR、GDP.pdf |
| Non-Hallucination | 10% | AA-Omniscience Law Non-Hallucination |
| Agentic Tool Use | 5% | AutomationBench-AA Support & Operations |
スコア
Artificial Analysis Legal Index
Artificial Analysis Legal Index:能力の内訳
能力の内訳
Artificial Analysis Legal Index:Legal Knowledge
代表的なワークフロー
Legal Indexで特に重視される能力を測る、実際の業務を想定したワークフローです。
例:Determine whether a non-compete is enforceable under controlling state precedent by gathering the relevant case law, applying each holding to the employee's facts, distinguishing unfavorable rulings, and synthesizing the analysis into a report.
例:Reconcile US and EU indemnity language in a cross-border M&A share purchase agreement 48 hours before signing to flag irreconcilable conflicts, propose drafting that satisfies both regimes where possible, and deliver a partner-ready redline.
例:Advise a startup shipping a feature that may trigger unsettled state privacy rules such as the CCPA to ask the clarifying questions, lay out the trade-offs by jurisdiction, surface open legal risks, and recommend a defensible launch posture.
例:Work a 200,000-document e-discovery production delivered ten days before trial to prioritise responsive material, flag likely privilege issues for attorney review, and draft a deposition outline tied to the strongest exhibits.
例:Rewrite internal policy for a new financial regulation taking effect in 90 days that clashes with procedures in three business units to produce a unified replacement policy, an implementation plan with named owners, and a training brief grounded in the statute and existing policy library.
例:Consolidate 40 active litigation matters tracked across three incompatible case-management systems to produce one unified docket, surface conflicting court deadlines, and propose a single workflow going forward.
コスト
Artificial Analysis Legal Index:タスク当たりのコスト
Artificial Analysis Legal Index vs. タスク当たりのコスト
速度
Artificial Analysis Legal Index:タスク当たりの時間
出力トークン
Artificial Analysis Legal Index:タスク当たりの出力トークン
リリース日
Artificial Analysis Legal Index vs. リリース日
よくある質問
Artificial AnalysisのLegal Indexによると、現在法務業務で最も高い性能を示すAIモデルはClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) (61)、Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (61)、GPT-6 Astra (max) (59)です。ランキングは新しいモデルのリリースに合わせて更新されます。
はい。Artificial AnalysisのLegal Indexは、法務業務におけるAIモデルの性能を測る独立したベンチマークです。法律知識、エージェント型ナレッジワーク、推論、長文文書の分析、ハルシネーション抑制、エージェント型ツール使用を評価します。
Legal Indexは、法務分野におけるモデルの性能を評価するArtificial Analysisの複合ベンチマークです。契約法、不法行為法、憲法などの専門知識に加え、法務調査と文書作成、訴訟支援、コンプライアンスレビューなどを評価します。
Legal Indexは、各能力のサブスコアを加重平均して算出します。サブスコアとその重みは、Legal Knowledge (35%)、Agentic Knowledge Work (25%)、Reasoning (15%)、Long-Context (10%)、Non-Hallucination (10%)、Agentic Tool Use (5%)です。
Legal IndexにはAA-Omniscience Law Accuracy、GDPval-AA v2、AA-Briefcase、HLE、LCR、GDP.pdf、AA-Omniscience Law Non-Hallucination、AutomationBench-AA Support & Operationsが含まれます。
結果が公開されているモデルの中では、現在Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)がLegal Indexで最高の61を記録しています。 モデルを見る
Legal Indexのスコアが高いほど、指数を構成するベンチマーク全体での性能が優れていることを示します。特定の用途では、複合スコアよりも個別のベンチマーク結果のほうが参考になる場合があります。