法律指数
评估模型在法律领域的表现。评估的能力包括合同法、侵权法、宪法等领域知识,以及法律研究与起草、诉讼支持、合规审查等。
查看代表性工作流The Artificial Analysis Legal Index combines performance across benchmarks chosen for legal practice. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across legal tasks. 所有底层基准测试均由 Artificial Analysis 独立运行。有关评测的实施方式,请参阅我们的智能基准测试方法论。
| 能力 | 权重 | 评测 |
|---|---|---|
| 法律知识 | 35% | AA-Omniscience 法律准确率 |
| 智能体知识工作 | 25% | GDPval-AA v2.1和AA-Briefcase v1.1 |
| 推理 | 15% | HLE |
| 长上下文 | 10% | LCR和GDP.pdf |
| 避免幻觉 | 10% | AA-Omniscience 法律避免幻觉 |
| 智能体工具使用 | 5% | AutomationBench-AA 支持和运营 |
得分
Artificial Analysis 法律指数
Artificial Analysis 法律指数:能力明细
能力明细
Artificial Analysis 法律指数:法律知识
代表性工作流
这些真实工作流重点检验 法律指数 中权重最高的能力。
示例:Determine whether a non-compete is enforceable under controlling state precedent by gathering the relevant case law, applying each holding to the employee's facts, distinguishing unfavorable rulings, and synthesizing the analysis into a report.
示例:Reconcile US and EU indemnity language in a cross-border M&A share purchase agreement 48 hours before signing to flag irreconcilable conflicts, propose drafting that satisfies both regimes where possible, and deliver a partner-ready redline.
示例:Advise a startup shipping a feature that may trigger unsettled state privacy rules such as the CCPA to ask the clarifying questions, lay out the trade-offs by jurisdiction, surface open legal risks, and recommend a defensible launch posture.
示例:Work a 200,000-document e-discovery production delivered ten days before trial to prioritise responsive material, flag likely privilege issues for attorney review, and draft a deposition outline tied to the strongest exhibits.
示例:Rewrite internal policy for a new financial regulation taking effect in 90 days that clashes with procedures in three business units to produce a unified replacement policy, an implementation plan with named owners, and a training brief grounded in the statute and existing policy library.
示例:Consolidate 40 active litigation matters tracked across three incompatible case-management systems to produce one unified docket, surface conflicting court deadlines, and propose a single workflow going forward.
成本
Artificial Analysis 法律指数:单任务成本
Artificial Analysis 法律指数 与单任务成本
速度
Artificial Analysis 法律指数:单任务耗时
输出 token
Artificial Analysis 法律指数:单任务输出 token
发布日期
Artificial Analysis 法律指数 与发布日期
常见问题
根据 Artificial Analysis 法律指数,目前在法律工作上表现最佳的 AI 模型是 Claude Opus 5.5 (Max, Default Fallback) (63)、Claude Fable 5.1 (Xhigh, Default Fallback) (61)和Claude Opus 5.5 (Xhigh, Default Fallback) (61)。新模型发布后,排行榜会随之更新。
有。Artificial Analysis 法律指数 是一项独立基准测试,用于衡量 AI 模型在法律工作上的表现。它评估法律知识、智能体知识工作、推理、长文档分析、避免幻觉和智能体工具使用等能力。
法律指数 是 Artificial Analysis 推出的综合基准测试,用于评估模型在法律领域的表现。评估能力包括合同法、侵权法、宪法等专业知识,以及法律研究与起草、诉讼支持、合规审查等。
法律指数 按各项能力子分数的加权平均值计算。各项子分数及其权重为:法律知识 (35%)、智能体知识工作 (25%)、推理 (15%)、长上下文 (10%)、避免幻觉 (10%)和智能体工具使用 (5%)。
法律指数 包含 AA-Omniscience 法律准确率、GDPval-AA v2.1、AA-Briefcase v1.1、HLE、LCR、GDP.pdf、AA-Omniscience 法律避免幻觉和AutomationBench-AA 支持和运营。
在已公布结果的模型中,Claude Opus 5.5 (Max, Default Fallback) 目前以 63 分位居 法律指数 榜首。 查看模型
法律指数 得分越高,表示模型在构成该指数的各项基准测试中整体表现越强。对于特定用例,单项基准测试结果可能比综合得分更具参考价值。