법률 지수
법률 분야 전반에서 모델 성능을 평가합니다. 계약법, 불법행위법, 헌법에 관한 전문 지식, 법률 조사 및 작성, 소송 지원, 규정 준수 검토 등을 평가합니다.
대표 워크플로 보기The Artificial Analysis Legal Index combines performance across benchmarks chosen for legal practice. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across legal tasks. 모든 기반 벤치마크는 Artificial Analysis가 독립적으로 실행합니다. 평가 수행 방식은 지능 벤치마킹 방법론을 참조하세요.
| 역량 | 가중치 | 평가 |
|---|---|---|
| 법률 지식 | 35% | AA-Omniscience 법률 정확도 |
| 에이전트형 지식 작업 | 25% | GDPval-AA v2.1 및 AA-Briefcase v1.1 |
| 추론 | 15% | HLE |
| 장문 맥락 | 10% | LCR 및 GDP.pdf |
| 환각 방지 | 10% | AA-Omniscience 법률 환각 방지 |
| 에이전트형 도구 사용 | 5% | AutomationBench-AA 지원 및 운영 |
점수
Artificial Analysis 법률 지수
Artificial Analysis 법률 지수: 역량 세부 분석
역량 세부 분석
Artificial Analysis 법률 지수: 법률 지식
대표 워크플로
법률 지수에서 가장 높은 가중치를 둔 역량을 평가하는 실제 워크플로입니다.
예시: Determine whether a non-compete is enforceable under controlling state precedent by gathering the relevant case law, applying each holding to the employee's facts, distinguishing unfavorable rulings, and synthesizing the analysis into a report.
예시: Reconcile US and EU indemnity language in a cross-border M&A share purchase agreement 48 hours before signing to flag irreconcilable conflicts, propose drafting that satisfies both regimes where possible, and deliver a partner-ready redline.
예시: Advise a startup shipping a feature that may trigger unsettled state privacy rules such as the CCPA to ask the clarifying questions, lay out the trade-offs by jurisdiction, surface open legal risks, and recommend a defensible launch posture.
예시: Work a 200,000-document e-discovery production delivered ten days before trial to prioritise responsive material, flag likely privilege issues for attorney review, and draft a deposition outline tied to the strongest exhibits.
예시: Rewrite internal policy for a new financial regulation taking effect in 90 days that clashes with procedures in three business units to produce a unified replacement policy, an implementation plan with named owners, and a training brief grounded in the statute and existing policy library.
예시: Consolidate 40 active litigation matters tracked across three incompatible case-management systems to produce one unified docket, surface conflicting court deadlines, and propose a single workflow going forward.
비용
Artificial Analysis 법률 지수: 작업당 비용
Artificial Analysis 법률 지수 vs. 작업당 비용
속도
Artificial Analysis 법률 지수: 작업당 시간
출력 토큰
Artificial Analysis 법률 지수: 작업당 출력 토큰
출시일
Artificial Analysis 법률 지수 vs. 출시일
자주 묻는 질문
Artificial Analysis 법률 지수에 따르면 현재 법률 업무에서 가장 뛰어난 AI 모델은 Claude Opus 5.5 (Max, Default Fallback) (63), Claude Fable 5.1 (Xhigh, Default Fallback) (61) 및 Claude Opus 5.5 (Xhigh, Default Fallback) (61)입니다. 순위는 새 모델 출시에 맞춰 업데이트됩니다.
네. Artificial Analysis의 법률 지수는 법률 업무에서 AI 모델의 성능을 측정하는 독립적인 벤치마크입니다. 법률 지식, 에이전트 지식 업무, 추론, 긴 문서 분석, 환각 방지, 에이전트 도구 사용을 평가합니다.
법률 지수는 법률 분야에서 모델의 성능을 평가하는 Artificial Analysis의 종합 벤치마크입니다. 계약법, 불법행위법, 헌법 관련 전문 지식과 법률 조사 및 문서 작성, 소송 지원, 규제 준수 검토 등을 평가합니다.
법률 지수는 각 역량 하위 점수의 가중 평균으로 계산합니다. 하위 점수와 가중치는 법률 지식 (35%), 에이전트형 지식 작업 (25%), 추론 (15%), 장문 맥락 (10%), 환각 방지 (10%) 및 에이전트형 도구 사용 (5%)입니다.
법률 지수에는 AA-Omniscience 법률 정확도, GDPval-AA v2.1, AA-Briefcase v1.1, HLE, LCR, GDP.pdf, AA-Omniscience 법률 환각 방지 및 AutomationBench-AA 지원 및 운영이 포함됩니다.
결과가 공개된 모델 중 현재 Claude Opus 5.5 (Max, Default Fallback)의 법률 지수 점수가 63로 가장 높습니다. 모델 보기
법률 지수 점수가 높을수록 지수를 구성하는 벤치마크 전반에서 더 뛰어난 성능을 보인다는 뜻입니다. 특정 사용 사례에서는 종합 점수보다 개별 벤치마크 결과가 더 유용할 수 있습니다.