Engineering Index
Assesses model performance across the engineering domain. Capabilities evaluated include domain-specific knowledge (civil, electrical, and mechanical engineering), design and analysis, tooling and automation, technical documentation, and more.
대표 워크플로 보기The Artificial Analysis Engineering Index combines performance across benchmarks chosen for engineering work, spanning engineering knowledge, reasoning, agentic execution, and terminal use. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across engineering tasks. 모든 기반 벤치마크는 Artificial Analysis가 독립적으로 실행합니다. 평가 수행 방식은 지능 벤치마킹 방법론을 참조하세요.
| 역량 | 가중치 | 평가 |
|---|---|---|
| Engineering Knowledge | 35% | AA-Omniscience Science, Engineering & Mathematics Accuracy |
| Reasoning | 30% | HLE 및 CritPt |
| Agentic Knowledge Work | 20% | GDPval-AA v2 및 AA-Briefcase |
| Agentic Terminal Use | 15% | Terminal-Bench v4.0 |
점수
Artificial Analysis Engineering Index
Artificial Analysis Engineering Index: 역량 세부 분석
역량 세부 분석
Artificial Analysis Engineering Index: Engineering Knowledge
대표 워크플로
Engineering Index에서 가장 높은 가중치를 둔 역량을 평가하는 실제 워크플로입니다.
예시: Design and analyze a wind turbine support structure to size components against fatigue and extreme-wind load cases, justify safety margins against a governing standard such as IEC 61400, and maximize power output while reliably withstanding environmental stress.
예시: Track an intermittent CFD pipeline failure through the CMake build and Conda environment on a Slurm cluster from the terminal, then ship a fix that spares adjacent batch jobs.
예시: Read a vendor package of CAD schematics and dimensioned drawings to extract GD&T callouts, materials, and interface dimensions per ASME Y14.5, reconcile conflicts across sheets, and draft a specification that cites each source drawing.
출시일
Artificial Analysis Engineering Index vs. 출시일
비용
Artificial Analysis Engineering Index: 작업당 비용
Artificial Analysis Engineering Index: 총비용
속도
Artificial Analysis Engineering Index: 작업당 시간
출력 토큰
Artificial Analysis Engineering Index: 작업당 출력 토큰
자주 묻는 질문
Artificial Analysis Engineering Index에 따르면 현재 엔지니어링 업무에서 가장 뛰어난 AI 모델은 Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) (57), Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (57) 및 GPT-6 Astra (max) (55)입니다. 순위는 새 모델 출시에 맞춰 업데이트됩니다.
네. Artificial Analysis의 Engineering Index는 엔지니어링 업무에서 AI 모델의 성능을 측정하는 독립적인 벤치마크입니다. 공학 지식, 정량적 추론, 에이전트 실행, 터미널 사용을 평가합니다.
Engineering Index는 엔지니어링 분야에서 모델의 성능을 평가하는 Artificial Analysis의 종합 벤치마크입니다. 토목, 전기, 기계 공학 관련 전문 지식과 설계 및 분석, 도구 및 자동화, 기술 문서 작성 등을 평가합니다.
Engineering Index는 각 역량 하위 점수의 가중 평균으로 계산합니다. 하위 점수와 가중치는 Engineering Knowledge (35%), Reasoning (30%), Agentic Knowledge Work (20%) 및 Agentic Terminal Use (15%)입니다.
Engineering Index에는 AA-Omniscience Science, Engineering & Mathematics Accuracy, HLE, CritPt, GDPval-AA v2, AA-Briefcase 및 Terminal-Bench v4.0이 포함됩니다.
결과가 공개된 모델 중 현재 Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)의 Engineering Index 점수가 57로 가장 높습니다. 모델 보기
Engineering Index 점수가 높을수록 지수를 구성하는 벤치마크 전반에서 더 뛰어난 성능을 보인다는 뜻입니다. 특정 사용 사례에서는 종합 점수보다 개별 벤치마크 결과가 더 유용할 수 있습니다.