Engineering Index
Assesses model performance across the engineering domain. Capabilities evaluated include domain-specific knowledge (civil, electrical, and mechanical engineering), design and analysis, tooling and automation, technical documentation, and more.
See representative workflowsThe Artificial Analysis Engineering Index combines performance across benchmarks chosen for engineering work, spanning engineering knowledge, reasoning, agentic execution, and terminal use. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across engineering tasks. All underlying benchmarks are run independently by Artificial Analysis. See our Intelligence Benchmarking Methodology for how evaluations are conducted.
| Capability | Weight | Evaluations |
|---|---|---|
| Engineering Knowledge | 35% | AA-Omniscience Science, Engineering & Mathematics Accuracy |
| Reasoning | 35% | HLE, GPQA Diamond, Crit-Pt |
| Agentic Knowledge Work | 25% | GDPval-AA v2 |
| Agentic Terminal Use | 5% | Terminal-Bench v2.1 |
Score
Artificial Analysis Engineering Index
Artificial Analysis Engineering Index: Capability Breakdown
Capability Breakdown
Artificial Analysis Engineering Index: Engineering Knowledge
Representative Workflows
Real-world workflows that exercise the capabilities the Engineering Index weights most heavily.
Example: Design and analyze a wind turbine support structure to size components against fatigue and extreme-wind load cases, justify safety margins against a governing standard such as IEC 61400, and maximize power output while reliably withstanding environmental stress.
Example: Track an intermittent CFD pipeline failure through the CMake build and Conda environment on a Slurm cluster from the terminal, then ship a fix that spares adjacent batch jobs.
Example: Read a vendor package of CAD schematics and dimensioned drawings to extract GD&T callouts, materials, and interface dimensions per ASME Y14.5, reconcile conflicts across sheets, and draft a specification that cites each source drawing.
Release Date
Artificial Analysis Engineering Index vs. Release Date
Cost
Artificial Analysis Engineering Index: Cost per Task
Artificial Analysis Engineering Index: Total Cost
Speed
Artificial Analysis Engineering Index: Time per Task
Output Tokens
Artificial Analysis Engineering Index: Output Tokens per Task
Frequently Asked Questions
Based on the Artificial Analysis Engineering Index, the top-performing AI models for engineering work are currently Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (63), GPT-5.6 Sol (max) (60), and GPT-5.6 Sol (xhigh) (58). Rankings are updated as new models are released.
Yes. The Engineering Index from Artificial Analysis is an independent benchmark of how AI models perform on engineering work. It measures performance on the engineering domain, including engineering knowledge, quantitative reasoning, agentic execution, and terminal use.
The Engineering Index is a composite benchmark from Artificial Analysis that assesses model performance across the engineering domain. Capabilities evaluated include domain-specific knowledge (civil, electrical, and mechanical engineering), design and analysis, tooling and automation, technical documentation, and more.
The Engineering Index is calculated as a weighted average of its capability sub-scores. The sub-scores and their weights are: Engineering Knowledge (35%), Reasoning (35%), Agentic Knowledge Work (25%), and Agentic Terminal Use (5%).
The Engineering Index includes AA-Omniscience Science, Engineering & Mathematics Accuracy, HLE, GPQA Diamond, Crit-Pt, GDPval-AA v2, and Terminal-Bench v2.1.
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) currently has the highest Engineering Index score, with a score of 63 among models with published results. View model
A higher Engineering Index score indicates stronger overall performance across the benchmarks that make up the index. For a specific use case, individual benchmark results may be more informative than the composite score.