All capability indexes

Engineering Index

Assesses model performance across the engineering domain. Capabilities evaluated include domain-specific knowledge (civil, electrical, and mechanical engineering), design and analysis, tooling and automation, technical documentation, and more.

See representative workflows

The Artificial Analysis Engineering Index combines performance across benchmarks chosen for engineering work, spanning engineering knowledge, reasoning, agentic execution, and terminal use. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.

This composite metric provides a single score for tracking model performance across engineering tasks. All underlying benchmarks are run independently by Artificial Analysis. See our Intelligence Benchmarking Methodology for how evaluations are conducted.

CapabilityWeightEvaluations
Engineering Knowledge35%AA-Omniscience Science, Engineering & Mathematics Accuracy
Reasoning35%HLE, GPQA Diamond, Crit-Pt
Agentic Knowledge Work25%GDPval-AA v2
Agentic Terminal Use5%Terminal-Bench v2.1

Score

Artificial Analysis Engineering Index

Incorporates 6 evaluations: AA-Omniscience, Humanity's Last Exam, GPQA Diamond, CritPt, GDPval-AA v2, Terminal-Bench v2.1 · Higher is better
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Engineering Index: Capability Breakdown

Incorporates 6 evaluations: AA-Omniscience, Humanity's Last Exam, GPQA Diamond, CritPt, GDPval-AA v2, Terminal-Bench v2.1 · Segmented by contribution
Reasoning models are indicated by a lightbulb icon

Capability Breakdown

Artificial Analysis Engineering Index: Engineering Knowledge

Incorporates 1 evaluation: AA-Omniscience · Higher is better
Reasoning models are indicated by a lightbulb icon

Representative Workflows

Real-world workflows that exercise the capabilities the Engineering Index weights most heavily.

Design & analysisEngineering KnowledgeReasoning

Example: Design and analyze a wind turbine support structure to size components against fatigue and extreme-wind load cases, justify safety margins against a governing standard such as IEC 61400, and maximize power output while reliably withstanding environmental stress.

Tooling & automationReasoningAgentic Knowledge WorkAgentic Terminal Use

Example: Track an intermittent CFD pipeline failure through the CMake build and Conda environment on a Slurm cluster from the terminal, then ship a fix that spares adjacent batch jobs.

Documentation & specificationEngineering KnowledgeAgentic Knowledge Work

Example: Read a vendor package of CAD schematics and dimensioned drawings to extract GD&T callouts, materials, and interface dimensions per ASME Y14.5, reconcile conflicts across sheets, and draft a specification that cites each source drawing.

Release Date

Artificial Analysis Engineering Index vs. Release Date

Most attractive region

Cost

Artificial Analysis Engineering Index: Cost per Task

Average cost per task (USD), broken down by input, cache hit, cache write, reasoning, and answer tokens

Average cost per task in the index. Costs are split by input, cache hit, cache write, reasoning, and answer token pricing where canonical token counts are available.

Artificial Analysis Engineering Index: Total Cost

Total cost (USD) to run the index

The cost to run the index, calculated using the model's input and output token pricing and the number of tokens used.

Speed

Artificial Analysis Engineering Index: Time per Task

Weighted average decode time (minutes) per task; excludes TTFT and overhead time · Lower is better

The weighted average time (minutes) per index task. This is calculated by dividing output tokens per task by output speed, weighted by the relative weights of each benchmark in the index.

Output Tokens

Artificial Analysis Engineering Index: Output Tokens per Task

Output tokens used to run one task, broken down by reasoning and answer tokens

The average number of answer and reasoning tokens produced per benchmark task in this index.

Frequently Asked Questions

Based on the Artificial Analysis Engineering Index, the top-performing AI models for engineering work are currently Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (63), GPT-5.6 Sol (max) (60), and GPT-5.6 Sol (xhigh) (58). Rankings are updated as new models are released.

Yes. The Engineering Index from Artificial Analysis is an independent benchmark of how AI models perform on engineering work. It measures performance on the engineering domain, including engineering knowledge, quantitative reasoning, agentic execution, and terminal use.

The Engineering Index is a composite benchmark from Artificial Analysis that assesses model performance across the engineering domain. Capabilities evaluated include domain-specific knowledge (civil, electrical, and mechanical engineering), design and analysis, tooling and automation, technical documentation, and more.

The Engineering Index is calculated as a weighted average of its capability sub-scores. The sub-scores and their weights are: Engineering Knowledge (35%), Reasoning (35%), Agentic Knowledge Work (25%), and Agentic Terminal Use (5%).

The Engineering Index includes AA-Omniscience Science, Engineering & Mathematics Accuracy, HLE, GPQA Diamond, Crit-Pt, GDPval-AA v2, and Terminal-Bench v2.1.

Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) currently has the highest Engineering Index score, with a score of 63 among models with published results. View model

A higher Engineering Index score indicates stronger overall performance across the benchmarks that make up the index. For a specific use case, individual benchmark results may be more informative than the composite score.