Engineering Index
Assesses model performance across the engineering domain. Capabilities evaluated include domain-specific knowledge (civil, electrical, and mechanical engineering), design and analysis, tooling and automation, technical documentation, and more.
代表的なワークフローを見るThe Artificial Analysis Engineering Index combines performance across benchmarks chosen for engineering work, spanning engineering knowledge, reasoning, agentic execution, and terminal use. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across engineering tasks. 基礎となるすべてのベンチマークは、Artificial Analysisが独立して実施しています。評価方法の詳細は、知能ベンチマークの方法論をご覧ください。
| 能力 | 重み | 評価 |
|---|---|---|
| Engineering Knowledge | 35% | AA-Omniscience Science, Engineering & Mathematics Accuracy |
| Reasoning | 30% | HLE、CritPt |
| Agentic Knowledge Work | 20% | GDPval-AA v2、AA-Briefcase |
| Agentic Terminal Use | 15% | Terminal-Bench v4.0 |
スコア
Artificial Analysis Engineering Index
Artificial Analysis Engineering Index:能力の内訳
能力の内訳
Artificial Analysis Engineering Index:Engineering Knowledge
代表的なワークフロー
Engineering Indexで特に重視される能力を測る、実際の業務を想定したワークフローです。
例:Design and analyze a wind turbine support structure to size components against fatigue and extreme-wind load cases, justify safety margins against a governing standard such as IEC 61400, and maximize power output while reliably withstanding environmental stress.
例:Track an intermittent CFD pipeline failure through the CMake build and Conda environment on a Slurm cluster from the terminal, then ship a fix that spares adjacent batch jobs.
例:Read a vendor package of CAD schematics and dimensioned drawings to extract GD&T callouts, materials, and interface dimensions per ASME Y14.5, reconcile conflicts across sheets, and draft a specification that cites each source drawing.
コスト
Artificial Analysis Engineering Index:タスク当たりのコスト
Artificial Analysis Engineering Index vs. タスク当たりのコスト
速度
Artificial Analysis Engineering Index:タスク当たりの時間
出力トークン
Artificial Analysis Engineering Index:タスク当たりの出力トークン
リリース日
Artificial Analysis Engineering Index vs. リリース日
よくある質問
Artificial AnalysisのEngineering Indexによると、現在エンジニアリング業務で最も高い性能を示すAIモデルはClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) (57)、Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (57)、GPT-6 Astra (max) (55)です。ランキングは新しいモデルのリリースに合わせて更新されます。
はい。Artificial AnalysisのEngineering Indexは、エンジニアリング業務におけるAIモデルの性能を測る独立したベンチマークです。工学知識、定量的推論、エージェント型の実行、ターミナル操作を評価します。
Engineering Indexは、エンジニアリング分野におけるモデルの性能を評価するArtificial Analysisの複合ベンチマークです。土木、電気、機械工学などの専門知識に加え、設計と分析、ツールと自動化、技術文書などを評価します。
Engineering Indexは、各能力のサブスコアを加重平均して算出します。サブスコアとその重みは、Engineering Knowledge (35%)、Reasoning (30%)、Agentic Knowledge Work (20%)、Agentic Terminal Use (15%)です。
Engineering IndexにはAA-Omniscience Science, Engineering & Mathematics Accuracy、HLE、CritPt、GDPval-AA v2、AA-Briefcase、Terminal-Bench v4.0が含まれます。
結果が公開されているモデルの中では、現在Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)がEngineering Indexで最高の57を記録しています。 モデルを見る
Engineering Indexのスコアが高いほど、指数を構成するベンチマーク全体での性能が優れていることを示します。特定の用途では、複合スコアよりも個別のベンチマーク結果のほうが参考になる場合があります。