エンジニアリング指数
エンジニアリング分野におけるモデルの性能を評価します。土木、電気、機械工学に関する専門知識、設計・分析、ツール・自動化、技術文書などを対象とします。
代表的なワークフローを見るThe Artificial Analysis Engineering Index combines performance across benchmarks chosen for engineering work, spanning engineering knowledge, reasoning, agentic execution, and terminal use. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across engineering tasks. 基礎となるすべてのベンチマークは、Artificial Analysisが独立して実施しています。評価方法の詳細は、知能ベンチマークの方法論をご覧ください。
| 能力 | 重み | 評価 |
|---|---|---|
| 工学知識 | 35% | AA-Omniscience 科学・工学・数学の正確性 |
| 推論 | 30% | HLE、CritPt |
| エージェント型知識作業 | 20% | GDPval-AA v2.1、AA-Briefcase v1.1 |
| エージェント型ターミナル利用 | 15% | Terminal-Bench 4.0 |
スコア
Artificial Analysis エンジニアリング指数
Artificial Analysis エンジニアリング指数:能力の内訳
能力の内訳
Artificial Analysis エンジニアリング指数:工学知識
代表的なワークフロー
エンジニアリング指数で特に重視される能力を測る、実際の業務を想定したワークフローです。
例:Design and analyze a wind turbine support structure to size components against fatigue and extreme-wind load cases, justify safety margins against a governing standard such as IEC 61400, and maximize power output while reliably withstanding environmental stress.
例:Track an intermittent CFD pipeline failure through the CMake build and Conda environment on a Slurm cluster from the terminal, then ship a fix that spares adjacent batch jobs.
例:Read a vendor package of CAD schematics and dimensioned drawings to extract GD&T callouts, materials, and interface dimensions per ASME Y14.5, reconcile conflicts across sheets, and draft a specification that cites each source drawing.
コスト
Artificial Analysis エンジニアリング指数:タスク当たりのコスト
Artificial Analysis エンジニアリング指数 vs. タスク当たりのコスト
速度
Artificial Analysis エンジニアリング指数:タスク当たりの時間
出力トークン
Artificial Analysis エンジニアリング指数:タスク当たりの出力トークン
リリース日
Artificial Analysis エンジニアリング指数 vs. リリース日
よくある質問
Artificial Analysisのエンジニアリング指数によると、現在エンジニアリング業務で最も高い性能を示すAIモデルはClaude Opus 5.5 (Max, Default Fallback) (60)、Claude Opus 5.5 (Xhigh, Default Fallback) (59)、Claude Sonnet 5.5 (Max, Default Fallback) (58)です。ランキングは新しいモデルのリリースに合わせて更新されます。
はい。Artificial Analysisのエンジニアリング指数は、エンジニアリング業務におけるAIモデルの性能を測る独立したベンチマークです。工学知識、定量的推論、エージェント型の実行、ターミナル操作を評価します。
エンジニアリング指数は、エンジニアリング分野におけるモデルの性能を評価するArtificial Analysisの複合ベンチマークです。土木、電気、機械工学などの専門知識に加え、設計と分析、ツールと自動化、技術文書などを評価します。
エンジニアリング指数は、各能力のサブスコアを加重平均して算出します。サブスコアとその重みは、工学知識 (35%)、推論 (30%)、エージェント型知識作業 (20%)、エージェント型ターミナル利用 (15%)です。
エンジニアリング指数にはAA-Omniscience 科学・工学・数学の正確性、HLE、CritPt、GDPval-AA v2.1、AA-Briefcase v1.1、Terminal-Bench 4.0が含まれます。
結果が公開されているモデルの中では、現在Claude Opus 5.5 (Max, Default Fallback)がエンジニアリング指数で最高の60を記録しています。 モデルを見る
エンジニアリング指数のスコアが高いほど、指数を構成するベンチマーク全体での性能が優れていることを示します。特定の用途では、複合スコアよりも個別のベンチマーク結果のほうが参考になる場合があります。