공학 지수
공학 전반에서 모델 성능을 평가합니다. 토목, 전기, 기계 공학에 관한 전문 지식, 설계 및 분석, 도구와 자동화, 기술 문서화 등을 평가합니다.
대표 워크플로 보기The Artificial Analysis Engineering Index combines performance across benchmarks chosen for engineering work, spanning engineering knowledge, reasoning, agentic execution, and terminal use. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across engineering tasks. 모든 기반 벤치마크는 Artificial Analysis가 독립적으로 실행합니다. 평가 수행 방식은 지능 벤치마킹 방법론을 참조하세요.
| 역량 | 가중치 | 평가 |
|---|---|---|
| 공학 지식 | 35% | AA-Omniscience 과학, 공학 및 수학 정확도 |
| 추론 | 30% | HLE 및 CritPt |
| 에이전트형 지식 작업 | 20% | GDPval-AA v2.1 및 AA-Briefcase v1.1 |
| 에이전트형 터미널 사용 | 15% | Terminal-Bench 4.0 |
점수
Artificial Analysis 공학 지수
Artificial Analysis 공학 지수: 역량 세부 분석
역량 세부 분석
Artificial Analysis 공학 지수: 공학 지식
대표 워크플로
공학 지수에서 가장 높은 가중치를 둔 역량을 평가하는 실제 워크플로입니다.
예시: Design and analyze a wind turbine support structure to size components against fatigue and extreme-wind load cases, justify safety margins against a governing standard such as IEC 61400, and maximize power output while reliably withstanding environmental stress.
예시: Track an intermittent CFD pipeline failure through the CMake build and Conda environment on a Slurm cluster from the terminal, then ship a fix that spares adjacent batch jobs.
예시: Read a vendor package of CAD schematics and dimensioned drawings to extract GD&T callouts, materials, and interface dimensions per ASME Y14.5, reconcile conflicts across sheets, and draft a specification that cites each source drawing.
비용
Artificial Analysis 공학 지수: 작업당 비용
Artificial Analysis 공학 지수 vs. 작업당 비용
속도
Artificial Analysis 공학 지수: 작업당 시간
출력 토큰
Artificial Analysis 공학 지수: 작업당 출력 토큰
출시일
Artificial Analysis 공학 지수 vs. 출시일
자주 묻는 질문
Artificial Analysis 공학 지수에 따르면 현재 엔지니어링 업무에서 가장 뛰어난 AI 모델은 Claude Opus 5.5 (Max, Default Fallback) (60), Claude Opus 5.5 (Xhigh, Default Fallback) (59) 및 Claude Sonnet 5.5 (Max, Default Fallback) (58)입니다. 순위는 새 모델 출시에 맞춰 업데이트됩니다.
네. Artificial Analysis의 공학 지수는 엔지니어링 업무에서 AI 모델의 성능을 측정하는 독립적인 벤치마크입니다. 공학 지식, 정량적 추론, 에이전트 실행, 터미널 사용을 평가합니다.
공학 지수는 엔지니어링 분야에서 모델의 성능을 평가하는 Artificial Analysis의 종합 벤치마크입니다. 토목, 전기, 기계 공학 관련 전문 지식과 설계 및 분석, 도구 및 자동화, 기술 문서 작성 등을 평가합니다.
공학 지수는 각 역량 하위 점수의 가중 평균으로 계산합니다. 하위 점수와 가중치는 공학 지식 (35%), 추론 (30%), 에이전트형 지식 작업 (20%) 및 에이전트형 터미널 사용 (15%)입니다.
공학 지수에는 AA-Omniscience 과학, 공학 및 수학 정확도, HLE, CritPt, GDPval-AA v2.1, AA-Briefcase v1.1 및 Terminal-Bench 4.0이 포함됩니다.
결과가 공개된 모델 중 현재 Claude Opus 5.5 (Max, Default Fallback)의 공학 지수 점수가 60로 가장 높습니다. 모델 보기
공학 지수 점수가 높을수록 지수를 구성하는 벤치마크 전반에서 더 뛰어난 성능을 보인다는 뜻입니다. 특정 사용 사례에서는 종합 점수보다 개별 벤치마크 결과가 더 유용할 수 있습니다.