Allen Institute for AIのモデル:知能・性能・料金
Artificial Analysis Intelligence Index
* 推定値
Intelligence Indexのタスクあたり
Artificial AnalysisはAllen Institute for AIの10個のモデルをベンチマークしています。以下では、これらのモデルの主要指標を比較します。
- 知能が最も高いAllen Institute for AIのモデルはOlmo 3.1 32B Think(7、推定)です。
知能
Artificial Analysis Intelligence Index
Intelligence Index vs. Cost per Intelligence Index Task
費用
Cost per Intelligence Index Task
速度と遅延
Output Speed
能力スコア
能力指数
Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better
Incorporates 6 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better
Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AA-LCR v1.1, GDP.pdf, AutomationBench-AA · Higher is better
Incorporates 6 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, MLCR-AA, Humanity's Last Exam, AutomationBench-AA · Higher is better
Incorporates 6 evaluations: AA-Omniscience, Humanity's Last Exam, CritPt, GDPval-AA v2.1, AA-Briefcase v1.1, Terminal-Bench 4.0 · Higher is better
Incorporates 5 evaluations: AA-Omniscience, Humanity's Last Exam, GDPval-AA v2.1, AA-Briefcase v1.1, AA-LCR v1.1 · Higher is better
Allen Institute for AIのすべてのリリース
詳細
ウェイト | プロバイダーのベンチマーク | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Llama 3.1 Tulu3 405B | 7 | 405B | 128k | - | - | - | |||
| Olmo 3.1 32B Think | 7 | 32.2B | 66k | $0.0 | - | ||||
| Olmo 3.1 32B Instruct | 6 | 32.2B | 66k | - | - | - | |||
| Olmo 3 32B Think | 6 | 32.2B | 66k | - | - | - | |||
| OLMo 2 32B | 6 | 32.2B | 4k | - | - | - | |||
| Olmo 3 7B Think | 6 | 7B | 66k | - | - | - | |||
| OLMo 2 7B | 6 | 7.3B | 4k | - | - | - | |||
| Molmo 7B-D | 6 | 8.0B | 4k | - | - | - | |||
| Olmo 3 7B Instruct | 5 | 7B | 66k | $0.1 | - | ||||
| Molmo2-8B | 5 | 8.7B | 37k | - | - | - |