Allen Institute for AI 模型:智能、性能与价格
Artificial Analysis Intelligence Index
* 估计值
每项 Intelligence Index 任务
Artificial Analysis 已对 Allen Institute for AI 的 10 个模型进行基准测试。下方对这些模型的关键指标进行了比较。
- 智能方面,Allen Institute for AI 得分最高的模型是 Olmo 3.1 32B Think(7,估算值)。
智能
Artificial Analysis Intelligence Index
Intelligence Index vs. Cost per Intelligence Index Task
成本
Cost per Intelligence Index Task
速度与延迟
Output Speed
能力得分
能力指数
Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better
Incorporates 6 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better
Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AA-LCR v1.1, GDP.pdf, AutomationBench-AA · Higher is better
Incorporates 6 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, MLCR-AA, Humanity's Last Exam, AutomationBench-AA · Higher is better
Incorporates 6 evaluations: AA-Omniscience, Humanity's Last Exam, CritPt, GDPval-AA v2.1, AA-Briefcase v1.1, Terminal-Bench 4.0 · Higher is better
Incorporates 5 evaluations: AA-Omniscience, Humanity's Last Exam, GDPval-AA v2.1, AA-Briefcase v1.1, AA-LCR v1.1 · Higher is better
Allen Institute for AI 的所有发布
更多详情
权重 | 服务商基准测试 | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Llama 3.1 Tulu3 405B | 7 | 405B | 128k | - | - | - | |||
| Olmo 3.1 32B Think | 7 | 32.2B | 66k | US$0.0 | - | ||||
| Olmo 3.1 32B Instruct | 6 | 32.2B | 66k | - | - | - | |||
| Olmo 3 32B Think | 6 | 32.2B | 66k | - | - | - | |||
| OLMo 2 32B | 6 | 32.2B | 4k | - | - | - | |||
| Olmo 3 7B Think | 6 | 7B | 66k | - | - | - | |||
| OLMo 2 7B | 6 | 7.3B | 4k | - | - | - | |||
| Molmo 7B-D | 6 | 8.0B | 4k | - | - | - | |||
| Olmo 3 7B Instruct | 5 | 7B | 66k | US$0.1 | - | ||||
| Molmo2-8B | 5 | 8.7B | 37k | - | - | - |