所有发布•

13 个模型

IBM 模型:智能、性能与价格

Artificial Analysis 已对 IBM 的 13 个模型进行基准测试。下方对这些模型的关键指标进行了比较。

  • 智能方面,IBM 得分最高的模型是 Granite 4.2 30B(15,估算值)。
  • 输出速度方面,最快的模型是 Granite 4.2 3B(216 token/秒)。
  • 延迟方面,Granite 4.2 3B(9.77 秒) 的首个答案 token 等待时间最短。
  • 价格方面,Granite 4.2 3B($0.01) 的单任务成本最低。不同模型的价格差异最高可达 3.9 倍。

智能

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Index vs. Cost per Intelligence Index Task

Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task
Most attractive quadrant
Pareto line

成本

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better

速度与延迟

Output Speed

Output tokens per second · Higher is better

能力得分

能力指数

衡量模型在特定能力与行业中的表现
Finance & Accounting Index

Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better

Strategy & Ops Index

Incorporates 6 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better

Legal Index

Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AA-LCR v1.1, GDP.pdf, AutomationBench-AA · Higher is better

Healthcare & Medical Index

Incorporates 6 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, MLCR-AA, Humanity's Last Exam, AutomationBench-AA · Higher is better

No data available
Engineering Index

Incorporates 6 evaluations: AA-Omniscience, Humanity's Last Exam, CritPt, GDPval-AA v2.1, AA-Briefcase v1.1, Terminal-Bench 4.0 · Higher is better

Economics Index

Incorporates 5 evaluations: AA-Omniscience, Humanity's Last Exam, GDPval-AA v2.1, AA-Briefcase v1.1, AA-LCR v1.1 · Higher is better

IBM 的所有发布

更多详情

权重
服务商基准测试
Granite 4.2 30B
IBM 标志IBM
15
30B
131k
US$0.1
77
DeepInfra
Granite 4.2 8B
IBM 标志IBM
11
8B
131k
US$0.0
76
CoreWeaveDeepInfra
Granite 4.2 3B
IBM 标志IBM
9
3B
131k
US$0.0
216
DeepInfra
Granite 4.1 30B
IBM 标志IBM
7
30B
131k
-
-
-
Granite 4.1 8B
IBM 标志IBM
7
8B
131k
US$0.1
126
CoreWeave
Granite 4.0 H Small
IBM 标志IBM
6
32B
推理时启用 9B 个参数
128k
US$0.1
15
Replicate
Granite 4.1 3B
IBM 标志IBM
6
3B
131k
-
-
-
Granite 4.0 H 1B
IBM 标志IBM
5
1.5B
128k
-
-
-
Granite 4.0 Micro
IBM 标志IBM
5
3B
128k
-
-
-
Granite 4.0 1B
IBM 标志IBM
5
1.6B
128k
-
-
-
Granite 3.3 8B (Non-reasoning)
IBM 标志IBM
5
8.2B
128k
US$0.1
15
Replicate
Granite 4.0 350M
IBM 标志IBM
5
0.3B
33k
-
-
-
Granite 4.0 H 350M
IBM 标志IBM
5
0.3B
33k
-
-
-