All Releases•

22 models

Meta Models: Intelligence, Performance & Price

Artificial Analysis has benchmarked 22 models from Meta. Below is a comparison of the key metrics across these models.

  • For intelligence, the top model from Meta is Muse Spark 1.3 (max) at 48.
  • For output speed, the fastest model is Muse Spark 1.3 (xhigh) at 221 t/s.
  • For latency, Llama 4 Scout at 0.84s offers the lowest time to first answer token.
  • For pricing, Muse Glimmer (high) at $0.06 offers the lowest cost per task. Prices vary up to 28x across models.

Intelligence

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Index vs. Cost per Intelligence Index Task

Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task
Most attractive quadrant
Pareto line

Cost

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better

Speed & Latency

Output Speed

Output tokens per second · Higher is better

Capability Scores

Capability Indexes

Measures the performance of models on specific capabilities and industries
Finance & Accounting Index

Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better

Strategy & Ops Index

Incorporates 6 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better

Legal Index

Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AA-LCR v1.1, GDP.pdf, AutomationBench-AA · Higher is better

Healthcare & Medical Index

Incorporates 6 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, MLCR-AA, Humanity's Last Exam, AutomationBench-AA · Higher is better

Engineering Index

Incorporates 6 evaluations: AA-Omniscience, Humanity's Last Exam, CritPt, GDPval-AA v2.1, AA-Briefcase v1.1, Terminal-Bench 4.0 · Higher is better

Economics Index

Incorporates 5 evaluations: AA-Omniscience, Humanity's Last Exam, GDPval-AA v2.1, AA-Briefcase v1.1, AA-LCR v1.1 · Higher is better

All Meta Releases

Further details

Weights
Provider Benchmarks
Muse Spark 1.3 (max)
Meta logoMeta
48
-
1M
$0.8
203
Not available
Meta
Muse Spark 1.3 (xhigh)
Meta logoMeta
45
-
1M
$0.8
221
Not available
Meta
Muse Spark 1.2 (xhigh)
Meta logoMeta
40
-
1M
$0.8
234
Not available
Meta
Muse Spark 1.1 (xhigh)
Meta logoMeta
34
-
1M
$0.8
-
Not available
Meta
Muse Spark
Meta logoMeta
31
-
262k
-
-
Not available
-
Muse Glimmer (high)
Meta logoMeta
17
30B
131k
$0.2
122
Together AIDeepInfraSystalyzeFireworks
Llama 4 Maverick
Meta logoMeta
10
402B
17B active at inference time
1M
$0.3
110
NovitaMicrosoft AzureAmazon Bedrock
+3
Llama 4 Scout
Meta logoMeta
8
109B
17B active at inference time
10M
$0.2
107
DeepInfraMicrosoft AzureGoogle
+3
Llama 3.3 Instruct 70B
Meta logoMeta
8
70B
128k
$0.7
90
FireworksGoogleGroq
+10
Llama 3.1 Instruct 405B
Meta logoMeta
7
405B
128k
-
-
-
Llama 3.1 Instruct 8B
Meta logoMeta
7
8B
128k
$0.0
136
GroqNovitaCoreWeave
+3
Llama 3.1 Instruct 70B
Meta logoMeta
7
70B
128k
$0.6
52
Amazon BedrockDeepInfra
Llama 3.2 Instruct 90B (Vision)
Meta logoMeta
6
90B
128k
-
-
-
Llama 2 Chat 7B
Meta logoMeta
6
7B
4k
$0.1
-
Replicate
Llama 3.2 Instruct 3B
Meta logoMeta
6
3B
128k
-
-
-
Llama 3 Instruct 70B
Meta logoMeta
5
70B
8k
$0.9
-
NovitaAmazon BedrockReplicate
Llama 3.2 Instruct 11B (Vision)
Meta logoMeta
5
11B
128k
$0.3
17
DeepInfra
Llama 2 Chat 70B
Meta logoMeta
5
70B
4k
-
-
-
Llama 2 Chat 13B
Meta logoMeta
5
13B
4k
-
-
-
Llama 65B
Meta logoMeta
5
65B
2k
-
-
-
Llama 3.2 Instruct 1B
Meta logoMeta
5
1B
128k
-
-
-
Llama 3 Instruct 8B
Meta logoMeta
5
8B
8k
$0.1
-
NovitaAmazon BedrockReplicateDeepInfra