Benchmarking pocket-scale inference
We benchmark small models on mobile phones. Artificial Analysis' testing covers model intelligence on a set of benchmarks chosen to represent real-world mobile device usage, and we partner with Liquid AI to gather real inference data measured on the devices themselves. Note: we have independently validated Liquid AI's inference measurement process.
“Small” models are all models that fit inside 8 GB of memory after quantization, including KV cache at 8K context. View all rules and our process in the methodology page.
Intelligence and Inference Performance Summary
Average Score (16K max context) vs. End-to-End Generation TimeiPhone 17 Pro
Inference Performance
End-to-End Generation TimeiPhone 17 Pro
Model Intelligence
Average Score (Mobile Device Benchmark Set, 16K max context)iPhone 17 Pro
Looking for the Artificial Analysis Intelligence Index scores for these models? The following models have been evaluated on our full index, and their scores are visible on their model pages:
Token Efficiency
Context Budget Overruns
Evaluation Breakdown
Mobile Device Benchmark Set Evaluations (16K max context)iPhone 17 Pro
BFCL
Tool calling (index subset)
IFBench
Instruction following
AA-Omniscience Accuracy
Knowledge
AA-Omniscience Non-Hallucination Rate
1 - hallucination rate
GPQA Diamond
Scientific reasoning
MATH-500
Quantitative reasoning