Benchmarking pocket-scale inference

We benchmark small models on mobile phones. Artificial Analysis' testing covers model intelligence on a set of benchmarks chosen to represent real-world mobile device usage, and we partner with Liquid AI to gather real inference data measured on the devices themselves. Note: we have independently validated Liquid AI's inference measurement process.

“Small” models are all models that fit inside 8 GB of memory after quantization, including KV cache at 8K context. View all rules and our process in the methodology page.

Intelligence and Inference Performance Summary

Average Score (16K max context) vs. End-to-End Generation TimeiPhone 17 Pro

Simple average of 5 evaluations chosen to represent real-world mobile device usage: BFCL (subset), IFBench, AA-Omniscience, GPQA Diamond, MATH-500 · Context limited to 16K tokens · E2E time is the seconds taken to process a 1024-token prompt and generate a 256-token response
Most attractive quadrant
Pareto line

Inference Performance

End-to-End Generation TimeiPhone 17 Pro

Seconds to process a 1024-token prompt and generate a 256-token response · Lower is better

Model Intelligence

Average Score (Mobile Device Benchmark Set, 16K max context)iPhone 17 Pro

Simple average of 5 evaluations chosen to represent real-world mobile device usage: BFCL (subset), IFBench, AA-Omniscience, GPQA Diamond, MATH-500 · Context limited to 16K tokens · Higher is better
Score at 64K max context
Performance measurements omitted as the model did not fit on this device or exceeded the time limit

Looking for the Artificial Analysis Intelligence Index scores for these models? The following models have been evaluated on our full index, and their scores are visible on their model pages:

Token Efficiency

Context Budget Overruns

Generations that stopped at the 16K-token limit instead of finishing, across every evaluation run on the model · Lower is better

Evaluation Breakdown

Mobile Device Benchmark Set Evaluations (16K max context)iPhone 17 Pro

Intelligence evaluations measured independently by Artificial Analysis · Context limited to 16K tokens · Higher is better

BFCL

Tool calling (index subset)

IFBench

Instruction following

AA-Omniscience Accuracy

Knowledge

AA-Omniscience Non-Hallucination Rate

1 - hallucination rate

GPQA Diamond

Scientific reasoning

MATH-500

Quantitative reasoning