All articles
August 24, 2026

Intelligence at pocket scale: Benchmarking small models and mobile phones

Today we are launching intelligence and inference benchmarking for models that fit on mobile phones. Intelligence is measured by a set of five evaluations selected to represent the types of function/tool calling, knowledge and reasoning tasks small models are often asked to accomplish. Our mobile phone inference benchmarking, operated in partnership with Liquid AI, measures speed, latency, and memory use on real phones in a standardized testing environment. Both the intelligence and inference benchmarks use the same builds, quantized to 4-bit or smaller, representing the precision at which small models are realistically run on mobile phones.

This is the original launch article and shows results as of launch. For live results across the latest models and devices, see the mobile phones page.

Why small models are harder to benchmark

Models under a few billion parameters can now follow instructions, call tools, and answer questions on mobile phones. But available benchmark results often describe neither the quantized model artifact nor the device and runtime combination a user will run. Some of the issues that add to the difficulty of benchmarking this space:

Frontier evaluations stop discriminating

Most headline benchmarks were built to separate frontier models, and small models cluster near the floor on them. Evaluations of large-scale agentic work, such as GDPval-AA v2 and Terminal-Bench v2.1, set long multi-step tasks that small models rarely finish at all, so they return near-zero scores that say little about how one small model compares to another. Benchmarking small models requires evaluations that produce meaningful score differences within this class.

The inference stack is younger

Mobile device runtimes are less mature than their datacenter-scale counterparts, and configuration details can change evaluation results: chat templates that do not render tool calls, parsers that drop parallel tool calls, and server flags that silently divide the context window.

Devices are moving targets

On-device performance depends on thermals, power modes, cooling, and how each operating system accounts for memory. A set of results is only interpretable in the context of its exact combination of model, quantization, runtime, flags, and level of physical temperature control.

One model name covers many builds

The same weights ship in several quantization formats, and the memory a build needs depends on its quantization, architecture, context length, and KV cache. These quantized model versions can be officially released by the model labs, or created by third-parties, with varying levels of transparency.

What we're launching

This launch consists of intelligence benchmarking, administered by Artificial Analysis, and inference benchmarking, developed and run by Liquid AI. We test intelligence and inference on the same build, quantized to 4-bit or smaller, that a device would load, served with llama.cpp, ensuring each side is measured under as similar conditions as possible. A model qualifies for testing if it fits inside 8 GB of memory after quantization, including the KV cache required at 8K context.

Mobile Device Benchmark Set

Intelligence is measured by a set of five evaluations, chosen to represent the range of capabilities a phone-scale model needs: instruction following, tool calling, knowledge/recall, and reasoning.

BFCL

Tool calling, on a 640-task subset across three categories: picking the right tool, multi-turn tool use, and declining when no tool fits.

IFBench

Instruction following with precise, verifiable output constraints.

AA-Omniscience

Knowledge accuracy and hallucination resistance, split evenly between accuracy and non-hallucination.

GPQA Diamond

Graduate-level scientific reasoning, multiple choice format.

MATH-500

Competition mathematics with symbolic answer checking.

Full details on our benchmarking process can be found on the Mobile Device Benchmark Set methodology page.

Mobile Phone Inference

Liquid AI has developed Pipette, a mobile device inference benchmarking product, and open-sourced it on GitHub. We have independently validated Liquid AI's inference benchmarking methodology, which includes a standardized testing harness and a climate-controlled physical testing facility for repeatable results. Read more about Liquid AI's methodology on the Pipette documentation.

Launch results for the iPhone 17 Pro

Nanbeige4.2-3B and LFM2.5-2.6B share the top average evaluation scores (with a 16K context limit) at launch, both scoring 63, ahead of Ornith-1.0-9B on 62 and Qwen3.5 9B (Reasoning) on 61. LFM2.5-2.6B achieves its score more efficiently: on an iPhone 17 Pro it answers a standard 1,024-token prompt in 8.0 seconds using 2.3 GB of memory, against 21.4 seconds and 4.0 GB for Nanbeige4.2-3B and more than 25 seconds and 6.9 GB for the two 9B models. The benchmark set scores 41 quantized builds, 33 of which successfully ran on the iPhone 17 Pro.

  • The speed-intelligence Pareto frontier is short. Six models are unbeaten on both axes on an iPhone 17 Pro: LFM2.5-230M (27 at 0.9 seconds), MiniCPM5-1B (45 at 2.9 seconds), LFM2.5-8B-A1B (58 at 5.7 seconds), Ling 3.0 Tiny (59 at 5.7 seconds), LFM2.5-2.6B (63 at 8.0 seconds) and Nanbeige4.2-3B (63 at 21.4 seconds). LFM2.5-8B-A1B and Ling 3.0 Tiny are mixture-of-experts models that activate around a billion parameters per token, which is how they answer in under six seconds with 8B-class weights.
  • The 16K context limit defines the leaderboard. Models like Qwen3.5 9B are highly intelligent but spend far more output tokens on evaluation tasks than is practical on a mobile phone: 74.5M tokens across one pass of the benchmark set, against 5.2M for Gemma 4 E4B (Reasoning), which scores within two points of it overall. Qwen3.5 9B (Reasoning) hits the 16K context limit on 29% of its generations, which reduces its scores. With the limit raised to 64K, Ling 3.0 Tiny takes first place with a score of 66, ahead of Nanbeige4.2-3B (65), with LFM2.5-2.6B and Qwen3.5 9B (Reasoning) both on 64. But a 64K window does not fit in mobile phone memory, and at 55 output tokens per second on an iPhone 17 Pro, generating 64K output tokens could mean a wait of more than twenty minutes and substantial battery use. This is why our primary results are capped at 16K, but we also publish results capped at 64K and results capped at one minute of generation time.
  • A one-minute answer budget inverts the order completely. Capping each answer at what a model can generate in 60 seconds on the iPhone 17 Pro (its output speed × 60, ignoring prompt processing time) puts LFM2.5-8B-A1B first on 47, followed by LFM2-2.6B-Exp on 45 and Gemma 4 E4B (Non-reasoning) on 44. LFM2.5-2.6B falls to 38, and Nanbeige4.2-3B to 18: at 14 output tokens per second it gets about 850 tokens per answer, so most of its reasoning never finishes. Qwen3.5 9B (Reasoning) scores 14. While we aren't using the one-minute limit as the primary filter for this benchmark data, it may be more in line with the maximum wait time mobile users are willing to tolerate.
  • The leading models have opposite strengths. Nanbeige4.2-3B is the most balanced: within a point of the best on BFCL (76%), third on MATH-500 (96%) and sixth on GPQA Diamond (67%). Qwen3.5 9B (Non-reasoning) is the strongest tool caller (77% on BFCL) and the strongest scientific reasoner (79% on GPQA Diamond), and Ornith-1.0-9B recalls the most facts (15% accuracy on AA-Omniscience). LFM2.5-2.6B follows instructions best of any model measured on the iPhone (59% on IFBench), clears 90% on MATH-500, and is far readier to decline than to guess: 79% non-hallucination against 33% for Nanbeige4.2-3B and 24% and 1% for the two Qwen3.5 9B variants. That component is what lifts it level with Nanbeige4.2-3B; on the other four evaluations alone, Nanbeige4.2-3B, Ornith-1.0-9B and Qwen3.5 9B (Reasoning) would all rank above it.
  • Generation time spans a factor of 30, from 0.9 seconds for LFM2.5-230M to 26.7 for Falcon-H1R-7B, measured end to end on an identical 1,024-token prompt and 256-token response.
  • Memory needs span an order of magnitude. Peak in-test memory at 4K context runs from 0.4 GB for the smallest models to 6.9 GB for Ornith-1.0-9B and Qwen3.5 9B, the heaviest here. On a 12 GB phone, this top end leaves limited room for the operating system and other apps.

The inference results below are measured on an iPhone 17 Pro. Results for the other devices we benchmark are on the mobile phones page, which is updated as new models and devices are added. Intelligence results are device-agnostic; every model is evaluated as the same build, quantized to 4-bit or smaller, that a device would load.

Average Score (16K max context) vs. End-to-End Generation TimeiPhone 17 Pro

Simple average of 5 evaluations chosen to represent real-world mobile device usage: BFCL (subset), IFBench, AA-Omniscience, GPQA Diamond, MATH-500 · Context limited to 16K tokens · E2E time is the seconds taken to process a 1024-token prompt and generate a 256-token response
Most attractive quadrant
Pareto line

A simple average of five evaluations chosen to represent real-world mobile device usage, run against models small enough to fit on portable hardware and each measured independently by Artificial Analysis. See the methodology for further details.

Total wall-clock time to process a 1,024-token prompt and generate a 256-token response. Note that this is fundamental to the hardware used and the model’s architecture, and does not include the effect of model verbosity or tendency to use more or fewer turns.

End-to-End Generation TimeiPhone 17 Pro

Seconds to process a 1024-token prompt and generate a 256-token response · Lower is better

Total wall-clock time to process a 1,024-token prompt and generate a 256-token response. Note that this is fundamental to the hardware used and the model’s architecture, and does not include the effect of model verbosity or tendency to use more or fewer turns.

Average Score (Mobile Device Benchmark Set, 16K max context)iPhone 17 Pro

Simple average of 5 evaluations chosen to represent real-world mobile device usage: BFCL (subset), IFBench, AA-Omniscience, GPQA Diamond, MATH-500 · Context limited to 16K tokens · Higher is better
Score at 64K max context

A simple average of five evaluations chosen to represent real-world mobile device usage, run against models small enough to fit on portable hardware and each measured independently by Artificial Analysis. See the methodology for further details.

Context Budget Overruns

Generations that stopped at the 16K-token limit instead of finishing, across every evaluation run on the model · Lower is better

Mobile Device Benchmark Set Evaluations (16K max context)

Intelligence evaluations measured independently by Artificial Analysis · Context limited to 16K tokens · Higher is better

BFCL

Tool calling (index subset)

IFBench

Instruction following

AA-Omniscience Accuracy

Knowledge

AA-Omniscience Non-Hallucination Rate

1 - hallucination rate

GPQA Diamond

Scientific reasoning

MATH-500

Quantitative reasoning

Next steps

We plan to expand both the intelligence and on-device inference benchmarks. Next steps include:

  • Comparing models by memory footprint rather than a single quantization, so a heavily quantized large model and a high-precision small model compete on equal terms
  • Adding new models as they are released
  • Incorporating results from more inference frameworks beyond llama.cpp, and more devices
  • Investigating evaluations purpose-built for small models and on-device use

Explore further