Laptop & workstation inference performance

We benchmark the local agent inference speeds that users can expect to get out of desk-scale hardware. Our official leaderboard consists of runs conducted in a controlled environment on Artificial Analysis' own hardware. In addition, our fully open-source AA-AgentPerf-Local tool enables any user to easily benchmark their own system's agentic performance.

GitHubRun your own local AI benchmarks using AA-AgentPerf-Local

Systems & models overview

We are currently testing on a limited set of systems and models, and expect these to expand significantly over time. A brief overview of each currently included system and model is included below.

Systems

Drawing of the MacBook Pro
Apple

MacBook Pro

M5 Pro · 64 GB · 20-core GPU

Launch MSRP
$3,000
Memory
64 GB unified
Memory bandwidth
307 GB/s
Memory type
LPDDR5X-9600
Chip
M5 Pro
Launched
Mar 2026
Drawing of the AMD Ryzen AI Halo
AMD

AMD Ryzen AI Halo

Launch MSRP
$4,000
Memory
128 GB unified
Memory bandwidth
256 GB/s
Memory type
LPDDR5X-8000
Chip
Ryzen AI Max+ 395
Launched
Jul 2026
Drawing of the NVIDIA DGX Spark
NVIDIA

NVIDIA DGX Spark

Launch MSRP
$4,000
Memory
128 GB unified
Memory bandwidth
273 GB/s
Memory type
LPDDR5X-8533
Chip
GB10
Launched
Oct 2025
Drawing of the NVIDIA GeForce RTX 5090
NVIDIA

NVIDIA GeForce RTX 5090

Launch MSRP (card)
$2,000
Memory
32 GB VRAM
Memory bandwidth
1,792 GB/s
Memory type
GDDR7
Chip
RTX 5090
Launched
Jan 2025

Models

Alibaba

Qwen3.5 9B (Reasoning)

Intelligence Index
11
Total parameters
9.65B
Active parameters
9.65B
Context window
256K
Released
Mar 2026
Builds tested
NVFP4, Q4_K_M
Alibaba

Qwen3.8 27B (xhigh)

Intelligence Index
34
Total parameters
27B
Active parameters
27B
Context window
256K
Released
Aug 2026
Builds tested
Q4_K_M, NVFP4
Alibaba

Qwen3.6 35B A3B (Reasoning)

Intelligence Index
18
Total parameters
36B
Active parameters
3B
Context window
256K
Released
Apr 2026
Builds tested
UD-Q4_K_M, NVFP4
InclusionAI

Ling 3.0 Flash

Intelligence Index
20
Total parameters
124B
Active parameters
5.1B
Context window
256K
Released
Aug 2026
Builds tested
Q4_K_M

Results from all system & model combinations

Completion Time by Model and System

Time to complete the AgentPerf-Local default workload (168 turns, 172 tool calls), with prefill and decode speeds below · Lower is better
AppleMacBook Pro (M5 Pro, 64 GB)64 GB unified
AMDAMD Ryzen AI Halo128 GB unified
NVIDIANVIDIA DGX Spark128 GB unified
NVIDIANVIDIA GeForce RTX 509032 GB VRAM
Alibaba
Qwen3.5 9B (Reasoning)9.65B parameters
15.7 min636 tok/s prefill
49 tok/s decode
13.0 min775 tok/s prefill
58 tok/s decode
13.0 min2,161 tok/s prefill
46 tok/s decode
2.2 min4,980 tok/s prefill
331 tok/s decode
Alibaba
Qwen3.8 27B (xhigh)27B parameters
37.6 min212 tok/s prefill
23 tok/s decode
34.5 min261 tok/s prefill
23 tok/s decode
24.2 min612 tok/s prefill
28 tok/s decode
4.9 min2,156 tok/s prefill
151 tok/s decode
Alibaba
Qwen3.6 35B A3B (Reasoning)36B, 3B active parameters
11.9 min718 tok/s prefill
70 tok/s decode
11.7 min687 tok/s prefill
74 tok/s decode
7.3 min1,308 tok/s prefill
119 tok/s decode
2.0 min5,101 tok/s prefill
381 tok/s decode
InclusionAI
Ling 3.0 Flash124B, 5.1B active parameters
Does not fitin 64 GB unified
25.1 min313 tok/s prefill
35 tok/s decode
14.9 min467 tok/s prefill
65 tok/s decode
Does not fitin 32 GB VRAM
FasterSlowerDoes not fit in memory

Inference details

Completion TimeQwen3.8 27B (xhigh)

Time to complete the AgentPerf-Local default workload (168 turns, 172 tool calls) · Lower is better

Decode SpeedQwen3.8 27B (xhigh)

Tokens generated per second after the first token on the AgentPerf-Local default workload · Higher is better

Time to First Token at Each Context LengthQwen3.8 27B (xhigh)

Median seconds to first token of the turns in each prompt-length band, with earlier turns already cached · Lower is better

Pricing

System prices

Systems are listed at their launch MSRP by default. Customize to better represent current prices available to you.

AppleMacBook Pro
AMDAMD Ryzen AI Halo
NVIDIANVIDIA DGX Spark
NVIDIANVIDIA GeForce RTX 5090

Completion Time vs. PriceQwen3.8 27B (xhigh)

Time to complete the AgentPerf-Local default workload against the price of the system as tested · Lower time and lower price is better
Most attractive quadrant
Pareto line
Prices represent launch MSRP, with most current prices significantly higher