Laptop & workstation inference performance
We benchmark the local agent inference speeds that users can expect to get out of desk-scale hardware. Our official leaderboard consists of runs conducted in a controlled environment on Artificial Analysis' own hardware. In addition, our fully open-source AA-AgentPerf-Local tool enables any user to easily benchmark their own system's agentic performance.
Systems & models overview
We are currently testing on a limited set of systems and models, and expect these to expand significantly over time. A brief overview of each currently included system and model is included below.
Systems

MacBook Pro
M5 Pro · 64 GB · 20-core GPU
- Launch MSRP
- $3,000
- Memory
- 64 GB unified
- Memory bandwidth
- 307 GB/s
- Memory type
- LPDDR5X-9600
- Chip
- M5 Pro
- Launched
- Mar 2026

AMD Ryzen AI Halo
- Launch MSRP
- $4,000
- Memory
- 128 GB unified
- Memory bandwidth
- 256 GB/s
- Memory type
- LPDDR5X-8000
- Chip
- Ryzen AI Max+ 395
- Launched
- Jul 2026

NVIDIA DGX Spark
- Launch MSRP
- $4,000
- Memory
- 128 GB unified
- Memory bandwidth
- 273 GB/s
- Memory type
- LPDDR5X-8533
- Chip
- GB10
- Launched
- Oct 2025

NVIDIA GeForce RTX 5090
- Launch MSRP (card)
- $2,000
- Memory
- 32 GB VRAM
- Memory bandwidth
- 1,792 GB/s
- Memory type
- GDDR7
- Chip
- RTX 5090
- Launched
- Jan 2025
Models
Qwen3.5 9B (Reasoning)
- Intelligence Index
- 11
- Total parameters
- 9.65B
- Active parameters
- 9.65B
- Context window
- 256K
- Released
- Mar 2026
- Builds tested
- NVFP4, Q4_K_M
Qwen3.8 27B (xhigh)
- Intelligence Index
- 34
- Total parameters
- 27B
- Active parameters
- 27B
- Context window
- 256K
- Released
- Aug 2026
- Builds tested
- Q4_K_M, NVFP4
Qwen3.6 35B A3B (Reasoning)
- Intelligence Index
- 18
- Total parameters
- 36B
- Active parameters
- 3B
- Context window
- 256K
- Released
- Apr 2026
- Builds tested
- UD-Q4_K_M, NVFP4
Ling 3.0 Flash
- Intelligence Index
- 20
- Total parameters
- 124B
- Active parameters
- 5.1B
- Context window
- 256K
- Released
- Aug 2026
- Builds tested
- Q4_K_M
Results from all system & model combinations
Completion Time by Model and System
Time to complete the AgentPerf-Local default workload (168 turns, 172 tool calls), with prefill and decode speeds below · Lower is better
Qwen3.5 9B (Reasoning)9.65B parameters
15.7 min636 tok/s prefill
49 tok/s decode
49 tok/s decode
13.0 min775 tok/s prefill
58 tok/s decode
58 tok/s decode
13.0 min2,161 tok/s prefill
46 tok/s decode
46 tok/s decode
2.2 min4,980 tok/s prefill
331 tok/s decode
331 tok/s decode
Qwen3.8 27B (xhigh)27B parameters
37.6 min212 tok/s prefill
23 tok/s decode
23 tok/s decode
34.5 min261 tok/s prefill
23 tok/s decode
23 tok/s decode
24.2 min612 tok/s prefill
28 tok/s decode
28 tok/s decode
4.9 min2,156 tok/s prefill
151 tok/s decode
151 tok/s decode
Qwen3.6 35B A3B (Reasoning)36B, 3B active parameters
11.9 min718 tok/s prefill
70 tok/s decode
70 tok/s decode
11.7 min687 tok/s prefill
74 tok/s decode
74 tok/s decode
7.3 min1,308 tok/s prefill
119 tok/s decode
119 tok/s decode
2.0 min5,101 tok/s prefill
381 tok/s decode
381 tok/s decode
Ling 3.0 Flash124B, 5.1B active parameters
Does not fitin 64 GB unified
25.1 min313 tok/s prefill
35 tok/s decode
35 tok/s decode
14.9 min467 tok/s prefill
65 tok/s decode
65 tok/s decode
Does not fitin 32 GB VRAM
FasterSlowerDoes not fit in memory
Inference details
Completion TimeQwen3.8 27B (xhigh)
Time to complete the AgentPerf-Local default workload (168 turns, 172 tool calls) · Lower is better
Decode SpeedQwen3.8 27B (xhigh)
Tokens generated per second after the first token on the AgentPerf-Local default workload · Higher is better
Time to First Token at Each Context LengthQwen3.8 27B (xhigh)
Median seconds to first token of the turns in each prompt-length band, with earlier turns already cached · Lower is better
Pricing
System prices
Systems are listed at their launch MSRP by default. Customize to better represent current prices available to you.
Completion Time vs. PriceQwen3.8 27B (xhigh)
Time to complete the AgentPerf-Local default workload against the price of the system as tested · Lower time and lower price is better
Most attractive quadrant
Pareto line
Prices represent launch MSRP, with most current prices significantly higher