All articles

September 29, 2026

AA-AgentPerf-Local: Benchmarking local AI agents on laptops and workstations

Announcing AA-AgentPerf-Local, our open-source inference testing tool for local AI models - test how fast agentic AI can run on your own laptop or workstation, and browse our list of serving configurations to plan your next agent setup

Key points:

➤ We’re open sourcing AA-AgentPerf-Local, which replays real agent trajectories on laptop & workstation hardware to test inference performance

➤ We’re releasing initial results for NVIDIA DGX Spark, NVIDIA GeForce RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro

➤ The tool and leaderboard will soon expand to cover more hardware, frameworks and models, and will stay updated over time as new releases launch

We already benchmark inference performance on mobile phones and datacenter-scale hardware, and are expanding our coverage to include laptops and workstations. Alongside hosting a leaderboard displaying results from popular model and hardware combinations, we are open-sourcing all code and data required to run AA-AgentPerf-Local to ensure that individuals and companies are able to run their own trials, informing their local AI serving decisions.

AA-AgentPerf-Local replays real agent sessions. Our default workload is 8 recorded agentic tasks, spanning 168 model turns. Each request carries the full conversation so far, as a real agent's would, so context grows to ~56K tokens. Every turn generates exactly its recorded number of tokens, so every system does identical work. Tool execution is skipped by default to isolate inference speed, but when benchmarking your own system, you can also replay the real recorded tool delays or run tool calls live on your CPU.

At launch, we are focusing on the performance a single agent can achieve when able to utilize the entire system; we plan to expand this coverage over time to cover multi-agent systems and scenarios where an agent must run alongside other regular processes.

The initial set of hardware covered on our official page is: NVIDIA DGX Spark (128 GB), AMD Ryzen AI Halo (128 GB), MacBook Pro M5 Pro (64 GB), and NVIDIA GeForce RTX 5090 (32 GB). These have been selected to cover a range of platforms (CUDA, ROCm, Vulkan, Metal), memory capacities, and bandwidths. We will be expanding the featured hardware to include x86 (and other) laptops, AI-focused graphics cards such as the RTX PRO 6000 Blackwell, and more, to give consumers a well-rounded view of performance across different hardware types.

The models featured at launch are: Qwen3.5-9B, Qwen3.8-27B, Qwen3.6-35B-A3B, and Ling 3.0 Flash (124B / 5B active). These initial models span a range of memory requirements and dense/MoE architectures, and are each benchmarked at 4-bit quantizations to reflect realistic serving conditions. The core set of models we feature will shift over time as new open weights models are released. Beyond the featured models, AA-AgentPerf-Local is able to benchmark performance of any OpenAI-compatible inference server the user runs, meaning that any model and config can be tested locally.

Choice of serving configuration is important, with the runtime, quantization and speculative decoding changing results substantially. Where an official off-the-shelf config was available for a system and model pair, we used the published config. We developed our own configs for all other cases. Every config uses speculative decoding (MTP, DFlash or DSpark), and all 14 are published in the repo and on the configs page on our website.

Initial results:

➤ Completion time mapped most closely to each model’s active parameter count: Qwen3.6-35B-A3B (3B active) was the fastest model on every system, e.g. 2.5-3.3x faster than the dense Qwen3.8-27B. However, active parameters are not the whole story, with Ling 3.0 Flash (124B total, 5B active) still finishing behind Qwen3.5-9B (nearly double the active parameter count) on all hardware that can support it.

➤ The GeForce RTX 5090 was the fastest system for every model that fits in its 32 GB, achieving completion times >3.5x faster than the other systems. Single-user decoding is heavily influenced by memory bandwidth, and the GeForce RTX 5090 has 1,792 GB/s against 256-307 GB/s for the unified-memory systems.

➤ The DGX Spark and Ryzen AI Halo are overall similar systems, with the same amount of unified memory, comparable memory bandwidth, and the same launch MSRP of $4,000. On our default trajectory set, the DGX Spark was 1.4-1.7x faster on three of four models, with the Ryzen AI Halo tying it on Qwen3.5-9B. The gap is far larger than their 7% bandwidth difference - the Spark has greater low-precision compute than the Ryzen AI Halo, enabling it to prefill faster, and its more mature CUDA software likely plays a role in enabling better MoE and speculative decoding performance.

➤ The MacBook Pro (M5 Pro, 64 GB, 20-core GPU) is the only laptop tested so far, and exhibited competitive results, finishing within 2–9% of the Ryzen AI Halo on Qwen3.6-35B-A3B and Qwen3.8-27B (though 21% slower on Qwen3.5-9B). It has the most memory bandwidth of the three unified-memory systems (307 GB/s) and its current price of $3,700 is the lowest of the systems tested so far. Its results are likely held back by software maturity and compute available for prefill.

➤ Despite the agentic trajectories serving 73-93% of prompt tokens from the KV cache, simply reading each turn’s new input (during prefill) used up substantial proportions of the end-to-end completion time, e.g. 22-41% for Qwen3.8-27B. This was especially impactful on systems with low compute FLOP/s relative to their memory bandwidth, such as the Ryzen AI Halo and potentially the MacBook Pro (MacBook FLOP/s are unpublished).

➤ Speculative decoding was implemented on all of the most successful configs so far, e.g. raising Qwen3.8-27B decode speeds ~30-120% above the bandwidth-constraint roofline.

Most hardware prices are significantly inflated vs. their launch MSRP. At current market prices, each of the initial four hardware types increase in price as they decrease in end-to-end completion time, with the MacBook Pro (M5 Pro, 64 GB) the current cheapest and the GeForce RTX 5090 system by far the most expensive. Our page defaults to launch MSRP but allows custom price entry. We’re showing below our best estimate of current market prices.

Decode speeds for Qwen3.8-27B were broadly similar on the three unified-memory devices, which have similar memory bandwidths at 256-307 GB/s. The GeForce RTX 5090 exhibited >5x faster decode speeds than these, owing mostly to its >5x faster memory, at 1,792 GB/s bandwidth.

Prefill speeds for Qwen3.8-27B followed a similar pattern to decode, but the DGX Spark’s disproportionately high compute/bandwidth ratio showed up in faster context reading speeds. This is important in our agentic trajectories as, even with most of each prompt served from cache (~93% on llama.cpp), each session still has ~196K new input tokens against ~31K output tokens. When a tool returns a large file, the agent can pause up to 24 seconds on the Halo and 39 seconds on the MacBook before replying, against about 2 seconds on the GeForce RTX 5090.

AA-AgentPerf-Local and the Laptops & Workstations page will be frequently updated as relevant new hardware and software becomes available. We are planning to implement a range of additional features to better help users run local AI optimally:

➤ A wider range of popular hardware and models

➤ A more comprehensive repository of configs for local AI deployment

➤ More custom inference frameworks, such as Inco Splash, which is currently being tested

➤ User-submitted result leaderboards

➤ Leaderboards featuring live CPU tool-calling

➤ Multi-agent scenarios

Try AA-AgentPerf-Local on your own hardware: https://github.com/ArtificialAnalysis/aa-agentperf-local

View the initial results here: https://artificialanalysis.ai/hardware-inference-stack/laptops-workstations and every serving configuration here: https://artificialanalysis.ai/hardware-inference-stack/laptops-workstations/configs