AA-AgentPerf: The Hardware Benchmark for the Agent Era

AA-AgentPerf has been shaped by our work with inference providers and engagement with AI accelerator companies, developers, and enterprise buyers over the past year. It tests inference hardware and software by replaying real coding agent trajectories, measuring how well each system handles long context, tool-calling delays, substantial KV cache reuse opportunities, and other features of agentic work.

AA-AgentPerf is open for submissions, and results will be published on a rolling basis.

View all serving configurations →

Total Throughput per MW vs. Output Speed

Total server throughput (input + output tokens per second) per megawatt of accelerator power vs. p25 output speed per request. Input tokens include KV-cache hits. Output speed excludes time to first token.
Serving configuration changes

Inside a single agent trajectory

Each simulated agent replays a real multi-turn coding session — reasoning, calling tools, and editing code. Context compounds turn over turn: prior turns are served from KV cache, only new tokens are prefilled, and trajectories grow beyond 100K tokens. This access pattern is what separates agent serving from synthetic benchmarks.