Datacenter Inference Systems

Inference Serving Configurations

Every serving configuration used in AA-AgentPerf hardware benchmarking, from June 2026 onwards. Each entry captures the model, accelerator system, precision, inference framework, parallelism topology, and full launch command.

Read the AA-AgentPerf methodology.

AA-AgentPerf

Serving configurations from AA-AgentPerf benchmark runs, with the full launch command for each. Read the methodology

Showing 47 of 47 configurations

Kimi K3 (max) — 13 configurations

GLM-5.2 (max) — 15 configurations

DeepSeek V4 Pro 0813 (max) — 10 configurations

DeepSeek V4 Pro (max) — 6 configurations

gpt-oss-120b (high) — 3 configurations