Laptop & Workstation Inference Systems
Laptop & Workstation Serving Configurations
We are building a library of the best serving configurations for popular models on laptop and workstation hardware. Each configuration below produced a result on the laptops & workstations page, on the AgentPerf-Local default workload.
MacBook Pro (M5 Pro, 64 GB)

M5 Pro · 64 GB · 20-core GPU
- Launch MSRP
- $3,000
- Memory
- 64 GB unified
- Memory bandwidth
- 307 GB/s
- Memory type
- LPDDR5X-9600
- Chip
- M5 Pro
- Launched
- Mar 2026
Qwen3.5 9B (Reasoning)
Launch command
llama-server \ --model "$MODEL_DIR"/Qwen3.5-9B-Q4_K_M.gguf \ --alias qwen35-9b-q4-k-m-mtp-m5-pro-8336903315cb89ff \ --host 127.0.0.1 \ --port 8080 \ --ctx-size 65536 \ --parallel 1 \ --gpu-layers all \ --jinja \ --batch-size 2048 \ --ubatch-size 512 \ --no-mmproj \ --device MTL0 \ --load-mode mmap \ --threads 6 \ --flash-attn on \ --fit off \ --cache-ram 0 \ --spec-type draft-mtp \ --spec-draft-n-max 2 \ --spec-draft-backend-sampling \ --no-webui \ --verbosity 4
$MODEL_DIR is a local copy of the model weights below.
- Runtime version
- version: 0.4.1-dev (build 11026, commit b49650adb)
- Model weights
- unsloth/Qwen3.5-9B-MTP-GGUF@9716a636
- Weights format
- GGUF, 5.9 GB
- Reasoning
- On
- Context window
- 65,536 tokens
- Sampling
- temperature 0.7, top_p 0.8, top_k 20, min_p 0
- AgentPerf-Local profile
qwen35-9b-q4-k-m-mtp-m5-pro- Benchmark client
- AgentPerf-Local 0.3.0
On the AgentPerf-Local default workload: 15.7 min to complete · 636 tok/s prefill · 49 tok/s decode · 0.8 s median time to first token
Qwen3.8 27B (xhigh)
Launch command
llama-server \ --model "$MODEL_DIR"/Qwen3.8-27B-Q4_K_M.gguf \ --alias qwen38-27b-q4-k-m-mtp-m5-pro-13b17d8ce67c773b \ --host 127.0.0.1 \ --port 8080 \ --ctx-size 65536 \ --parallel 1 \ --gpu-layers all \ --jinja \ --batch-size 2048 \ --ubatch-size 512 \ --no-mmproj \ --device MTL0 \ --load-mode mmap \ --threads 6 \ --flash-attn on \ --fit off \ --cache-ram 0 \ --spec-type draft-mtp \ --spec-draft-n-max 4 \ --spec-draft-backend-sampling \ --no-webui \ --verbosity 4
$MODEL_DIR is a local copy of the model weights below.
- Runtime version
- version: 0.4.1-dev (build 11026, commit b49650adb)
- Model weights
- unsloth/Qwen3.8-27B-GGUF@f1bfb127
- Weights format
- GGUF, 17.1 GB
- Reasoning
- On
- Context window
- 65,536 tokens
- Sampling
- temperature 0.7, top_p 0.8, top_k 20, min_p 0
- AgentPerf-Local profile
qwen38-27b-q4-k-m-mtp-m5-pro- Benchmark client
- AgentPerf-Local 0.3.0
On the AgentPerf-Local default workload: 37.6 min to complete · 212 tok/s prefill · 23 tok/s decode · 2.5 s median time to first token
Qwen3.6 35B A3B (Reasoning)
Launch command
llama-server \ --model "$MODEL_DIR"/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf \ --alias qwen36-35b-a3b-q4-k-m-mtp-m5-pro-01d9da65529d7211 \ --host 127.0.0.1 \ --port 8080 \ --ctx-size 65536 \ --parallel 1 \ --gpu-layers all \ --jinja \ --batch-size 2048 \ --ubatch-size 512 \ --no-mmproj \ --device MTL0 \ --load-mode mmap \ --threads 6 \ --flash-attn on \ --fit off \ --cache-ram 0 \ --spec-type draft-mtp \ --spec-draft-n-max 2 \ --spec-draft-backend-sampling \ --no-webui \ --verbosity 4
$MODEL_DIR is a local copy of the model weights below.
- Runtime version
- version: 0.4.1-dev (build 11026, commit b49650adb)
- Model weights
- unsloth/Qwen3.6-35B-A3B-MTP-GGUF@5bc3e238
- Weights format
- GGUF, 22.7 GB
- Reasoning
- On
- Context window
- 65,536 tokens
- Sampling
- temperature 0.7, top_p 0.8, top_k 20, min_p 0
- AgentPerf-Local profile
qwen36-35b-a3b-q4-k-m-mtp-m5-pro- Benchmark client
- AgentPerf-Local 0.3.0
On the AgentPerf-Local default workload: 11.9 min to complete · 718 tok/s prefill · 70 tok/s decode · 0.7 s median time to first token
AMD Ryzen AI Halo

- Launch MSRP
- $4,000
- Memory
- 128 GB unified
- Memory bandwidth
- 256 GB/s
- Memory type
- LPDDR5X-8000
- Chip
- Ryzen AI Max+ 395
- Launched
- Jul 2026
Qwen3.5 9B (Reasoning)
Launch command
llama-server \ --model "$MODEL_DIR"/Qwen3.5-9B-Q4_K_M.gguf \ --alias qwen35-9b-q4-k-m-mtp-strix-halo-7c80d3dc9520e8d4 \ --host 127.0.0.1 \ --port 8080 \ --ctx-size 65536 \ --parallel 1 \ --gpu-layers all \ --jinja \ --batch-size 2048 \ --ubatch-size 2048 \ --no-mmproj \ --device ROCm0 \ --load-mode none \ --threads 16 \ --flash-attn on \ --fit off \ --cache-ram 0 \ --spec-type draft-mtp \ --spec-draft-n-max 2 \ --spec-draft-backend-sampling \ --no-webui \ --verbosity 4
$MODEL_DIR is a local copy of the model weights below.
- Runtime version
- version: 0.4.1-dev (build 11026, commit b49650adb)
- Model weights
- unsloth/Qwen3.5-9B-MTP-GGUF@9716a636
- Weights format
- GGUF, 5.9 GB
- Reasoning
- On
- Context window
- 65,536 tokens
- Sampling
- temperature 0.7, top_p 0.8, top_k 20, min_p 0
- AgentPerf-Local profile
qwen35-9b-q4-k-m-mtp-strix-halo- Benchmark client
- AgentPerf-Local 0.3.0
On the AgentPerf-Local default workload: 13.0 min to complete · 775 tok/s prefill · 58 tok/s decode · 0.7 s median time to first token
Qwen3.8 27B (xhigh)
Launch command
llama-server \ --model "$MODEL_DIR"/Qwen3.8-27B-Q4_K_M.gguf \ --alias qwen38-27b-q4-k-m-mtp-strix-halo-800b9cd4aaefab6d \ --host 127.0.0.1 \ --port 8080 \ --ctx-size 65536 \ --parallel 1 \ --gpu-layers all \ --jinja \ --batch-size 2048 \ --ubatch-size 512 \ --no-mmproj \ --device ROCm0 \ --load-mode none \ --threads 16 \ --flash-attn on \ --fit off \ --cache-ram 0 \ --spec-type draft-mtp \ --spec-draft-n-max 2 \ --backend-sampling \ --spec-draft-backend-sampling \ --no-webui \ --verbosity 4
$MODEL_DIR is a local copy of the model weights below.
- Runtime version
- version: 0.4.1-dev (build 11026, commit b49650adb)
- Model weights
- unsloth/Qwen3.8-27B-GGUF@f1bfb127
- Weights format
- GGUF, 17.1 GB
- Reasoning
- On
- Context window
- 65,536 tokens
- Sampling
- temperature 0.7, top_p 0.8, top_k 20, min_p 0
- AgentPerf-Local profile
qwen38-27b-q4-k-m-mtp-strix-halo- Benchmark client
- AgentPerf-Local 0.3.0
On the AgentPerf-Local default workload: 34.5 min to complete · 261 tok/s prefill · 23 tok/s decode · 2.0 s median time to first token
Qwen3.6 35B A3B (Reasoning)
Launch command
llama-server \ --model "$MODEL_DIR"/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf \ --alias qwen36-35b-a3b-q4-k-m-mtp-strix-halo-53cb946e97294e4b \ --host 127.0.0.1 \ --port 8080 \ --ctx-size 65536 \ --parallel 1 \ --gpu-layers all \ --jinja \ --batch-size 2048 \ --ubatch-size 512 \ --no-mmproj \ --device Vulkan0 \ --load-mode none \ --threads 16 \ --flash-attn on \ --fit off \ --cache-ram 0 \ --spec-type draft-mtp \ --spec-draft-n-max 2 \ --backend-sampling \ --spec-draft-backend-sampling \ --no-webui \ --verbosity 4
$MODEL_DIR is a local copy of the model weights below.
- Runtime version
- version: 0.4.1-dev (build 11026, commit b49650adb)
- Model weights
- unsloth/Qwen3.6-35B-A3B-MTP-GGUF@5bc3e238
- Weights format
- GGUF, 22.7 GB
- Reasoning
- On
- Context window
- 65,536 tokens
- Sampling
- temperature 0.7, top_p 0.8, top_k 20, min_p 0
- AgentPerf-Local profile
qwen36-35b-a3b-q4-k-m-mtp-strix-halo- Benchmark client
- AgentPerf-Local 0.3.0
On the AgentPerf-Local default workload: 11.7 min to complete · 687 tok/s prefill · 74 tok/s decode · 0.8 s median time to first token
Ling 3.0 Flash
Launch command
llama-server \ --model "$MODEL_DIR"/Q4_K_M/Ling-3.0-flash-Q4_K_M-00001-of-00002.gguf \ --alias ling3-flash-q4-k-m-dspark-strix-halo-4316d051c13f4422 \ --host 127.0.0.1 \ --port 8080 \ --ctx-size 65536 \ --parallel 1 \ --gpu-layers all \ --jinja \ --batch-size 2048 \ --ubatch-size 512 \ --no-mmproj \ --device ROCm0 \ --flash-attn on \ --fit off \ --ctx-checkpoints 2 \ --model-draft "$MODEL_DIR"/Ling-3.0-flash-DSpark-bf16.gguf \ --gpu-layers-draft all \ --spec-type draft-dspark \ --spec-draft-n-max 8 \ --no-webui \ --verbosity 4
$MODEL_DIR is a local copy of the model weights below.
- Runtime version
- version: 0.4.1-dev (build 11114, commit 7dc9a32df)
- Model weights
- inclusionAI/Ling-3.0-flash-GGUF@29edce51
- Weights format
- GGUF, 77.0 GB
- Reasoning
- On
- Context window
- 65,536 tokens
- Sampling
- temperature 0.7, top_p 0.8, top_k 20, min_p 0
- AgentPerf-Local profile
ling3-flash-q4-k-m-dspark-strix-halo- Benchmark client
- AgentPerf-Local 0.3.0
On the AgentPerf-Local default workload: 25.1 min to complete · 313 tok/s prefill · 35 tok/s decode · 1.7 s median time to first token
NVIDIA DGX Spark

- Launch MSRP
- $4,000
- Memory
- 128 GB unified
- Memory bandwidth
- 273 GB/s
- Memory type
- LPDDR5X-8533
- Chip
- GB10
- Launched
- Oct 2025
Qwen3.5 9B (Reasoning)
Launch command
vllm serve "$MODEL_DIR" \
--served-model-name qwen35-9b-nvfp4-mtp-dgx-spark-b05e797fea417478 \
--host 127.0.0.1 \
--port 8080 \
--max-model-len 65536 \
--tensor-parallel-size 1 \
--max-num-seqs 1 \
--enable-prompt-tokens-details \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--seed 0 \
--trust-remote-code \
--kv-cache-dtype fp8 \
--attention-backend flashinfer \
--gpu-memory-utilization 0.5 \
--max-num-batched-tokens 8192 \
--enable-chunked-prefill \
--async-scheduling \
--enable-prefix-caching \
--speculative-config '{"method":"mtp","num_speculative_tokens":2}' \
--load-format fastsafetensors$MODEL_DIR is a local copy of the model weights below.
- Runtime version
- 0.28.0
- Model weights
- AxionML/Qwen3.5-9B-NVFP4@97aef923
- Weights format
- ModelOpt, 9.4 GB
- Reasoning
- On
- Context window
- 65,536 tokens
- Sampling
- temperature 0.7, top_p 0.8, top_k 20, min_p 0
- AgentPerf-Local profile
qwen35-9b-nvfp4-mtp-dgx-spark- Benchmark client
- AgentPerf-Local 0.3.2
On the AgentPerf-Local default workload: 13.0 min to complete · 2,161 tok/s prefill · 46 tok/s decode · 0.4 s median time to first token
Qwen3.8 27B (xhigh)
Launch command
vllm serve "$MODEL_DIR" \
--served-model-name qwen38-27b-nvfp4-dgx-spark-fa8cb342e4620146 \
--host 127.0.0.1 \
--port 8080 \
--max-model-len 65536 \
--tensor-parallel-size 1 \
--max-num-seqs 1 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--seed 0 \
--enable-prompt-tokens-details \
--trust-remote-code \
--kv-cache-dtype fp8 \
--gpu-memory-utilization 0.8 \
--max-num-batched-tokens 16384 \
--enable-chunked-prefill \
--async-scheduling \
--enable-prefix-caching \
--speculative-config '{"method":"dflash","model":"z-lab/Qwen3.8-27B-DFlash2","num_speculative_tokens":8}' \
--load-format safetensors$MODEL_DIR is a local copy of the model weights below.
- Runtime version
- 0.28.0
- Model weights
- RadixArk/Qwen3.8-27B-NVFP4@319f741c
- Weights format
- ModelOpt, 21.9 GB
- Reasoning
- On
- Context window
- 65,536 tokens
- Sampling
- temperature 0.7, top_p 0.8, top_k 20, min_p 0
- AgentPerf-Local profile
qwen38-27b-nvfp4-dgx-spark- Benchmark client
- AgentPerf-Local 0.3.0
On the AgentPerf-Local default workload: 24.2 min to complete · 612 tok/s prefill · 28 tok/s decode · 1.6 s median time to first token
Qwen3.6 35B A3B (Reasoning)
Launch command
vllm serve "$MODEL_DIR" \
--served-model-name qwen36-35b-a3b-nvfp4-mtp-dgx-spark-ac98a02d42f01169 \
--host 127.0.0.1 \
--port 8080 \
--max-model-len 65536 \
--tensor-parallel-size 1 \
--max-num-seqs 1 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--seed 0 \
--enable-prompt-tokens-details \
--trust-remote-code \
--kv-cache-dtype fp8 \
--attention-backend flashinfer \
--moe-backend marlin \
--gpu-memory-utilization 0.8 \
--max-num-batched-tokens 8192 \
--enable-chunked-prefill \
--async-scheduling \
--enable-prefix-caching \
--speculative-config '{"method":"mtp","num_speculative_tokens":3,"moe_backend":"triton"}' \
--load-format safetensors$MODEL_DIR is a local copy of the model weights below.
- Runtime version
- 0.28.0
- Model weights
- nvidia/Qwen3.6-35B-A3B-NVFP4@1355db6a
- Weights format
- ModelOpt, 23.5 GB
- Reasoning
- On
- Context window
- 65,536 tokens
- Sampling
- temperature 0.7, top_p 0.8, top_k 20, min_p 0
- AgentPerf-Local profile
qwen36-35b-a3b-nvfp4-mtp-dgx-spark- Benchmark client
- AgentPerf-Local 0.3.0
On the AgentPerf-Local default workload: 7.3 min to complete · 1,308 tok/s prefill · 119 tok/s decode · 0.8 s median time to first token
Ling 3.0 Flash
Launch command
llama-server \ --model "$MODEL_DIR"/Q4_K_M/Ling-3.0-flash-Q4_K_M-00001-of-00002.gguf \ --alias ling3-flash-q4-k-m-dspark-dgx-spark-1db997c2387108af \ --host 127.0.0.1 \ --port 8080 \ --ctx-size 65536 \ --parallel 1 \ --gpu-layers all \ --jinja \ --batch-size 2048 \ --ubatch-size 512 \ --no-mmproj \ --flash-attn on \ --fit off \ --ctx-checkpoints 2 \ --model-draft "$MODEL_DIR"/Ling-3.0-flash-DSpark-bf16.gguf \ --spec-type draft-dspark \ --spec-draft-n-max 8 \ --no-webui \ --verbosity 4
$MODEL_DIR is a local copy of the model weights below.
- Runtime version
- version: 0.4.1-dev (build 11114, commit 7dc9a32d)
- Model weights
- inclusionAI/Ling-3.0-flash-GGUF@29edce51
- Weights format
- GGUF, 77.0 GB
- Reasoning
- On
- Context window
- 65,536 tokens
- Sampling
- temperature 0.7, top_p 0.8, top_k 20, min_p 0
- AgentPerf-Local profile
ling3-flash-q4-k-m-dspark-dgx-spark- Benchmark client
- AgentPerf-Local 0.3.0
On the AgentPerf-Local default workload: 14.9 min to complete · 467 tok/s prefill · 65 tok/s decode · 1.6 s median time to first token
NVIDIA GeForce RTX 5090

- Launch MSRP (card)
- $2,000
- Memory
- 32 GB VRAM
- Memory bandwidth
- 1,792 GB/s
- Memory type
- GDDR7
- Chip
- RTX 5090
- Launched
- Jan 2025
Qwen3.5 9B (Reasoning)
Launch command
CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=0 llama-server \ --model "$MODEL_DIR"/Qwen3.5-9B-Q4_K_M.gguf \ --alias qwen35-9b-q4-k-m-mtp-rtx5090-086b2463c77e3274 \ --host 127.0.0.1 \ --port 8080 \ --ctx-size 65536 \ --parallel 1 \ --gpu-layers all \ --jinja \ --batch-size 4096 \ --ubatch-size 4096 \ --no-mmproj \ --spec-type draft-mtp \ --spec-draft-n-max 2 \ --no-webui \ --verbosity 4
$MODEL_DIR is a local copy of the model weights below.
- Runtime version
- version: 0.4.1-dev (build 11026, commit b49650adb)
- Model weights
- unsloth/Qwen3.5-9B-MTP-GGUF@9716a636
- Weights format
- GGUF, 5.9 GB
- Reasoning
- On
- Context window
- 65,536 tokens
- Sampling
- temperature 0.7, top_p 0.8, top_k 20, min_p 0
- AgentPerf-Local profile
qwen35-9b-q4-k-m-mtp-rtx5090- Benchmark client
- AgentPerf-Local 0.3.0
On the AgentPerf-Local default workload: 2.2 min to complete · 4,980 tok/s prefill · 331 tok/s decode · 0.2 s median time to first token
Qwen3.8 27B (xhigh)
Launch command
CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=0 llama-server \ --model "$MODEL_DIR"/Qwen3.8-27B-Q4_K_M.gguf \ --alias qwen38-27b-q4-k-m-mtp-488339a4d3e0e3af \ --host 127.0.0.1 \ --port 8080 \ --ctx-size 65536 \ --parallel 1 \ --gpu-layers all \ --jinja \ --batch-size 4096 \ --ubatch-size 2048 \ --no-mmproj \ --spec-type draft-mtp \ --spec-draft-n-max 2 \ --backend-sampling \ --spec-draft-backend-sampling \ --no-webui \ --verbosity 4
$MODEL_DIR is a local copy of the model weights below.
- Runtime version
- version: 0.4.1-dev (build 11026, commit b49650adb)
- Model weights
- unsloth/Qwen3.8-27B-GGUF@f1bfb127
- Weights format
- GGUF, 17.1 GB
- Reasoning
- On
- Context window
- 65,536 tokens
- Sampling
- temperature 0.7, top_p 0.8, top_k 20, min_p 0
- AgentPerf-Local profile
qwen38-27b-q4-k-m-mtp- Benchmark client
- AgentPerf-Local 0.3.0
On the AgentPerf-Local default workload: 4.9 min to complete · 2,156 tok/s prefill · 151 tok/s decode · 0.3 s median time to first token
Qwen3.6 35B A3B (Reasoning)
Launch command
llama-server \ --model "$MODEL_DIR"/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf \ --alias qwen36-35b-a3b-q4-k-m-mtp-rtx5090-27e26f3617e55468 \ --host 127.0.0.1 \ --port 8080 \ --ctx-size 65536 \ --parallel 1 \ --gpu-layers all \ --jinja \ --batch-size 2048 \ --ubatch-size 2048 \ --no-mmproj \ --spec-type draft-mtp \ --spec-draft-n-max 2 \ --backend-sampling \ --no-webui \ --verbosity 4
$MODEL_DIR is a local copy of the model weights below.
- Runtime version
- version: 0.4.1-dev (build 11026, commit b49650adb)
- Model weights
- unsloth/Qwen3.6-35B-A3B-MTP-GGUF@5bc3e238
- Weights format
- GGUF, 22.7 GB
- Reasoning
- On
- Context window
- 65,536 tokens
- Sampling
- temperature 0.7, top_p 0.8, top_k 20, min_p 0
- AgentPerf-Local profile
qwen36-35b-a3b-q4-k-m-mtp-rtx5090- Benchmark client
- AgentPerf-Local 0.3.0
On the AgentPerf-Local default workload: 2.0 min to complete · 5,101 tok/s prefill · 381 tok/s decode · 0.2 s median time to first token