Technical Writing
Benchmark results for Qwen/Qwen3.6-35B-A3B-FP8 (NVIDIA DGX Spark, GB10) serving via vLLM
Benchmark results for Qwen/Qwen3.6-35B-A3B-FP8 (NVIDIA DGX Spark, GB10) serving via vLLM.
DGX Benchmark Report — Qwen3.6-35B-A3B-FP8
Benchmark results for Qwen/Qwen3.6-35B-A3B-FP8 on (NVIDIA DGX Spark, GB10) serving via vLLM.
TL;DR
- Single-user steady-state: ~28–30 tok/s. At or near the memory-bandwidth ceiling for this model on this hardware.
- Aggregate under load: up to 156 tok/s at c=32. Scales sub-linearly (memory-bandwidth bound, not KV-bound).
- Interactive UX sweet spot: c=1–8. Beyond c=16, per-user TPOT exceeds 130 ms/tok — good for agentic/batch, sluggish for chat.
- No failed requests at any tested concurrency up to 32.
- Alternative quantizations tested (AWQ-Int4) and speculative decoding (MTP-2) were rejected — details in appendix.
Environment
Hardware
- GPU: NVIDIA GB10 (Grace Blackwell), capability (12, 1) / sm_121
- Memory: 121 GiB unified LPDDR5x, 273 GB/s bandwidth
- CPU: 20-core ARM Cortex-A725, arm64
- Storage: 4 TB NVMe
Software
- OS: Ubuntu 24.04 LTS, kernel 6.17.x-nvidia
- NVIDIA driver: 580.142 (pinned via
apt-mark hold) - CUDA: 13.0 host / 13.2 container via forward-compat
- Container runtime: Docker CE 29.2.1 + nvidia-container-toolkit 1.19.0 (legacy
--gpus allpath) - Image:
nvcr.io/nvidia/vllm:26.03.post1-py3 - vLLM: 0.17.1+bd67d66a.nvinternal.26.03.post1.48207566
- PyTorch: 2.11.0a0+a6c236b9fd.nv26.03.46836102
- Model: Qwen3.6-35B-A3B-FP8 (35B total / 3B active, MoE 256 experts top-8+1, hybrid attention + Mamba, quantization
fp8block-128) - MoE backend selected by vLLM: TRITON (FLASHINFER_TRTLLM/CUTLASS/DEEPGEMM all unavailable on SM121 as of this date)
vLLM configuration
Flags in use at measurement time:
--served-model-name qwen3.6-35b
--host 0.0.0.0 --port 8000
--reasoning-parser qwen3
--enable-auto-tool-choice --tool-call-parser qwen3_coder
--max-model-len 262144
--gpu-memory-utilization 0.85
--enable-prefix-caching
Runtime container flags: --gpus all --ipc=host --ulimit memlock=-1 --ulimit stack=67108864 --restart unless-stopped
Env: VLLM_FLASHINFER_MOE_BACKEND=latency (set but ignored by vLLM 0.17.1 — preserved for future NGC upgrades).
Engine-reported capacity at boot
- Available KV cache memory: 64.53 GiB
- GPU KV cache size: 846,912 tokens
- Maximum concurrency at full 262K context: 12.56×
- Init engine time (profile + KV + warmup): 130.87 s
- First inference after container start takes ~60 s (kernel autotune). Discard from benchmarks.
Methodology
Measurements produced via vllm bench serve running inside the container (removes network leg from timing). Executed from the host via docker exec.
- Dataset:
random(random token IDs generated by the bench harness). Adversarial for MoE routing — real-world natural-text prompts typically realize +10–20% above these numbers. - Prompt shape:
--random-input-len 512 --random-output-len 256 --random-range-ratio 0.1(small variation in both input and output length, median ~512 in / 256 out). - Warmup: 2 prompts per run, discarded from stats.
- Main sample size:
max(8, concurrency × 4)prompts. Small-sample variance is non-trivial at c=1 and c=2; trust the curve shape, not individual points. - Output generation:
--ignore-eosso the model always emits exactlyoutput_lentokens (stable generation work per request). - Mode:
chat_template_kwargs: {enable_thinking: false}— no-think mode. Reasoning-on is a different workload and not tested here. - Auth: Bearer token, identical to production access.
Results
Primary concurrency sweep
| c | Output tok/s | Peak tok/s | TTFT p50 | TTFT p95 | TPOT p50 | TPOT p95 | E2E p50 | E2E p95 | Failed |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 21.0 | 33 | 363 ms | 421 ms | 48 ms | 49 ms | 13.1 s | 13.7 s | 0 |
| 2 | 31.5 | 34 | 396 ms | 649 ms | 62 ms | 62 ms | 16.7 s | 17.6 s | 0 |
| 4 | 51.5 | 60 | 510 ms | 917 ms | 75 ms | 76 ms | 19.9 s | 21.7 s | 0 |
| 8 | 70.2 | 96 | 617 ms | 12,266* | 98 ms | 101 ms | 26.3 s | 37.9 s | 0 |
| 16 | 115.8 | 160 | 861 ms | 3,339 ms | 132 ms | 138 ms | 34.4 s | 38.7 s | 0 |
| 32 | 155.6 | 291 | 1,121 ms | 6,236 ms | 195 ms | 207 ms | 51.6 s | 59.4 s | 0 |
*c=8 p95 is a small-sample outlier (n=32 prompts). Resolves in c=16 as vLLM’s continuous-batching scheduler finds steady state.
Scaling efficiency
| From → To | Throughput ratio | Expected if linear | Efficiency |
|---|---|---|---|
| c=1 → c=2 | 1.5× | 2× | 75% |
| c=1 → c=4 | 2.5× | 4× | 62% |
| c=1 → c=8 | 3.3× | 8× | 42% |
| c=1 → c=16 | 5.5× | 16× | 34% |
| c=1 → c=32 | 7.4× | 32× | 23% |
Scaling flattens significantly past c=8 — MoE routing cost rises with batch width (more experts activated per step across the batch), which caps aggregate throughput.
Real-prompt sanity (non-random inputs)
Small-sample spot checks show natural-text prompts realize ~30 tok/s single-user, consistent with community reports of 31–32 tok/s as the observed Spark ceiling for this model class.
Example warm-path measurements:
- “Count to 10” → 31 tokens in 1.03 s = 30.1 tok/s
- “Say 42 only” → 3 tokens in 0.16 s = TTFT-dominated
- Best c=1 benchmark run observed: 28.7 tok/s (variance across runs; compile state and MoE routing luck matter)
Interpretation
Capacity guidance for orchestrator tuning
| User load | Per-user experience | Suitable workloads |
|---|---|---|
| 1–4 concurrent | Snappy — hosted-API feel (sub-second TTFT, sub-100ms/token) | Interactive chat, real-time coding assist |
| 5–8 concurrent | Fine for work — slight lag noticeable | Interactive mixed-user team |
| 9–16 concurrent | Works; per-user feels slow for interactive chat | Agentic/batch, orchestrator internal calls, long-context summarization |
| 17+ concurrent | Agentic/batch only; don’t put a human in the loop | Bulk generation, background jobs |
Recommended LiteLLM / orchestrator settings against this backend:
- Default soft ceiling: c=8 (balance of latency + throughput)
- Burst allowance up to c=16 before returning 429 or spilling to fallback
- Aggressive prompt-prefix caching at LiteLLM (we observed 0% prefix cache hit in random bench; agentic traffic with shared system prompts will benefit substantially)
Not KV-bound anywhere in the sweep
Engine reports max concurrency 12.73× at full 262K context; bench at 512-token prompts never approached that ceiling. The bottleneck we hit is memory bandwidth on the forward pass, not KV cache exhaustion. Lowering --max-model-len would free KV cache we aren’t using — no throughput benefit.
Bottleneck analysis
Spark is memory-bandwidth bound on decode. For Qwen3.6-35B-A3B-FP8:
theoretical ceiling = bandwidth / bytes_per_active_params_per_token
= 273 GB/s / ~3 GB
= 91 tok/s (single-stream theoretical)
Real-world typically realizes 25–40% of the theoretical ceiling (attention, norms, scheduling, Python overhead). Measured single-stream of 28–30 tok/s = ~32% realization, which is normal for this class of workload.
What would move the ceiling up (not available today):
- NVFP4 quantization: ~2× theoretical, ~55–60 tok/s projected. Kernels not yet mature on Spark (NVIDIA dev forum consensus Apr 2026). Revisit quarterly via NGC release notes.
- Native sm_121 kernels (instead of sm_120 forward-compat): +10–20% expected. Will arrive in a future NGC tag.
- TP=2 across two Sparks via ConnectX-7: ~1.8× aggregate throughput per community benchmarks. Helps aggregate, not single-stream.
- MXFP4 community patches: community reports ~35 tok/s single-node. Marginal gain, requires running a forked vLLM — operational risk.
What does not move it:
- Any Qwen dense model of comparable quality — dense reads full weights per token, ~10× more bandwidth cost, substantially slower on Spark.
- Kernel autotune, cudagraph, compile flags — already on by default, no room.
- Custom MoE configs via
vllm/tools/benchmark_moe.py— community reports mixed results (30.5 vs 32 tok/s). Not worth the effort.
Appendix A — Experiments rejected
AWQ-Int4 (QuantTrio/Qwen3.6-35B-A3B-AWQ, vLLM nightly 0.19.2)
Same sweep, same prompts, same image baseline, awq_marlin quantization backend.
| c | FP8 tok/s | AWQ tok/s | Δ | FP8 TPOT p50 | AWQ TPOT p50 |
|---|---|---|---|---|---|
| 1 | 21.0 | 23.9 | +14% | 48ms | 41ms |
| 2 | 31.5 | 39.7 | +26% | 62ms | 48ms |
| 4 | 51.5 | 65.3 | +27% | 75ms | 59ms |
| 8 | 70.2 | 92.2 | +31% | 98ms | 74ms |
| 16 | 115.8 | 147.3 | +27% | 132ms | 103ms |
| 32 | 155.6 | 192.1 | +23% | 195ms | 158ms |
Rejected for production. Quality regression reproduced (Alice/Bob age problem — AWQ said 18, FP8 said 20) and matches Tencent AngelSlim’s published Qwen3-32B data (INT4-AWQ drops ~3 HumanEval points vs FP8-Static). Speed win modest (~25%), not worth the correctness trade for agentic/code workloads.
MTP speculative decoding
MTP-2 (num_speculative_tokens: 2)
Per Qwen recipe on recipes.vllm.ai. At launch, vLLM printed the warning: Enabling num_speculative_tokens > 1 will run multiple times of forward on same MTP layer, which may result in lower acceptance rate.
| c | MTP-2 tok/s | MTP-2 TPOT p50 | Baseline tok/s | Baseline TPOT p50 | Note |
|---|---|---|---|---|---|
| 1 | 26.1 | 38ms | 28.7 | 30ms | −9% tok/s, +27% TPOT |
| 4 | 44.5 | 76ms | 51.5 | 75ms | −14% tok/s |
| 8 | — | — | 70.2 | 98ms | 17/32 requests FAILED |
Rejected. Strictly worse than baseline, and breaks under c=8 concurrency. Hypothesis: MTP-2 on a model with a single MTP layer runs that layer twice, drafting the second token from the (unverified) first draft, tanking acceptance. Also incompatible with --enable-prefix-caching.
MTP-1 (num_speculative_tokens: 1)
Per community Qwen3.5 latency recipe. Disables --enable-prefix-caching (mutually exclusive). Max concurrency at full 262K drops to 11.03× (vs 12.56× for baseline).
| c | MTP-1 tok/s | MTP-1 TPOT p50 | Baseline tok/s | Baseline TPOT p50 | Δ throughput | Δ TPOT | Failed |
|---|---|---|---|---|---|---|---|
| 1 | 28.4 | 37ms | 28.7 | 30ms | ≈ | +23% | 0 |
| 4 | 63.8 | 59ms | 51.5 | 75ms | +24% | −21% | 0 |
| 8 | 85.7 | 76ms | 70.2 | 98ms | +22% | −22% | 0 |
Surprising result: MTP is typically pitched as a single-stream latency win. On our model+hardware it’s actually a mid-concurrency throughput win — c=1 is flat (speculative overhead eats the small acceptance benefit), but c=4–8 gain ~22% throughput and ~20% lower TPOT. Hypothesis: moderate MTP acceptance rate + amortization of draft computation across concurrent requests.
Trade vs --enable-prefix-caching: MTP-1 is mutually exclusive with prefix caching. Under random-prompt bench (0% prefix hit rate), MTP-1 wins clearly. Under real agentic traffic with long shared system prompts, prefix caching may win. Worth re-measuring once real orchestrator traffic exists.
Current production config uses MTP-1 (decision 2026-04-23). Rollback path: relaunch with --enable-prefix-caching instead of the speculative config flag.
Appendix B — Reproducing this benchmark
From the DGX:
ssh dgx-NN
KEY=<VLLM_API_KEY>
for C in 1 2 4 8 16 32; do
N=$(( C * 4 < 8 ? 8 : C * 4 ))
echo "=== concurrency=$C num_prompts=$N ==="
docker exec vllm vllm bench serve \
--backend openai-chat \
--base-url http://localhost:8000 \
--endpoint /v1/chat/completions \
--header "Authorization=Bearer $KEY" \
--model qwen3.6-35b \
--tokenizer Qwen/Qwen3.6-35B-A3B-FP8 \
--dataset-name random --random-input-len 512 --random-output-len 256 --random-range-ratio 0.1 \
--num-prompts "$N" --num-warmups 2 --max-concurrency "$C" --ignore-eos \
--extra-body '{"chat_template_kwargs": {"enable_thinking": false}}' \
--percentile-metrics ttft,tpot,itl,e2el --metric-percentiles 50,95,99 \
--disable-tqdm 2>&1 | grep -E "^(Successful|Failed|Benchmark duration|Output token throughput|Peak|Mean TTFT|P50 TTFT|P95 TTFT|Mean TPOT|P50 TPOT|P95 TPOT|Mean E2EL|P50 E2EL|P95 E2EL)"
echo
done
Note on --tokenizer: required because --served-model-name qwen3.6-35b aliases the model; without --tokenizer, bench tries to fetch qwen3.6-35b from HuggingFace and errors with OSError: not a local folder and is not a valid model identifier.
Discard the first bench run after any container restart (kernel autotune dominates).
References
- adadrag/qwen3.5-dgx-spark (GitHub) — community guide confirming 31–32 tok/s is observed Spark ceiling for Qwen3.5-35B-A3B
- Tencent/AngelSlim (GitHub) — Qwen3-32B cross-format CEVAL/MMLU/GSM8K/HumanEval numbers (AWQ quality regression data)
- vLLM Qwen3.5/3.6 Recipe — MTP-1 latency guidance
- PSA: State of FP4/NVFP4 Support for DGX Spark in VLLM (NVIDIA Developer Forums) — NVFP4 Spark status, track quarterly