Technical Writing

Benchmark results for Qwen/Qwen3.6-35B-A3B-FP8 (NVIDIA DGX Spark, GB10) serving via vLLM

· 11 min read · Updated

Benchmark results for Qwen/Qwen3.6-35B-A3B-FP8 (NVIDIA DGX Spark, GB10) serving via vLLM.

DGX Benchmark Report — Qwen3.6-35B-A3B-FP8

Benchmark results for Qwen/Qwen3.6-35B-A3B-FP8 on (NVIDIA DGX Spark, GB10) serving via vLLM.

TL;DR

  • Single-user steady-state: ~28–30 tok/s. At or near the memory-bandwidth ceiling for this model on this hardware.
  • Aggregate under load: up to 156 tok/s at c=32. Scales sub-linearly (memory-bandwidth bound, not KV-bound).
  • Interactive UX sweet spot: c=1–8. Beyond c=16, per-user TPOT exceeds 130 ms/tok — good for agentic/batch, sluggish for chat.
  • No failed requests at any tested concurrency up to 32.
  • Alternative quantizations tested (AWQ-Int4) and speculative decoding (MTP-2) were rejected — details in appendix.

Environment

Hardware

  • GPU: NVIDIA GB10 (Grace Blackwell), capability (12, 1) / sm_121
  • Memory: 121 GiB unified LPDDR5x, 273 GB/s bandwidth
  • CPU: 20-core ARM Cortex-A725, arm64
  • Storage: 4 TB NVMe

Software

  • OS: Ubuntu 24.04 LTS, kernel 6.17.x-nvidia
  • NVIDIA driver: 580.142 (pinned via apt-mark hold)
  • CUDA: 13.0 host / 13.2 container via forward-compat
  • Container runtime: Docker CE 29.2.1 + nvidia-container-toolkit 1.19.0 (legacy --gpus all path)
  • Image: nvcr.io/nvidia/vllm:26.03.post1-py3
  • vLLM: 0.17.1+bd67d66a.nvinternal.26.03.post1.48207566
  • PyTorch: 2.11.0a0+a6c236b9fd.nv26.03.46836102
  • Model: Qwen3.6-35B-A3B-FP8 (35B total / 3B active, MoE 256 experts top-8+1, hybrid attention + Mamba, quantization fp8 block-128)
  • MoE backend selected by vLLM: TRITON (FLASHINFER_TRTLLM/CUTLASS/DEEPGEMM all unavailable on SM121 as of this date)

vLLM configuration

Flags in use at measurement time:

--served-model-name qwen3.6-35b
--host 0.0.0.0 --port 8000
--reasoning-parser qwen3
--enable-auto-tool-choice --tool-call-parser qwen3_coder
--max-model-len 262144
--gpu-memory-utilization 0.85
--enable-prefix-caching

Runtime container flags: --gpus all --ipc=host --ulimit memlock=-1 --ulimit stack=67108864 --restart unless-stopped

Env: VLLM_FLASHINFER_MOE_BACKEND=latency (set but ignored by vLLM 0.17.1 — preserved for future NGC upgrades).

Engine-reported capacity at boot

  • Available KV cache memory: 64.53 GiB
  • GPU KV cache size: 846,912 tokens
  • Maximum concurrency at full 262K context: 12.56×
  • Init engine time (profile + KV + warmup): 130.87 s
  • First inference after container start takes ~60 s (kernel autotune). Discard from benchmarks.

Methodology

Measurements produced via vllm bench serve running inside the container (removes network leg from timing). Executed from the host via docker exec.

  • Dataset: random (random token IDs generated by the bench harness). Adversarial for MoE routing — real-world natural-text prompts typically realize +10–20% above these numbers.
  • Prompt shape: --random-input-len 512 --random-output-len 256 --random-range-ratio 0.1 (small variation in both input and output length, median ~512 in / 256 out).
  • Warmup: 2 prompts per run, discarded from stats.
  • Main sample size: max(8, concurrency × 4) prompts. Small-sample variance is non-trivial at c=1 and c=2; trust the curve shape, not individual points.
  • Output generation: --ignore-eos so the model always emits exactly output_len tokens (stable generation work per request).
  • Mode: chat_template_kwargs: {enable_thinking: false} — no-think mode. Reasoning-on is a different workload and not tested here.
  • Auth: Bearer token, identical to production access.

Results

Primary concurrency sweep

cOutput tok/sPeak tok/sTTFT p50TTFT p95TPOT p50TPOT p95E2E p50E2E p95Failed
121.033363 ms421 ms48 ms49 ms13.1 s13.7 s0
231.534396 ms649 ms62 ms62 ms16.7 s17.6 s0
451.560510 ms917 ms75 ms76 ms19.9 s21.7 s0
870.296617 ms12,266*98 ms101 ms26.3 s37.9 s0
16115.8160861 ms3,339 ms132 ms138 ms34.4 s38.7 s0
32155.62911,121 ms6,236 ms195 ms207 ms51.6 s59.4 s0

*c=8 p95 is a small-sample outlier (n=32 prompts). Resolves in c=16 as vLLM’s continuous-batching scheduler finds steady state.

Scaling efficiency

From → ToThroughput ratioExpected if linearEfficiency
c=1 → c=21.5×2×75%
c=1 → c=42.5×4×62%
c=1 → c=83.3×8×42%
c=1 → c=165.5×16×34%
c=1 → c=327.4×32×23%

Scaling flattens significantly past c=8 — MoE routing cost rises with batch width (more experts activated per step across the batch), which caps aggregate throughput.

Real-prompt sanity (non-random inputs)

Small-sample spot checks show natural-text prompts realize ~30 tok/s single-user, consistent with community reports of 31–32 tok/s as the observed Spark ceiling for this model class.

Example warm-path measurements:

  • “Count to 10” → 31 tokens in 1.03 s = 30.1 tok/s
  • “Say 42 only” → 3 tokens in 0.16 s = TTFT-dominated
  • Best c=1 benchmark run observed: 28.7 tok/s (variance across runs; compile state and MoE routing luck matter)

Interpretation

Capacity guidance for orchestrator tuning

User loadPer-user experienceSuitable workloads
1–4 concurrentSnappy — hosted-API feel (sub-second TTFT, sub-100ms/token)Interactive chat, real-time coding assist
5–8 concurrentFine for work — slight lag noticeableInteractive mixed-user team
9–16 concurrentWorks; per-user feels slow for interactive chatAgentic/batch, orchestrator internal calls, long-context summarization
17+ concurrentAgentic/batch only; don’t put a human in the loopBulk generation, background jobs

Recommended LiteLLM / orchestrator settings against this backend:

  • Default soft ceiling: c=8 (balance of latency + throughput)
  • Burst allowance up to c=16 before returning 429 or spilling to fallback
  • Aggressive prompt-prefix caching at LiteLLM (we observed 0% prefix cache hit in random bench; agentic traffic with shared system prompts will benefit substantially)

Not KV-bound anywhere in the sweep

Engine reports max concurrency 12.73× at full 262K context; bench at 512-token prompts never approached that ceiling. The bottleneck we hit is memory bandwidth on the forward pass, not KV cache exhaustion. Lowering --max-model-len would free KV cache we aren’t using — no throughput benefit.

Bottleneck analysis

Spark is memory-bandwidth bound on decode. For Qwen3.6-35B-A3B-FP8:

theoretical ceiling = bandwidth / bytes_per_active_params_per_token
                    = 273 GB/s  /  ~3 GB
                    = 91 tok/s (single-stream theoretical)

Real-world typically realizes 25–40% of the theoretical ceiling (attention, norms, scheduling, Python overhead). Measured single-stream of 28–30 tok/s = ~32% realization, which is normal for this class of workload.

What would move the ceiling up (not available today):

  • NVFP4 quantization: ~2× theoretical, ~55–60 tok/s projected. Kernels not yet mature on Spark (NVIDIA dev forum consensus Apr 2026). Revisit quarterly via NGC release notes.
  • Native sm_121 kernels (instead of sm_120 forward-compat): +10–20% expected. Will arrive in a future NGC tag.
  • TP=2 across two Sparks via ConnectX-7: ~1.8× aggregate throughput per community benchmarks. Helps aggregate, not single-stream.
  • MXFP4 community patches: community reports ~35 tok/s single-node. Marginal gain, requires running a forked vLLM — operational risk.

What does not move it:

  • Any Qwen dense model of comparable quality — dense reads full weights per token, ~10× more bandwidth cost, substantially slower on Spark.
  • Kernel autotune, cudagraph, compile flags — already on by default, no room.
  • Custom MoE configs via vllm/tools/benchmark_moe.py — community reports mixed results (30.5 vs 32 tok/s). Not worth the effort.

Appendix A — Experiments rejected

AWQ-Int4 (QuantTrio/Qwen3.6-35B-A3B-AWQ, vLLM nightly 0.19.2)

Same sweep, same prompts, same image baseline, awq_marlin quantization backend.

cFP8 tok/sAWQ tok/sΔFP8 TPOT p50AWQ TPOT p50
121.023.9+14%48ms41ms
231.539.7+26%62ms48ms
451.565.3+27%75ms59ms
870.292.2+31%98ms74ms
16115.8147.3+27%132ms103ms
32155.6192.1+23%195ms158ms

Rejected for production. Quality regression reproduced (Alice/Bob age problem — AWQ said 18, FP8 said 20) and matches Tencent AngelSlim’s published Qwen3-32B data (INT4-AWQ drops ~3 HumanEval points vs FP8-Static). Speed win modest (~25%), not worth the correctness trade for agentic/code workloads.

MTP speculative decoding

MTP-2 (num_speculative_tokens: 2)

Per Qwen recipe on recipes.vllm.ai. At launch, vLLM printed the warning: Enabling num_speculative_tokens > 1 will run multiple times of forward on same MTP layer, which may result in lower acceptance rate.

cMTP-2 tok/sMTP-2 TPOT p50Baseline tok/sBaseline TPOT p50Note
126.138ms28.730ms−9% tok/s, +27% TPOT
444.576ms51.575ms−14% tok/s
8——70.298ms17/32 requests FAILED

Rejected. Strictly worse than baseline, and breaks under c=8 concurrency. Hypothesis: MTP-2 on a model with a single MTP layer runs that layer twice, drafting the second token from the (unverified) first draft, tanking acceptance. Also incompatible with --enable-prefix-caching.

MTP-1 (num_speculative_tokens: 1)

Per community Qwen3.5 latency recipe. Disables --enable-prefix-caching (mutually exclusive). Max concurrency at full 262K drops to 11.03× (vs 12.56× for baseline).

cMTP-1 tok/sMTP-1 TPOT p50Baseline tok/sBaseline TPOT p50Δ throughputΔ TPOTFailed
128.437ms28.730ms≈+23%0
463.859ms51.575ms+24%−21%0
885.776ms70.298ms+22%−22%0

Surprising result: MTP is typically pitched as a single-stream latency win. On our model+hardware it’s actually a mid-concurrency throughput win — c=1 is flat (speculative overhead eats the small acceptance benefit), but c=4–8 gain ~22% throughput and ~20% lower TPOT. Hypothesis: moderate MTP acceptance rate + amortization of draft computation across concurrent requests.

Trade vs --enable-prefix-caching: MTP-1 is mutually exclusive with prefix caching. Under random-prompt bench (0% prefix hit rate), MTP-1 wins clearly. Under real agentic traffic with long shared system prompts, prefix caching may win. Worth re-measuring once real orchestrator traffic exists.

Current production config uses MTP-1 (decision 2026-04-23). Rollback path: relaunch with --enable-prefix-caching instead of the speculative config flag.

Appendix B — Reproducing this benchmark

From the DGX:

ssh dgx-NN
KEY=<VLLM_API_KEY>
for C in 1 2 4 8 16 32; do
  N=$(( C * 4 < 8 ? 8 : C * 4 ))
  echo "=== concurrency=$C num_prompts=$N ==="
  docker exec vllm vllm bench serve \
    --backend openai-chat \
    --base-url http://localhost:8000 \
    --endpoint /v1/chat/completions \
    --header "Authorization=Bearer $KEY" \
    --model qwen3.6-35b \
    --tokenizer Qwen/Qwen3.6-35B-A3B-FP8 \
    --dataset-name random --random-input-len 512 --random-output-len 256 --random-range-ratio 0.1 \
    --num-prompts "$N" --num-warmups 2 --max-concurrency "$C" --ignore-eos \
    --extra-body '{"chat_template_kwargs": {"enable_thinking": false}}' \
    --percentile-metrics ttft,tpot,itl,e2el --metric-percentiles 50,95,99 \
    --disable-tqdm 2>&1 | grep -E "^(Successful|Failed|Benchmark duration|Output token throughput|Peak|Mean TTFT|P50 TTFT|P95 TTFT|Mean TPOT|P50 TPOT|P95 TPOT|Mean E2EL|P50 E2EL|P95 E2EL)"
  echo
done

Note on --tokenizer: required because --served-model-name qwen3.6-35b aliases the model; without --tokenizer, bench tries to fetch qwen3.6-35b from HuggingFace and errors with OSError: not a local folder and is not a valid model identifier.

Discard the first bench run after any container restart (kernel autotune dominates).

References