Technical Writing
Benchmark results for AEON-7/Qwen3.6-35B-A3B-heretic-NVFP4 + DFlash (NVIDIA DGX Spark, GB10)
The previous post rejected MTP-1/2 and projected ~2× from NVFP4. We measured +152% aggregate at c=10 and discovered that published benchmark numbers are only honest when paired with the workload that produced them.
DGX Benchmark Report — Qwen3.6-35B-A3B-NVFP4 + DFlash
Benchmark results for AEON-7/Qwen3.6-35B-A3B-heretic-NVFP4 + z-lab/Qwen3.6-35B-A3B-DFlash drafter, served on a DGX Spark (GB10) via the AEON-7 Spark-tuned vLLM image.
TL;DR
- DFlash (k=5) replaces MTP-1 as the spec-decode path on this model: +23.2% c=10 aggregate throughput, +3.8% c=1, measured against the same hardware. The previous post’s MTP-1 +22% at c=8 was the previous best; DFlash beats it on every metric that matters.
- Single-stream NVFP4 + DFlash hits 77 tok/s on math reasoning, 31 tok/s on mixed traffic — a 2.5× spread from prompt class alone. The drafter’s lifetime 9.6% acceptance is real but misleading; the drafter is tuned for math/code reasoning and is doing its job there.
num_speculative_tokens=15(the AEON-7 default) is mostly wasted compute: pos-0 acceptance 78%, pos-8 acceptance 2.3%. Reducing draft length to 5 captures the high-acceptance head and reclaims the rest as forward-pass budget. Cold-start unchanged.- MTP-1 / MTP-2 from the previous post remain rejected. DFlash is the spec-decode choice that survives measurement on this hardware.
- The 4×4 natural-prompt matrix matches the AEON-7 published numbers within 3% (code/reasoning) and 13% (prose) at c=1. At c=16, our 35B saturates 27–29% earlier than the published 27B reference — the expected memory-budget effect (more weights, same 128 GB).
Environment
Hardware
- GPU: NVIDIA GB10 (Grace Blackwell), capability (12, 1) / sm_121
- Memory: 121 GiB unified LPDDR5x, 273 GB/s bandwidth
- CPU: 20-core ARM Cortex-A725, arm64
- Storage: 4 TB NVMe
Software
- OS: Ubuntu 24.04 LTS, kernel 6.17.x-nvidia
- NVIDIA driver: 580.142 (pinned via
apt-mark hold) - CUDA: 13.0 host / 13.2 container via forward-compat (
cuda-compat-13-2package) - Container runtime: Docker CE 29.2.1 + nvidia-container-toolkit 1.19.0
- Image:
ghcr.io/aeon-7/vllm-spark-omni-q36:v1.2(Spark-tuned, hardcoded MARLIN MoE backend, DFlash support) - vLLM: 0.1.dev1+gbfde49e28.d20260418 (custom AEON-7 build, ~9 days before upstream v0.20.0)
- PyTorch: 2.12.0.dev20260408+cu130
- Model:
AEON-7/Qwen3.6-35B-A3B-heretic-NVFP4(35B total / 3B active, MoE 256 experts top-8+1, hybrid attention + Mamba, NVFP4 quantization) - Drafter:
z-lab/Qwen3.6-35B-A3B-DFlash(~905 MB) - Weights on disk:
/home/system/qwen36/(22 GB + 905 MB)
vLLM configuration
Flags in use at measurement time (see Appendix B for the reproducible command):
--model AEON-7/Qwen3.6-35B-A3B-heretic-NVFP4
--speculative-model z-lab/Qwen3.6-35B-A3B-DFlash
--num-speculative-tokens 5 # AEON-7 default is 15; see DFlash k=5 vs k=15
--moe-backend marlin # forced; auto-select picks wrong on GB10
--max-model-len 262144
--gpu-memory-utilization 0.85
--enable-prefix-caching
--chat-template-kwargs '{"enable_thinking": false}'
Container flags: --gpus all --network host --ipc host --ulimit memlock=-1.
Engine-reported capacity at boot
- Available KV cache memory: 49.25 GiB
- GPU KV cache size: 846,912 tokens
- Maximum concurrency at full 262K context: 12.56×
- Cold start: 10–15 minutes end-to-end, dominated by
torch.compileon the main backbone (~200 s). Mount/root/.cache/vllmto a persistent host volume to preserve the AOT cache acrossdocker rm. - First inference after container start takes ~60 s (kernel autotune). Discard from benchmarks.
Methodology
This post replaces the previous post’s random-token benchmark with the AEON-7 natural-prompt matrix (their bench_natural.py, mirrored into the local harness) so results are comparable to AEON-7’s published numbers.
- Endpoint: LiteLLM → Nginx → vLLM on the DGX (production traffic path). The previous post measured inside the container; the headline numbers here are through the proxy and include a gateway tax.
- Prompt set: 4 classes (code, reasoning, dialogue, prose) from the AEON-7 BENCHMARKS.md, with the exact prompt strings. Reasoning is the DFlash sweet spot; prose is the worst case.
- Per-class × per-concurrency: 16 requests, max-tokens 512, temperature 0, no-thinking mode.
- Concurrency levels: 1, 4, 8, 16.
- Warmup: 3 serial requests per class, discarded from stats.
- Cache-busting: each request prefixed with
[bench-N]to defeat shared-prefix caching. - Tool: a stdlib-only Python harness that streams
/v1/chat/completionsSSE and captures TTFT, TPOT, aggregate throughput, and median output tokens.
The previous post’s random-token dataset is adversarial for MoE routing AND for DFlash acceptance. The natural-prompt matrix is the right shape for measuring a spec-decoded stack on real workloads.
Results
Primary: 4×4 natural-prompt matrix
| class | c | agg tok/s | per-req p50 | TTFT p50 | TTFT p95 | TPOT p50 | TPOT p95 | med out | wall |
|---|---|---|---|---|---|---|---|---|---|
| code | 1 | 62.0 | 61.8 | 557 ms | 607 ms | 15.1 ms | 16.0 ms | 512 | 128.0 s |
| code | 4 | 135.1 | 35.0 | 796 ms | 872 ms | 27.0 ms | 31.1 ms | 512 | 59.9 s |
| code | 8 | 194.1 | 25.8 | 883 ms | 1126 ms | 36.8 ms | 40.5 ms | 512 | 40.8 s |
| code | 16 | 241.0 | 16.2 | 1654 ms | 1685 ms | 58.7 ms | 64.4 ms | 512 | 32.9 s |
| reasoning | 1 | 51.7 | 52.3 | 605 ms | 797 ms | 17.9 ms | 21.8 ms | 512 | 158.5 s |
| reasoning | 4 | 125.2 | 31.2 | 765 ms | 891 ms | 30.5 ms | 33.3 ms | 512 | 65.4 s |
| reasoning | 8 | 181.2 | 23.4 | 1019 ms | 1096 ms | 41.0 ms | 44.4 ms | 512 | 45.2 s |
| reasoning | 16 | 222.9 | 14.3 | 1527 ms | 1757 ms | 67.2 ms | 68.7 ms | 512 | 36.8 s |
| dialogue | 1 | 12.0 | 12.0 | 549 ms | 646 ms | 50.3 ms | 59.4 ms | 15 | 21.4 s |
| dialogue | 4 | 31.6 | 8.3 | 756 ms | 879 ms | 81.1 ms | 90.0 ms | 19 | 8.7 s |
| dialogue | 8 | 47.4 | 6.1 | 919 ms | 1021 ms | 105.6 ms | 121.8 ms | 14 | 5.0 s |
| dialogue | 16 | 55.7 | 4.0 | 1167 ms | 1399 ms | 184.7 ms | 195.0 ms | 14 | 4.2 s |
| prose | 1 | 25.7 | 26.3 | 559 ms | 642 ms | 37.0 ms | 43.1 ms | 512 | 318.9 s |
| prose | 4 | 62.5 | 15.7 | 791 ms | 1201 ms | 62.2 ms | 64.7 ms | 512 | 131.1 s |
| prose | 8 | 87.2 | 11.3 | 978 ms | 1025 ms | 87.0 ms | 93.8 ms | 512 | 93.9 s |
| prose | 16 | 107.8 | 6.8 | 1240 ms | 1279 ms | 143.9 ms | 146.2 ms | 512 | 76.0 s |
Dialogue prompt is broken for steady-state benchmarking. Median output is 14–19 tokens (well below the 512 cap) — the Alice/Bob prompt elicits a short reply and EOS, so DFlash never warms up. Treat dialogue numbers as prompt-processing cost, not decode throughput. Flagged for the next matrix revision.
Headline: NVFP4 + DFlash vs FP8 baseline (deck numbers)
These are the numbers that drove the production cutover. Measured through the LiteLLM proxy on math-reasoning prompts at max-tokens=200, single fresh container restart per condition. The previous post’s FP8 numbers were measured direct inside the container, so the apples-to-oranges delta understates the real gain.
| Concurrency | FP8 + MTP-1 (prev. post, direct) | NVFP4 + DFlash k=5 (this post, proxied) | Delta |
|---|---|---|---|
| 1 | 28–30 tok/s | 77.7 tok/s | +94% |
| 5 | 106 tok/s | 122 tok/s | +15% |
| 10 | 96 tok/s | 242 tok/s | +152% |
| 20 | 212 tok/s | 222 tok/s | +5% |
The previous post projected “NVFP4 quantization: ~2× theoretical, 55–60 tok/s projected.” Measured: 77 tok/s single-stream, 242 tok/s aggregate at c=10. The projection was conservative. NVFP4 works on GB10 once the right MoE backend is selected.
DFlash k=5 vs k=15
The AEON-7 image defaults to num_speculative_tokens=15. Per-position acceptance falls off the cliff past position 4: pos-0 51.8%, pos-1 32.3%, pos-2 20.4%, pos-3 12.4%, pos-4 8.1%, pos-5 5.4%, pos-8 2.3%, pos-14 0.6%. Positions 0–4 capture most of the high-acceptance head; positions 5–14 contribute < 1 accepted token combined. Reducing draft length to 5 cuts drafter forward-pass cost by 67% per verify step.
| k=15 (AEON-7 default) | k=5 | Delta | |
|---|---|---|---|
| Draft-token acceptance ratio | 9.6% | 30.8% | +21.2 pp |
| Mean accepted / draft | 1.44 | 1.54 | +7% |
| c=1 agg tok/s | 31.9 | 33.1 | +3.8% |
| c=10 agg tok/s | 103.4 | 127.4 | +23.2% |
| c=10 TPOT p50 | 86.7 ms | 66.0 ms | −24% |
| c=10 TTFT p50 | 1018 ms | 944 ms | −7% |
| c=10 TTFT p95 | 1134 ms | 3067 ms | +170% ⚠️ |
⚠️ The c=10 TTFT p95 widening is at N=30, where p95 = sample 28–29 of 30 — extremely noise-sensitive. Could be real or sample variance. Need N≥100 to confirm before promoting to production. (Status: working theory, not validated.)
k=5 is the production recommendation for math/code reasoning workloads. For mixed-traffic workloads with effective acceptance closer to 1.5 tokens/draft, k=4 may be optimal — not yet swept.
Interpretation
Capacity guidance for orchestrator tuning
| User load | Per-user experience | Suitable workloads |
|---|---|---|
| 1–4 concurrent | Snappy — sub-second TTFT, sub-30 ms/token on code, sub-70 ms on reasoning | Interactive chat, real-time coding assist |
| 5–8 concurrent | Fine for work — TTFT stays under 1 s, per-req TPOT grows linearly | Mixed-user team, agentic coding |
| 9–16 concurrent | Per-req feels slow for chat; fine for agentic | Long-context summarization, orchestrator internal calls |
| 17+ concurrent | Agentic only | Bulk generation, batch jobs |
Recommended orchestrator settings:
- Default soft ceiling: c=8 (balance of latency + throughput)
- Burst allowance up to c=16 before returning 429 or spilling to fallback
- For math/code reasoning workloads, set
num_speculative_tokens=5in the request body — k=5 wins by 23% on the workloads the drafter was tuned for - Aggressive prompt-prefix caching at the orchestrator: cache-busted benches above show 0% hit; real traffic with shared system prompts will benefit substantially
Workload-class reality check
The single most important finding from this work is prompt-class sensitivity. The 77 tok/s headline is real but only on math/code reasoning; the drafter’s lifetime 9.6% acceptance is averaged across all production traffic, and most production traffic is not math/code reasoning.
| Workload | c=1 tok/s | DFlash pos-0 accept |
|---|---|---|
| Math reasoning, max-tokens=200, fresh container | 77.7 | 78% (deck headline) |
| Math reasoning, repeat (same prompt, req 7) | 25.3 | 41% (first-request advantage degrades) |
| Code prompt, max-tokens=200 | 58.6 | ~75% |
| Mixed bench (5 prompt rotation) | 31.9 | ~52% (realistic production) |
| Prose, max-tokens=512 | 25.7 | ~30% (DFlash weak case) |
Published benchmark numbers without the workload that produced them are not honest. Future benchmark claims should pair every tok/s number with the prompt class and max_tokens that produced it. The harness now supports a --workload {math-reasoning, generic-explain, code, creative} flag for this reason.
Bottleneck analysis
Spark is memory-bandwidth bound on decode. For this NVFP4 model:
theoretical ceiling = bandwidth / bytes_per_active_param_per_token
= 273 GB/s / ~1.5 GB (NVFP4 active params, ~3B active @ 4-bit)
≈ 182 tok/s (single-stream theoretical)
FP8 was 91 tok/s theoretical (3B active @ 8-bit, ~3 GB). NVFP4 should roughly double it.
Measured single-stream math-reasoning throughput: 77 tok/s. Realization: ~42% of theoretical, vs the 32% on FP8 in the previous post. DFlash closes the gap to theoretical by drafting cheaply — every accepted draft token saves a full main-model forward pass.
The bottleneck has moved: with FP8 alone, we were bandwidth-bound on the main-model forward pass. With NVFP4 + DFlash, we are bandwidth-bound on drafter scheduling — the drafter’s forward pass and the verify step compete for the same memory bus. The k=15 → k=5 reduction is a drafter-scheduling optimization, not a kernel one.
What would move the ceiling up (not available today):
- sm_121-native FP4 kernels (instead of MARLIN W4A16 fallback): expected +15–25%. Waiting on upstream FlashInfer 0.6.8+ for SM120 kernels.
- A purpose-trained 35B-A3B drafter: the published drafter is a generic 27B-style checkpoint. A drafter trained against this exact main model would likely extend the high-acceptance head by 1–2 positions, worth another ~10% at c=10.
- TP=2 across two Sparks via ConnectX-7: helps aggregate, not single-stream.
What does not move it:
num_speculative_tokens=15(AEON-7 default): +0% at this drafter. The high-acceptance head is positions 0–4.- Kernel autotune, cudagraph, compile flags: already on by default, no room.
The previous post projected +2× from NVFP4. We measured +152% on aggregate, 94% on single-stream. The remaining gap to theoretical is drafter scheduling, not the kernel.
Appendix A — Experiments rejected
MTP-1 / MTP-2 (re-stated from the previous post)
The previous post rejected MTP-1 (best +22% c=8 over FP8 baseline) and MTP-2 (strictly worse, broke under c=8). Both remain rejected. For the NVFP4 stack, DFlash is the spec-decode path that survives measurement — pos-0 acceptance 78% (vs MTP’s typical 30–50%), drafter overhead amortized across concurrent requests, no --enable-prefix-caching mutual exclusion.
Stock vLLM 0.20.0 with default flags
Upstream vllm/vllm-openai:v0.20.0 with identical serve flags underperforms the AEON-7 image on this hardware. Two independent regressions: (1) MoE backend auto-selection picks FLASHINFER_CUTLASS instead of MARLIN — force --moe-backend marlin to recover. (2) DFlash spec-decode is reinterpreted as EAGLE with auxiliary layers (1, 10, 19, 28, 37), which at c=10 triples TTFT p95 from 1344 ms → 4191 ms. The AEON-7 image runs the same drafter against the same flags without the c=10 tail blowup. Measured delta vs baseline (c=1 / c=10 / c=10 p95): +(-7.2% / -14.6% / +212%) stock, +(+7.6% / -7.7% / +198%) with MARLIN forced, +(-18.6% / +6.7% / -4%) with MARLIN and no spec. Stay on the AEON-7 image until those patches land upstream — its value is the Spark-specific patches, not the vLLM revision.
num_speculative_tokens=15 (AEON-7 default)
Tried as the production default. See DFlash k=5 vs k=15 above. Net result: k=5 wins by 23% c=10 on the workloads the drafter was tuned for. Rollback path is a one-line edit to the compose file.
Appendix B — Reproducing this benchmark
From the DGX (commands assume the AEON-7 image is already running on port 8000 with the drafter loaded):
KEY=<VLLM_API_KEY>
for p in code reasoning dialogue prose; do
for c in 1 4 8 16; do
python3 scripts/bench/run.py \
--base-url http://localhost:8000 \
--model AEON-7/Qwen3.6-35B-A3B-heretic-NVFP4 \
--prompt-class "$p" \
--concurrency "$c" --num-requests 16 --max-tokens 512 \
--temperature 0 --warmup 3 --cache-bust \
--extra-body '{"chat_template_kwargs": {"enable_thinking": false}}' \
--output "results/${p}-c${c}.json"
done
done
Discard the first bench run after any container restart (kernel autotune + DFlash warmup dominate). Pin num_speculative_tokens=5 in the serve command; pin --moe-backend marlin to survive a future vLLM upgrade. The full natural-prompt matrix takes ~25 minutes wall-clock. The exact prompt strings are mirrored from the AEON-7 BENCHMARKS.md so any rerun is comparable to AEON-7’s published numbers.
References

- AEON-7/vllm-dflash (GitHub) — DFlash drafter and
bench_natural.pymethodology used in this post - AEON-7/Qwen3.6-35B-A3B-heretic-NVFP4 (HuggingFace) — NVFP4 model weights
- z-lab/Qwen3.6-35B-A3B-DFlash (HuggingFace) — DFlash drafter checkpoint
ghcr.io/aeon-7/vllm-spark-omni-q36:v1.2— Spark-tuned vLLM image used- Previous post: Benchmark results for Qwen/Qwen3.6-35B-A3B-FP8 (NVIDIA DGX Spark, GB10) serving via vLLM — the FP8 baseline