Technical Writing

Benchmark results for AEON-7/Qwen3.6-35B-A3B-heretic-NVFP4 + DFlash (NVIDIA DGX Spark, GB10)

· 14 min read

The previous post rejected MTP-1/2 and projected ~2× from NVFP4. We measured +152% aggregate at c=10 and discovered that published benchmark numbers are only honest when paired with the workload that produced them.

DGX Benchmark Report — Qwen3.6-35B-A3B-NVFP4 + DFlash

Benchmark results for AEON-7/Qwen3.6-35B-A3B-heretic-NVFP4 + z-lab/Qwen3.6-35B-A3B-DFlash drafter, served on a DGX Spark (GB10) via the AEON-7 Spark-tuned vLLM image.

TL;DR

  • DFlash (k=5) replaces MTP-1 as the spec-decode path on this model: +23.2% c=10 aggregate throughput, +3.8% c=1, measured against the same hardware. The previous post’s MTP-1 +22% at c=8 was the previous best; DFlash beats it on every metric that matters.
  • Single-stream NVFP4 + DFlash hits 77 tok/s on math reasoning, 31 tok/s on mixed traffic — a 2.5× spread from prompt class alone. The drafter’s lifetime 9.6% acceptance is real but misleading; the drafter is tuned for math/code reasoning and is doing its job there.
  • num_speculative_tokens=15 (the AEON-7 default) is mostly wasted compute: pos-0 acceptance 78%, pos-8 acceptance 2.3%. Reducing draft length to 5 captures the high-acceptance head and reclaims the rest as forward-pass budget. Cold-start unchanged.
  • MTP-1 / MTP-2 from the previous post remain rejected. DFlash is the spec-decode choice that survives measurement on this hardware.
  • The 4×4 natural-prompt matrix matches the AEON-7 published numbers within 3% (code/reasoning) and 13% (prose) at c=1. At c=16, our 35B saturates 27–29% earlier than the published 27B reference — the expected memory-budget effect (more weights, same 128 GB).

Environment

Hardware

  • GPU: NVIDIA GB10 (Grace Blackwell), capability (12, 1) / sm_121
  • Memory: 121 GiB unified LPDDR5x, 273 GB/s bandwidth
  • CPU: 20-core ARM Cortex-A725, arm64
  • Storage: 4 TB NVMe

Software

  • OS: Ubuntu 24.04 LTS, kernel 6.17.x-nvidia
  • NVIDIA driver: 580.142 (pinned via apt-mark hold)
  • CUDA: 13.0 host / 13.2 container via forward-compat (cuda-compat-13-2 package)
  • Container runtime: Docker CE 29.2.1 + nvidia-container-toolkit 1.19.0
  • Image: ghcr.io/aeon-7/vllm-spark-omni-q36:v1.2 (Spark-tuned, hardcoded MARLIN MoE backend, DFlash support)
  • vLLM: 0.1.dev1+gbfde49e28.d20260418 (custom AEON-7 build, ~9 days before upstream v0.20.0)
  • PyTorch: 2.12.0.dev20260408+cu130
  • Model: AEON-7/Qwen3.6-35B-A3B-heretic-NVFP4 (35B total / 3B active, MoE 256 experts top-8+1, hybrid attention + Mamba, NVFP4 quantization)
  • Drafter: z-lab/Qwen3.6-35B-A3B-DFlash (~905 MB)
  • Weights on disk: /home/system/qwen36/ (22 GB + 905 MB)

vLLM configuration

Flags in use at measurement time (see Appendix B for the reproducible command):

--model AEON-7/Qwen3.6-35B-A3B-heretic-NVFP4
--speculative-model z-lab/Qwen3.6-35B-A3B-DFlash
--num-speculative-tokens 5                  # AEON-7 default is 15; see DFlash k=5 vs k=15
--moe-backend marlin                        # forced; auto-select picks wrong on GB10
--max-model-len 262144
--gpu-memory-utilization 0.85
--enable-prefix-caching
--chat-template-kwargs '{"enable_thinking": false}'

Container flags: --gpus all --network host --ipc host --ulimit memlock=-1.

Engine-reported capacity at boot

  • Available KV cache memory: 49.25 GiB
  • GPU KV cache size: 846,912 tokens
  • Maximum concurrency at full 262K context: 12.56×
  • Cold start: 10–15 minutes end-to-end, dominated by torch.compile on the main backbone (~200 s). Mount /root/.cache/vllm to a persistent host volume to preserve the AOT cache across docker rm.
  • First inference after container start takes ~60 s (kernel autotune). Discard from benchmarks.

Methodology

This post replaces the previous post’s random-token benchmark with the AEON-7 natural-prompt matrix (their bench_natural.py, mirrored into the local harness) so results are comparable to AEON-7’s published numbers.

  • Endpoint: LiteLLM → Nginx → vLLM on the DGX (production traffic path). The previous post measured inside the container; the headline numbers here are through the proxy and include a gateway tax.
  • Prompt set: 4 classes (code, reasoning, dialogue, prose) from the AEON-7 BENCHMARKS.md, with the exact prompt strings. Reasoning is the DFlash sweet spot; prose is the worst case.
  • Per-class × per-concurrency: 16 requests, max-tokens 512, temperature 0, no-thinking mode.
  • Concurrency levels: 1, 4, 8, 16.
  • Warmup: 3 serial requests per class, discarded from stats.
  • Cache-busting: each request prefixed with [bench-N] to defeat shared-prefix caching.
  • Tool: a stdlib-only Python harness that streams /v1/chat/completions SSE and captures TTFT, TPOT, aggregate throughput, and median output tokens.

The previous post’s random-token dataset is adversarial for MoE routing AND for DFlash acceptance. The natural-prompt matrix is the right shape for measuring a spec-decoded stack on real workloads.

Results

Primary: 4×4 natural-prompt matrix

classcagg tok/sper-req p50TTFT p50TTFT p95TPOT p50TPOT p95med outwall
code162.061.8557 ms607 ms15.1 ms16.0 ms512128.0 s
code4135.135.0796 ms872 ms27.0 ms31.1 ms51259.9 s
code8194.125.8883 ms1126 ms36.8 ms40.5 ms51240.8 s
code16241.016.21654 ms1685 ms58.7 ms64.4 ms51232.9 s
reasoning151.752.3605 ms797 ms17.9 ms21.8 ms512158.5 s
reasoning4125.231.2765 ms891 ms30.5 ms33.3 ms51265.4 s
reasoning8181.223.41019 ms1096 ms41.0 ms44.4 ms51245.2 s
reasoning16222.914.31527 ms1757 ms67.2 ms68.7 ms51236.8 s
dialogue112.012.0549 ms646 ms50.3 ms59.4 ms1521.4 s
dialogue431.68.3756 ms879 ms81.1 ms90.0 ms198.7 s
dialogue847.46.1919 ms1021 ms105.6 ms121.8 ms145.0 s
dialogue1655.74.01167 ms1399 ms184.7 ms195.0 ms144.2 s
prose125.726.3559 ms642 ms37.0 ms43.1 ms512318.9 s
prose462.515.7791 ms1201 ms62.2 ms64.7 ms512131.1 s
prose887.211.3978 ms1025 ms87.0 ms93.8 ms51293.9 s
prose16107.86.81240 ms1279 ms143.9 ms146.2 ms51276.0 s

Dialogue prompt is broken for steady-state benchmarking. Median output is 14–19 tokens (well below the 512 cap) — the Alice/Bob prompt elicits a short reply and EOS, so DFlash never warms up. Treat dialogue numbers as prompt-processing cost, not decode throughput. Flagged for the next matrix revision.

Headline: NVFP4 + DFlash vs FP8 baseline (deck numbers)

These are the numbers that drove the production cutover. Measured through the LiteLLM proxy on math-reasoning prompts at max-tokens=200, single fresh container restart per condition. The previous post’s FP8 numbers were measured direct inside the container, so the apples-to-oranges delta understates the real gain.

ConcurrencyFP8 + MTP-1 (prev. post, direct)NVFP4 + DFlash k=5 (this post, proxied)Delta
128–30 tok/s77.7 tok/s+94%
5106 tok/s122 tok/s+15%
1096 tok/s242 tok/s+152%
20212 tok/s222 tok/s+5%

The previous post projected “NVFP4 quantization: ~2× theoretical, 55–60 tok/s projected.” Measured: 77 tok/s single-stream, 242 tok/s aggregate at c=10. The projection was conservative. NVFP4 works on GB10 once the right MoE backend is selected.

DFlash k=5 vs k=15

The AEON-7 image defaults to num_speculative_tokens=15. Per-position acceptance falls off the cliff past position 4: pos-0 51.8%, pos-1 32.3%, pos-2 20.4%, pos-3 12.4%, pos-4 8.1%, pos-5 5.4%, pos-8 2.3%, pos-14 0.6%. Positions 0–4 capture most of the high-acceptance head; positions 5–14 contribute < 1 accepted token combined. Reducing draft length to 5 cuts drafter forward-pass cost by 67% per verify step.

k=15 (AEON-7 default)k=5Delta
Draft-token acceptance ratio9.6%30.8%+21.2 pp
Mean accepted / draft1.441.54+7%
c=1 agg tok/s31.933.1+3.8%
c=10 agg tok/s103.4127.4+23.2%
c=10 TPOT p5086.7 ms66.0 ms−24%
c=10 TTFT p501018 ms944 ms−7%
c=10 TTFT p951134 ms3067 ms+170% ⚠️

⚠️ The c=10 TTFT p95 widening is at N=30, where p95 = sample 28–29 of 30 — extremely noise-sensitive. Could be real or sample variance. Need N≥100 to confirm before promoting to production. (Status: working theory, not validated.)

k=5 is the production recommendation for math/code reasoning workloads. For mixed-traffic workloads with effective acceptance closer to 1.5 tokens/draft, k=4 may be optimal — not yet swept.

Interpretation

Capacity guidance for orchestrator tuning

User loadPer-user experienceSuitable workloads
1–4 concurrentSnappy — sub-second TTFT, sub-30 ms/token on code, sub-70 ms on reasoningInteractive chat, real-time coding assist
5–8 concurrentFine for work — TTFT stays under 1 s, per-req TPOT grows linearlyMixed-user team, agentic coding
9–16 concurrentPer-req feels slow for chat; fine for agenticLong-context summarization, orchestrator internal calls
17+ concurrentAgentic onlyBulk generation, batch jobs

Recommended orchestrator settings:

  • Default soft ceiling: c=8 (balance of latency + throughput)
  • Burst allowance up to c=16 before returning 429 or spilling to fallback
  • For math/code reasoning workloads, set num_speculative_tokens=5 in the request body — k=5 wins by 23% on the workloads the drafter was tuned for
  • Aggressive prompt-prefix caching at the orchestrator: cache-busted benches above show 0% hit; real traffic with shared system prompts will benefit substantially

Workload-class reality check

The single most important finding from this work is prompt-class sensitivity. The 77 tok/s headline is real but only on math/code reasoning; the drafter’s lifetime 9.6% acceptance is averaged across all production traffic, and most production traffic is not math/code reasoning.

Workloadc=1 tok/sDFlash pos-0 accept
Math reasoning, max-tokens=200, fresh container77.778% (deck headline)
Math reasoning, repeat (same prompt, req 7)25.341% (first-request advantage degrades)
Code prompt, max-tokens=20058.6~75%
Mixed bench (5 prompt rotation)31.9~52% (realistic production)
Prose, max-tokens=51225.7~30% (DFlash weak case)

Published benchmark numbers without the workload that produced them are not honest. Future benchmark claims should pair every tok/s number with the prompt class and max_tokens that produced it. The harness now supports a --workload {math-reasoning, generic-explain, code, creative} flag for this reason.

Bottleneck analysis

Spark is memory-bandwidth bound on decode. For this NVFP4 model:

theoretical ceiling = bandwidth / bytes_per_active_param_per_token
                    = 273 GB/s  /  ~1.5 GB            (NVFP4 active params, ~3B active @ 4-bit)
                    ≈ 182 tok/s (single-stream theoretical)

FP8 was 91 tok/s theoretical (3B active @ 8-bit, ~3 GB). NVFP4 should roughly double it.

Measured single-stream math-reasoning throughput: 77 tok/s. Realization: ~42% of theoretical, vs the 32% on FP8 in the previous post. DFlash closes the gap to theoretical by drafting cheaply — every accepted draft token saves a full main-model forward pass.

The bottleneck has moved: with FP8 alone, we were bandwidth-bound on the main-model forward pass. With NVFP4 + DFlash, we are bandwidth-bound on drafter scheduling — the drafter’s forward pass and the verify step compete for the same memory bus. The k=15 → k=5 reduction is a drafter-scheduling optimization, not a kernel one.

What would move the ceiling up (not available today):

  • sm_121-native FP4 kernels (instead of MARLIN W4A16 fallback): expected +15–25%. Waiting on upstream FlashInfer 0.6.8+ for SM120 kernels.
  • A purpose-trained 35B-A3B drafter: the published drafter is a generic 27B-style checkpoint. A drafter trained against this exact main model would likely extend the high-acceptance head by 1–2 positions, worth another ~10% at c=10.
  • TP=2 across two Sparks via ConnectX-7: helps aggregate, not single-stream.

What does not move it:

  • num_speculative_tokens=15 (AEON-7 default): +0% at this drafter. The high-acceptance head is positions 0–4.
  • Kernel autotune, cudagraph, compile flags: already on by default, no room.

The previous post projected +2× from NVFP4. We measured +152% on aggregate, 94% on single-stream. The remaining gap to theoretical is drafter scheduling, not the kernel.

Appendix A — Experiments rejected

MTP-1 / MTP-2 (re-stated from the previous post)

The previous post rejected MTP-1 (best +22% c=8 over FP8 baseline) and MTP-2 (strictly worse, broke under c=8). Both remain rejected. For the NVFP4 stack, DFlash is the spec-decode path that survives measurement — pos-0 acceptance 78% (vs MTP’s typical 30–50%), drafter overhead amortized across concurrent requests, no --enable-prefix-caching mutual exclusion.

Stock vLLM 0.20.0 with default flags

Upstream vllm/vllm-openai:v0.20.0 with identical serve flags underperforms the AEON-7 image on this hardware. Two independent regressions: (1) MoE backend auto-selection picks FLASHINFER_CUTLASS instead of MARLIN — force --moe-backend marlin to recover. (2) DFlash spec-decode is reinterpreted as EAGLE with auxiliary layers (1, 10, 19, 28, 37), which at c=10 triples TTFT p95 from 1344 ms → 4191 ms. The AEON-7 image runs the same drafter against the same flags without the c=10 tail blowup. Measured delta vs baseline (c=1 / c=10 / c=10 p95): +(-7.2% / -14.6% / +212%) stock, +(+7.6% / -7.7% / +198%) with MARLIN forced, +(-18.6% / +6.7% / -4%) with MARLIN and no spec. Stay on the AEON-7 image until those patches land upstream — its value is the Spark-specific patches, not the vLLM revision.

num_speculative_tokens=15 (AEON-7 default)

Tried as the production default. See DFlash k=5 vs k=15 above. Net result: k=5 wins by 23% c=10 on the workloads the drafter was tuned for. Rollback path is a one-line edit to the compose file.

Appendix B — Reproducing this benchmark

From the DGX (commands assume the AEON-7 image is already running on port 8000 with the drafter loaded):

KEY=<VLLM_API_KEY>
for p in code reasoning dialogue prose; do
  for c in 1 4 8 16; do
    python3 scripts/bench/run.py \
      --base-url http://localhost:8000 \
      --model AEON-7/Qwen3.6-35B-A3B-heretic-NVFP4 \
      --prompt-class "$p" \
      --concurrency "$c" --num-requests 16 --max-tokens 512 \
      --temperature 0 --warmup 3 --cache-bust \
      --extra-body '{"chat_template_kwargs": {"enable_thinking": false}}' \
      --output "results/${p}-c${c}.json"
  done
done

Discard the first bench run after any container restart (kernel autotune + DFlash warmup dominate). Pin num_speculative_tokens=5 in the serve command; pin --moe-backend marlin to survive a future vLLM upgrade. The full natural-prompt matrix takes ~25 minutes wall-clock. The exact prompt strings are mirrored from the AEON-7 BENCHMARKS.md so any rerun is comparable to AEON-7’s published numbers.

References

AEON-7 HuggingFace org page — 30 models, 5 collections, with NVFP4 quantization variants prominent. The Qwen3.6-35B-A3B-heretic-NVFP4 model benchmarked in this post sits in the right column of the Models grid.