The NVIDIA RTX PRO 6000 Blackwell combines 96 GB of GDDR7 memory with native FP4 support, offering considerable headroom for serving large language models and handling more concurrent requests on a single GPU. When speculative decoding enters the equation, however, choosing the right format becomes more than a matter of model size or theoretical compute performance.
FP8 offers higher precision and keeps the model closer to its original behavior, while NVFP4 reduces the memory footprint and leaves more VRAM available for the KV cache. So which format makes speculative decoding more effective on the NVIDIA RTX PRO 6000?
In this post, we will compare FP8 and NVFP4 under a real-world serving workload to determine which delivers better throughput and scalability – and ultimately, greater practical value.
Why Speculative Decoding Exists
Generating text one token at a time wastes a GPU. Every step drags the model’s entire weight set out of memory to produce a single token — decoding is bound by memory bandwidth, not math, so the compute units sit mostly idle. Speculative decoding spends that idle capacity: a small drafter proposes k tokens ahead, and the big target model verifies all k in one forward pass. Matching guesses cost the price of one token; at the first mismatch, everything after it is discarded. The output is mathematically identical to normal decoding — a pure speed optimization, not a quality trade-off.

That idle capacity shrinks as weights get faster, and batches get bigger — exactly what format choice and concurrency change.
Three Ways to Buy Free Tokens
- MTP ships inside the checkpoint (mtp.* tensors) — no second model, no extra VRAM. Free, in every sense.
- DSPARK is an external 1.36B BF16 drafter with a confidence head that decides how many tokens to propose per request.
- DFlash2 drafts by block diffusion — predicting 8 positions at once and tracing the best coherent path through them.
On paper, all three sit within a few percent of each other. That parity is what makes format the more interesting variable.
The Test Setup
One NVIDIA RTX PRO 6000 Blackwell (96 GB, sm_120, CUDA 13.2) per arm, three arms in parallel, SGLang with FlashInfer attention and an FP8 KV cache. Workload: NVIDIA’s SPEED-Bench (throughput_8k, low_entropy), 300 prompts of ~8.8K tokens, swept across concurrency 10 → 150. –ignore-eos forces a fixed 1024-token decode, so decode is a real share of the work; replicas restart between levels for a cold cache each time. Two formats ran the identical sweep: Qwen3.8-27B-FP8 (~27 GB) and RadixArk/Qwen3.8-27B-NVFP4 (~21 GB, mixed NVFP4 W4A4 with an FP8 KV scheme).
VRAM Headroom Beats Drafter Choice


Same three arms, same workload, only the weight format changes — throughput jumps roughly a third across the board:
| CCU | MTP: FP8 → NVFP4 | DSPARK: FP8 → NVFP4 | DFlash2: FP8 → NVFP4 |
| 10 | 452 → 592 (+31%) | 460 → 581 (+26%) | 519 → 658 (+27%) |
| 20 | 569 → 765 (+34%) | 545 → 709 (+30%) | 612 → 793 (+30%) |
| 40 | 563 → 817 (+45%) | 546 → 728 (+33%) | 604 → 806 (+34%) |
| 100 | 570 → 816 (+43%) | 549 → 729 (+33%) | 609 → 812 (+33%) |
| 150 | 566 → 820 (+45%) | 548 → 730 (+32%) | 607 → 806 (+33%) |
Dropping ~6 GB of weights feeds straight back into the two pools that gate throughput:
| Arm | FP8: cap / KV pool | NVFP4: cap / KV pool |
| MTP | 104 / 216,291 tok | 64, uncapped / 253,621 tok |
| DSPARK | 43 / 177,952 tok | 50 / 216,289 tok |
| DFlash2 | 42 / 175,645 tok | 49 / 212,500 tok |
Request caps grow 16–17%; the KV pool grows 21–22%. MTP gains most because it’s the only arm that reaches its configured cap outright — external drafters carry their own state on top of the target, so their caps stay lower even with the extra room.
The Latency Payoff

At concurrency 20 the run needs ~198K KV tokens in flight. FP8’s pool can’t hold it, so requests queue and TTFT is 4.0 s. NVFP4’s larger pool holds it — TTFT is 0.78 s, 5.2x better, purely from capacity. Past each format’s own saturation point, extra concurrency buys latency and nothing else: by concurrency 150, TTFT is 239 s on FP8 vs 174 s on NVFP4. The practical operating knee sits around 18 concurrent requests on FP8 and 22–25 on NVFP4 at this context length.
The Ranking Flip

On FP8, DFlash2 leads every level by ~+7%. On NVFP4 that lead evaporates — +11% at concurrency 10, then below the free MTP baseline from 40 upward (0.98–0.99x). DSPARK, already net loss on FP8 (~0.97x), drops further on NVFP4 (~0.89x).
The mechanism: faster weights saturate the compute units sooner, and once saturated, a drafter’s extra forward passes compete for FLOPs that are no longer free. Speculative decoding fades with achieved batch size and hardware speed — not with the concurrency number in your config.
The Acceptance-Length Trap

DSPARK accepts 3.08 tokens per verification step against MTP’s 2.67 — a better prediction rate on paper — yet delivers 3–4% less throughput on FP8 and 11% less on NVFP4. Acceptance length is only the revenue side of speculative decoding; the drafter’s own compute is the cost, and here the cost wins outright. Acceptance barely moves between weight formats at all (MTP 2.67→2.65, DSPARK 3.08→3.12, DFlash2 3.72→3.79): the drafters themselves stay BF16 either way, so the entire NVFP4 throughput gain is coming from raw speed, not from any change in how well the drafters predict.
Configuration and Limits
The NVFP4 + MTP arm ran SGLang with numGPUs: 1, gpuMemoryUtilization: 0.85, attentionBackend: flashinfer, kvCacheDtype: fp8_e4m3, and speculative settings speculativeAlgorithm: EAGLE, speculativeNumSteps: 3, speculativeEagleTopk: 1, speculativeNumDraftTokens: 4. Two settings were deliberately left unset: maxTotalTokens — the FP8 arms pinned it to 175,000 tokens for cross-arm parity, while NVFP4 was left to size naturally — and quantization, since the checkpoint’s hf_quant_config.json declares MIXED_PRECISION and SGLang auto-detects it. Benchmarking used vllm bench serve with –ignore-eos, –seed 42, and the FP8 checkpoint tokenizer pinned for both formats, so FP8 and NVFP4 are compared over identical token streams.
Limits worth knowing only low_entropy (code completion) were measured, the drafter-friendly end of the spectrum. Single GPU, no tensor parallelism. Neither external-drafter arm could reach the requested 64-concurrent cap — a drafter’s own KV allocation grows with the request cap, which is a second-GPU or larger-pool problem, not a tuning problem.
What This Means, and How to Set It Up
Put together, the sweep answers the opening question more decisively than expected: format, not drafter choice, is the lever that matters on this card. Across every drafter tested, moving from FP8 to NVFP4 was worth 4–6x more throughput than picking the best of the three speculative-decoding strategies.
For anyone sizing a similar deployment on an NVIDIA RTX PRO 6000:
- Quantize first, pick a drafter second. NVFP4 raised the request cap 16–17% and the KV pool 21–22%, which is what turned into the ~30–45% throughput gain and 5.2x TTFT improvement at concurrency 20.
- Default to the built-in MTP head. No extra VRAM, reaches its full cap, and fastest on NVFP4 from concurrency 40 up. DFlash2 only wins at low concurrency; DSPARK never won anywhere in this sweep.
- Size admission control to achieve concurrency, not your config number. The knee here sat around 18 requests on FP8 and 22–25 on NVFP4 — measure your own rather than importing this one.
- Let the KV pool size itself on NVFP4, and budget drafter VRAM separately if a drafter is in play — that’s a capacity problem to plan for, not a setting to tune away.
A single NVIDIA RTX PRO 6000 held this entire comparison — three drafting strategies, two weight formats, a concurrency sweep from 10 to 150 — inside 96 GB, without ever touching a second GPU.
Author:
An Hoang Minh is an AI Infrastructure Engineer at FPT Smart Cloud, focusing on LLM infrastructure and production of AI systems. His work involves designing and deploying distributed inference systems, optimizing model serving and KV-cache performance, and managing large-scale GPU infrastructure. He also provides technical consulting on AI architecture, infrastructure sizing, deployment strategies, and performance optimization for enterprise AI workloads.
***
Enjoy FPT Smart Cloud’s technical blogs from these authors? Get to know the products they are building!
- FPT GPU Cloud
- FPT Token Factory
- FPT AI Agents
- FPT.AI
All details are at: https://factory.fpt.ai/
