llama.cppQwenMTPRTX 5090
Qwen 3.8 Server Profiles: MTP Depth 8 & N-Gram Chain (RTX 5090)
Both profiles share the same base: Qwen 3.8 27B UD-Q6_K_XL, q4_0 K/V cache, Flash Attention, one slot, 160K capacity. Swap only the speculative-decoding flags.
Default profile — uniform speed (155 tok/s everywhere):
/opt/llama-sm120/bin/llama-server \
-hf unsloth/Qwen3.8-27B-GGUF:UD-Q6_K_XL \
-ngl 999 \
--ctx-size 163840 \
--cache-type-k q4_0 --cache-type-v q4_0 \
--flash-attn on \
--parallel 1 \
--spec-type draft-mtp --spec-draft-n-max 8 \
--batch-size 2048 --ubatch-size 1024 \
--main-gpu 0 --jinja \
--host 0.0.0.0 --port 8080
Repetition-heavy profile — code refactors, boilerplate, templates (290–380 tok/s warm):
# identical except the spec line:
--spec-type ngram-mod,draft-mtp --spec-draft-n-max 8 \
Measured trade-offs (same binary, 3 warm runs each):
| Configuration | Repetitive/warm | Novel content |
|---|---|---|
| MTP n=4 (old default) | 148 tok/s | 148 tok/s |
| MTP n=8 | 155 tok/s | 155 tok/s |
| ngram-mod,draft-mtp n=8 | 290–380 tok/s | 100–120 tok/s |
Key findings:
--spec-draft-n-max 8is the decode knee for this dense 27B: n=10 declines, n=12 collapses to 139 tok/s (acceptance fall-off beats draft savings)- The n-gram pool persists across requests and accumulates during a session — first request runs at plain-MTP speed, later requests hit the 2.4x numbers once patterns exist
- On novel text the chain loses ~25% vs plain MTP (draft acceptance drops to ~45–50%, wasted verification batches), so keep both profiles and switch per workload