Skip to main content
NJannasch.Dev
llama.cppQwenMTPRTX 5090

Qwen 3.8 Server Profiles: MTP Depth 8 & N-Gram Chain (RTX 5090)

Both profiles share the same base: Qwen 3.8 27B UD-Q6_K_XL, q4_0 K/V cache, Flash Attention, one slot, 160K capacity. Swap only the speculative-decoding flags.

Default profile — uniform speed (155 tok/s everywhere):

/opt/llama-sm120/bin/llama-server \
  -hf unsloth/Qwen3.8-27B-GGUF:UD-Q6_K_XL \
  -ngl 999 \
  --ctx-size 163840 \
  --cache-type-k q4_0 --cache-type-v q4_0 \
  --flash-attn on \
  --parallel 1 \
  --spec-type draft-mtp --spec-draft-n-max 8 \
  --batch-size 2048 --ubatch-size 1024 \
  --main-gpu 0 --jinja \
  --host 0.0.0.0 --port 8080

Repetition-heavy profile — code refactors, boilerplate, templates (290–380 tok/s warm):

# identical except the spec line:
  --spec-type ngram-mod,draft-mtp --spec-draft-n-max 8 \

Measured trade-offs (same binary, 3 warm runs each):

ConfigurationRepetitive/warmNovel content
MTP n=4 (old default)148 tok/s148 tok/s
MTP n=8155 tok/s155 tok/s
ngram-mod,draft-mtp n=8290–380 tok/s100–120 tok/s

Key findings:

  • --spec-draft-n-max 8 is the decode knee for this dense 27B: n=10 declines, n=12 collapses to 139 tok/s (acceptance fall-off beats draft savings)
  • The n-gram pool persists across requests and accumulates during a session — first request runs at plain-MTP speed, later requests hit the 2.4x numbers once patterns exist
  • On novel text the chain loses ~25% vs plain MTP (draft acceptance drops to ~45–50%, wasted verification batches), so keep both profiles and switch per workload