Skip to main content
NJannasch.Dev
llama.cppQwenMTPRTX 5060 Ti

Qwen 3.8 27B MTP Server (RTX 5060 Ti)

This is the measured default profile for the Qwen 3.8 27B Dense IQ3_XXS GGUF on a 16 GB RTX 5060 Ti. It uses the model’s built-in Multi-Token Prediction (MTP) head rather than n-gram speculation.

llama-server \
  -m /path/to/Qwen3.8-27B-UD-IQ3_XXS.gguf \
  -ngl 99 \
  --ctx-size 98304 \
  -fa on -ctk q4_0 -ctv q4_0 \
  --parallel 1 \
  --spec-type draft-mtp --spec-draft-n-max 3 \
  --jinja \
  --host 0.0.0.0 --port 11433

Use the OpenAI-compatible endpoint with thinking disabled for ordinary coding and interactive HTML tasks:

{
  "temperature": 0.7,
  "top_p": 0.8,
  "top_k": 20,
  "chat_template_kwargs": { "enable_thinking": false }
}

Key flags:

  • -ngl 99: offloads all supported layers to CUDA.
  • --ctx-size 98304: allows a real 96K-token prompt on this 16 GB card, though it leaves little VRAM margin. Use 32K or 64K for comfortable day-to-day sessions.
  • -fa on: enables Flash Attention.
  • -ctk q4_0 -ctv q4_0: quantizes the KV cache so the model and useful context fit together.
  • --parallel 1: reserves VRAM for one responsive session.
  • --spec-type draft-mtp --spec-draft-n-max 3: the fastest tested configuration: 63.89 tok/s at 4K, 45.48 tok/s at 32K, and 39.06 tok/s at 64K.
  • --jinja: enables the model’s Qwen chat template and request-level template options.

Do not add --spec-default to this command. In current upstream llama.cpp it enables the ngram-mod speculative preset, not MTP. This server already selects MTP explicitly; the measured n-gram-mod gain was only 1.7% on the coding workload, so combining the two is not a justified default.

For a little more VRAM headroom and higher draft acceptance at long contexts, change only --spec-draft-n-max 3 to 2; it was slower but still strong. Do not use ngram-mod as the default: it gained only 1.7% on the measured 512-token coding workload.

Thinking profiles should be selected at request time. Medium and Xhigh used the entire 2,048-token budget on a visual-code test without returning code, so start with thinking off and reserve those profiles for bounded hard tasks with a larger output budget.