Qwen 3.8 27B MTP Server (RTX 5060 Ti)
This is the measured default profile for the Qwen 3.8 27B Dense IQ3_XXS GGUF on a 16 GB RTX 5060 Ti. It uses the model’s built-in Multi-Token Prediction (MTP) head rather than n-gram speculation.
llama-server \
-m /path/to/Qwen3.8-27B-UD-IQ3_XXS.gguf \
-ngl 99 \
--ctx-size 98304 \
-fa on -ctk q4_0 -ctv q4_0 \
--parallel 1 \
--spec-type draft-mtp --spec-draft-n-max 3 \
--jinja \
--host 0.0.0.0 --port 11433
Use the OpenAI-compatible endpoint with thinking disabled for ordinary coding and interactive HTML tasks:
{
"temperature": 0.7,
"top_p": 0.8,
"top_k": 20,
"chat_template_kwargs": { "enable_thinking": false }
}
Key flags:
-ngl 99: offloads all supported layers to CUDA.--ctx-size 98304: allows a real 96K-token prompt on this 16 GB card, though it leaves little VRAM margin. Use 32K or 64K for comfortable day-to-day sessions.-fa on: enables Flash Attention.-ctk q4_0 -ctv q4_0: quantizes the KV cache so the model and useful context fit together.--parallel 1: reserves VRAM for one responsive session.--spec-type draft-mtp --spec-draft-n-max 3: the fastest tested configuration: 63.89 tok/s at 4K, 45.48 tok/s at 32K, and 39.06 tok/s at 64K.--jinja: enables the model’s Qwen chat template and request-level template options.
Do not add --spec-default to this command. In current upstream llama.cpp it enables the ngram-mod speculative preset, not MTP. This server already selects MTP explicitly; the measured n-gram-mod gain was only 1.7% on the coding workload, so combining the two is not a justified default.
For a little more VRAM headroom and higher draft acceptance at long contexts, change only --spec-draft-n-max 3 to 2; it was slower but still strong. Do not use ngram-mod as the default: it gained only 1.7% on the measured 512-token coding workload.
Thinking profiles should be selected at request time. Medium and Xhigh used the entire 2,048-token budget on a visual-code test without returning code, so start with thinking off and reserve those profiles for bounded hard tasks with a larger output budget.