Snippets
Short, practical code references. Copy-paste ready.
Qwen 3.8 Server Profiles: MTP Depth 8 & N-Gram Chain (RTX 5090)
Two llama.cpp server profiles for Qwen 3.8 27B on a RTX 5090: uniform MTP depth 8 (+5%) and chained n-gram/MTP drafting (2.4x on repetitive code, slower on novel text).
Qwen 3.8 27B MTP Server (RTX 5060 Ti)
Measured llama.cpp configuration for Qwen 3.8 27B IQ3_XXS on an RTX 5060 Ti 16 GB.
OpenCode with Local llama.cpp (Qwen 3.6)
Connect OpenCode to a local llama.cpp server running Qwen 3.6 MTP. Zero API costs, 90K context, local-first AI coding.
Gemma 4 MTP Server (ik_llama.cpp)
Run Gemma 4 26B-A4B with MTP speculative decoding using ik_llama.cpp. Separate drafter model, 133 t/s on an NVIDIA RTX 5060 Ti 16 GB.
Qwen 3.6 MTP Server (llama.cpp)
Run Qwen 3.6 35B-A3B with MTP speculative decoding on llama.cpp. 144 t/s on an NVIDIA RTX 5060 Ti 16 GB.
Gemma 4 256K Context Server (llama.cpp)
llama-server config for Gemma 4 26B-A4B MoE with full 256K context on an NVIDIA RTX 5060 Ti 16 GB. The key: do NOT use --swa-full.