Skip to main content
NJannasch.Dev

Snippets

Short, practical code references. Copy-paste ready.

llama.cppQwenMTPRTX 5090

Qwen 3.8 Server Profiles: MTP Depth 8 & N-Gram Chain (RTX 5090)

Two llama.cpp server profiles for Qwen 3.8 27B on a RTX 5090: uniform MTP depth 8 (+5%) and chained n-gram/MTP drafting (2.4x on repetitive code, slower on novel text).

llama.cppQwenMTPRTX 5060 Ti

Qwen 3.8 27B MTP Server (RTX 5060 Ti)

Measured llama.cpp configuration for Qwen 3.8 27B IQ3_XXS on an RTX 5060 Ti 16 GB.

OpenCodellama.cppQwenLocal-First

OpenCode with Local llama.cpp (Qwen 3.6)

Connect OpenCode to a local llama.cpp server running Qwen 3.6 MTP. Zero API costs, 90K context, local-first AI coding.

llama.cppGemmaMTPRTX 5060 Ti

Gemma 4 MTP Server (ik_llama.cpp)

Run Gemma 4 26B-A4B with MTP speculative decoding using ik_llama.cpp. Separate drafter model, 133 t/s on an NVIDIA RTX 5060 Ti 16 GB.

llama.cppQwenMTPRTX 5060 Ti

Qwen 3.6 MTP Server (llama.cpp)

Run Qwen 3.6 35B-A3B with MTP speculative decoding on llama.cpp. 144 t/s on an NVIDIA RTX 5060 Ti 16 GB.

llama.cppGemmaRTX 5060 Ti

Gemma 4 256K Context Server (llama.cpp)

llama-server config for Gemma 4 26B-A4B MoE with full 256K context on an NVIDIA RTX 5060 Ti 16 GB. The key: do NOT use --swa-full.