Skip to content

Repository files navigation

Local LLM Guides

Team reference for running Qwen3.5, Qwen3.6, and Gemma 4 models locally across NVIDIA (CUDA), AMD (ROCm), and Apple Silicon (MLX) hardware using llama.cpp, Ollama, vLLM, and LM Studio.


Quick Decision Tree

Backend — what do I want to do?

Backend Selector — Decision Schematic

Hardware — what do I have?

Hardware Selector — Decision Schematic


Model Selection Matrix

Gemma 4 (Apache 2.0, released Apr 2026)

Variant Arch Total Params Active Params Context Modalities 4-bit RAM Best For
E2B Dense+PLE ~9.6 GB BF16 all 128K Text, Image, Audio ~3 GB Phone / edge / Raspberry Pi
E4B Dense+PLE ~15 GB BF16 all 128K Text, Image, Audio ~5 GB Laptops, 8 GB RAM machines
26B-A4B MoE 26B 4B active 256K Text, Image ~16 GB Best speed/quality balance
31B Dense 31B all 256K Text, Image ~17 GB Highest quality at 31B scale

Qwen3.5 (Apache 2.0, released Feb 2026)

Variant Arch Total Params Active Params Context 4-bit RAM Best For
27B Dense 27B all 262K ~14 GB Strong dense, fits 16 GB VRAM
35B-A3B MoE 35B 3B active 262K ~18 GB Fast MoE, coding + reasoning
122B-A10B MoE 122B 10B active 262K ~62 GB Multi-GPU powerhouse
397B-A17B MoE 397B 17B active 262K ~200 GB Data center / multi-node

All Qwen3.5 models have native vision support and hybrid thinking.

Qwen3.6 (Apache 2.0, released Apr 2026)

Variant Arch Total Params Active Params Context 4-bit RAM Best For
27B Dense 27B all 262K (1M YaRN) ~17 GB Flagship coding — beats 397B MoE; fits 16 GB VRAM
35B-A3B MoE 35B 3B active 256K (1M YaRN) ~23 GB Agentic coding, best single-GPU MoE

Qwen3.6 supports 201 languages, hybrid thinking, and tool calling. The 27B dense surpasses Qwen3.5-397B-A17B (807 GB) across all major coding benchmarks at 15× lower memory.


Backend Comparison

Backend Best Use Case GPU Support Format API Difficulty
LM Studio Local exploration, GUI CUDA, ROCm, Metal GGUF OpenAI-compat Beginner
Ollama Fast local API, teams CUDA, ROCm, Metal GGUF OpenAI-compat Beginner
llama.cpp Full control, CPU+GPU, servers CUDA, ROCm, Metal, CPU GGUF OpenAI-compat Intermediate
vLLM Production serving, multi-GPU CUDA, ROCm (MI300+) BF16/FP8/AWQ OpenAI-compat Advanced

Hardware Compatibility Matrix

Model 8 GB VRAM 12 GB VRAM 16 GB VRAM 24 GB VRAM 48 GB VRAM 80 GB VRAM
Gemma 4 E2B Q8 BF16 BF16 BF16 BF16 BF16
Gemma 4 E4B Q4 Q8 BF16 BF16 BF16 BF16
Gemma 4 26B-A4B Q4 Q6 BF16 BF16
Gemma 4 31B Q4 Q6 BF16 BF16
Qwen3.5 27B Q4 Q6 BF16 BF16
Qwen3.6 27B Q4 Q6 BF16 BF16
Qwen3.5/3.6 35B-A3B Q4 Q6-Q8 BF16
Qwen3.5 122B-A10B Q4 (2×80)

For CPU-only or partial GPU offload, llama.cpp can layer-split across VRAM + system RAM. Expect slower generation.


Recommended Quantizations (GGUF)

Use Case Quantization Notes
Maximum quality Q8_0 Near-lossless; largest file
Best quality/size balance Q4_K_M or UD-Q4_K_XL Unsloth Dynamic quants preferred
Memory constrained Q3_K_M Slight quality drop
Ultra-low RAM Q2_K / UD-Q2_K_XL Noticeable degradation

Unsloth Dynamic (UD) quants are strongly preferred — they use higher precision for critical layers (embeddings, attention) while compressing others, achieving better quality at the same file size vs. standard GGUF quants.


Cost Comparison

  • cost-comparison.md — Hardware costs, electricity, 3-year TCO, cloud API pricing (DeepSeek, Kimi, GLM, Mistral), benchmark comparison, and break-even analysis

Model Guides

Backend Guides

Hardware Guides

Agent Integration


Common Gotchas

  • CUDA 13.2 + GGUF: Do not use the CUDA 13.2 runtime with any GGUF model — it produces garbage output. Use CUDA 12.x or 13.0. NVIDIA is tracking this.
  • Qwen3.x thinking mode: Must be explicitly toggled per use case. Wrong mode = wrong sampling params = bad output. See model guides.
  • Gemma 4 EOS token: The end-of-sentence token is <turn|>, not </s>. Some older inference wrappers may not stop correctly.
  • MoE VRAM: All MoE model weights load into VRAM even though only a fraction is active per token. Plan for full model size in VRAM.
  • KV cache overhead: Hardware tables show weight-only memory. Add 10–30% for KV cache at typical context lengths. Reduce --ctx-size if OOM.
  • ROCm consumer GPUs: vLLM officially only supports MI300X+ for AMD. Use llama.cpp for RX 6000/7000 series.

About

Team guides for running Qwen3.5/3.6 and Gemma 4 locally (llama.cpp, Ollama, vLLM, LM Studio · CUDA, ROCm, MLX)

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages