Skip to content

← All series & guides

Meta

llama.cpp from scratch

0 of 7 parts live

This path is for you if you want Georgi Gerganov's llama.cpp engine: a local run, a GGUF file, and a small API on your own machine. No part is live on this page yet. Start with Run open models, which walks a desktop runner first, and come back when part 1 is up.

  1. 1 What Is llama.cpp? A Program That Runs Models Locally, Not an AI Company Treat llama.cpp as a building block, not a software vendor. It powers Ollama and LM Studio behind the scenes. Compile once, run offline, no per-seat bill. Scheduled · November 9, 2026
  2. 2 Install Paths for Normal People (And When to Stop) Prefer package managers or official release zips for llama.cpp. Compile from source only when CUDA on Linux or bleeding-edge models force it. Stop when tokens outpace reading. Scheduled · November 10, 2026
  3. 3 Loading a GGUF Model and First Reply Load a verified GGUF, set model path, temperature, and GPU layers, then read llama.cpp timings. Interactive -cnv mode keeps the KV cache warm for follow-ups. Scheduled · November 11, 2026
  4. 4 Context Size, Threads, and Why is This Slow? Diagnose llama.cpp slowness with KV cache size, physical-core thread counts, and memory bandwidth limits. Quantize the KV cache before you buy new silicon. Scheduled · November 12, 2026
  5. 5 Server Mode: Sharing Your Local Model With Other Programs Run llama-server with host, port, parallel slots, and cont-batching. Point the official OpenAI client at localhost. Keep long system prompts warm with prefix cache. Scheduled · November 13, 2026
  6. 6 Frontends That Wrap llama.cpp Put a real UI on llama-server: Open WebUI for teams, Continue for editors, desktop apps for zero Docker. Use host.docker.internal, not localhost, from containers. Scheduled · November 14, 2026
  7. 7 Troubleshooting the Usual Fails in llama.cpp Triage llama.cpp fails: read stderr, check GGUF magic (curl -L), match CPU flags, shrink offload or quantize KV, fix chat templates. Add a self-healing wrapper for edge boxes. Scheduled · November 15, 2026