Skip to content
,
Run open models from scratch · Part 6

How to update local AI models without breaking your setup

13 min read
How to update local AI models without breaking your setup

Updating a local AI model on your own computer is not like updating a phone app. The short name you type can quietly point at a new file overnight, so the answers change while your prompts stay exactly the same. The safe routine has three steps. Keep a known-good copy, test one prompt you already trust on the new file, and then delete whichever copy lost so your disk and your results stay predictable.

Say it is Monday morning and a coworker tells you the newest version of the model you run on your laptop is much better. You run ollama pull, the command that downloads the latest version of a model in Ollama, a free app for running models on your own computer. Friday’s file was 4.7 gigabytes (GB), and Monday’s file is also 4.7 GB, but the list shows a different digest, which is a fingerprint of the exact file. You keep 12 trusted test prompts on a sticky note on the laptop lid, and their answers have started to ramble. The morning delay note that used to come out at 90 words now invents a dock 4 reroute that the warehouse floor never ran.

This post shows how to update on purpose: pin the old file, test the new one, and delete the loser. Size alone is a weak fingerprint, because two different files can have the same size. The digest is the fingerprint you can trust.

Why latest is not a pin

A moving tag such as latest, or a short size tag like :8b, can change quality and size under a name you already use. To protect yourself, pin a known-good copy with a dated alias or by keeping the original GGUF filename (GGUF is the file format most local model apps load). Run one canned prompt before you point daily jobs at a new file. Leftover files fill the disk after a few days of “just update it,” so clean up as you go. After an app update, also check the cloud toggles and which model is selected.

A tag is a pointer

Four update risks: a moving tag, an in-place pull, leftover blobs on disk, and an app default that can include a cloud toggle
Four update risks: a moving tag, an in-place pull, leftover blobs on disk, and an app default that can include a cloud toggle

Ollama model names look like model:tag. If you leave out the tag, you get latest. That word is a pointer, and the library maintainers can aim it at a new file whenever they publish one. Your chat window still shows the same short name, but the file behind it is not the same file.

Open a tags page before you treat a name as a fixed fact. When it was checked in August 2026, the Llama 3.2 tags list showed 63 tags. Sizes in that one family ran from a 581 megabyte (MB) q2_K 1B file to a 6.4 GB 3B fp16 file. On that page llama3.2:latest matched llama3.2:3b, which matched llama3.2:3b-instruct-q4_K_M, about 2.0 GB, with the digest prefix a80c4f17acd5. There is one family and many files, and latest is an alias, not a freeze.

In the story above, you are not on Llama 3.2. You are on an 8B-class tag that lists at 4.7 GB, the same size class covered in the earlier post on hosted versus downloaded models. On Friday you wrote the ID from ollama show on the sticky note under prompt 12. On Monday the ID no longer matches, even though the list still says 4.7 GB. Size is a weak fingerprint and the digest is the real one.

The web check is GET /api/tags on http://localhost:11434, which asks the local Ollama server (the background program that answers requests) for its model list. Each local model comes back with a name, a size in bytes, and a digest. If the digest moved and you did not mean to change jobs, you already have the Monday from the story.

What an update can change

A new file under an old name can shift more than a changelog post admits. The quantization can change, which means the way the model is squeezed to fit in memory, so the memory use and heat covered in the earlier hardware post move even when the marketing name stays “8B.” The chat template in the Modelfile can change too, so the model talks more, or hedges, or adds “next steps” that nobody asked for. A different instruct tune can even treat your canned prompt as a request for a plan. That is how a 90-word delay note becomes 340 words with a dock 4 story.

Ollama’s own project page says ollama pull can update a local model and that only the changed layers are downloaded. That is handy on a slow connection, and it is easy to miss because the command looks the same as the first download. If your daily name is qwen3:8b and you pull that string, the daily name now points at the new layers. Your 12 prompts did not get a vote.

LM Studio (a desktop app; LM stands for language model) and llama.cpp work more bluntly. They load a GGUF file, or a file in Apple’s own MLX format, from a path on your disk, and a new download gets a new filename if you let it. The failure mode is loading “the Qwen 8B” from a dropdown that now prefers the fresh file. Hugging Face model pages list hashes for those files. Keep the old path until the new one has earned the job.

Update typeRiskHow to pin
ollama pull on :latest or a short size tag like :8bSame name, new digest. Quality, size, and ramble length can all move.Copy to a dated local name first. Compare digest. Switch the daily name only after one canned prompt wins.
New GGUF in LM Studio or a Hugging Face saveTwo files on disk. The Discover list likes the new one. You load it by habit.Keep the old filename. Load by path. Delete the extra quant only after the test.
Runner or engine bump (llama.cpp inside the app)Same file, different runtime. Answers can still shift.Note the engine version in Settings. Rerun the one prompt on the pinned file.
Desktop app update (Ollama, LM Studio, friends)Selected model, context size, and cloud toggles can move. The 4.7 GB file may sit unused.Read the model line after install. If it says cloud or remote, you changed location. The file on disk did not move.

Pin the file that already does the job. Test the candidate on one canned prompt. Delete the loser so the disk matches the sticky note.

Pin a known-good name

Four steps: copy the known-good file, pull the candidate beside it, run one canned prompt, delete the loser
Four steps: copy the known-good file, pull the candidate beside it, run one canned prompt, delete the loser

The cheap pin in Ollama is a copy. The command ollama cp and the matching POST /api/copy create a new name that points at the layers you already like. The docs show the shape ollama cp llama3.2 my-model, and a copy made through the API (the way a program sends commands to Ollama) from gemma4 to gemma4-backup. The names in those docs change, but the idea does not. On Friday you should have run a copy into something like ops-2026-08-15 before anyone typed pull.

A longer library tag is a softer pin. llama3.2:3b-instruct-q4_K_M moves less than llama3.2:latest, but maintainers can still retarget a long tag. The local copy is the one that cannot move unless you pull over it. Write the size and the digest on the same sticky note as the 12 prompts. That is two extra lines, and it costs far less than a 340-word dock story.

LM Studio and a raw GGUF

LM Studio keeps ordinary files under a models directory, and the My Models screen can move it (this reflects the docs in August 2026). The llama.cpp program takes a -m path to a file. Pinning here just means do not delete Friday’s .gguf. Load the old path until the new one beats the 90-word delay note without inventing dock 4. If the Hugging Face card lists a SHA256 hash, paste it next to the filename.

Test one prompt, then delete

Do not test the whole sticky note on a live warehouse thread. Pick the one prompt that already has a known good answer. In this story it is the morning delay note, which runs 90 words, names the lane, and does not invent a reroute. Run it on the pinned copy and then on the candidate, with the same temperature (the setting that controls how random the wording is) and no extra instructions. Read the length and check the facts. If the candidate is tighter and still true, point the daily name at it. If it rambles, keep the pin. Only then delete the file you will not use, because deleting first is how people “update” into a hole with 4 GB free and no copy of Friday’s model file.

Here is a checklist you can paste. Flags and model names change, so read the help text the week you run it. This is a shape to follow and not a promise that every runner uses the same subcommand.

# 1. Snapshot what you already run (size + digest).
ollama ls
curl -s http://localhost:11434/api/tags
# 2. Pin Friday before anyone pulls.
# Replace the source name with whatever `ollama ls` showed.
ollama cp qwen3:8b ops-2026-08-15
# 3. Pull the candidate. The short tag may move. The copy should not.
ollama pull qwen3:8b
ollama show qwen3:8b
ollama show ops-2026-08-15
# 4. One canned prompt only. Same text you already trust.
# Example: "Write the 7am delay note: Lane B late 40 min, no reroute."
# If it invents dock 4, keep ops-2026-08-15 as the daily name.
# 5. After the new file wins, free the loser. Check disk.
# ollama rm qwen3:8b
# Then confirm the models folder actually shrank.

The numbers in the table below are toy numbers for teaching, and they are not from a live library listing.

NameDigest (short)SizeRole
ops-2026-08-157a3c… (Friday)4.7 GBKnown-good. 90-word delay note.
qwen3:8b after pull9f21… (Monday)4.7 GBCandidate. 340 words, dock 4 fiction.
Daily chat nameMust match the pinPoint this at Friday until the test passes.

Disk fills with leftover files

Ollama stores layers instead of one tidy icon per model. According to the public FAQ, the default locations are ~/.ollama/models on macOS, often /usr/share/ollama/.ollama/models on Linux, and .ollama\models inside the user profile on Windows. The OLLAMA_MODELS setting moves that whole folder. Pulling a moving tag can add new files while the old ones stay until no name points at them. On Friday you had 11 GB free, and after two more days of “just pull” you had 4 GB.

Shared layers are the reason a delete sometimes frees less space than you expected. List first, and remove only the name that lost the test, and then look at free space. Do not empty the whole models directory to “be sure.” In LM Studio the leftovers are extra .gguf files, so keep the quantization you tested. One path per job keeps things clear.

App updates change defaults

The weights (the model file itself) are not the only thing that moves. The desktop app can ship a new default model, a new context size (how much text the model reads at once), or a cloud control. The post on hosted versus downloaded models already showed the one-app, two-mode trap, because Ollama’s public site advertises both computer and cloud. An installer that “just updates the app” can leave you on a remote model while the 4.7 GB file sits on disk looking unused. Your legal team heard local, but the data went to a host.

After every app update, read the model line before you paste in a vendor PDF. If it says cloud or remote, you changed location, so turn that off or stop pasting. Then rerun the one canned prompt on the pinned local name, because an engine update can still change the tone on the same GGUF. Confirm the selected model and the models directory as well. If LM Studio reset My Models to the home folder, you may have loaded a different file.

A worked example: 12 prompts and one bad Monday

On Friday you pencil the ollama show ID under prompt 12. On Monday you pull the short tag, and prompt 1 comes back at 340 words and names dock 4, a real dock the floor did not use. You almost send it. The other 11 prompts add “next steps” to jobs that were one paragraph.

CheckFridayMonday after the pull
Listed size4.7 GB4.7 GB (useless as a fingerprint)
Digest from show / /api/tagsThe ID on the sticky noteA different ID
Delay note length90 words, lane named, no reroute340 words, invented dock 4
Disk free11 GBOn the way to 4 GB after two more pulls
FixShould have cp’d firstIf the old blobs still exist, copy is still possible. If not, you wait on a re-download of the Friday digest, if the registry still has it.

The boring ending goes like this. You copy whatever still matches Friday’s digest to ops-2026-08-15, point the daily name there, and treat the Monday file as a candidate. You run one prompt. If it still invents docks, you run ollama rm on the candidate when disk is tight. The 12 prompts stay on the pin until a candidate beats 90 words without fiction.

Do not update because the badge is green

  • Pulling :latest (or a short :8b) on the same name your canned jobs already use.
  • Trusting size alone, when two 4.7 GB files are not the same file.
  • Deleting Friday’s copy before the one-prompt test.
  • Leaving every old file “in case,” until the drive fills and you delete the pin in a panic.
  • Updating the app and missing a cloud toggle, so the location changed even though the weights did not.
  • Testing on a live vendor email instead of the delay note you already know.
  • Assuming a longer tag can never move, when only a local cp or a kept GGUF path is a true freeze.

Write the current tag before you pull

List what you actually run and copy it to a dated name. Write the size and digest on the sticky note. Pull only if you have a reason, and pull into a name you can throw away. Run one canned prompt and keep the winner, then use rm on the loser and watch the free space. The next post, on sharing a local setup with family or a team, covers other people using the same machine, which is how a moving tag becomes everyone else’s problem. For hardware limits, see the post on memory, graphics, heat, and battery. For file size and memory, see the post on model size and quantization.

Update on purpose, keep a rollback name

  • A tag is a pointer, and latest can change the 4.7 GB file under you.
  • Copy a known-good name, or keep the GGUF, before you pull, because a pull can quietly replace the model you tested with a different one.
  • Test one canned prompt, then delete the loser so the disk matches the pin.
  • After an app update, read the model line for cloud use and for a different file.

Series notes

This is Part 6 of Run open models from scratch (series code OS13). Previous: hardware reality. Next: sharing a local setup. Related: hosted vs download, Open-source AI explained, and Learn.

Sources

Written by

Jose S

Founder & Lead Analyst · Analytics Made Simple

Hands-on data strategist, analytics engineering lead, and educator. Writing practical, no-fluff guides to help everyday teams, analysts, and engineers master SQL, AI systems, and modern data architectures.

Keep going

Same lessons in your feed

Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.

Google Search Prefer our practical guides in Google Search & Top Stories: