“Open” does not mean one thing when people talk about AI models. It can mean you can download the finished model, it can mean you can see everything used to build it, or it can just mean the chat is free. Mixing those up is how a confident slide becomes a legal problem later, so this post separates the three claims and shows how to check which one you are actually getting.
Picture an architecture slide that says “we’ll use open source Llama” in 28-point type. Your legal team asks three questions. Can we download it? Can we use it at a hospital with 2,118 patient rows? Can we modify it and share it? Engineering answers “yes” to the first and then stares at the other two. Downloading the model is like picking up a crate of metal plates, which people call open weights. Other crates stay locked, such as the training data, the full recipe, and a license that is not the permissive MIT kind. Calling both crates “open source” is how a slide becomes an incident later.
This is the first post in the Open-source AI explained series. It gives the chooser in Which AI product should I use? a place to send readers who said “I care about privacy” or “I want to run things myself.” Open is at least three separate claims, so re-check licenses the week you download anything. They change, and marketing pages describe them in a friendly way that is not always accurate.
Open is three claims

| Claim | Plain English | Typical example | Does not mean |
|---|---|---|---|
| Open weights | You can download the finished model numbers, called parameters, and run them | Many Llama, DeepSeek, Qwen, and Gemma releases | You got the training set, or a license that lets you do anything with it |
| Open source AI | The Open Source Initiative (OSI) bar: you can use, study, modify, and share it, with the preferred form for making changes | Rarer than the marketing suggests, such as some fully open research stacks | A chat box with a cute llama icon |
| Free to use | No invoice today | A vendor’s free chat tier | The weights are yours, or your prompts stay on your own machine |
The Open Source Initiative (OSI) is blunt on this point: open weights are not the same as open source AI. Weights are only the final numbers. Open source AI, as the OSI defines it, also cares about the training code and the data, or a full description of the data when it cannot legally be shared. It also requires a license that lets you use the system for any purpose. Cards on Hugging Face (a site where people share models) that say “open” often mean only “downloadable under this license,” so read the card and do not trust the adjective in the title.
Weights, in one picture
A weight is a number that the training run learned. A file in the GGUF format (GGUF means a popular way to pack a model to run on a laptop) is a pile of those numbers plus just enough structure to run them. When you “download Llama,” you are usually grabbing that pile and not the internet it was trained on. You can run it, you can often fine-tune it, and you can sometimes ship it, depending on the license. What you cannot do is reproduce the original training from that file alone, and that gap is why auditors and the OSI keep saying “not quite.”
# Not a real loader. A picture of the split.
# weights.bin -> numbers you can download (open weights, if licensed that way)
# train.py -> recipe (often missing)
# data/ -> the internet-shaped pile (almost always missing)
def can_i_call_this_open_source(card):
if not card.has_weights:
return "closed or hosted-only"
if not card.license_allows_any_purpose:
return "open weights with strings; not OSI open source"
if not (card.has_training_code and card.has_data_or_disclosure):
return "open weights, still not the OSI AI bar"
return "closer to open source AI; still read the card"That code sketches a decision you can make in your head when you look at a model card. If legal asks “is it open source,” the honest answer is often “open weights under license X.” That sentence is longer, but it is true.
Licenses are the plot
The Massachusetts Institute of Technology (MIT) license and Apache 2.0 are the licenses people mean when they say “do what you want, and keep the notice.” Some model families are closer to that, and DeepSeek has shipped certain releases under MIT, though you should re-check the card in your hand. Llama’s community license has historically added acceptable-use rules and a threshold for very large user counts. The OSI and others have said that this is not open source. Google’s Gemma license includes clauses that protect Google’s products. None of this is a reason to panic. It is a reason to stop saying “open source” as a general good feeling.
If you work at a hospital, a bank, or a school, your legal team should read the license and not the meme. When a license limits monthly active users or the field of use, you are not in MIT territory. Print the license and put it next to the slide.
Not a free lunch

- Hardware: you need memory (RAM), a graphics chip (GPU) or Apple Silicon, and a way to handle the heat. A 7B model and a 70B model are very different animals, where B stands for billions of parameters.
- License: watch for limits on field of use, caps on users, attribution rules, and “no competing with our product.”
- Operations: you patch it, you shrink it to fit your machine (a step called quantizing), and you fix the desktop app when it breaks on a Tuesday.
- Quality: a $20 closed chat plan still wins a lot of writing tasks, and running locally is a choice about control.
Rule of thumb: If you cannot name the license file, you cannot call the model “open source” on a slide.
Worked example: the 28-point slide
| Legal question | What the slide said | The accurate sentence |
|---|---|---|
| Can we download it? | Open source Llama | Yes, the weights are available under the Llama community license (check the current PDF) |
| Can we use it on 2,118 patient rows locally? | Implied yes | Maybe, if the license and hospital policy allow it, because running locally is a deployment choice and not a permission |
| Can we modify and share it? | Open source always means that | Only if that license says so, and the OSI would still want the recipe |
| Is a free ChatGPT tab the same thing? | Also “open” in hallway talk | No, because free to use is a price, and closed weights stay closed |
The chooser’s earlier post on privacy and running AI yourself is about location, meaning where your text goes. This post is about the label on the model, and you need both. Hosted Llama on a cheap API is still someone else’s GPU, while local Llama means your fan and your license homework.
When a closed chat is simply easier
If you only need a polished thank-you email, you do not need a GGUF file at all. If your repository cannot leave the building, though, a local model might be the right answer for a coding helper. Most readers should stay on a work chat plan until they have a real reason to leave, such as working offline, cost at scale, a license they have actually read, or a research need. Shame is not a reason, and neither is a conference slide.
How to read a model card without getting lost
A useful card answers four questions on its first screen. What is the file: an instruct model, a base model, or a fine-tune? What is the license name, with a link to the full text? What was it trained for? Is there an acceptable-use policy on top? If the card is a poem about “democratizing AI” and buries the license, that is a smell. If it says “open source” and the linked license is a custom community agreement, believe the PDF.
Quantization, meaning shrinking a model into fewer bits so it fits on a smaller machine, is the next post’s job. The only preview you need is that smaller files run on smaller machines and get worse at hard tasks in ugly, quiet ways. Do not promise legal that a 4-bit file is “the same as the closed model, but local” just because you found one. Try it on a toy prompt first. Hardware reality and laptop choices come later in this series, and today you only need to stop the 28-point slide.
Families you will hear next
You will hear about Llama from Meta, DeepSeek, Qwen from Alibaba, Kimi from Moonshot, and GLM (General Language Model) from Zhipu, plus a zoo of fine-tunes. Each will get its own map on this site, and this page gives them shared vocabulary so those maps do not repeat the OSI lecture. Hosted chat pages for open models, such as a Groq playground, a Together chat, or a vendor’s “open” tab, are still hosted. Local runners such as Ollama, llama.cpp, and the LM Studio desktop app (LM stands for language model) are local. Keep pointing at the crate that is actually open.
PBS and the OSI both stress the same three-way split that this post uses: closed, open weight, and open source. This site is not a law firm. If your counsel says the current Llama PDF is fine for your use, that wins over any blog. If your slide still says “open source” only because it sounded generous, change the slide.
Common mistakes
- Saying “open source” when you mean “I downloaded a file.”
- Saying “private” because the model is Llama, while you actually sent your text to a host.
- Skipping the license PDF because the card had a green badge.
- Promising legal reproducibility that you cannot deliver without the data.
- Assuming free ChatGPT is the same category as a GGUF file on your disk.
A hallway sentence you can reuse
When someone says “we should go open source,” answer with a short script and not a feeling. Ask, “Do you mean downloadable weights, an OSI-qualified project, or a free chat tier?” Then ask, “Where will the text live: our laptop, a host, or a closed vendor?” Then ask, “Which license PDF is on the slide?” That is three questions and two minutes, with no big speech. If they cannot answer, the slide is not ready. If they can, you may still pick Claude for writing and a local 7B model for a private notes experiment, because mixed stacks are allowed. Mixed nouns are not.
Cost stories get slippery when nobody says which cost they mean. Local is “free” until you count the machine, the electricity, and the afternoon you lost to a CUDA error (CUDA means Nvidia’s software for running work on its graphics chips). Hosted open models can cost cents per thousand words and still be a privacy miss. Closed work plans are a line item that finance understands. Pick the cost you actually mean. The next posts in this series map hosted pages, desktop runners, and hardware, and none of them will make the license file optional.
If you arrived from the chooser because you care about privacy, go back to the privacy post once you can name the three claims. Look at location first and license second. If you arrived because a conference said Llama is the future, stay here until the 28-point slide is rewritten. The future can wait for a PDF.
How to practice this week
Open one model card on Hugging Face or the vendor’s site. Write four lines: whether the weights are available, the license name, whether the training data is available, and whether you can use it at work. Bring that paper to the next slide review. The next posts in this series cover hosted versus laptop use, then size and quantization. The chooser defaults remain in the post on which AI product to try first.
Quick recap
- Open weights, open source AI, and free to use are different claims.
- Read the license, because the OSI’s bar is higher than a download button.
- Running locally is homework, and a closed work chat is allowed to be the right tool.
Series notes
This is post OS1, the first in Open-source AI explained. It links back to the privacy post (P4) and the chooser defaults (P6) in the products chooser series.
Sources
- OSI: Open Weights
- OSI: Open Source AI Definition
- OSI on Llama’s license (not open source)
- PBS NewsHour explainer on closed, open source, and open weight
- Analytics Made Simple: privacy lanes for running AI yourself
Keep going
Same lessons in your feed
Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.
