Skip to content
,
Open-source AI explained · Part 1

What open means: weights, license, not a free lunch

10 min read
Featured image: What open means, open crate of weight plates versus a locked crate

“Open” does not mean one thing when people talk about AI models. It can mean you can download the finished model, it can mean you can see everything used to build it, or it can just mean the chat is free. Mixing those up is how a confident slide becomes a legal problem later, so this post separates the three claims and shows how to check which one you are actually getting.

Picture an architecture slide that says “we’ll use open source Llama” in 28-point type. Your legal team asks three questions. Can we download it? Can we use it at a hospital with 2,118 patient rows? Can we modify it and share it? Engineering answers “yes” to the first and then stares at the other two. Downloading the model is like picking up a crate of metal plates, which people call open weights. Other crates stay locked, such as the training data, the full recipe, and a license that is not the permissive MIT kind. Calling both crates “open source” is how a slide becomes an incident later.

This is the first post in the Open-source AI explained series. It gives the chooser in Which AI product should I use? a place to send readers who said “I care about privacy” or “I want to run things myself.” Open is at least three separate claims, so re-check licenses the week you download anything. They change, and marketing pages describe them in a friendly way that is not always accurate.

Open is three claims

Open is three different claims
Open is three claims: open weights, open source AI under the OSI bar, and free to use
ClaimPlain EnglishTypical exampleDoes not mean
Open weightsYou can download the finished model numbers, called parameters, and run themMany Llama, DeepSeek, Qwen, and Gemma releasesYou got the training set, or a license that lets you do anything with it
Open source AIThe Open Source Initiative (OSI) bar: you can use, study, modify, and share it, with the preferred form for making changesRarer than the marketing suggests, such as some fully open research stacksA chat box with a cute llama icon
Free to useNo invoice todayA vendor’s free chat tierThe weights are yours, or your prompts stay on your own machine

The Open Source Initiative (OSI) is blunt on this point: open weights are not the same as open source AI. Weights are only the final numbers. Open source AI, as the OSI defines it, also cares about the training code and the data, or a full description of the data when it cannot legally be shared. It also requires a license that lets you use the system for any purpose. Cards on Hugging Face (a site where people share models) that say “open” often mean only “downloadable under this license,” so read the card and do not trust the adjective in the title.

Weights, in one picture

A weight is a number that the training run learned. A file in the GGUF format (GGUF means a popular way to pack a model to run on a laptop) is a pile of those numbers plus just enough structure to run them. When you “download Llama,” you are usually grabbing that pile and not the internet it was trained on. You can run it, you can often fine-tune it, and you can sometimes ship it, depending on the license. What you cannot do is reproduce the original training from that file alone, and that gap is why auditors and the OSI keep saying “not quite.”

# Not a real loader. A picture of the split.
# weights.bin  -> numbers you can download (open weights, if licensed that way)
# train.py     -> recipe (often missing)
# data/        -> the internet-shaped pile (almost always missing)

def can_i_call_this_open_source(card):
    if not card.has_weights:
        return "closed or hosted-only"
    if not card.license_allows_any_purpose:
        return "open weights with strings; not OSI open source"
    if not (card.has_training_code and card.has_data_or_disclosure):
        return "open weights, still not the OSI AI bar"
    return "closer to open source AI; still read the card"

That code sketches a decision you can make in your head when you look at a model card. If legal asks “is it open source,” the honest answer is often “open weights under license X.” That sentence is longer, but it is true.

Licenses are the plot

The Massachusetts Institute of Technology (MIT) license and Apache 2.0 are the licenses people mean when they say “do what you want, and keep the notice.” Some model families are closer to that, and DeepSeek has shipped certain releases under MIT, though you should re-check the card in your hand. Llama’s community license has historically added acceptable-use rules and a threshold for very large user counts. The OSI and others have said that this is not open source. Google’s Gemma license includes clauses that protect Google’s products. None of this is a reason to panic. It is a reason to stop saying “open source” as a general good feeling.

If you work at a hospital, a bank, or a school, your legal team should read the license and not the meme. When a license limits monthly active users or the field of use, you are not in MIT territory. Print the license and put it next to the slide.

Not a free lunch

Not a free lunch
Not a free lunch: hardware, license, ops, and quality
  • Hardware: you need memory (RAM), a graphics chip (GPU) or Apple Silicon, and a way to handle the heat. A 7B model and a 70B model are very different animals, where B stands for billions of parameters.
  • License: watch for limits on field of use, caps on users, attribution rules, and “no competing with our product.”
  • Operations: you patch it, you shrink it to fit your machine (a step called quantizing), and you fix the desktop app when it breaks on a Tuesday.
  • Quality: a $20 closed chat plan still wins a lot of writing tasks, and running locally is a choice about control.

Rule of thumb: If you cannot name the license file, you cannot call the model “open source” on a slide.

Worked example: the 28-point slide

Legal questionWhat the slide saidThe accurate sentence
Can we download it?Open source LlamaYes, the weights are available under the Llama community license (check the current PDF)
Can we use it on 2,118 patient rows locally?Implied yesMaybe, if the license and hospital policy allow it, because running locally is a deployment choice and not a permission
Can we modify and share it?Open source always means thatOnly if that license says so, and the OSI would still want the recipe
Is a free ChatGPT tab the same thing?Also “open” in hallway talkNo, because free to use is a price, and closed weights stay closed

The chooser’s earlier post on privacy and running AI yourself is about location, meaning where your text goes. This post is about the label on the model, and you need both. Hosted Llama on a cheap API is still someone else’s GPU, while local Llama means your fan and your license homework.

When a closed chat is simply easier

If you only need a polished thank-you email, you do not need a GGUF file at all. If your repository cannot leave the building, though, a local model might be the right answer for a coding helper. Most readers should stay on a work chat plan until they have a real reason to leave, such as working offline, cost at scale, a license they have actually read, or a research need. Shame is not a reason, and neither is a conference slide.

How to read a model card without getting lost

A useful card answers four questions on its first screen. What is the file: an instruct model, a base model, or a fine-tune? What is the license name, with a link to the full text? What was it trained for? Is there an acceptable-use policy on top? If the card is a poem about “democratizing AI” and buries the license, that is a smell. If it says “open source” and the linked license is a custom community agreement, believe the PDF.

Quantization, meaning shrinking a model into fewer bits so it fits on a smaller machine, is the next post’s job. The only preview you need is that smaller files run on smaller machines and get worse at hard tasks in ugly, quiet ways. Do not promise legal that a 4-bit file is “the same as the closed model, but local” just because you found one. Try it on a toy prompt first. Hardware reality and laptop choices come later in this series, and today you only need to stop the 28-point slide.

Families you will hear next

You will hear about Llama from Meta, DeepSeek, Qwen from Alibaba, Kimi from Moonshot, and GLM (General Language Model) from Zhipu, plus a zoo of fine-tunes. Each will get its own map on this site, and this page gives them shared vocabulary so those maps do not repeat the OSI lecture. Hosted chat pages for open models, such as a Groq playground, a Together chat, or a vendor’s “open” tab, are still hosted. Local runners such as Ollama, llama.cpp, and the LM Studio desktop app (LM stands for language model) are local. Keep pointing at the crate that is actually open.

PBS and the OSI both stress the same three-way split that this post uses: closed, open weight, and open source. This site is not a law firm. If your counsel says the current Llama PDF is fine for your use, that wins over any blog. If your slide still says “open source” only because it sounded generous, change the slide.

Common mistakes

  • Saying “open source” when you mean “I downloaded a file.”
  • Saying “private” because the model is Llama, while you actually sent your text to a host.
  • Skipping the license PDF because the card had a green badge.
  • Promising legal reproducibility that you cannot deliver without the data.
  • Assuming free ChatGPT is the same category as a GGUF file on your disk.

A hallway sentence you can reuse

When someone says “we should go open source,” answer with a short script and not a feeling. Ask, “Do you mean downloadable weights, an OSI-qualified project, or a free chat tier?” Then ask, “Where will the text live: our laptop, a host, or a closed vendor?” Then ask, “Which license PDF is on the slide?” That is three questions and two minutes, with no big speech. If they cannot answer, the slide is not ready. If they can, you may still pick Claude for writing and a local 7B model for a private notes experiment, because mixed stacks are allowed. Mixed nouns are not.

Cost stories get slippery when nobody says which cost they mean. Local is “free” until you count the machine, the electricity, and the afternoon you lost to a CUDA error (CUDA means Nvidia’s software for running work on its graphics chips). Hosted open models can cost cents per thousand words and still be a privacy miss. Closed work plans are a line item that finance understands. Pick the cost you actually mean. The next posts in this series map hosted pages, desktop runners, and hardware, and none of them will make the license file optional.

If you arrived from the chooser because you care about privacy, go back to the privacy post once you can name the three claims. Look at location first and license second. If you arrived because a conference said Llama is the future, stay here until the 28-point slide is rewritten. The future can wait for a PDF.

How to practice this week

Open one model card on Hugging Face or the vendor’s site. Write four lines: whether the weights are available, the license name, whether the training data is available, and whether you can use it at work. Bring that paper to the next slide review. The next posts in this series cover hosted versus laptop use, then size and quantization. The chooser defaults remain in the post on which AI product to try first.

Quick recap

  • Open weights, open source AI, and free to use are different claims.
  • Read the license, because the OSI’s bar is higher than a download button.
  • Running locally is homework, and a closed work chat is allowed to be the right tool.

Series notes

This is post OS1, the first in Open-source AI explained. It links back to the privacy post (P4) and the chooser defaults (P6) in the products chooser series.

Sources

Written by

Jose S

Founder & Lead Analyst · Analytics Made Simple

Hands-on data strategist, analytics engineering lead, and educator. Writing practical, no-fluff guides to help everyday teams, analysts, and engineers master SQL, AI systems, and modern data architectures.

Keep going

Same lessons in your feed

Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.

Google Search Prefer our practical guides in Google Search & Top Stories: