Skip to content
,
Meta Llama from scratch · Part 6

Fine-tunes and community variants without the zoo

15 min read
Featured image: Llama variants, not a zoo, with the official Llama (LLaMA) lockup. Editorial illustration for Analytics Made Simple.

Strangers online publish thousands of modified versions of Llama, Meta’s free-to-download AI model, usually by giving it extra training on their own examples (called a fine-tune), and downloading a pile of them makes your setup harder to trust, not better. Pick one official version from Meta and stick with it for about ninety days. Every modified copy is still built from Meta’s model, so it still carries Meta’s license, and a file labeled “uncensored” does not come with fewer rules.

Imagine your team wants a live demo at lunch, so the night before you download fourteen modified Llama versions from community sites, each with a name promising something extra, like “uncensored.” By midnight they fill 83 gigabytes (GB) of your work laptop. The next morning your legal team asks one simple question: which license came with those files? You open the folder and find no license and no file explaining where they came from, just 83 GB of models nobody can vouch for.

Eighty-three gigabytes of llama stickers

The Finder window looked busy, which is exactly how a zoo like this hides. Fourteen rows, one column of llama names, and a disk number that felt like real work. You can scroll through it, and you can sort by file size, but none of that answers the compliance check waiting for you.

Here is the dump, using local names only, since these are files on a local disk rather than Hub ids worth hunting down. This is not a list of an “official uncensored” repo, because Meta does not ship one. Community drops reuse the word Llama the same way third-party mobile apps reuse a mascot they never licensed.

  • Llama-4-uncensored-Q2_K.gguf, 8.40 GB on disk, with no card URL attached anywhere.
  • Llama-4-uncensored-Q3_K_M.gguf, 11.20 GB, same missing card URL.
  • Llama-4-uncensored-Q4_K_M.gguf, 14.10 GB, again no card URL.
  • Llama-4-uncensored-Q5_K_M.gguf, 16.10 GB, no card URL either.
  • llama4-hermes-style-Q3.gguf 4.90 GB, Discord filename, no Hub page.
  • llama4-hermes-style-Q4.gguf 6.20 GB, the one used for the leadership demo.
  • llama4-hermes-style-Q5.gguf 7.40 GB, “better quality” of the same unknown cook.
  • scout-chat-mystery-Q4.gguf 4.20 GB, Scout in the name, parent blank.
  • llama-4-ablit-Q3.gguf 2.10 GB and llama-4-ablit-Q4.gguf 2.80 GB, a community nickname, still no meta-llama link.
  • maverick-lite-mystery-Q2.gguf 1.90 GB. Maverick on the tin is not Maverick on the card.
  • llama4-roleplay-Q4.gguf 1.40 GB, llama4-system-stripped-Q3.gguf 1.30 GB, llama-sticker-merged-Q3.gguf 1.00 GB. The last one is the sticker. Add them up and you get 83.00 GB.

What it felt like was a buffet of Llama 4 variants, so leadership could test an uncensored build if official Instruct felt too restrained. What actually existed on disk was fourteen blobs whose parents nobody could name. The lunch chat sounded fluent, but fluency is cheap. The compliance question is the one that survives the demo.

Hermes-class, in this post, means a community chat fine-tune style. Nous Research and other groups have shipped Hermes-named chats on several bases over the years. This is not pointing you at a Llama 4 Hermes Hub id, and none of those fourteen files are official. If a cook is real and useful later, it will have a card, a parent, and a license line you can read without digging through a Discord scrollback. The leadership demo had a sticker instead.

Rule of thumb: If you cannot paste a Hugging Face URL that lists the parent and the license, you do not have a Llama stack. You have a folder.

Base, Instruct, and “someone cooked this”

Four cards for Llama variants: official Base, official Instruct, a community fine-tune, and a random GGUF with no matching card
Four cards for Llama variants: official Base, official Instruct, a community fine-tune, and a random GGUF with no matching card

Four different objects share the word Llama on disk, and only two of them are actually Meta’s official Scout line on meta-llama. The other two are how you end up with 83 GB of mystery files.

Base

Base is the pretrained next-token model. Official Scout Base lives at meta-llama/Llama-4-Scout-17B-16E, at least according to vendor pages checked in September 2026. You feed it a prefix, and it continues the prefix; it is not a finished chat product on its own. A runner can wrap Base in a generic chat template, and the window will still look like a chatbot, but that wrapper is your own UI, not Meta’s actual Instruct training. If the job from the earlier post on first useful tasks is to write, summarize, or do light code on a named path, Base is the wrong first pull. Keep Base for people who are pretraining or doing research completions. You do not need an unaligned Base model for an operational demo.

Instruct

Instruct is Meta’s chat-tuned official tag. For this series, pick one and stick with it: Llama-4-Scout-17B-16E-Instruct, at the Hub path meta-llama/Llama-4-Scout-17B-16E-Instruct. That is the file, or the host id copied from it, you actually want for the jobs covered in the earlier post on first useful tasks. Maverick has its own separate Instruct tag. Do not collect both “just in case” on a laptop, since that exact hoarding habit already proved expensive in the earlier post on sizing. Hosted Groq, Together, Fireworks, and the others each list their own Instruct-class Llama 4 ids, so copy the host’s string directly. Groq, spelled with a Q, is still not Grok, the xAI chatbot.

Gated Hub repos are a feature, not an obstacle. You accept the Llama 4 Community License on meta-llama before the weights even start downloading. Skipping that gate by grabbing a Discord GGUF instead does not skip the license itself. It only skips the moment where you would have actually read it.

Someone cooked this

A community fine-tune is a further training run on top of Base or Instruct, or sometimes a mix of both. The cook changes behavior: more chat, more roleplay, more refusal stripping, more tool syntax, whatever their training data happened to emphasize. Hermes-class chats fall into this same category, and so do merges, “abliterated” nicknames, and roleplay dumps. Some of those cooks publish a real Hub repo with a proper card. Some only publish a .gguf file inside a zip. The cook does not become Meta by doing this work, and the cook does not mint a brand new license out of thin air just because the parent model was Llama 4.

A GGUF is a file format, usually a quantized, compressed shrink meant for llama.cpp, LM Studio, or Ollama. Community quantizers do genuinely useful work when the card says clearly: this Q4_K_M is a quant of meta-llama/Llama-4-Scout-17B-16E-Instruct, license llama4. That file is still Instruct underneath, just smaller on disk. A random GGUF titled Llama is only a filename, and a filename is not parentage. Those fourteen files sitting in the zoo represent the worst category of all: random GGUF files, llama labels slapped on, and zero verifiable provenance behind any of them.

If you need a local shrink of official Instruct, wait for a card that names that Instruct model as base_model, with base_model_relation: quantized set explicitly. Confirm the card the week you actually pull it, since quantizers change tags over time, and hardware numbers still disagree across different blogs, the same warning covered in the earlier post on Llama sizes.

The card that has to match the file

Five rows that must match a GGUF to its model card: repo, license, parent, quant, and why you kept this file
Five rows that must match a GGUF to its model card: repo, license, parent, quant, and why you kept this file

A model card on the Hugging Face Hub is really the repo’s README (README stands for read-me), the file formatted as README.md, with a structured data block at the top, in a format called YAML (YAML stands for a plain-text data layout), that the Hub itself parses. Hugging Face documents these fields in its own Model Cards page. What you are actually looking for is a match between the bytes on disk and that YAML block, not a loose vibe match between the filename and the word Llama.

Open the card, and read five specific rows. If any single row is blank, you do not load that file for a leadership demo.

YAML fieldWhat you readFail if
licenseHub identifier such as llama4Empty, unknown, or mit on a Llama 4 derivative
base_modelParent Hub id, often meta-llama/Llama-4-Scout-17B-16E-InstructMissing, or a name you cannot open
base_model_relationquantized, finetune, adapter, or mergeYou cannot tell shrink vs new cook
library_namegguf, transformers, or the real loaderFile is .gguf and the card never says so
license_name / license_linkFull text when license is otherother with no link and no LICENSE file
pipeline_tagUsually text-generation for this workBlank on a chat demo you plan to ship

The Hub also uses base_model to draw out the full model tree: Base, then Instruct, then quants and fine-tunes hanging off Instruct beneath that. Official Scout Instruct shows that whole tree clearly. The download folder in this story showed nothing verifiable at all, because none of the files tracked back to a clean repo anywhere. A shared drive folder of GGUFs has no YAML attached to it. Discord has no base_model field either. “The filename says Scout” is simply not a field in any of this.

Match the quant too, not just the license. If the card lists Q4_K_M and you actually downloaded Q2_K because it was the first link you clicked, you did not download the card you just read. Multi-part GGUFs need every single part present. A 14.10 GB file named Q4 is not proof of Q4 on its own. The card’s file table is the real proof. If the uploader omitted sizes entirely, you do not guess with du -h and a shrug.

Write the five rows into the demo notes before the lunch starts, the same way the earlier post wanted a named path before you touched a vendor PDF. Repo, license, parent, quant, and why this one. “Why this one” needs to be a full sentence, not a sticker. A good sentence looks like: “Official Scout Instruct, Q4_K_M, card lists meta-llama/Llama-4-Scout-17B-16E-Instruct, license llama4, so legal or compliance can verify the LICENSE and the Notice line.” A bad sentence looks like: “Uncensored Llama, tested great late last night.”

License follows the weights

The compliance question turns out to be the right one to ask. The Llama 4 Community License, effective April 2025, follows the Llama 4 weights (the billions of numbers the model learned in training) into every derivative built on top of them. A fine-tune, a merge, a GGUF quant, a Hermes-class chat layered onto Scout: if the parent was Llama 4, you still have that same agreement, the same Acceptable Use Policy, and the same notice rules attached. You do not get MIT (a very permissive open source license) just because a Discord title said uncensored. You do not get Apache 2.0 just because the cook added a chat template on top. Apache 2.0 on Meta’s own side, according to vendor pages checked in September 2026, belongs to Muse Glimmer under meta-models, an entirely different family. The earlier post on licensing already walked through those gates in full; this part is really just the derivative reminder.

Section 1(b) of the Llama 4 Community License is the redistribution rule you need on a demo day. If you distribute Llama materials or a derivative, including another AI model built from them, you have to provide a copy of the agreement, and you have to prominently display “Built with Llama” on a related website, interface, blog post, about page, or piece of product documentation. If you use the Llama materials or their outputs to create, train, fine-tune, or otherwise improve an AI model that you then distribute, that new model’s name has to include “Llama” at the beginning. You also have to keep this exact Notice line in a plain-text notice file alongside the copies: Llama 4 is licensed under the Llama 4 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved. Meta’s own Llama FAQ restates that same attribution pattern. Renaming a file is really just putting on a sticker, and it does not satisfy any actual license notice requirement.

The 700-million-monthly-user extra license clause still cares about the licensee and its affiliates on the Llama 4 release date, not about whether your GGUF happens to be Q3 or Q5. The Acceptable Use Policy still applies in full. Multimodal Llama 4 rights are still not granted at all if you are an individual living in the EU, or a company whose main office sits in the EU. An uncensored filename does not open a side path around any of this. The earlier post on licensing covers it in full.

Community naming is genuinely messy in the wild. You will see files that drop Llama from the front entirely, files that keep Llama and add a cook name on top, and files that only carry a llama pictogram somewhere in the folder. For a model you distribute, the license asks for the word Llama at the beginning of the name. For a model you only run internally, you still need to know exactly which agreement you accepted in the first place. Compliance asks which license file shipped because a public demo is a distribution-shaped event even when you think of it as just a screenshot. Put LICENSE, USE_POLICY, and the Notice file next to the GGUF you actually load. If you cannot do that, do not load it.

One honest edge worth naming: a documented community fine-tune really can be the right tool later on. A support-tone chat, a tighter SQL dialect (SQL is the language for asking databases questions), a tool-call wrapper your team specifically measured, all of these are legitimate. That decision wants a card, a parent id, a written reason, and a legal sign-off on the thread, though. It does not want fourteen overnight downloads and a sticker slapped on top. The cook is allowed under the license in many cases. The zoo itself is simply not a process.

One default stack for the next 90 days

Pick one Instruct tag and stick with it. For this series that tag is Llama-4-Scout-17B-16E-Instruct. Hosted work reuses the host’s own Scout Instruct id from the earlier post on hosted versus self-host, so copy it rather than inventing your own version. Local work is either the official Hub repo or one GGUF whose card lists that Instruct model as parent, license llama4, relation quantized. Confirm the card the week you actually pull it. Do not add Maverick, Base, a Hermes-class cook, or a second quant “for backup” until the 90 days end, or until real hardware needs from the sizing post make a second file genuinely necessary.

The proper workflow here is an audit, a delete, and one verified pull from official channels. Put all fourteen rows in a table you can sort through. “Keep” is a high bar: the card URL actually opens, the license is llama4 (or the parent’s own Llama 4 agreement), the parent is meta-llama/Llama-4-Scout-17B-16E-Instruct or Scout Base if Base is truly needed, the quant matches the filename, and you can say why this file in one plain sentence. In this story, all fourteen files fail that bar. Delete the whole folder.

What that script does: it prints a clean keep-or-delete audit sheet that legal or compliance can inspect without ever opening a terminal (the text window for typing commands) themselves. Run it against the actual folder on disk, not against a fuzzy memory of what you downloaded one night. Then pull Instruct fresh from meta-llama, accepting the gate as you go, or from a quant card you just personally verified. Put LICENSE, the use policy link, and the Notice line in the same directory as the GGUF you keep. Take the llama sticker off the slide entirely. The model name on the slide should be the actual Hub id.

# zoo_audit.py: Audit local folder for unverified model weights
files = [
    {"path": "Llama-4-uncensored-Q2_K.gguf", "size_gb": 8.40, "card_url": "", "license": "", "parent": "", "keep": "delete"},
    {"path": "Llama-4-uncensored-Q3_K_M.gguf", "size_gb": 11.20, "card_url": "", "license": "", "parent": "", "keep": "delete"},
    {"path": "Llama-4-uncensored-Q4_K_M.gguf", "size_gb": 14.10, "card_url": "", "license": "", "parent": "", "keep": "delete"},
    {"path": "Llama-4-uncensored-Q5_K_M.gguf", "size_gb": 16.10, "card_url": "", "license": "", "parent": "", "keep": "delete"},
    {"path": "llama4-hermes-style-Q3.gguf", "size_gb": 4.90, "card_url": "", "license": "", "parent": "", "keep": "delete"},
    {"path": "llama4-hermes-style-Q4.gguf", "size_gb": 6.20, "card_url": "", "license": "", "parent": "", "keep": "delete"},
    {"path": "llama4-hermes-style-Q5.gguf", "size_gb": 7.40, "card_url": "", "license": "", "parent": "", "keep": "delete"},
    {"path": "scout-chat-mystery-Q4.gguf", "size_gb": 4.20, "card_url": "", "license": "", "parent": "", "keep": "delete"},
    {"path": "llama-4-ablit-Q3.gguf", "size_gb": 2.10, "card_url": "", "license": "", "parent": "", "keep": "delete"},
    {"path": "llama-4-ablit-Q4.gguf", "size_gb": 2.80, "card_url": "", "license": "", "parent": "", "keep": "delete"},
    {"path": "maverick-lite-mystery-Q2.gguf", "size_gb": 1.90, "card_url": "", "license": "", "parent": "", "keep": "delete"},
    {"path": "llama4-roleplay-Q4.gguf", "size_gb": 1.40, "card_url": "", "license": "", "parent": "", "keep": "delete"},
    {"path": "llama4-system-stripped-Q3.gguf", "size_gb": 1.30, "card_url": "", "license": "", "parent": "", "keep": "delete"},
    {"path": "llama-sticker-merged-Q3.gguf", "size_gb": 1.00, "card_url": "", "license": "", "parent": "", "keep": "delete"},
]
print("path\tsize_gb\tcard_url\tlicense\tparent\tkeep")
for f in files:
    print("{path}\t{size_gb}\t{card_url}\t{license}\t{parent}\t{keep}".format(**f))
kept = [f for f in files if f["keep"] == "keep"]
print("keep {0} of {1}; disk {2:.2f} GB".format(len(kept), len(files), sum(f["size_gb"] for f in files)))
# keep 0 of 14; disk 83.00 GB
# pull instead: meta-llama/Llama-4-Scout-17B-16E-Instruct
# Notice: Llama 4 is licensed under the Llama 4 Community License, Copyright (c) Meta Platforms, Inc. All Rights Reserved.

After the delete, the 90-day stack is boring on purpose. Hosted Scout Instruct handles the write-and-summarize loop from the first-tasks post whenever the prompt may leave the building. Local Scout Instruct, or one matching GGUF, handles it whenever the prompt has to stay put. Use the exact same tag across every team environment so test prompts and operational demos cannot quietly drift apart from each other. No second cook gets added until someone writes down why, pastes the card URL in #llama-demo, and compliance actually signs off on it.

This week, run the audit on your real folder, delete every row with a blank card URL, pull Llama-4-Scout-17B-16E-Instruct once, drop LICENSE plus the Notice file beside it, and take the sticker off the slide for good. Then read the next post in this series, on when Claude or ChatGPT is simply easier, before anyone orders GPUs, so the next vendor lunch can be a $20 seat instead of 83 GB of unlabeled files.

Series notes

This is Part 6 of Learn Llama (LL6). The earlier post covered first useful tasks; the next one covers when Claude or ChatGPT is simply easier.

Sources

Research and further reading used for this article:

Written by

Jose S

Founder & Lead Analyst · Analytics Made Simple

Hands-on data strategist, analytics engineering lead, and educator. Writing practical, no-fluff guides to help everyday teams, analysts, and engineers master SQL, AI systems, and modern data architectures.

Keep going

Same lessons in your feed

Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.

Google Search Prefer our practical guides in Google Search & Top Stories: