Skip to content
,
Open-source AI explained · Part 2

Hosted open models vs download to your machine

9 min read
Featured image: Hosted vs on your machine, server rack on one side and a laptop on the other

An open AI model is really just a large file, and you can keep that file in two different places. Hosted means your prompt travels to a company that owns powerful computers, such as Groq, Together, Fireworks, or a chat website that lists Llama as one of its models. Local is the opposite. The file sits on your own disk, and your prompt can stay there too, as long as you do not switch on a cloud option inside the same app. Having “Llama” in the name tells you nothing about where the model actually runs.

Say you are asked to summarize an 86-page contract packet from a vendor, and you promise your legal team that the documents will stay on your laptop. You spend 41 minutes watching a progress bar while a 4.7 gigabyte file called qwen3:8b downloads. Then the first local reply takes 28 seconds per sentence, so you flip a cloud switch in the same app to speed things up. The whole packet goes out on your very next prompt. Your legal team heard “local,” but the packet went to a host.

This post separates those two places, names the common hosts and desktop programs, and shows the trap where one app offers both. Groq (the company that sells fast model hosting, spelled with a q) is not Grok, which is the chat product from xAI. Names matter here. Double-check the product names the week you touch any setting that says Cloud.

Location is the noun

Two places a model can live
Hosted open models send the prompt off-device. Local downloads can stay on disk if you skip the cloud toggle. Groq is not Grok. A Llama n…

An open-weights model is a file, or a pile of file pieces, plus a license that says what you may do with it. You can send that file to a company that already owns the computers to run it, or you can load it on the machine in front of you. The license does not change when your laptop fan spins up, but the privacy story changes completely. When a model is hosted, your prompt leaves your device. When it is local, the prompt can stay put, provided you do not turn on a cloud mode and the disk is really yours.

Closed products such as ChatGPT, Claude, Gemini, and Grok are hosted too, and they are simply not the subject of this series. If you only needed a summary of a public blog post, the chooser for writing and everyday questions was enough. It is easy to reach for an open model because the packet feels “too sensitive for ChatGPT,” and then end up sending it to a different host with a friendlier logo.

PlaceWhat you runWhere the prompt goesYou pay with
Hosted open modelGroq, Together, Fireworks, OpenRouter, or a vendor chat that lists Llama, Qwen, or DeepSeekTheir data centerTokens, credits, or a seat
On your machineOllama, LM Studio, or llama.cpp, using a downloaded model file such as a GGUFYour own memory, if you stay offlineHardware, heat, and your time
Closed chatThe ChatGPT, Claude, Gemini, or Grok appsTheir data centerFree caps or plans around $20 (see the earlier post on free versus paid plans)
Same app, cloud toggleOllama cloud models, or any “run on our servers” optionNot your laptop, even if you already downloaded 4.7 gigabytesWhatever plan that toggle belongs to

Hosted: fast graphics chips, still a vendor

A handful of companies rent out the computers that run open models, which the trade calls inference. Groq runs models on its own custom chips and is known for speed on a short list of open models. Together and Fireworks rent graphics processing units (GPU for short, the chips that do most AI work) and tend to list more model families, custom tunes, and sometimes a dedicated machine just for you. OpenRouter is a switchboard, meaning one account key that reaches many different backends. Prices vary widely per million tokens (a token is a small piece of a word), so do not copy a number from a blog post into a budget. Open the host’s pricing page the week you start.

Hosted is the right choice when you want a very large model that will not fit on a laptop, or when you want programmatic access without buying a GPU. It is the wrong choice when the sentence you told your legal team was “nothing leaves.” Read each host’s retention and training policy, because Fireworks has advertised options that keep no data at all, while others differ. Until a contract says otherwise, assume logs exist. The privacy paste test still applies here, and it asks whether you would email this PDF to that vendor.

The playground chats on those sites feel like ChatGPT with a Llama badge. They are not local, and they are not the Grok chatbot either. If someone says “we moved to Groq,” ask whether they mean the hosting company or whether they misspelled Grok. That mix-up is how a slide ends up with the wrong vendor and the wrong privacy story in a single breath.

On your machine: the 4.7 gigabytes is only the beginning

Ollama is the one-command route that many people actually finish. LM Studio gives you a friendly window instead of a terminal, and llama.cpp is the engine underneath many of those programs. You download a compressed model file (a common packaging format is called GGUF, short for a model file format), load it, and start chatting. The next post on model size, compression, and whether your laptop can run it covers memory in detail. For now, know that a model with about 8 billion settings is a laptop-sized animal and one with 70 billion is a different animal entirely. That 4.7 gigabyte download was the easy half, and the 28-second sentences were your hardware talking.

Local wins when the file must never leave and you are willing to live with lower quality and a hot laptop. It loses when you expected polished prose from a small compressed model, or when the same app offers a cloud shortcut and you take it. Ollama’s homepage currently sells both a computer version and a cloud version, which is useful, and it is also exactly how the toggle trap happens. Read the model line before you paste anything, because if it says cloud, your 41-minute download did not apply.

Pick a location first

Pick a location before you pick a model
Pick a location before you pick a model: allowed to leave, need a big model, need offline, or just writing
  1. Ask whether this text may leave the building. If the answer is no, use a local model or use nothing.
  2. If it may leave, ask whether you need a large open model today. If so, choose a hosted one and read the host’s policy.
  3. If you need to work offline, or want something cheap when idle, download a size your machine can hold, and feel the fan on a toy prompt first.
  4. If the job is a thank-you note, use the chooser and a closed chat plan, because a paragraph does not justify building a local setup.

Once you think you are local, you can run a quick check in the terminal. Treat it as a shape to copy, not a promise that every program uses the same command.

# After a local pull, stay explicit.
# Ollama-style (flags and names move; read `ollama --help` this week):
ollama list
# If a model line says cloud / remote, do not paste NDAs into it. # Toy prompt only:
# "Summarize this public sentence: The warehouse delay was ours."
# Time it. If you then switch to a cloud id because it was slow,
# you changed location. Tell legal the new location or do not paste.

The point of that check is that a model list shows what is really configured. In the packet story, listing the models afterward would have revealed a cloud entry sitting right beside the 4.7 gigabyte file. The list is the map. The download progress bar is not.

Worked example: 41 minutes, 86 pages, one toggle

Claim you might makeWhat was trueWhat legal needed
“We run Llama locally”You downloaded a Qwen file with about 8 billion settings, then used a cloud entryThe host name and the retention policy
“The 4.7 gigabyte file is the privacy control”The file sat unused after the toggleWhether the prompt left the laptop
“Open source, so we’re fine”The model is weights plus a licenseThe license document, plus the location
28 seconds per sentenceThe cost of running an 8-billion-setting model on a laptopA decision: wait, use a smaller model, or sign a contract with a host

The fix is boring. Turn cloud off, and summarize a two-page public PDF first. If 28 seconds per sentence is unacceptable, either keep the packet away from models entirely or sign up with a host that has a real agreement. Do not split the difference with a toggle that you never mention on the slide.

When hosted is the honest cheaper path

Builders sometimes rent from Together or Fireworks because a weekend GPU quote is worse than a token bill under 100 million tokens a month. That math belongs to a later post about running costs. For an analyst who wanted a private summary, hosted open models are still vendors, so they can be cheaper than Claude’s paid interface but they are not free just because the model is Qwen. Local is cheaper when idle, if the machine already exists and you accept the quality. A closed plan around $20 is cheaper than buying a GPU if the job is writing.

Later posts cover size and compression, and an easy hosted chat window. Do not skip this location split to chase a model card. It comes first. The families of models (Llama, DeepSeek, Qwen) make more sense once you can say “hosted” or “local” without hesitating.

Common mistakes

  • Saying Groq when you mean Grok, or the reverse.
  • Calling a Groq or Together playground “local” because the model is Llama.
  • Downloading 4.7 gigabytes and then using the cloud entry in the same window.
  • Promising legal an “open source, local” setup without the license document or a named location.
  • Building a local setup just to write the thank-you note covered in the earlier guide.
  • Forgetting that Ollama and similar programs may offer both modes.

How to practice this week

Write one sentence that says hosted or local, plus the app name. Run only a toy prompt at first. If you turn on cloud, write the host’s name on the same sticky note. After that, work out whether your laptop can run the model you want. The privacy chooser is still the privacy paste test, and the setup guide is accounts and free versus paid plans.

Quick recap

  • Hosted open weights still leave the building, while local files can stay home if you do not toggle cloud.
  • Groq is not Grok, and having Llama in the name does not mean the model is on your disk.
  • Pick the location first and the model second, because the download bar alone is not a privacy control.

Series notes

This is Part 2 of Open-source AI explained (series code OS2). Previous: what open means. Next: model size, quantization, and will my laptop run it. Related: privacy paste test and Run open models from scratch.

Sources

Written by

Jose S

Founder & Lead Analyst · Analytics Made Simple

Hands-on data strategist, analytics engineering lead, and educator. Writing practical, no-fluff guides to help everyday teams, analysts, and engineers master SQL, AI systems, and modern data architectures.

Keep going

Same lessons in your feed

Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.

Google Search Prefer our practical guides in Google Search & Top Stories: