Mei pasted a 12-page draft offer letter into the desktop runner she had labeled local. Base salary $118,500, $8,000 sign-on, 0.35% equity, start date 8 September, candidate full name on page 1. The Word file still sat in Downloads as Offer_Samir_DRAFT_v7.docx, next to a weekend packing list. The model picker still showed last Tuesday’s cloud id because the 8B file on her laptop had taken 22 seconds per paragraph. She wanted the letter warmer and the numbers untouched. The reply landed in 4 seconds. That speed was the giveaway. The 12 pages had already left the building.
This is OS4, Part 4 of Open-source AI explained. After OS3 on size and RAM, this page asks whether the prompt is allowed to leave, even when the icon lives on the dock. OS2 split hosted from a file on disk. The chooser’s privacy door is still P4. Here we name what travels (text, files, metadata), when local is worth the fan, and when a ~$20 closed chat on a work account is the grown-up move. Next is OS5, quality without benchmark theater.
What you’ll learn
- What leaves on send: prompt, files, metadata, and the transcript after
- Why a download plus a dock icon is not a privacy control if cloud mode is on
- How Groq, Together, Fireworks, and Ollama cloud still put the text in someone else’s building
- When local is worth it: health notes, unreleased numbers, kids on a shared house chat
- When a $20 work plan is the honest privacy move, and how Mei’s 12-page letter leaked
What leaves when you hit send

People argue about brands. The useful object is the packet. Mei’s “make this warmer, keep the numbers” line did not hide $118,500. The model has to see the letter to rewrite it. Custom instructions ride along too. Files are the same payload with a different button: a PDF, a CSV, a Slack screenshot, a 12-page paste. Hidden sheet rows and Word comments still count. The filename on your disk is a local label. The other side sees content.
Even when a host says it does not keep prompt text, the call still has to be routed. That usually means an account, an IP address, a model id, a timestamp, and a token count. Ollama’s FAQ, as of writing (August 2026): local runs stay on your machine; cloud-hosted models are processed to provide the service, with a claim of no storage or training on that content, plus “basic account info and limited usage metadata.” Together’s docs say inputs and outputs are not stored by default (zero data retention), with training opt-in. Those are policy sentences. They are not “nothing left the laptop.”
After the reply, history exists. A ChatGPT or Claude thread sits in a conversation list. A Groq playground is the same shape: a product with an account. A local app writes chats to a folder on disk. Time Machine, OneDrive, and the other login on the living-room laptop can all read that folder. Delete in the UI is a button plus a vendor policy. Legal holds and abuse-log clauses are why you should not treat the trash-can as proof.
Rule of thumb: If the model id contains the word cloud, or the reply was faster than your fan, treat the paste as a document you emailed to that host.
Local is not a privacy spell
A file on disk can stay. llama.cpp, a GGUF you pulled, Ollama or LM Studio with a local id and cloud features off: the tokens run in RAM you paid for. Ollama’s FAQ states they do not see prompts when you run locally. LM Studio’s app privacy page (effective June 2026) says local messages and documents stay on the device.
That prize has conditions. Full-disk encryption is on. The app is not signed into a cloud id for this chat. Web search add-ons send query text to a search vendor. Serving the model on the house network (OLLAMA_HOST=0.0.0.0, LM Studio “serve on network”) means a sibling’s laptop can hit the same endpoint. OS14 will spend a page on sharing a box. A stolen laptop is still a privacy problem. So is the partner who uses the same macOS user. Local means “not that vendor.” It does not mean invisible.
Hosted open weights still leave
Llama or Qwen on a playground is a location story, not a license story. Groq is a hosted inference company (LPU hardware, spelling with a q). It is not xAI’s Grok chatbot. Together and Fireworks rent GPUs and list many families. OpenRouter is a switchboard: one key, many backends, so the privacy story is whichever backend you hit. As of writing, Together describes zero data retention by default and a separate toggle for passthrough models that forward traffic to an upstream provider. Fireworks has advertised zero-retention options. Groq publishes a trust center and a privacy policy. Do not paste Mei’s letter because a blog said “ZDR.” Open the host’s current privacy pages the week you send traffic.
Hosted is the right door when you want a 70B-class model your laptop cannot hold, and the text is allowed to leave. It is the wrong door when you told people-ops “we run this locally.” The license PDF on Hugging Face does not travel to the host’s logs. If you picked “open” because you wanted privacy, a Groq or Together call missed the requirement. If you picked “open” because you wanted a model family and a token bill, write that in the team note.
The desktop app with a cloud toggle
This is the trap that ate Mei’s letter, and Ken’s NDA in OS2. Ollama’s public site and docs, as of writing, sell models on your computer and cloud models offloaded to Ollama’s service. Cloud ids look like the local ones. Docs show names such as gpt-oss:120b-cloud. You stay in the same CLI and the same window. The request does not stay in RAM. LM Studio has grown a similar split: local, remote on hardware you own (LM Link), and a Secure Cloud path with a zero-retention claim. One icon, two buildings.
You can turn the cloud path off. Ollama documents local-only mode: set disable_ollama_cloud to true in ~/.ollama/server.json, or export OLLAMA_NO_CLOUD=1, then restart. Logs should show that cloud is disabled. Flags move. Read ollama --help and the FAQ the week you lock this down. Until that switch is on, a leftover :cloud id in the picker is enough. Speed is a cheap detector. Mei’s local 8B needed 22 seconds per paragraph. Four seconds is another computer. If you felt the fan on a toy prompt yesterday and you do not feel it today, look at the model line before you paste the real file.
When local is worth the heat

You do not need a threat model named after a nation-state. You need three boring piles of text that should not become someone else’s ticket, autocomplete, or training set.
Health notes
A discharge summary, a therapy journal, a list of meds, a school nurse email: that is a person. Your own notes still identify you. P4’s cafeteria paste (2,118 rows, dates of birth in column D) was this pile. Local with cloud off is a fit if you accept slower models. A contracted work workspace is a fit only if policy says health data may go there. A free personal tab on a phone is not. If you will not maintain the local stack, write the summary yourself.
Unreleased numbers
Offer letters, a forecast that is not public, headcount you have not announced, a pipeline sheet with customer names. Mei’s $118,500 plus 0.35% is this pile. So is “Q3 is going to miss by $140k” in a Slack screenshot. If you would not attach the file to an email addressed to that vendor’s privacy inbox, do not paste it. Local, or no model. A work ChatGPT or Claude workspace can be allowed when legal has a data processing agreement and an admin who turned off training-on-chats (when the product offers that control). Re-read the admin docs this month.
Kids on a shared house chat
The kitchen iPad is signed into one ChatGPT or Claude. Memory is on. A child pastes a homework prompt that includes a full name, a school, and a teacher. The next adult asks for a packing list and the model still has last night’s names in reach. That is a house policy problem, not an open-weights problem. A local app on the living-room laptop leaks the same way if everyone shares one OS user. Separate profiles. Cloud off. Or keep homework in a notebook.
When a $20 work chat is honest
Local is a control you operate. If you will not operate it, skip the trophy. A closed chat plan around $20 a month (ChatGPT Plus, Claude Pro, and cousins; as of writing August 2026, confirm the vendor page the week you buy) on a work email is often the honest privacy move for allowed text. You get an admin, a merchant finance recognizes, and, on many work tiers, a training-off control. P7 covered the $20 rung. P9 is the identity split: the email on the account is the legal person in the room.
Use that plan for a job post with no candidate names, a blog outline, a thank-you, a public FAQ. Do not stand up Ollama for a paragraph you would have pasted into Claude at work anyway. Do not put Mei’s 12 pages in Plus on a personal Gmail either. The work identity and the policy are the control. If work has no approved vendor, you still have two adult options: local with cloud off, or no model. “I used Llama so it was fine” is not a third option when Llama ran on Groq.
Worked example: 12 pages, one cloud id
Mei is people-ops at a 40-person analytics shop. Legal stuffed a 6-page IP assignment into the offer, so the draft hit 12 pages. The picker still had a cloud id from the slow 8B run. She pasted. Four seconds later she had warmer prose and a vendor document.
| Lane | What leaves | Who can see | Delete path |
|---|---|---|---|
| Local (cloud off, your disk) | Prompt stays in RAM and the chat folder. Update checks may send app version, not the letter. | You, other logins, backups, a stolen disk. | Delete the chat file. Empty backups. Encrypt the disk. |
| Hosted open (Groq, Together, Fireworks, Ollama cloud) | Prompt, files, and usage metadata go to that host. ZDR, if true this week, is a no-keep claim, not a no-send claim. | The host, subprocessors, anyone with your API key. | Account settings, org privacy toggles, key rotation. ZDR can mean nothing listed to erase. |
| Closed chat (~$20 work plan) | Text plus a stored thread. Work tiers may keep it inside a workspace. | The vendor, your workspace admin, possibly legal hold. Personal Gmail Plus is not work SSO. | In-app delete, Data controls, admin retention. Read this week’s help page. |
Mei’s row was hosted open wearing a local icon. The fix is dull. Disable cloud. Confirm the model line has no :cloud, no remote host, no Groq or Together id. Paste a two-paragraph public sample first and time it. If 22 seconds is unacceptable for a 12-page offer, either keep the offer off models or put it in a contracted work workspace legal already approved. Do not split the difference with a toggle you omit from the hiring checklist.
A checklist you can keep next to the app. Names and flags move; fill the blanks the week you run it.
# Leave-the-building checklist (before a real paste)
# Date: [today]
# App: [Ollama / LM Studio / Groq playground / ChatGPT work / Claude work]
# Mode I think I am in: [local file / :cloud id / hosted API / closed chat]
1. Copy the model line from the picker:
[ ]
STOP if it contains: cloud, remote, groq, together, fireworks, openrouter, or a URL.
2. What I am about to send:
Prompt text: yes/no
File or page count: [ ]
Names of people: yes/no
Money numbers not public: yes/no
Health, school, or kids: yes/no
3. If any names / money / health / kids is yes:
Allowed to leave the building? yes/no
If no: cloud off, confirm local-only, or do not paste.
4. Local-only (Ollama as of writing, August 2026; re-read the FAQ):
# ~/.ollama/server.json -> { "disable_ollama_cloud": true }
# or: export OLLAMA_NO_CLOUD=1
# restart, then:
ollama list
# Use a line with no :cloud suffix.
5. After the session, write the delete path I can click:
Chat lives in: [app folder / vendor history / I do not know]
I clicked: [ ]What that checklist is for: forcing a location sentence before the paste. Mei could have filled line 1, seen :cloud, and stopped. The 4-second reply would have been the warning instead of the postmortem.
Local is not the same as private
- Calling a desktop runner private because a 4.7 GB file once downloaded, then using the cloud id in the same window.
- Saying Groq when you mean Grok, or treating a Groq Llama playground as a laptop.
- Deleting a ChatGPT thread and filing that as “the vendor never got it.”
- Pasting an offer, a forecast, or a child’s homework into a shared house account with memory on.
- Building a local stack for a public blog outline that belonged on the $20 work plan.
- Leaving Ollama reachable on the house LAN and calling that offline.
Write where the prompt goes
Pick one app you already have. Write one sentence on a sticky note: local, hosted open, or closed work chat, plus the model line. Run the checklist on a public two-paragraph sample, never on an offer letter. If you use Ollama, set local-only and confirm the log line. If the job is allowed writing, pay on the website with the work email (P7) and skip the fan. Next: OS5, quality without benchmark theater. The chooser door remains P4. Hands-on local setup starts in Run open models from scratch. More paths: Learn.
Three privacy reasons that matter
- Read the model line before you paste. Cloud in the id means the text left.
- Count prompt, files, metadata, and where the transcript will sit.
- Keep health notes, unreleased numbers, and kids’ named homework off shared and hosted chats unless policy says otherwise.
- If the text is allowed to leave and you want better prose, use the $20 work plan. If it is not allowed, local with cloud off, or no model.
Sources
Names, flags, and retention claims move. Re-open these the week you toggle anything.
- Ollama: Cloud models (cloud ids, same CLI, offload to Ollama’s service)
- Ollama FAQ (local prompts stay local; cloud processing plus metadata;
disable_ollama_cloud/OLLAMA_NO_CLOUD) - Ollama (product home: computer and cloud)
- LM Studio: Desktop app privacy (local stays on device; cloud under a zero-retention claim)
- LM Studio and llama.cpp
- Groq (hosted inference; not xAI Grok) and Groq privacy policy
- Together AI: Privacy and security (default ZDR, training opt-in, passthrough toggle)
- Fireworks AI (hosted open models; confirm retention on their current legal pages)
- Hugging Face: Model cards
- AMS P4: Privacy / run it yourself (three lanes chooser)
Keep going
Same lessons in your feed
Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.
