When you use AI on private data, you have three places to send it: a closed chat app with strong controls, a hosted open-weight model on someone else’s server, or a model running on your own machine. Pick the place on purpose before you paste anything.
Say you have a file called patients_q2.csv with 2,118 rows and dates of birth in column D. You paste it into a free personal chat on your phone in the cafeteria to “just get a summary.” The model is helpful, and the problem is the room, the account, and the file. Legal asks later whether the vendor ever had a contract with your company, and nobody has a good answer.
Three places your text can go

| Lane | What you get | Where the text goes | Pick it when |
|---|---|---|---|
| Closed chat (Claude, ChatGPT, Gemini, Grok) | Best everyday quality for most people | Vendor cloud, under that product’s policy | The content is allowed in that product, ideally a paid work plan |
| Hosted open model (Groq, Together, a cloud Llama, etc.) | Often cheaper APIs, downloadable cousins exist | Still someone else’s GPU | You want a model family, not a laptop project |
| Local (Ollama-style, llama.cpp, a GUI) | Control, offline, no cafeteria upload | Your disk, your RAM, your fan | The file must not leave, and you will maintain the stack |
A file like that belongs in the third lane, on your own machine, or in a contracted work workspace. It does not belong in a personal free tab. A hosted Llama model is still run by a vendor, and downloading the weights later does not un-send Tuesday’s CSV. For the vocabulary of weights versus licenses, read the post on what “open” means for AI models after this page.
Before you paste

- Whose account is it? A personal free chat is not your hospital, your school, or your company.
- Would you email this CSV to a random vendor? If not, do not paste it.
- Can you strip names and dates of birth first? A 20-row toy file beats 2,118 real people.
- Do you need a local model, or do you need a contract? Comfort is not a threat model.
Rule of thumb: Privacy is a location and a contract, not a brand personality. “It feels trustworthy” is not a control.
When local is the wrong trophy
Local models cost hardware, heat, and your own time as the person who has to keep them patched. A 7-billion-parameter model on a laptop can summarize a meeting note, but it will not reliably match a top paid chat plan on hard writing. If your real requirement is good prose from data you are allowed to share, a work plan for Claude, ChatGPT, Gemini, or Grok with admin controls may beat a weekend spent setting up Ollama. If the real requirement is that this CSV never leaves your computer, you accept the drop in quality or you do not use a model at all.
The Practical AI series on this site already covers what you should not paste for analysis. It is the same muscle through a different door.
A 15-minute local smoke test, if you want one
Try this only if you have already decided that the file must stay on your machine. It is not a full Ollama tutorial, and that series comes later.
# Example shape only. Package names and flags move.
# 1) Install a friendly runner (Ollama-style) from its current site.
# 2) Pull a small model you can actually run (think 7B-class, not a 70B hope).
# 3) Chat with a toy paragraph, not patients_q2.csv.
# Fake check: if Activity Monitor shows the fan and the reply is slow,
# you learned the hardware tax. That is the point of the smoke test.The smoke test exists so you can feel the fan spin before you promise Legal a local setup. If you will not keep it updated, stop. Use a contracted work chat, or do the summary yourself.
A worked example: 2,118 rows in a cafeteria
| Choice | What leaves the laptop | Quality of the summary | Incident risk |
|---|---|---|---|
| Personal free chat | The CSV | High | High |
| Work workspace, allowed data | The CSV, under a contract | High | Lower if policy matches |
| Toy 20-row strip, then any chat | Fake names | Good enough to learn | Low |
| Local small model on the real file | Nothing | Uneven | Low on network, still a stolen-laptop problem |
The cafeteria paste is the first row of that table. The better first try for a regulated file is the third or fourth row, and never the first row on public Wi-Fi.
Work plans are a privacy tool too
Teams sometimes treat “we refuse all vendors” as the only moral position, and then someone pastes anyway from a phone. A paid work workspace has a data processing agreement, single sign-on (SSO, one company login for many tools), and an admin who can turn off training on chats when the vendor offers that control. That is a real control. It is not as strong as a machine with no network connection, but it is much stronger than a personal tab in a cafeteria. Read the current admin docs, and do not quote a social media thread from 2024 as your policy.
Personal projects are allowed to be sloppy in a different way, since your recipe blog is not 2,118 dates of birth. The chooser still asks you to pick a lane on purpose. If you want local for learning, that is great, so feel the fan and then decide. The later local-model series will spend pages on Ollama, llama.cpp, and point-and-click apps, and this page only has to stop the unexamined paste.
Hosted open models, without the halo
A DeepSeek or Llama endpoint at Groq, Together, Fireworks, or a cloud you already pay for can be cheap and fast. You still sent the text off your device, and the license on the weights does not travel to the host’s logs. If your reason for wanting “open” was privacy, a hosted open model missed the point. If your reason was cost, or a model family you like, it may be the right lane. The post on what “open” means is where the license vocabulary lives, so keep the lanes honest here: first the location of the bytes, then the license of the file.
Mistakes that come up again and again
- Calling a free ChatGPT tab “private” because you deleted the thread. Deletion is not a legal theory you should invent.
- Assuming Llama on a host is local, when the graphics processor (GPU) is still in someone else’s building.
- Promising Legal a local setup that you will not keep patched.
- Pasting dates of birth to “make the summary more personal.”.
- Skipping the license vocabulary and calling every downloadable model “open source.”.
A copy-paste policy you can actually use
Teams stall because they wait for a 40-page AI policy. You can ship a 12-line one this week and let Legal thicken it later.
Paste policy (v0)
Allowed in personal free chats: public text, toy data, your own shopping list.
Allowed in work workspace only: internal drafts without regulated fields.
Never: dates of birth, patient/student IDs, payroll, unpublished financials,
other people's inboxes, secrets, API keys.
If you cannot strip it, do not paste it.
Local models: only on company laptops if IT says so; still encrypted disk.
When unsure: ask [name], do not test in the cafeteria.This policy names the lanes without pretending that a Llama icon is a control. Put it in the same place as the 10-line chooser note from the post on choosing an AI product, and review it every quarter because vendors change their training-on-chat settings. The 2,118-row file fails every line after “Never.” That is the point of writing the policy down before the export.
Travel and coffee shops make people sloppy. If you must work on a train, work on the toy file and not the real one. A privacy screen helps against people looking over your shoulder, but it does nothing for the upload request that already left the phone. The location of the human and the location of the bytes are different maps, and this post is about the bytes. Stolen-laptop risk still exists for local models, so encrypt the disk, use a firmware password if you know how, and do not leave patients_q2.csv on an unlocked cafe table next to the “private” model. Local is not magic. It is a smaller room with a lock that you have to remember to use.
How to practice this week
Take one file you almost pasted and strip it down to a toy. If you cannot strip it, you cannot paste it. After that, read the post on images and live answers, or the post on what “open” means if you want the license vocabulary. The decision tree is in the post on which AI product to try first.
Quick recap
- Closed chat, hosted open models, and local models are different places for your text to live.
- Ask whether you would email this to a vendor, and if not, do not paste it.
- Local is control plus homework, and it is not a free lunch.
Series notes
This is Part 4 of Which AI product should I use?. Related: privacy when pasting data into chat tools.
Sources
- Vendor privacy / terms pages for Claude, ChatGPT, Gemini, Grok (re-check; they change).
- Open Source Initiative: Open Weights
- Open Source Initiative: Open Source AI Definition
- Analytics Made Simple: Practical AI
Keep going
Same lessons in your feed
Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.
