Skip to content
,
Practical AI for analytics people · Part 5

RAG in plain English

11 min read
Editorial featured image for RAG in plain English. Title text reads RAG in plain English.

RAG stands for retrieval-augmented generation. It means the system finds the useful pieces of your company’s documents at the moment someone asks a question, puts those pieces into the prompt, and lets the model answer from that open book. It helps with “what did we write down?” questions, and it does not replace a warehouse query or a human owner for numbers that matter.

Say someone asks your internal chatbot, “What is the official definition of active customer?” The model answers fluently, and the answer sounds like last year’s all-hands meeting. It is also wrong, because Finance updated the rule on a Confluence page in March and the model never saw March.

RAG in plain words

Retrieval-augmented generation works like this. Before the model answers, a system searches a document store for passages that look relevant to the question. It then adds those passages to the prompt so the model can use them. The generation step is still a language model, and the new idea is the open-book exam. A closed-book answer comes from training alone, while an open-book answer comes with these pages on the desk.

Researchers popularized the name RAG for combining two kinds of memory. One is what the model learned in training, and the other is what you retrieve at the moment of the question. You do not need to read the paper to use the idea, but you do need some humility. Retrieved text can be wrong, out of date, or about the wrong product line, and the model will still sound sure of itself.

Rule of thumb: RAG upgrades the prompt with company context. It does not replace data quality, metric ownership, or reading the source page when the number matters.

The flow from documents to answer

Strip away the product names, and most company knowledge assistants do some version of these six steps:

  1. Collect docs: wiki pages, metric cards, runbooks, exported chat threads (handled carefully), PDFs, and ticket templates.
  2. Chunk: split long pages into smaller pieces, so that search can rank single paragraphs and not a 40-page manual as one blob.
  3. Index: store the chunks so you can find them fast. This often uses embeddings (numbers that capture what a passage is about) and a vector index, but it can also be plain keyword search, or both.
  4. Retrieve: turn the user’s question into a search and pull the top chunks.
  5. Stuff, or select: put the best chunks into the model’s prompt, within the context window (the amount of text the model can read at once) budget covered earlier in this series.
  6. Generate: the model writes an answer, ideally with citations to the chunks.
RAG in four phases: ingest docs, retrieve chunks, augment the prompt, generate an answer. Grounding is not a warranty.
RAG in four phases: ingest docs, retrieve chunks, augment the prompt, generate an answer. Grounding is not a warranty.

Vector databases show up in the indexing step when you want similarity search over embeddings. That is useful for a request such as “find paragraphs about refund exclusions even if the wording differs,” but it is not required for every internal FAQ. Keyword search still works when people share the same vocabulary. Our vector post covers the storage side, and this post cares about the decision itself: do you need retrieval at all, and what fails when you have it?

Why analytics people should care

Analytics work is full of half-structured knowledge that never lands cleanly in a warehouse column:

  • Metric definitions and edge cases, such as “active means logged in once in 28 days, except enterprise trials.”.
  • Instructions for the scheduled jobs that move data between systems (often called pipelines), plus the “do not use this table on Mondays” folklore that should be written down.
  • Notes on what each dashboard, a page of charts that updates from your data, is for, and known caveats.
  • Access policies and notes on handling personal data.
  • Old incident writeups that explain strange spikes.

A database query, written in SQL (the standard language for asking a database questions), does not answer “what does Finance mean by net.” A model alone invents a plausible Finance. RAG plus a maintained definition page can quote the real rule. Catalogs and stewardship habits still matter, because a bot that retrieves a stale card is just a faster way to spread the wrong definition. Pair this thinking with data stewardship and with the metric specs from the metrics series when those paths are on your Learn map.

Grounding is not a warranty

Vendors love the word “grounded.” In practice it only means that the model saw some retrieved text. It does not mean any of these things:

  • The retrieved text was the latest approved version.
  • The chunk was complete, and half a definition can be worse than none.
  • The model used the chunk instead of its old habits.
  • Two conflicting docs were resolved correctly.
  • The answer is safe to put on a board slide without a human check.

Good systems show citations and let you open the source. Great teams treat citations as the start of an audit and not as a stamp of approval. The test suite from the earlier post on human evals should include at least one RAG case, such as “answer only from these docs, and if the topic is missing, say you do not know.”

When RAG helps and when it does not

Use this as a conversation card for your team, and not as a rulebook.

When RAG helps versus when it does not table
When RAG helps versus when it does not table
SituationRAG often helpsPrefer something else
Policy and definition Q&AYes, if pages are current and citedIf definitions live only in Slack lore, fix docs first
Onboarding “where is X?”Yes, for runbooks and linksStill maintain a human map for critical systems
Exact metric number for last weekOnly if retrieval hits a trusted report store and you verifyQuery the warehouse or certified dashboard; do not invent from prose
SQL generation against live schemaMaybe, if schema cards are retrievedSchema tools + human SQL check usually beat random wiki dumps
Conflicting sourcesOnly if the system surfaces conflictOwner decision, not a chatbot vote
PII or secret runbooksDangerous without access control on the indexPermissioned systems; least privilege (stewardship habits)
Fast-changing ops statusWeak if index is hours staleStatus page, monitoring, on-call
Teaching concepts (what is a join?)Optional; general models already knowTutorials and practice; see SQL series

Notice the pattern. RAG is strong for “what did we write down?” It is weak for “what is true in the data right now?” and weak for “what should the company decide?” Those questions need queries and owners.

A worked example: what is an active customer?

An active customer is whatever your company’s current, certified definition says, and RAG only gets that right if it finds the current document instead of an old one. Suppose three documents exist in your company:

  • Wiki page A (2024): active means any login in 90 days.
  • Metric card B (March 2026, certified): active means a paid seat with at least one intentional session in 28 days, with internal tenants excluded and trials counted separately.
  • Chat thread C: a vice president saying “just use 30 days for the board, whatever.”.

Closed-book chat (no RAG): the model invents a plausible software-company definition, maybe 30 days and maybe 90, with nothing specific to your company. It sounds fine and is wrong for your Finance process.

Naive RAG: the search returns page A because it is longer and older and ranks high on the phrase “active customer.” The model quotes 90 days with a citation, which is still wrong for 2026 reporting.

Better RAG design: the index prefers certified metric cards, stores effective dates, and boosts anything tagged “certified.” The prompt says to prefer metric cards over informal pages, and to list the sources and stop if they conflict. The answer cites card B, notes page A as outdated if that information exists, and ignores the chat thread unless policy allows it.

A human still owns the number: if you need a headcount for the board, you run the certified query or open the certified dashboard. The RAG answer helps you remember the rule, and it is not a substitute for the measurement.

Here is a tiny prompt pattern that helps when you control the assistant’s instructions:

You answer using ONLY the passages provided below.
If the passages are missing, incomplete, or conflict, say so explicitly.
Quote the definition and name the source title and date when available.
Do not invent SQL table names. Do not invent numeric results.
Passages:
{{retrieved_chunks}}
Question:
{{user_question}}

That pattern is only as good as the retrieval behind it. If card B never enters {{retrieved_chunks}}, the model cannot quote it, so garbage in still means fluent garbage out.

Chunks, context windows, and the suitcase problem

The earlier post on tokens treated them like suitcase space. RAG fills that suitcase with other people’s packing. If you retrieve ten noisy chunks, you crowd out the question and the table layout. If you retrieve one tiny chunk that cuts a definition in the middle of a sentence, you get half a rule. These habits help:

  • Prefer fewer high-quality chunks over many that are only “sort of related.”.
  • Keep metric cards short and self-contained, so a single chunk carries the full rule.
  • Store titles, owners, and dates with each chunk, so answers can cite them.
  • Re-index when certified pages change, and not “someday.”.

Long PDFs of board decks are often poor material to retrieve from, while clean metric cards and runbooks are better. That is a documentation problem wearing an AI hat.

RAG is not a warehouse, and not an agent

Two confusions come up again and again in meetings:

  • “We will put the data warehouse into RAG.” Usually people mean “embed dashboard PDFs” or “index table descriptions.” That can help people find things, but it does not turn the model into a correct calculator for weekly revenue. Aggregations belong in SQL and certified transforms. See the pipelines and metrics series on Learn when you need that stack.
  • “RAG is our agent.” RAG retrieves and generates. Agents, which the next post covers, choose tools, take several steps, and can write back to systems. That is a different risk class. You can put RAG inside an agent, but you should not pretend that a document chatbot is already safe automation.

What to put in a RAG eval, sized for humans

Borrow the approach from the earlier post on human evals, and add cases like these:

CasePass looks like
Known definitionQuotes current certified card; cites source
Deprecated page still in indexDoes not prefer old page; or flags conflict
Missing topicSays unknown; does not invent policy
Numeric askRefuses to invent; points to dashboard or SQL path
Access-sensitive docNot returned to unauthorized user (system test)

If your bot cannot pass the “missing topic” case, it is a fiction generator with footnotes.

Mistakes that come up again and again

  • Indexing everything and curating nothing. Volume without ownership makes retrieval confidently wrong.
  • No effective dates. Definitions with no “as of” date turn into landfill.
  • Treating citations as proof. A citation proves that a chunk was present, not that the answer is ready for a decision.
  • Skipping access control on the index. A search index can become a side door to personal data.
  • Replacing the catalog with chat. Catalogs and stewards still need named humans, because chat is a screen and not an owner.
  • Expecting RAG to fix bad joins. Wrong SQL is still wrong, so check it.

Practice this week

  1. List five questions your team asks a chatbot, or wishes they could.
  2. For each one, name the single document that should answer it. If none exists, write a half-page card before you buy any tools.
  3. Mark which questions need a warehouse query instead of a document answer.
  4. If you already have an internal assistant, run three of those questions and open every citation, noting any stale or missing sources.
  5. Add one “missing topic” case and one “conflicting docs” case to your eval sheet.

The next post in the series covers agents, tools, and harnesses. It looks at what happens when the model can not only read docs but also call systems, and how limits and human review keep that from becoming an automated incident. For vector storage detail without redoing this conceptual map, stay with the vector databases post.

Quick recap

  • RAG retrieves company passages, puts them in the prompt, and then generates an answer.
  • It helps with “what did we write down?” more than with “what is true in the data right now?”.
  • Vectors and indexes are implementation details, so link out for that depth and do not confuse them with judgment.
  • Grounding is not a warranty, citations help an audit, and humans still own the numbers.
  • Curate docs, dates, and access, test the missing and conflicting cases, and pair RAG with SQL checks for anything that a query can answer.

Series notes

This is Part 5 of Practical AI for analytics people. The previous post covered human evals, and the next covers agents, tools, and harnesses.

Sources

Sources include work from the US National Institute of Standards and Technology (NIST) and from the Open Worldwide Application Security Project (OWASP).

Written by

Jose S

Founder & Lead Analyst · Analytics Made Simple

Hands-on data strategist, analytics engineering lead, and educator. Writing practical, no-fluff guides to help everyday teams, analysts, and engineers master SQL, AI systems, and modern data architectures.

Keep going

Same lessons in your feed

Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.

Google Search Prefer our practical guides in Google Search & Top Stories: