When someone asks “where is the truth?”, the problem is usually too many copies of the data and no short card that says which one to trust. The fix is to map three things on one page: the system where each fact is first recorded, the cleaned-up copy analysts should use, and the dashboards people actually open. Then write a minimum catalog entry for the copy you trust, so people stop arguing from memory.
Say you ask three people where the “official” customer table lives. You get four answers, two chat threads, and a spreadsheet named customers_FINAL_v7_use_this.xlsx. Someone points at a dashboard dataset, and someone else swears the CRM (the sales team’s customer system) is the only truth. Engineering says “the warehouse summary table,” then admits it is three weeks behind a hotfix that never got a ticket.
This is not a tooling failure first. It is a documentation failure wearing tooling as a costume. A data catalog, which is a searchable list of your data sets and what they mean, exists to answer “where is the truth?” in under two minutes. It fails when it turns into a museum, with every table tagged, no human keeping it up, and nobody trusting the badges.
“Source of truth” is three different sentences
People use one phrase for three different jobs. Split them apart and the arguments get shorter.
- The official record: where a business event is created and corrected in day-to-day operations. Think of the CRM for opportunity stage, the orders database for checkout, and the HR system for employment status.
- Curated analytics truth: the table or model built for analysis at a defined grain, meaning a stated answer to “what is one row?” It has tested transforms, documented exclusions, and a steward, which is the person who looks after its meaning and quality. Most “dashboard truth” should live here.
- The serving product: the dashboard dataset, sync job, API, or export that consumers actually use. It should point at the curated truth and not invent a fourth definition.
If finance says “truth is NetSuite” and growth says “truth is the revenue dashboard,” both may be right about different sentences. NetSuite can be the official record for invoices. The certified warehouse model can be the curated truth for weekly net revenue, and the Looker explore can be the serving product. Your job as a steward is to map the chain, not to declare a winner based on who sends the most email.
This sits on the path described in the series on data pipelines: sources, then landing, then transform, then serve. A catalog is how you label which object is certified at which stage, and which is raw scrap you should not join into board metrics.
The truth map: one page, not a novel
A truth map is a sketch of one domain. The official records go on the left, then the landing and raw areas, then the curated models in the middle, and the serving products on the right. Add the owners and stewards, and mark what is certified and what is experimental. You are not drawing every column. You are avoiding the “which customers table?” tax.

You can build one without a week-long workshop:
- Pick one domain, such as orders, customers, or support tickets.
- List the systems that create events, with five at most to start.
- List the curated tables or models people should use for decisions, again five at most.
- List the serving products that read them, such as dashboards, extracts, and apps.
- Draw arrows only for official paths, and put “do not use for X” notes on the dangerous clones.
- Attach owner and steward names, using the roles from the earlier post on stewardship.
If a path is unofficial but heavily used, draw it as a dashed risk line and not as certified. Shadow truth is still truth in practice until you replace it, and pretending it does not exist only hides the load on the wrong tables.
Document without a 200-page wiki
Wikis die when every field is mandatory and nobody has time, and catalogs die the same way. Aim for minimum honest documentation. That means enough that a competent new hire does not ship a wrong join, but not so much that updating the page takes longer than fixing the bug.
What “minimum” means
For a certified dataset, you need answers to six questions:
- What is one row?
- Which official record sits upstream?
- Who is the owner and who is the steward?
- What must consumers never do with this table?
- How fresh is it, and what do we do when freshness fails?
- Where is the definition of the key metrics that read it?
Column-level notes can wait until the columns people actually join on. A full data dictionary is great when it is generated automatically from code, but a hand-written one for 400 columns goes stale by Thursday. The U.S. National Institute of Standards and Technology (NIST) makes the same point in its privacy framework: keep an inventory of your data, but keep it light enough to maintain.
Where the words should live
| Home | Best for | Watch-outs |
|---|---|---|
| Repo README / dbt docs next to models | Engineers and analysts who already live in git | Business users may never open it; link from the catalog |
| Data catalog tool (any vendor or open source) | Search, lineage UI, access requests, tags | Empty badges; buy-in dies if stewards do not own updates |
| One-pager wiki / Notion / Confluence | Executive-friendly narrative and domain map | Drifts from code; treat as an overview, not the last word on SQL |
| Metric catalog / metrics layer | Definitions for KPIs (ties to the metrics series) | Do not redefine the same KPI in three tools |
Pick one primary home for each type of asset. Secondary homes should link to it and not copy long definitions that will drift apart. The series on metric definitions already pushes written specs for KPIs, so catalog cards should point at those specs instead of inventing a second formula in a description field.
The minimum catalog card
Here is a card format you can paste into a catalog tool, a plain text settings file, or a wiki template. Fill in one certified dataset this week, and ignore everything else until that one is honest.

The same fields appear below as a checklist table:
| Field | Example (orders domain) | Why it exists |
|---|---|---|
| Name / ID | analytics.certified.weekly_order_facts | Unambiguous handle |
| Status | Certified / Experimental / Deprecated | Stops silent trust of junk |
| Grain | One row per order per week snapshot | Prevents duplicated rows from bad joins |
| Official record | Checkout database orders | Upstream truth for corrections |
| Owner | Finance analytics lead | Accountable for publish risk |
| Steward | Orders domain analyst | Day-to-day honesty |
| Custodian contact | Platform data engineering on-call | Jobs and grants |
| Key consumers | Executive pack, weekly growth review | Who to tell when something breaks |
| Freshness target | By 08:00 local each Monday | When late becomes an incident |
| Quality checks | Link to test suite or scorecard | Pairs with the series on data quality |
| Personal data / sensitivity | No email; hashed customer_id only | Access and clean-room choices |
| Do not use for | Real-time fraud; same-day partial weeks | Explicit anti-use cases |
| Definition links | URL to weekly net revenue spec | One home for the formula |
| Last reviewed | 2026-03-01 by steward | Staleness signal |
Here is a machine-readable sketch that you can adapt to your catalog’s API or to dbt’s meta settings:
# catalog_card.weekly_order_facts.yaml
id: analytics.certified.weekly_order_facts
status: certified
grain: "one row per order_id per week_end_date"
system_of_record:
system: checkout_db
object: public.orders
owner: finance_analytics_lead
steward: orders_domain_analyst
custodian_oncall: platform_data_eng
consumers:
- exec_weekly_pack
- growth_weekly_review
freshness:
slo: "Monday 08:00 America/New_York"
check: job_success_and_max_week_end_date
sensitivity:
pii: false
notes: "customer_id is hashed; no email or phone"
do_not_use_for:
- real_time_fraud
- incomplete_current_week_as_final
definitions:
weekly_net_revenue: "https://wiki.example/metrics/weekly-net-revenue"
quality:
suite: "tests/weekly_order_facts.yml"
scorecard: "https://wiki.example/quality/orders"
last_reviewed: "2026-03-01"
reviewed_by: orders_domain_analystNotice what is missing: 80 free-text paragraphs about company history. History can live in a linked architecture decision record, a short document that explains why a design choice was made. The card stays easy to scan.
Certification is a promise, not a sticker
“Certified” only means something if you can take it away. Pair the status with the habits from the series on data quality: quality dimensions, automated checks, and a scorecard that someone actually reads. If freshness fails three Mondays in a row and the status stays Certified, consumers learn to ignore badges.
A simple ladder of statuses works well:
- Experimental: may break, is not for board use, needs a steward, and does not need an owner yet.
- Certified: has a documented grain, an owner, a steward, checks, a freshness target, and a list of consumers.
- Deprecated: still readable for history, blocked for new builds, and given a sunset date.
Deprecation without a replacement path creates shadow clones. When you deprecate a table, name its successor on the card and in the serving product’s description in the same week.
A mini catalog for the orders domain
Imagine a mid-size company. The checkout database is the official record, and raw landing is raw.checkout_orders. The curated truth is analytics.certified.weekly_order_facts. The serving products are a finance dashboard and a growth spreadsheet extract that should die but has not. Support also keeps a personal CSV of “top customers” that must never be treated as certified.
These are the notes you would write on the truth map:
- CRM account names are not the official record for order amounts.
- The growth extract must lag the certified table by design, or it will invent its own week boundaries.
- The top-customers CSV is explicitly uncertified, and access to the personal data in it may still need controls, which the next post in this series covers.
Here are sample “do not use” rules worth putting on the certified card:
-- Example guardrail comment in a BI-facing view
-- Certified grain: one row per order_id x week_end_date
-- Do NOT join to raw.checkout_orders for board metrics
-- Do NOT filter week_end_date = CURRENT_DATE for "final" weekly revenue
-- Definition: wiki/metrics/weekly-net-revenue
CREATE VIEW analytics.certified.v_weekly_order_facts AS
SELECT *
FROM analytics.certified.weekly_order_facts;Comments like these help, but they are not a catalog. The catalog card is still the search entry for people who never open SQL.
Lineage: enough to debug, not enough to impress
Lineage means the trail showing where a table came from and where it goes. Automated lineage is wonderful when it works, but a hand-written trail of three hops (source, curated, serve) is enough for many incidents. Capture these three things:
- The upstream object name.
- The job or model name that builds the curated table.
- The downstream dashboards or extracts.
If your catalog harvests lineage automatically, still verify the certified links by hand once a quarter. Tools sometimes label temporary tables as beloved ones, so stewards confirm which links matter for decisions.
When master data and clean rooms show up
Master data management programs often own the golden customer or product record, meaning the one agreed version. Your analytics catalog should say whether a customer table is fed by that program or comes straight from the CRM. Do not quietly rebuild a second golden customer in the warehouse without linking to the program’s decision.
Clean rooms are a different serving pattern. They let several companies measure something together while sharing only limited identifiers. Document them as serving products with strict do-not-use rules, not as general-purpose customer tables. If someone asks whether the clean room is the customer truth, the answer is usually no, because it is a constrained space for working with partners.
Common mistakes
- Cataloging everything in month one: you get empty fields and cynicism. Start with certified assets only.
- A status without teeth: Certified forever, with no review date and no checks.
- Duplicate definitions: the metric formula sits in the catalog, the wiki, and a chat canvas, all slightly different.
- Ignoring serving products: you document only warehouse tables while executives live in a spreadsheet extract.
- No owner or steward names: a card without humans is a brochure.
- Confusing the official record with curated truth: dashboards get forced to query the live production database “because truth.”
- Wiki novels: 200 pages nobody updates. Prefer short cards plus links to code and metric specs.
Practice: 45 minutes this week
Choose one certified dataset, or one that should be. Draw the three-column truth map on a whiteboard and take a photo. Fill in the minimum card fields, including “do not use for” and “last reviewed,” and put the card where people already search, whether that is the catalog tool or a wiki. Then send one message: “If you use another table for this decision, reply with why.” Collect the shadow paths without shaming anyone, and schedule deprecation or documentation next.
The next post in this series covers access, least privilege, and the eternal “just give me prod” request. A catalog without access rules becomes a treasure map for the wrong adventure. For more structured learning paths, see the Learn hub.
Quick recap
- Split the official record, the curated analytics truth, and the serving product.
- Draw a one-page truth map for each domain, with owners and stewards.
- Use a minimum catalog card: grain, status, humans, freshness, sensitivity, do-not-use, and definition links.
- Document lightly, link to metric specs and quality checks, and review on a set date.
- Governance and master data programs set the standards, and catalogs make them findable in weekly work.
Series notes
This is Part 2 of Data stewardship at work. The previous post covered steward roles, and the next covers access and least privilege.
Sources
- DAMA International, data governance and metadata management concepts in DMBOK: https://www.dama.org/cpages/body-of-knowledge
- DataHub (open source metadata platform docs, vendor-neutral catalog patterns): https://datahubproject.io/docs/
- OpenMetadata documentation (catalog, glossary, and data quality integration patterns): https://docs.open-metadata.org/
- dbt Docs (documentation co-located with transforms): https://docs.getdbt.com/docs/collaborate/documentation
- NIST Privacy Framework (inventory and data processing documentation as risk practice): https://www.nist.gov/privacy-framework
- ISO/IEC 11179 (metadata registries concepts, high-level background): https://www.iso.org/standard/78914.html
Keep going
Same lessons in your feed
Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.
