This path is for analysts who inherited a pipeline and were never in the room when it was designed. Parts walk the path from sources to a table someone queries, batch versus streaming, where the data lives, orchestration, what dbt is for, dev versus prod, and how you notice a break. Start with the first part on this page.
- 1 Sources, land, transform, serve Inherited a broken dashboard number? Map data from sources through land, transform, and serve so you can debug the path, not the logos. Part 1 of How data actually moves for analysts who inherit pipelines.
- 2 Batch vs streaming in plain English Real time is a latency budget, not a vibe. Learn batch, streaming, and micro-batch in plain English so you know when yesterday is fine, when minutes matter, and when delay truly costs money.
- 3 Warehouses vs lakes vs just a database App database, warehouse, or lake? Walk a plain decision tree for where data should live so product traffic, cheap history, and certified analytics stop fighting over one system.
- 4 Orchestration in one metaphor Orchestration is kitchen management for data: schedules, dependencies, and retries. Learn the Airflow and dbt mental models so broken morning numbers point to a clear task, not a vague outage.
- 5 What is dbt (conceptually)? dbt turns SQL transforms into versioned models with tests and docs. Learn staging, intermediate, and marts so you can read a change even if you do not own the repo.
- 6 Environments: dev, stage, and prod Laptop truth is not production truth. Learn what dev, stage, and prod mean for data work, a simple promotion path from notebook to trusted table, secrets hygiene, and who should be allowed to write where.
- 7 Observability for pipelines Pipeline observability means early signals: freshness, volume, schema, and a business check. Red, yellow, green without theater, plus a human who gets paged.
