Skip to content
,
Airflow hello DAG · Part 1

Airflow hello: your first DAG as a ticket rail

11 min read
Airflow hello: your first DAG as a ticket rail, with the official product logo. Editorial illustration for Analytics Made Simple.

Airflow is a tool that runs your data jobs in the right order, on a schedule, and shows you which one broke. That job order is easy to hold in your head on a good day and easy to lose on a bad one. A first Airflow workflow, called a DAG, works best as a ticket rail: a visible track of small tasks with clear handoffs, and not a dumping ground for every script you own.

Say you get a ticket that reads “run the nightly load.” In practice that means you wait for the file drop, load a staging table (a holding area for raw data), run the transforms, ping the warehouse tests, and only then refresh the dashboard extract. On a good week one person remembers the order. On a bad week two jobs start at once, a transform reads half a file, and the dashboard looks confident while it is empty. Orchestration, which just means coordinating jobs in order, is how teams stop treating someone’s memory as infrastructure.

This post is the first in the Airflow hello DAG series. It keeps the install light and the ideas heavy. For broader pipeline context see data pipelines, for transform habits see the dbt lab and data quality, and for the full map open Learn.

Airflow in one honest sentence

Apache Airflow is an open source platform to author, schedule, and monitor workflows written as code. You define a DAG, short for directed acyclic graph, which is a set of tasks with dependencies and no loops. Airflow’s scheduler decides when a run of that DAG should start, and workers (also called task runners, depending on how you set it up) carry out the tasks. The web screen then shows what is running, what failed, and what is waiting.

Airflow is not your warehouse, it is not dbt, and it is not a replacement for good SQL. It also does not make a bad task reliable. It is an orchestrator, which means it puts work in order, records the history, and gives people a place to look when the night went wrong.

Rule of thumb: Put in Airflow the steps you would write on a runbook sticky note. Keep heavy business logic in the tools those steps call, such as SQL, dbt, Spark jobs, or APIs, and avoid packing it into giant Python operators if you can help it.

The ticket rail metaphor

Think of the physical ticket rail in a restaurant kitchen or at a support desk, where tickets move left to right through stations. You can see at a glance which station is stuck. Nobody cooks the entire meal inside the rail itself, because the rail only coordinates the stations.

A good first DAG looks like that, with five stations.

  • Station 1 confirms that the raw file or API extract arrived.
  • Station 2 loads it into staging.
  • Station 3 runs the transforms, for example dbt build.
  • Station 4 runs data tests or a row-count check.
  • Station 5 sends a notice or triggers a downstream refresh.

Each station is a task, and the connections between stations are dependencies. If station 2 fails, station 3 should not pretend it succeeded. That is the whole point of a rail, compared with five unrelated scheduled lines that only work if everything happens to finish in the right order.

Filled Airflow ticket rail: extract, transform, test, notify
Filled Airflow ticket rail: extract, transform, test, notify

Core terms without the fog

TermPlain meaningKitchen analogy
DAGWorkflow definition (the menu and order of stations)The full ticket path for one dish type
DAG runOne execution of that workflow for a logical dateTonight’s tickets for that dish
TaskOne unit of work in the DAGOne station on the rail
OperatorTemplate for how a task runs (Bash, Python, empty, etc.)The type of station equipment
ScheduleWhen new DAG runs are createdWhen the kitchen opens a new batch
DependencyTask B waits for task ANo plating before cooking
XComSmall messages between tasks (use sparingly)A short note pinned on the ticket

You will also hear the word executor, which is the machinery that actually runs tasks (on one machine, across many, or inside containers). For a first mental model, ignore the executor debate and focus on clear tasks with honest dependencies. The next post in this series covers retries and sensors (tasks that wait for something to appear), and both sit on top of this rail.

Why “as code” matters

Airflow DAGs are Python files in a code repository, or a managed equivalent. That gives you three practical benefits.

  • You can have teammates review the runbook the same way they review a dbt model.
  • History lives in git, and not in a click-through console that someone inherited.
  • Test and production environments can use the same DAG definition with different connections.

The flip side is that the Python which builds the DAG runs in a special context, and heavy data work inside the DAG file can slow down the whole scheduler each time it re-reads the file. Keep DAG files thin. Declare the structure, and call out to jobs that do the real work. If you already check SQL carefully, the same skepticism applies to DAG code drafted by an AI tool, and the AI SQL checks carry over to generated operators.

A minimal first DAG (read-along)

Below is a teaching sketch, and you should not copy it into production as is. Names and imports vary slightly by Airflow version, so the shape is what matters: default settings, a DAG block, tasks, and then dependencies.

from datetime import datetime, timedelta
from airflow import DAG
from airflow.operators.bash import BashOperator
from airflow.operators.empty import EmptyOperator

default_args = {
    "owner": "analytics",
    "depends_on_past": False,
    "email_on_failure": False,
    "retries": 1,
    "retry_delay": timedelta(minutes=5),
}

with DAG(
    dag_id="hello_orders_rail",
    default_args=default_args,
    description="Ticket rail: extract check -> stage -> dbt -> test gate -> notify",
    schedule="0 6 * * *",  # 06:00 UTC daily; set timezone policy with your team
    start_date=datetime(2026, 1, 1),
    catchup=False,
    tags=["hello", "orders"],
) as dag:

    start = EmptyOperator(task_id="start")

    check_extract = BashOperator(
        task_id="check_extract",
        bash_command="test -f /data/incoming/orders_{{ ds }}.csv",
    )

    load_staging = BashOperator(
        task_id="load_staging",
        bash_command="python /opt/jobs/load_orders.py --date {{ ds }}",
    )

    run_dbt = BashOperator(
        task_id="run_dbt",
        bash_command="cd /opt/dbt && dbt build --select tag:orders",
    )

    quality_gate = BashOperator(
        task_id="quality_gate",
        bash_command="python /opt/jobs/check_orders_counts.py --date {{ ds }}",
    )

    notify = BashOperator(
        task_id="notify_success",
        bash_command="echo 'orders rail ok for {{ ds }}'",
    )

    end = EmptyOperator(task_id="end")

    start >> check_extract >> load_staging >> run_dbt >> quality_gate >> notify >> end

Read the code as a ticket rail, one line at a time.

  • start and end bookend the run so it is easier to read in the graph view.
  • check_extract fails right away if the file is missing, and the next post in the series covers sensors for cases where you would rather wait.
  • load_staging and run_dbt do the real work by calling scripts, so you do not have 200 lines of code sitting inline.
  • quality_gate is a deliberate station, because green transforms do not prove the counts make sense.
  • {{ ds }} is the date string Airflow fills in for each run, so every run is tied to one specific day.

Dependencies: the rail edges

The line start >> check_extract >> ... sets a straight rail, where each task waits for the one before it. Real jobs often branch, so here is how you run two independent transforms in parallel and then wait for both.

# After load, run two independent transforms, then join for tests
load_staging >> [run_dbt_orders, run_dbt_customers]
[run_dbt_orders, run_dbt_customers] >> quality_gate

Four rules keep a rail sane.

  • No loops. Task A cannot wait for task B while task B waits for task A.
  • Stop when something upstream fails. Downstream tasks should not run if the task before them failed, and that is the default behavior for normal dependencies.
  • Draw explicit edges, and avoid hidden coupling through shared files that have no task connecting them.
  • Never fake success. A notify task that always shows green after a failure trains people to ignore the rail.

Schedule, start_date, and catchup (the confusing trio)

Three settings confuse nearly every first-time user, so it helps to say what each one controls.

  • schedule (or the older name schedule_interval) sets how often new runs are created.
  • start_date is the earliest logical date the DAG is allowed to consider.
  • catchup decides whether Airflow should create backfill runs from the start date up to now.

For a hello DAG, set catchup=False so that turning the DAG on does not fire a year of historical runs by surprise. When you do need history, run a deliberate backfill with your eyes open. Also agree on a time zone policy with your warehouse and product teams, so “daily 6am” means the same morning for everyone.

Worked example: mapping a real ticket to tasks

Here is a ticket from your analytics operations team.

Every morning we need yesterday’s orders in the mart before 8am local. File lands in S3 around 5:30. Load, dbt orders models, confirm row count within 5% of yesterday, then Slack #data-status.

Each phrase in that ticket turns into one task, and each task gets a clear meaning of success.

Ticket phraseTask idWhat success means
File lands in S3check_extract (or a sensor later)Object exists for the logical date
Loadload_stagingStaging table replaced or appended correctly
dbt orders modelsrun_dbtSelected models build without error
Row count within 5%quality_gateGate script exits 0 only if rule holds
Slack #data-statusnotify_successMessage sent only after gate passes
Example DAG run result strip: check_extract success, load_staging success, run_dbt success, quality_gate success, notify_success success for logical date 2026-11-03
Example DAG run result strip: check_extract success, load_staging success, run_dbt success, quality_gate success, not…

That green strip is what stakeholders actually want. They do not care that Airflow is installed, they care that yesterday’s orders cleared the rail. When something fails, the strip shows which station broke, so you do not have to restart everything from the beginning.

What to put in a task and what to keep outside Airflow

  • Inside the task, put the call that starts a job, the date you pass to it, timeouts, and the rule that turns exit codes into success or failure.
  • Keep transform logic of several hundred lines, secret material, and ad hoc analysis notebooks outside the task body when possible.
  • Store warehouse credentials, bucket names, and environment-specific paths in Airflow’s connections and variables.
  • Never commit passwords or tokens to git. Use a secrets backend or environment injection instead.

This mirrors the environment discipline of modern data teams. Orchestration coordinates the work, and the warehouse and the transform tools still own the heavy lifting. dbt stays a great transform station on the rail, so Airflow should call it and not rebuild what it does.

Observability for humans

On day one, agree on three operational basics so nobody is guessing when something breaks.

  • Decide who owns the DAG, and put a real team name in owner and in the tags.
  • Decide where failures go: Slack, email, PagerDuty, or a ticket queue.
  • Decide the SLO, meaning the promise you make to the business. “Mart ready by 8am local” is better than “the job usually finishes.”

Name tasks after business stations, and not after temporary scripts. Whoever is on call in the middle of the night will thank you.

Common mistakes

  • One giant task. A single Bash blob that does load, transform, and notify hides the place where things failed.
  • Scheduled lines instead of dependencies. Five independent schedules that “usually” finish in order are not a rail.
  • Catchup surprises. Turning on an old start date with catchup set to true can flood the warehouse with old runs.
  • Business logic buried in operators. Python that nobody can review grows quickly, so prefer callable jobs and dbt models.
  • Alerts that lie. Some teams alert only on success, and others send a success alert even when the gate failed.
  • No quality station. “dbt finished” does not mean “the numbers are plausible.”
  • Hard-coded dates. Always prefer the run’s logical date templates over a constant like yesterday.
  • Treating Airflow as a notebook host. Long interactive analysis does not belong on the scheduler path.

How to practice this week

  • Write a five-station rail on paper for one real daily job that you already run by hand or on a timer.
  • Translate that paper rail into a DAG sketch with task ids and one line each for the success criteria.
  • If you have a sandbox Airflow, build a hello DAG that uses only EmptyOperator and Bash echo, so you learn the screens without touching production data.
  • Add a deliberately failing task and practice reading the graph and the logs.
  • List which station should call dbt and which should call custom Python in your setup, and keep that boundary clean.

The next post in the series covers retries, sensors, and when Airflow is the wrong tool. Do not skip it if you are about to wait on flaky files or wrap every spreadsheet refresh in a DAG.

Quick recap

  • Airflow orchestrates workflows written as code, with DAGs, tasks, schedules, and visible history.
  • Treat a first DAG as a ticket rail of stations with honest dependencies.
  • Keep DAG files thin, and put heavy logic in jobs and transform tools you already trust.
  • Learn schedule, start_date, and catchup before you enable anything near production.
  • Map real tickets to task ids and success criteria so the graph matches the business.
  • Next up are retries, sensors, and knowing when not to use Airflow at all.

Series notes

This is Part 1 of the Airflow hello DAG series. Related: data pipelines and data quality.

Sources

Written by

Jose S

Founder & Lead Analyst · Analytics Made Simple

Hands-on data strategist, analytics engineering lead, and educator. Writing practical, no-fluff guides to help everyday teams, analysts, and engineers master SQL, AI systems, and modern data architectures.

Keep going

Same lessons in your feed

Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.

Google Search Prefer our practical guides in Google Search & Top Stories: