Skip to content
,
Airflow hello DAG · Part 2

How to use retries and sensors in Apache Airflow, and when to skip Airflow

11 min read
How to use retries and sensors in Apache Airflow, and when to skip Airflow

Airflow is a tool that runs your data jobs in the right order on a schedule. A DAG (directed acyclic graph) is simply one of those jobs written out as a chain of steps. Retries give a failed step another chance, and sensors make a step wait for something, such as a file, before it starts. Knowing when Airflow is the wrong tool is the harder skill, and this post covers that too.

Say you are waiting on a file from a vendor, and it is late. Your DAG starts on schedule and the first step fails. Your team chat lights up, and you click “clear and rerun.” By the third morning, you have taken on a new job title: human sensor. Retries and sensors exist so that people stop playing that role. Used badly, they also hide real outages under a blanket of “it will work next try.”

Picture what happens without limits. Your sensor checks an empty folder in cloud storage (Amazon S3) every minute all night, and it ties up worker capacity because nobody set a timeout. A hard deadline and a clear failure path would raise one alert with a useful message instead of hundreds of empty checks.

Retries: recover from flukes, not from design bugs

A retry runs a failed task again, a limited number of times, after a delay. Good retries assume that the failure might be temporary, such as a blip in the network, a short queue at the warehouse or a rate limit (a cap on how many requests you may send) that clears on its own.

Bad retries assume that waiting five minutes will make a missing file appear or fix a wrong join. It will not, and all you have done is delay the alert and pile up overlapping work.

default_args = {
    "owner": "analytics",
    "retries": 2,
    "retry_delay": timedelta(minutes=10),
    # optional: retry_exponential_backoff=True in many setups
}

These guidelines keep retries honest.

  • Low counts for loaders and APIs (the connections one program uses to ask another for data) that can flake, where 1 to 3 retries is common.
  • Zero or one retry for logic bugs that you expect to fail the same way every time.
  • Idempotent tasks, which means running them twice gives the same result, so a second try does not insert rows twice.
  • Alert after the final failure and not after every attempt, unless you really want the noise.

Suppose a task fails three days in a row and always succeeds on the second retry. You do not have a resilient system. You have a hidden dependency on timing, meaning the step only works if something else happens to finish first. Fix that dependency, often with a sensor or a delivery deadline agreed with the upstream team, instead of raising the retry count forever.

Retry rule: Retries buy time for things to settle. They do not replace a correct wait condition or a correct join.

Sensors: waiting as a first-class station

A sensor is a task that waits for a condition. The condition might be a file in cloud storage, a new partition in a table, a successful run of an upstream DAG, or a row appearing in a control table. A sensor turns “a person checks the storage folder” into a step on the rail with a timeout.

Filled Airflow wait path: schedule, sensor, retries, alert
Filled Airflow wait path: schedule, sensor, retries, alert

Here is the classic example, which does not start the load until yesterday’s extract file exists. The imports below use the Airflow 2 paths. In Airflow 3, the file sensor lives in the standard provider package, so newer code imports it from airflow.providers.standard.sensors.filesystem, as the Airflow docs show.

from airflow.sensors.filesystem import FileSensor

wait_for_orders = FileSensor(
    task_id="wait_for_orders_file",
    filepath="/data/incoming/orders_{{ ds }}.csv",
    poke_interval=60,   # seconds between checks
    timeout=60 * 60 * 2,  # fail after 2 hours
    mode="reschedule",  # free the worker between pokes when supported
)

The table below lists the key settings.

KnobMeaningPractical tip
poke_intervalHow often to re-check the conditionToo aggressive hammers storage; too slow delays the rail
timeoutMax wait before the sensor failsMatch it to the business deadline and not to hope
mode=pokeHolds a worker slot while waitingFine for short waits; dangerous at scale
mode=rescheduleReleases the worker between checksPrefer for long waits when available
soft_failSkip downstream instead of failing hardUse rarely; easy to hide missing data

Sensors are not free. Hundreds of long sensors in poke mode can starve real work of capacity. Prefer reschedule for waits that last hours, or use event-driven patterns, such as messages or deferrable operators in newer Airflow versions, when your platform supports them. Read the official sensor docs for the version you run. The settings change over time, but the difference between waiting and working stays the same.

Retries vs sensors vs SLAs

People often mix up these three ideas.

  • Sensor: wait until a precondition is true, and then proceed.
  • Retry: after a task fails, try the same task again.
  • Deadline thinking, often called an SLA or service level agreement: if success arrives late, alert a person even if the task eventually finishes.

Many teams handle a late file with the following pattern.

  • A sensor waits for the extract, up to a defined timeout.
  • The load and transform steps use a small retry count for temporary warehouse errors.
  • If the sensor times out, the DAG fails loudly, and it does not load partial or made-up data.
  • Optionally, a separate monitoring check alerts someone if the finished table is still empty at the business deadline.

That last item matters, because a DAG can look like it is running fine while the business deadline has already passed. A successful run of the scheduler is not the same as a successful result for the people who use the data.

A worked example: an orders pipeline with a late file

Here is the business rule. The finished orders table must be ready by 08:00 local time. The extract usually lands by 05:30, although vendor delays sometimes push it to 07:00. After 07:30 waiting is pointless, because finance will switch to its fallback process.

from datetime import datetime, timedelta
from airflow import DAG
from airflow.operators.bash import BashOperator
from airflow.sensors.filesystem import FileSensor

default_args = {
    "owner": "analytics",
    "retries": 1,
    "retry_delay": timedelta(minutes=5),
}

with DAG(
    dag_id="orders_rail_with_sensor",
    default_args=default_args,
    schedule="0 5 * * *",
    start_date=datetime(2026, 1, 1),
    catchup=False,
    tags=["orders", "sensor"],
) as dag:

    wait_for_file = FileSensor(
        task_id="wait_for_orders_file",
        filepath="/data/incoming/orders_{{ ds }}.csv",
        poke_interval=120,
        timeout=60 * 90,  # 90 minutes after the DAG starts
        mode="reschedule",
    )

    load_staging = BashOperator(
        task_id="load_staging",
        bash_command="python /opt/jobs/load_orders.py --date {{ ds }}",
        retries=2,  # warehouse blips only
    )

    run_dbt = BashOperator(
        task_id="run_dbt",
        bash_command="cd /opt/dbt && dbt build --select tag:orders",
        retries=1,
    )

    quality_gate = BashOperator(
        task_id="quality_gate",
        bash_command="python /opt/jobs/check_orders_counts.py --date {{ ds }}",
        retries=0,  # a bad count is not transient
    )

    notify = BashOperator(
        task_id="notify_success",
        bash_command="echo 'orders ready {{ ds }}'",
    )

    wait_for_file >> load_staging >> run_dbt >> quality_gate >> notify

Notice how the responsibilities are split on purpose.

  • The sensor owns the lateness of the vendor file.
  • The load step gets more retries than the quality gate.
  • The quality gate fails closed, which means it stops the whole chain of steps, because bad counts should not be retried until they look right.
Result panel: sensor timed out at 90 minutes for missing orders file, downstream tasks skipped, alert sent; contrast success path when file arrived at minute 12
Result panel: sensor timed out at 90 minutes for missing orders file, downstream tasks skipped, alert sent; contrast …

That timeout panel is a feature and not a flaw. A clear red message that the file never arrived costs less than a green load of empty data that poisons dashboards until lunch.

Sensor anti-patterns

  • Infinite patience. With no timeout, you teach the vendor that being late is fine, and you burn capacity.
  • Soft fail everywhere. Downstream skips look healthy in the interface while the finished tables go stale.
  • Poke mode for waits of several hours across dozens of DAGs, with no capacity planning.
  • Sensing the wrong thing. The file exists but has zero bytes or is yesterday’s copy, so pair sensors with size checks or content checks.
  • Using a sensor as business logic. Complex rules belong in a small script (a short program file) or query and not in a tangle of nested sensors.

When not to use Airflow

Airflow shines at batch workflows that have many steps, touch many systems and need dependencies, history and people who operate them. It is a poor default for every small automation idea.

SituationPreferWhy not Airflow first
One SQL transform on a scheduledbt Cloud schedule, warehouse tasks, or a single orchestrated jobDAG overhead for one station
True streaming / sub-minute eventsStream processors, queues, CDC toolsAirflow is batch-oriented at heart
Ad hoc analyst notebooksNotebooks, scheduled notebook services carefully scopedNot a substitute for exploration UX
Simple app cron inside one serviceThe app’s own scheduler or platform cronCross-system orchestration not needed
CI for code testsGitHub Actions / GitLab CIDifferent lifecycle and secrets model
You have no operators on callManaged simpler tools until ownership existsAirflow without owners becomes a museum of failed tasks

Also pause if the real problem is about definitions. If two teams disagree on what one row of revenue means, a more sophisticated DAG will only deliver the wrong number faster. Fix the agreements and the tests first, and the review habits from the dbt lab carry over. A scheduler cannot create agreement about what the words mean.

Lighter alternatives to know about

Depending on the stack, teams also use Prefect, Dagster, or the schedulers built into cloud platforms. Examples are Google Cloud’s managed Airflow service, AWS (Amazon Web Services) Step Functions and Glue workflows, and Azure Data Factory. Some teams use the schedules built into their extract, load and transform (ELT) tools. The questions you ask stay the same, which are how many dependencies you have, how easily you can see what is happening, how skilled your team is, and whether you need a general graph of steps or a narrow managed job.

For many analytics teams the winning setup is boring. It combines a managed extract tool, dbt for transformations, a thin scheduler such as Airflow to connect the tools, and tests in the warehouse. Do not adopt Airflow just to feel like a big company. Adopt it when the chain of data steps is real and runs again and again.

Operational checklist for retries and sensors

  • Document for each task which failures are temporary and which are permanent.
  • Set retries only where a second try can succeed without doing damage.
  • Make loads safe to run twice, or safe to merge, before you raise retry counts.
  • Give every sensor a timeout that matches a business deadline.
  • Prefer reschedule or deferrable waits for conditions that take a long time.
  • Treat a sensor timeout as a vendor or upstream incident, and do not report it as “Airflow is broken.”.
  • Review the number and mode of sensors in capacity planning, and not only in reviews of DAG changes.

Common mistakes

  • Retries on bad data. Wrong logic fails the same way every time, so fix the model.
  • No timeout on waits. Tasks that wait forever hide outages.
  • Human sensors. If someone must click rerun every day, build the wait into the DAG or escalate to the vendor.
  • Alert fatigue. Storms of retries and soft fails teach teams to mute their channels.
  • Airflow for one SQL (the standard language for asking a database for data) file. Too much tooling leaves you owning platform pain with no value in return.
  • Ignoring idempotency. A double load after a retry creates duplicate rows, and the uniqueness test fires in the morning.
  • Skipping quality checks because the sensor already showed that the file exists.
  • Confusing green tasks with on-time data. Track business deadlines separately when you need to.

Quick recap

  • Retries help with temporary failures on tasks that are safe to repeat, and they do not fix wrong logic or missing files.
  • Sensors build waiting into the DAG with check intervals, timeouts and modes, and you should prefer waits that are gentle on capacity.
  • Treat late upstream data with a sensor, flaky execution with a retry and bad counts by failing closed.
  • Match timeouts to business deadlines, because a green run that finishes after the deadline can still be a miss for the business.
  • Skip Airflow when you need streaming, only a single scheduled step or automatic code tests, or when nobody owns the tool.
  • Together with the earlier post on the shape of a first DAG, you now have a starting path: first the shape of the chain of steps, then waiting and retries that survive problems, with honest limits.

How to practice this week

  • Pick one flaky job and sort last month’s failures into temporary glitches, late upstream data and logic bugs.
  • For each group, decide whether it needs a retry, a sensor or a code fix, and write that decision in the description of the change, as in the dbt lab.
  • Add or tighten one sensor timeout so it matches a real stakeholder deadline.
  • Remove one unjustified retry from a task that always fails for the same reason.
  • List two automations on your team that should not move to Airflow, and say out loud why.

If AI tools draft sensors and retry settings for you, still verify the timeouts, the modes and the idempotency by hand. Generated scheduling code fails the way generated SQL does, because it looks confident and plausible and is sometimes disastrous. The verification mindset in how to check AI-written SQL applies to DAG reviews too.

Series notes

This is Part 2 of Airflow hello DAG. The previous post covered the mental model of a pipeline as a ticket rail and the shape of a first DAG.

Sources

Written by

Jose S

Founder & Lead Analyst · Analytics Made Simple

Hands-on data strategist, analytics engineering lead, and educator. Writing practical, no-fluff guides to help everyday teams, analysts, and engineers master SQL, AI systems, and modern data architectures.

Keep going

Same lessons in your feed

Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.

Google Search Prefer our practical guides in Google Search & Top Stories: