Picture a nightly dbt run with 60 models. Model 41 fails at 3am: a transient BigQuery error, an upstream hiccup, whatever. Airflow does what Airflow does, it retries the task. And the pipeline rebuilds all 60 models from scratch, including the 40 that had already succeeded.

On BigQuery that is not just slow, it is billed. Every replayed model burns slots again, and the SLA slips while dbt rebuilds tables that were perfectly fine.

The interesting part is that nothing here is Airflow’s fault. Airflow knows when to retry. It has no idea what to retry, and it was never meant to.

dbt already knows what failed

dbt writes two artifacts on every invocation:

  • manifest.json: the compiled project, every model, test and dependency edge.
  • run_results.json: the status of each node in the last run, success, error or skipped.

Feed them back with state selection and dbt will re-run the failures, and only them:

dbt run \
  --select result:error+ \
  --state /artifacts/last_run

The + suffix matters. When a model fails, its children are skipped rather than run, so they still owe you an execution. result:error+ picks up the failure and everything downstream of it.

The catch: pods forget

With KubernetesPodOperator, every attempt is a fresh pod. The artifacts of attempt 1 die with its container, so attempt 2 opens on an empty directory and has no run_results.json to select against. State selection without state selects nothing useful.

The fix is unglamorous: write the artifacts to GCS at the end of every run, under a path that belongs to that specific DAG run, and read them back at the start of the next attempt.

attempt 1, pod A dbt run, fails at model 41 of 60 attempt 2, pod B dbt run --select result:error+ gs://state/DAG_ID/RUN_ID manifest.json run_results.json upload, even on failure download at startup
The state outlives the pod, so the second attempt knows what the first one broke.

All of it fits in the image entrypoint:

#!/usr/bin/env bash
# entrypoint.sh, runs inside the dbt pod
set -uo pipefail

STATE_URI="gs://my-dbt-state/${DAG_ID}/${RUN_ID}"

if gsutil -q stat "${STATE_URI}/run_results.json"; then
  # a previous attempt left state behind, so replay only what it broke
  mkdir -p /artifacts/last_run
  gsutil cp "${STATE_URI}/manifest.json" "${STATE_URI}/run_results.json" /artifacts/last_run/
  dbt run --select result:error+ --state /artifacts/last_run
else
  dbt run
fi
rc=$?

# the next attempt reads these, so upload them on failure too
gsutil cp target/manifest.json target/run_results.json "${STATE_URI}/" || true
exit $rc

The Airflow side stays a plain KubernetesPodOperator with retries set. All the intelligence lives in the entrypoint, which keeps the two responsibilities where they belong: Airflow decides when to retry, dbt decides what to re-run.

The details that bite

Scope the state to the run, not to the DAG. Key the GCS path on the DAG alone and tonight’s run will happily “retry” against last night’s failures. The RUN_ID in the path is what guarantees each scheduled run starts clean.

Upload on failure, not only on success. The whole mechanism is driven by the run_results.json of the run that failed. An if-success guard around the upload quietly disables the feature. Hence the || true: the upload must never mask the real exit code either.

Keep the same selector universe on both paths. If the full run is dbt run --select staging, the retry has to be --select staging,result:error+. The comma is an intersection. Without it, state selection can resurrect nodes that belong to another task.

Know what you are comparing against. --state compares to the manifest of the failed run. If the project changed between two attempts, which is rare at 3am and common at 3pm, state:modified semantics can surprise you. For a retry you want the manifest of that same run, which is exactly what this gives you.

What it buys

The arithmetic is simple, and it gets better the later the failure happens. A failure at model 55 of 60 replays 5 models instead of 60. A failure at model 3 replays almost everything, and the pattern buys you nothing that night.

I think the more useful effect is the second-order one. Once a retry costs a handful of models instead of a full rebuild, raising retries on a flaky external dependency stops being a conversation about the BigQuery bill. Retries become cheap enough to be generous with, and a class of 3am pages turns into a line in the next morning’s logs.

That is the quiet goal of most data platform work. Not more dashboards, fewer pages.

The pattern is implemented and tested in airflow-cloud-data-platform, alongside a few other Airflow patterns.