You're on Airflow: Keep Running the Schedule, and Let Data Workers Run the Operations When a DAG Run Goes Wrong
Airflow schedules and runs the work. Data Workers catches a green DAG run with wrong data, finds the cause, drafts the approved rerun and verifies the result.
Your team lives in Airflow. Hundreds of DAGs sit in a Git repository, deployed to Airflow 3 (3.3.2 since Sept 17, 2026) on Astronomer's Astro, on Google Cloud's Managed Service for Apache Airflow (formerly Cloud Composer), on Amazon MWAA or on your own Kubernetes cluster. Producer DAGs update assets, and consumer DAGs scheduled on those assets run the moment their inputs say they're ready. Deferrable sensors wait on the triggerer, and the Grid view holds every DAG run. Airflow is where your team schedules and runs the work. Data Workers is the operator when a run goes wrong: it diagnoses the cause, proposes the fix as a diff, drafts the rerun and verifies the result, all behind your approvals.
The run that hurts most is rarely the red one. A failed task pages someone. The expensive run goes green on the wrong data: a sensor that released early, a file that was half there, an asset event that told the next DAG to start. Someone still has to notice, trace, fix and rerun it before the business opens.
Key takeaways
- •Airflow keeps the schedule. DAGs, assets, sensors, retries and the run record stay in Airflow and your repository. Data Workers works through Airflow's REST API.
- •A green run gets checked against the data. Volume, freshness and balance checks on the tables your DAGs build catch a successful run on half the data overnight.
- •The rerun is planned, approved and checked. Data Workers drafts the run (which DAG, which
conf, in what order) and holds the approval. On Airflow 3 the owner triggers it; on Airflow 2 Data Workers queues it through the REST API. Airflow schedules, runs and records it like any other. - •Fixes arrive as diffs. A sensor or DAG change goes to its owner to merge, with the blast radius attached.
- •One loop across every Airflow you run. Self-managed, Astro, Managed Airflow or MWAA: one approval flow, one audit trail.
Airflow is the schedule your team runs on. Data Workers is the operator when a run goes wrong.
Airflow's contract ends at task state, and that is the right contract for a scheduler. Its asset documentation says so plainly: "Airflow marks an asset as updated only if the task completes successfully." Whether the file that task loaded was complete is a question about the data.
Here is a month-end night on Airflow 3.3.1, which Amazon MWAA has offered since Sept 1. This is an illustration, not a customer case. The producer DAG gl_landing waits for the September general ledger export from NetSuite with a deferrable S3KeySensor matching gl_20260930_*.csv, loads it into a Unity Catalog table in Databricks and updates the asset gl_entries. The consumer DAG month_end_close is scheduled on that asset and builds the close models with dbt, run as an Astronomer Cosmos task group. On Sept 29 the ERP admin changed the export to split large files into parts.
| Time (Thu Oct 1, ET) | System | What happens |
|---|---|---|
| Wed 23:50 | NetSuite | The month-end GL export starts; for the first time it writes two files |
| 00:02 | S3 | gl_20260930_part0001.csv lands |
| 00:03 | Airflow (gl_landing) | The deferrable sensor's trigger fires on the first matching file; the sensor succeeds |
| 00:09 | Airflow + Databricks | The load task writes the Sept 30 partition of finance.bronze.gl_entries and succeeds; Airflow updates the asset gl_entries |
| 00:10 | Airflow (month_end_close) | The consumer DAG run starts on the asset event |
| 00:34 | S3 | gl_20260930_part0002.csv lands; no DAG is waiting for it |
| 00:41 | dbt on Databricks | stg_gl_entries, int_gl_by_entity and fct_trial_balance build; every not_null, unique and relationships test passes; the DAG run is green |
| 00:45 | Data Workers | The volume check on the bronze partition reads 58% of the trailing six month-ends; freshness is green. Data Workers opens an incident |
| 00:52 | Data Workers | It reads both DAG runs and their task instances through the Airflow REST API, follows lineage from the Tableau "Close: Trial Balance" workbook back to the load task, and queries the partition in Databricks: every row comes from part 1, and posting dates stop at Sept 19 |
| 01:05 | Data Workers + Slack | It names the cause (a split export and a sensor that releases on the first match), maps the blast radius (three dbt models, two Tableau workbooks, the 09:00 close review) and proposes two changes: a new run of gl_landing for Sept 30, and a sensor diff that waits for the export's done-marker file. The approval request goes to the finance data on-call in Slack |
| 06:02 | Spellbook | The on-call reviews the plan in Spellbook and approves the rerun; the sensor diff waits for the DAG owner |
| 06:03 | Airflow | The on-call triggers the gl_landing run Data Workers drafted, with {"export_date": "2026-09-30"}. The load task replaces that date's partition, as it always does, and Airflow updates the asset |
| 06:48 | Airflow (month_end_close) | The asset event starts the close DAG again; dbt builds green |
| 06:55 | Data Workers | It verifies: row count matches both export files, debits equal credits for every entity in fct_trial_balance, freshness inside the SLA. The receipt lands in Spellbook |
| 07:00 | Tableau | The scheduled extract refresh reads the verified tables; the controller opens the right trial balance at 09:00 |
| 10:30 | GitHub | The DAG owner reviews and merges the sensor diff; next month's export waits for its done marker |

Every component did its job, and dbt's tests passed on the rows they were given. The night needed an operator that checks the data behind a green run, puts the exact rerun in front of a person who can say yes, and checks the result. On this Airflow 3 estate the on-call presses run; on an Airflow 2 estate, Data Workers queues the approved run itself.
| Job | What Airflow does | What Data Workers does |
|---|---|---|
| Waiting for data | Deferrable sensors poll on the triggerer and release when their condition is met | Checks the data the release let through: volume, freshness, balance against history |
| Starting the next step | Updates the asset when the producer task succeeds and schedules every consumer DAG | Reads the producer and consumer DAG runs and their task instances, and ties each task to the tables it builds |
| Noticing a problem | Retries, marks upstream_failed, fires deadline alerts and callbacks | Opens an incident when a green run's output is wrong, with the cause traced across systems |
| Fixing it | Runs whatever code is in the DAG | Proposes the sensor or DAG change as a diff for the owner to merge, with its blast radius |
| Running it again | Schedules and executes the new DAG run, and keeps its record | Drafts the run for the owner to trigger on Airflow 3, or queues it through the REST API on Airflow 2, with the undo written into the plan |
| Proving it | Shows the run as successful | Verifies the downstream tables and leaves a receipt: cause, approver, run, checks, undo |
Why doesn't Airflow just do this itself?
Because Airflow is built to schedule and run work reliably at scale, and a scheduler that judged every file's completeness would be slower and tied to formats it can't know. So Airflow offers checks as tasks, check_fn hooks, human-in-the-loop tasks and the common.ai provider for LLM steps, and keeps its contract clean: task state in, schedule out.
The managed Airflow vendors have added agents, each sensibly scoped to its platform. Since Sept 30, 2026, Astronomer's Otto is available across Astro; Otto investigations (in Labs) look into DAG failures "using proprietary Airflow, Astro, and Observe context" and suggest up to three ranked fixes, such as a code patch, clearing a task or triggering a new run, for a person to apply. Google's Managed Airflow Agent, in the Cloud Console since Aug 25, 2026, helps diagnose failed tasks and DAG runs, and Google's Managed Airflow remote MCP server went GA the same day. Both start from a failure inside the deployment. A green run on half a file never shows up as a failure.
Fixing it needs the data side (the Databricks partition, dbt models, Tableau workbook), the blast radius, a named approver, a recorded rerun, an undo, checks afterwards and an auditable receipt. That cross-system operations product, with approvals and rollback built in, is what Data Workers, the agentic data platform, is built to run.
Every tool owns a slice. Data Workers covers the whole lifecycle
Airflow owns its slice deeply: scheduling and running the work across thousands of DAGs. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the Airflow already running your schedule.

| Stage | Data Workers | Airflow | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 3 | Assets and the DAG graph say which task updates what, and OpenLineage can export it. Data Workers keeps one governed context graph of tables, models, owners and lineage across every system. |
| Analytics & Insights | 8 | 1 | Not Airflow's job: the UI reports on runs and tasks, not on the business numbers they build. Data Workers answers data questions from governed definitions with lineage behind every number. |
| Data Quality | 8 | 3 | Checks run as tasks you write, and an asset counts as updated when its task succeeds. Data Workers holds the tables the DAGs build to baselines the team records and drafts the owner's checks. |
| Observability & Incidents | 8.5 | 5 | Retries, deadline alerts, callbacks and the record of every DAG run and task instance. Data Workers diagnoses why a run went wrong across systems, proposes the fix behind an approval and verifies the result. |
| Pipelines & Ingestion | 8.5 | 9 | Airflow's home stage: schedules, assets, sensors, deferrable operators, retries and backfills across thousands of DAGs. Data Workers drafts approved reruns, checks them and keeps what they load right. |
| Schema & Migration | 8 | 2 | Not Airflow's job: a sensor sees files and a task sees success, not a changed export or schema. Data Workers catches schema and shape changes and assesses blast radius. |
| Governance & Access | 8.5 | 2 | Auth managers and roles control who can trigger, clear or edit DAGs. Data Workers routes every data change to a named approver with a receipt. |
| Security & Privacy | 8 | 2 | Connections, secrets backends and role-based access protect Airflow itself. Data Workers' pull request review flags new columns whose names or annotations look sensitive before the DAGs move them. |
| Cost / FinOps | 8 | 2 | Not Airflow's job: worker and triggerer capacity, not the warehouse bill a rerun creates. Data Workers traces Snowflake spend to the dbt model behind it and drafts the fix for its owner. |
| MLOps & Models | 7.5 | 4 | Training and scoring pipelines are often DAGs, so Airflow schedules the steps. Data Workers keeps the data under your models healthy. |
If your monitoring lives in an observability tool, Data Workers vs data observability covers how its signals feed the same loop. For the dbt side of a close like this one, read you're on dbt; teams on another orchestrator can compare you're on Prefect.
How Airflow and Data Workers work together
Engineers ask from their coding agent; approvers decide in Spellbook Data Catalog (in preview), where people see each proposal, approver, run and undo. Underneath, Data Context Wizard keeps one governed context graph across Airflow, Databricks, dbt and Tableau, the Data-Agents Swarm does the work, and the Autonomous Data-Conductor runs each incident end to end.

What Data Workers reads and triggers in Airflow. The native Airflow connector talks to the Airflow REST API, detecting /api/v2 on Airflow 3 or /api/v1 on Airflow 2. It reads DAG runs and their state, and the task instances in each run. Task-to-table lineage comes from the dbt manifest and the team's context-graph notes. On Airflow 2, the one thing it writes is a new DAG run, with an optional conf payload, queued through trigger_airflow_dag after a named approval; Airflow acknowledges it as queued, then schedules, runs and records it. On Airflow 3, the owner triggers the approved run from the plan Data Workers drafted, and Data Workers verifies the result. Clearing task instances, Airflow backfills, pausing DAGs and deploying DAG code stay with your team, by design. Pin the connection read-only and every write is refused. Astro, Managed Airflow and MWAA each expose that REST API per environment, so one pattern covers every Airflow you run. Databricks, dbt and Tableau (read) are native too; NetSuite and S3 connect over their APIs today. The wiring is in the Airflow integration guide.
Backfills, precisely. When the fix is a reload, Data Workers drafts it as a rerun (the owner triggers it on Airflow 3; Data Workers queues it on Airflow 2), reads its status and re-checks the quality assertions afterwards. The undo is written into the plan before the run, and the owner runs it if it's ever needed.
Setup over MCP today. Every Data Workers agent is an MCP server. The client setup guide documents the path: clone the repo and add one start-agent.sh entry per agent to your client.
# Example: Data Workers agents in Claude Code, from a clone of the open-source repo
claude mcp add --scope user dw-incidents -- "$(pwd)/start-agent.sh" dw-incidents
claude mcp add --scope user dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
claude mcp add --scope user dw-quality -- "$(pwd)/start-agent.sh" dw-quality
claude mcp add --scope user dw-connectors -- "$(pwd)/start-agent.sh" dw-connectorsList each agent's tools with your client's own command, such as /mcp. In this story: monitor_metrics and diagnose_incident catch and explain the break; trace_cross_platform_lineage and blast_radius_analysis map what it touches; the monitor_metrics baseline verifies the fix; remediate runs playbooks such as backfill_data; send_slack_alert tells the channel. In this Airflow 3 story the on-call triggers the run; on Airflow 2, trigger_airflow_dag queues it. Astronomer's open-source astro-airflow-mcp can sit in the same client.
Where things run. The agents run in your infrastructure on every tier and hold the Airflow, Databricks and model credentials. Your data stays in your systems; the hosted Conductor sees workflow metadata only. The production remote endpoint serves /mcp and takes an API key or OAuth tokens from your identity provider, such as Okta or Entra ID, verified through JWKS.
One incident, L0 to L4, set per domain:

- •L0 manual. The controller finds the trial balance off at 09:00.
- •L1 observe. The incident opens at 00:45 with cause and blast radius. Nothing changes.
- •L2 propose. Nothing runs until a named person approves; an unanswered request expires and escalates, never auto-grants.
- •L3 act reversibly. For a class with a clean record, such as reloading one date's partition, the plan goes out at 01:05 with the undo written first; Data Workers queues the run itself on Airflow 2, and on Airflow 3 the on-call triggers it.
- •L4 autonomous. A scoped domain handles that class end to end; people read receipts.
More in who owns the agents, how approvals work, is it safe to let AI agents change production data and where does our data go.
What changes for your team
Airflow keeps the schedule, and your DAG authors keep authoring. What changes is who does the overnight operations work: the 6 a.m. page becomes a diagnosed incident with a proposed fix and rerun waiting for one approval, and the audit trail records who said yes.

The Orchestration agent and Incident Debugging agent posts go deeper. For practice notes, see automating Airflow task failure analysis and reducing data on-call burden; Data Workers for Airflow has the overview.
Keep Airflow, or consolidate?
Keep Airflow if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
For nearly every Airflow team the answer is keep it. What teams consolidate is the scaffolding around it: "check the row count" tasks copied into every DAG, a separate freshness monitor, runbooks that say "clear, rerun, check the dashboard". Those move into one loop that ends with a verified Airflow run and a receipt. Weighing a build on a coding agent and MCP servers? Read build it ourselves with Claude Code and MCP servers: wiring the servers is the easy part. On Astro, you're on Astronomer covers how this sits next to Otto.
The case for your CFO
The outcome: the numbers your DAGs build arrive right, on time, more often. When a pipeline goes green on the wrong data, the error is caught overnight, fixed behind one approval and verified before the business reads it.
The risk story is plain. Agents decide nothing you haven't delegated, and autonomy is set per domain from L0 manual to L4 autonomous. At L2 a named person approves every rerun and fix, and no agent can promote its own work. In Airflow, the only write Data Workers makes is a new DAG run on Airflow 2, and only for the domains you open; on Airflow 3 your owner triggers the approved run. Every change carries a receipt with the cause, approver, run, checks and undo, and an org-wide stop halts all autonomous dispatch. Zero migration: Airflow, your DAGs, Databricks, dbt and Tableau stay as they are.
Why now: Airflow 3 puts assets at the center of scheduling, so more DAGs start on another DAG's success with nobody looking in between. The first win is L1 on one critical asset chain such as the close. What stays the same: your schedules, repository, review process, auth manager and on-call rota. For the numbers, see the ROI of agentic data operations.
The sentence to repeat upstairs: "Airflow keeps running our schedule; Data Workers makes sure a green run means right data, and fixes it behind an approval when it doesn't."
Getting started
Start with a pilot. Pick one asset chain that matters, such as the month-end close, connect Data Workers read-only to that Airflow environment, the warehouse and the dbt project, and run at L1 so every suspect run arrives with its cause and blast radius. Then turn on L2 for that domain; on Airflow 2, give Data Workers a token allowed to trigger only the DAGs it covers. Plans are on the pricing page, and the pilot is credited in full against the first year. Data leaders can read the Airflow guide for data leaders.
FAQ
What exactly does Data Workers read from Airflow, and what does it trigger? It reads DAG runs and their state, and the task instances in each run, through the Airflow REST API (/api/v2 or /api/v1, detected automatically). On Airflow 2 it triggers one thing: a new DAG run, with an optional conf payload, after a named approval. On Airflow 3 the owner triggers the approved run from Data Workers' plan, and Data Workers verifies it. Clears, Airflow backfills, pauses and DAG deploys stay with your team.
Our DAG run was green. How does Data Workers know the data is wrong? It checks the tables your DAGs build: volume against history, freshness, and rules such as debits equal credits.
Does it work on Astro, Managed Airflow and MWAA? Yes. Each exposes the Airflow REST API per environment, and teams running more than one Airflow get one approval flow and one audit trail across them.
Will Data Workers change our DAG code? It proposes the change as a diff for the DAG owner to merge, with the blast radius attached, and opens a pull request only when your team turns on the GitHub pull-request target.
Can a rerun make things worse? At L2 every rerun waits for approval, and the plan records the undo before it runs; the owner runs the undo if it's needed. Failed checks escalate to a person.
We already use Otto or the Managed Airflow Agent. Do we need both? Keep them for authoring and investigating DAGs inside their platform. Data Workers works beside them on the same REST API and takes the incidents whose cause or impact sits outside the DAG, including green runs with wrong data.
Sources
- •Apache Airflow releases: 3.3.2 (Sept 17, 2026), 3.3.1 (Aug 12, 2026), 3.3.0 (July 6, 2026), https://github.com/apache/airflow/releases (checked Oct 3, 2026)
- •Apache Airflow 3.3.2 docs, Asset-aware scheduling: asset updated only when the task completes successfully; consumer Dag scheduling, https://airflow.apache.org/docs/apache-airflow/stable/authoring-and-scheduling/asset-scheduling.html (checked Oct 3, 2026)
- •Apache Airflow 3.3.2 docs, Deferrable operators and triggers, https://airflow.apache.org/docs/apache-airflow/stable/authoring-and-scheduling/deferring.html (checked Oct 3, 2026)
- •Apache Airflow Amazon provider, S3KeySensor (
wildcard_match,deferrable,check_fn) and S3KeysUnchangedSensor, https://airflow.apache.org/docs/apache-airflow-providers-amazon/stable/operators/s3/s3.html (checked Oct 3, 2026) - •Apache Airflow common.ai provider, "Route pipeline failures to a fix or a person", https://airflow.apache.org/docs/apache-airflow-providers-common-ai/stable/use_cases/route_pipeline_failures.html (checked Oct 3, 2026)
- •Astronomer, Astro release notes (Sept 30, 2026: new Astro experience live for all Organizations with Otto available across Astro; Otto investigations list up to three ranked suggested fixes, Labs; Astro IDE GA), https://www.astronomer.io/docs/astro/release-notes (checked Oct 3, 2026)
- •Astronomer, Otto investigations (Labs; Airflow, Astro and Observe context; runs in the Astro control plane), https://www.astronomer.io/docs/astro/otto/otto-investigate (checked Oct 3, 2026)
- •Astronomer, Otto overview (labelled Labs), https://www.astronomer.io/docs/astro/otto (checked Oct 3, 2026)
- •Astronomer, astronomer/agents (astro-airflow-mcp,
afCLI, skills), https://github.com/astronomer/agents (checked Oct 3, 2026) - •Google Cloud, Managed Service for Apache Airflow release notes (rename from Cloud Composer Apr 15, 2026; remote MCP server GA and Managed Airflow Agent in Cloud Console Aug 25, 2026; Airflow 3.3.1 builds in Gen 3, Aug 26 and Sept 2, 2026), https://docs.cloud.google.com/composer/docs/release-notes (checked Oct 3, 2026)
- •AWS, Amazon Managed Workflows for Apache Airflow versions (Airflow 3.3.1 available Sept 1, 2026), https://docs.aws.amazon.com/mwaa/latest/userguide/airflow-versions.html (checked Oct 3, 2026)
- •AWS, Using the Apache Airflow REST API on Amazon MWAA, https://docs.aws.amazon.com/mwaa/latest/userguide/access-mwaa-apache-airflow-rest-api.html (checked Oct 3, 2026)
- •Data Workers open-source repository (tool registrations:
trigger_airflow_dag,send_slack_alertin dw-connectors; dw-incidents, dw-context-catalog, dw-quality), https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026) - •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)