Your Pipelines Run on Airflow. Now Let Data Workers Own What They Produce: A Guide for Data Leaders
A guide for VPs and heads of data platform whose company runs on Apache Airflow, Astro, Managed Airflow or MWAA: what they cover, and how Data Workers diagnoses failed runs, proposes fixes and verifies approved reruns.
Airflow runs the steps on time. Data Workers owns whether the numbers they produce are right, and fixes the cause when they aren't: it diagnoses the failed DAG run, proposes the fix to its owner as a diff, drafts the rerun for approval and checks the table afterwards. This guide is for the VP of data or head of data platform whose company runs on Airflow, self-managed or on a managed service.
Your platform team already lives in Airflow, where hundreds or thousands of DAGs schedule everything the business reads. Some run on Apache Airflow 3, current at 3.3.2 since September 17. Some run on Astronomer's Astro, where Otto writes and investigates DAGs and Astro Observe tracks data products against their SLAs. Some run on Google Cloud's Managed Service for Apache Airflow (formerly Cloud Composer), where the Managed Airflow Agent diagnoses failed tasks, or on Amazon MWAA, which has offered Airflow 3.3.1 since September 1.
Then the estate grows: twelve hundred DAGs, eight teams, one rotation paged for every failed task. Most failures start outside the DAG: a renamed source column, a late sync, a retyped field. A green DAG run can still load an empty partition, and nobody hears until a dashboard looks wrong. Meanwhile Airflow 2 reached end of life on April 22, and the move to Airflow 3 is on your roadmap or halfway done.
Your team has become the glue, carrying context by hand from the task log to the source, the dbt model and the dashboard. Airflow made running pipelines cheap. The expensive part is operating them: the triage, the fixing, the rerunning and the record of what changed.
What Airflow solved, and what it didn't
Airflow solved running. It turned pipelines into reviewed Python, with dependencies, retries and a record of every DAG run and task instance. Airflow 3 made that record firmer with DAG versioning, deadline alerts, human-in-the-loop tasks and the stable /api/v2 REST API. The managed services add AI aimed at the deployment they run. Otto, available across Astro since September 30, authors DAGs, diagnoses failures and reviews code, and its investigations (in Labs) suggest up to three ranked fixes, such as a code patch or clearing a run, for a person to apply. The Managed Airflow Agent, launched on August 25, diagnoses failed tasks and DAG runs on Google Cloud and recommends the steps.
What Airflow doesn't own is the work that starts and ends outside the DAG. Its model ends at the task boundary: a task succeeds, fails, retries or waits. A source team retypes a column, a task quietly loads nothing, a finance number is wrong at 8 a.m. Tracing that across the source, the warehouse, dbt and BI, getting the owner to approve a fix, rerunning the right DAG, checking the table and keeping a receipt is not a job Airflow was built for. That is the right focus for a scheduler that runs ten thousand tasks a day without caring why a Postgres column changed.
Airflow runs the steps on time. Data Workers is the agentic data platform that owns whether their output is right, and runs the rest of your data back office.
Or, in one sentence for your team: Airflow tells us what ran and what failed; Data Workers finds out why, takes the fix through the owner's approval and proves the data is right again.
What goes on autopilot
Each of these is a queue your team works by hand today, behind signals Airflow already records.

None of this requires replacing anything. Data Workers connects to Airflow natively through its REST API, detects Airflow 2 or Airflow 3, and works the same on self-managed Airflow, Astro and Managed Airflow; on MWAA it connects through the environment's Airflow REST API today. It reads DAG runs and task instances. Fixes are proposed as diffs the owner merges through your normal review. Clears, pauses and DAG deploys stay with your team, by design. When a rerun is approved, on Airflow 2 Data Workers queues the run through Airflow's REST API; on Airflow 3 the owner triggers the approved run. Airflow schedules and records it like any other, and Data Workers checks the result.
What changes for your organization
Access and spend follow the same pattern: the change is drafted for its owner, approved and recorded.

- •On-call starts from a cause. The engineer paged at 3 a.m. opens a failed task already traced to its cause, with its downstream reach and a proposed fix. The rotation reviews instead of digging.
- •Every failure arrives with its owner. The approval request goes to the named owner recorded in Data Workers' context graph, so nobody has to claim it in a channel first. An unanswered request expires and escalates.
- •Green means right. A DAG run that succeeds on an empty partition is caught on the table it builds and handled like any other incident.
- •Reruns and backfills become deliberate. After approval, the one rerun the fix needs goes through Airflow, instead of a whole DAG cleared to be safe. A backfill is drafted with its undo written down before it starts; the owner triggers it, holds that undo and runs it if it's ever needed, and Data Workers re-checks the quality assertions afterwards.
- •Audits get easier. Airflow records what ran. Data Workers records why it ran again, who approved it, what it touched and how to undo it.
One quarter on a 1,200-DAG estate
This is an illustration, not a customer case.
- •A company runs 1,200 DAGs across two estates: an older self-managed Airflow 2 cluster and a newer one on MWAA, feeding Snowflake, dbt and Tableau for finance, product and marketing.
- •The shared rotation is paged on most nights. Most causes start outside the DAG, and several green runs load bad data that reaches a board-deck number before anyone notices.
- •The Airflow 3 move stalls. SLAs become deadline alerts, catchup is off by default and removed context variables break templates. MWAA upgrades minor versions in place, but the major-version move means a new environment. Every migrated DAG is a risk nobody has time to watch.
- •At the quarterly review, the CFO asks why on-call keeps growing and which numbers were wrong for a morning. The answer takes a week of task logs and Slack threads.
With Data Workers, each failed or quietly wrong run arrives traced to its cause, and the fix goes to its owner as a diff. After approval, the one rerun the fix needs goes through Airflow, queued by Data Workers on the Airflow 2 cluster and triggered by the owner on Airflow 3, and Data Workers checks the table it builds. During the Airflow 3 move, the same connection reads both estates, and each moved DAG's tables keep their freshness, volume and quality checks, so a changed behaviour shows up as an incident on its next run, traced to the DAG that moved. At the quarterly review, the receipts answer: which runs broke and why, who approved each fix, and which moved DAGs raised an incident.
You choose how far and how fast
You choose the altitude, one domain at a time.

L0 manual is the work your team does by hand today. Every domain starts at L1 observe: each failed DAG run gets a diagnosis and a blast radius, and nothing changes. At L2 propose, agents draft the fix and the rerun, and a named owner approves. At L3 act reversibly, change classes with a proven record, such as rerunning a DAG after an approved upstream fix, run on their own with a receipt (on Airflow 3 the owner still presses run). At L4 autonomous, a domain that has earned it runs end to end, and any domain can be dialled back. An unanswered approval request expires and escalates, never auto-grants, and no agent can promote its own work.
The ladder in detail is in autonomy levels L0 to L4 explained, and how approvals work covers who signs off on what.
Why doesn't Airflow just do this itself?
Focus and risk. Airflow earns trust by running every task correctly and recording what happened: the task is the unit, retries and deadlines are the tools, and a person decides what to clear. Otto and the Managed Airflow Agent keep that scope, suggesting fixes inside their own deployment for a person to apply. The incident that matters to a leader usually starts in a source system and ends in a finance number. Fixing it means scoping what else a change touches across systems the orchestrator doesn't run, getting the right owner to approve, queuing the right rerun, proving the number and owning the record. That is a different product with a different liability, and the product we built.
How it fits with Airflow
- •Airflow keeps its job. Schedules, retries, deadline alerts, human-in-the-loop tasks, DAG versions, clears and backfill controls stay in Airflow.
- •Your managed service keeps its agent. Otto and the Managed Airflow Agent keep authoring and investigating DAGs. Data Workers takes the incidents whose cause sits in the source, dbt or the warehouse, across every estate.
- •Your review stays the gate. DAG and model changes are diffs your reviewers merge, behind your CI and branch protection.
- •Your permissions stay the lock. Data Workers starts with a viewer-role token in Airflow, and gets a token allowed to trigger runs, limited to named DAGs where your auth manager supports it, only once you open writes for a domain.
- •Your data stays in your systems. The agents run in your infrastructure and hold the warehouse credentials, and the hosted coordination service sees workflow metadata only. Our security and deployment guide has the detail.
When you don't need Data Workers
- •Your DAGs rarely fail, and when they do, the cause is almost always inside the DAG.
- •You run one small Airflow estate and one team owns every DAG.
- •Nobody upstairs is asking what on-call costs or which numbers were wrong last quarter.
If all three are true, keep your budget.
The case for your CFO
The outcome. The company already pays for Airflow and a platform team to run it. Data Workers makes that investment hold up where the business notices: failed runs are fixed at the cause, green runs that load bad data are caught, and each rerun is checked. The on-call rotation stops scaling with the DAG count, and platform engineers' hours go to the platform and the Airflow 3 move instead of night triage.
The risk story. Autonomy is set per domain, and nothing changes at L2 until a named person approves. Fixes go through your reviewers and CI; reruns go through Airflow, which records them. Every change carries a receipt: the cause, the diff, who approved it, what it touched downstream, how it was verified and how to undo it. An org-wide stop halts all autonomous dispatch. Our safety guide goes deeper, and who owns the agents covers accountability.
Why now. Airflow 2 is past end of life, so every estate still on it has the Airflow 3 move ahead, and Astro and Managed Airflow each ship their own agent in their own console. The choice is one approval flow and one record across every estate, or a different answer per deployment and each engineer's own wiring: see build it ourselves with Claude Code and MCP servers.
The first win. Read-only: every failed DAG run in one domain, such as the finance pipelines, arrives with a diagnosis, its owner and its blast radius.
What stays the same. Airflow, your managed service, your DAG repository, reviewers, on-call rotation and warehouse permissions. The path is a pilot on your own DAGs, and the ROI guide shows how to size it.
The sentence to repeat upstairs: Airflow runs our pipelines on time; Data Workers makes sure what they produce is right, fixes the cause when it isn't, and keeps a receipt for every change.
Where to start
Pick one domain: the finance DAGs behind the board deck, the DAGs fed by one source system, or the DAGs moving to Airflow 3 next.
Start with a pilot. A forward-deployed engineer connects Data Workers to your Airflow estates, warehouse, dbt project and repositories, and runs the first domain with your team, read-only first. See pricing for how the pilot works.
For the detail, send your platform engineers you're on Airflow, Data Workers + Airflow for the wiring, you're on Astronomer, you're on Managed Service for Apache Airflow, the dbt leaders guide and automated Airflow task failure root cause analysis. They cover the four products: the Autonomous Data-Conductor picks up Airflow's failed runs and runs each fix from detection to a verified result, the Data-Agents Swarm does the work, Data Context Wizard keeps one governed graph of definitions, lineage and owners across every platform, and Spellbook Data Catalog (in preview) is where your team approves, audits and rolls back. The thinking behind them is in our thesis.
Sources
- •Apache Airflow, releases on GitHub (3.3.2, Sep 17, 2026; 3.3.1, Aug 12, 2026), checked Oct 3, 2026: https://github.com/apache/airflow/releases
- •Apache Airflow blog, Airflow 3.3.0 (Jul 6, 2026), checked Oct 2, 2026: https://airflow.apache.org/blog/airflow-3.3.0/
- •Apache Airflow Docs, Supported versions (Airflow 2 EOL Apr 22, 2026), checked Oct 3, 2026: https://airflow.apache.org/docs/apache-airflow/stable/installation/supported-versions.html
- •Apache Airflow Docs, Upgrading to Airflow 3 (SLAs replaced by deadline alerts, catchup off by default, removed context variables,
/api/v2), checked Oct 3, 2026: https://airflow.apache.org/docs/apache-airflow/stable/installation/upgrading_to_airflow3.html - •Apache Airflow Docs, REST API security, checked Oct 2, 2026: https://airflow.apache.org/docs/apache-airflow/stable/security/api.html
- •Amazon MWAA User Guide, Apache Airflow versions (v3.3.1 available Sep 1, 2026; minor-version upgrades, no major-version upgrade), checked Oct 3, 2026: https://docs.aws.amazon.com/mwaa/latest/userguide/airflow-versions.html
- •Amazon MWAA User Guide, Using the Apache Airflow REST API (InvokeRestApi or web session token;
/api/v2on Airflow 3), checked Oct 3, 2026: https://docs.aws.amazon.com/mwaa/latest/userguide/access-mwaa-apache-airflow-rest-api.html - •Google Cloud, Managed Service for Apache Airflow release notes (rename from Cloud Composer Apr 15, 2026; Managed Airflow Agent Aug 25, 2026; Airflow 3.3.1 in Gen 3 Aug 20, 2026), checked Oct 3, 2026: https://docs.cloud.google.com/composer/docs/release-notes
- •Google Cloud, Troubleshoot tasks and DAG runs with the Managed Airflow Agent, checked Oct 3, 2026: https://docs.cloud.google.com/composer/docs/composer-3/troubleshoot-tasks-and-dag-runs-with-agent
- •Astronomer, Astro release notes (Otto across Astro, Astro IDE GA and ranked suggested fixes, Sep 30, 2026), checked Oct 3, 2026: https://www.astronomer.io/docs/astro/release-notes
- •Astronomer Docs, Investigate Dag failures with Otto (Labs; up to three suggested fixes), checked Oct 3, 2026: https://www.astronomer.io/docs/astro/otto/otto-investigate
- •Astronomer, Astro Observe product page (lineage, data product SLAs, alerts), checked Oct 3, 2026: https://www.astronomer.io/product/observe/
- •Astronomer blog, Astro Observe GA (Feb 13, 2025), checked Oct 3, 2026: https://www.astronomer.io/blog/observe-ga/