Data Workers for data engineers
What Data Workers changes in a data engineer's week: pipeline failures arrive with a root cause and a proposed fix, backfills run through a playbook with a recorded undo, source schema changes come with rollback SQL, while you keep the pipelines and every approval.
Data Workers takes the triage and the cleanup out of a data engineer's week: a failed DAG run arrives with its root cause and a proposed fix, a backfill runs through a playbook that records its own undo, and a source team's schema change comes with a migration and rollback SQL for you to apply. You keep owning the pipelines, the approvals and the autonomy settings: each agent action runs at the level your team sets, waits for a named approver wherever that level asks for one, and leaves a receipt.
Data Workers is the agentic data platform. For a data engineer, it works from the coding agent you already use, from the page that wakes you up, and from Spellbook, its catalog and approvals app.
Key takeaways
- •On-call starts with a diagnosis.
diagnose_incidentandget_root_causetrace a failed load to the change that caused it; the page carries the blast radius and a proposed fix. - •Backfills with a recorded undo. The backfill playbook finds the gap, starts the backfill as a task in your orchestrator, re-checks the data against your quality assertions, and records the step that removes the backfilled rows before it starts.
- •Schema changes caught in review. The dbt manifest diff, pull request review and
check_compatibilityflag a breaking change;generate_migrationandrollback_migrationwrite the forward and rollback SQL for you to apply. - •Fixes arrive as diffs. You review and merge them; reruns are queued through Airflow, Dagster, Prefect or ADF.
- •You keep the job. Pipelines, design, approvals and autonomy levels stay yours, set per domain from L0 manual to L4 autonomous.
Your week today
The planned work is new sources, a Spark job that needs tuning, and the migration everyone agreed to finish this quarter. The unplanned work is most of the week:
- •The 2am page. A DAG run fails, PagerDuty fires, and the first hour goes to finding out whether the cause is the source, the load, the warehouse or the code.
- •Schema changes from source teams. A service team ships a Postgres migration on Tuesday night. You find out when the Wednesday load breaks.
- •Backfills. Every fix ends with one, hand-written, row counts checked by hand, and no clean way back if the window was wrong.
- •Alerts and tickets. Monte Carlo, Great Expectations and the orchestrator each alert on the same incident, and analysts ask "is this table fresh?"
The coding agent already helps with the code. Stack Overflow's 2025 Developer Survey (third-party; 49,009 responses, published July 2025) found that "51% of professional developers use AI tools daily," yet "more developers actively distrust the accuracy of AI tools (46%) than trust it (33%)," and on AI agents "87% of all respondents agree they are concerned about the accuracy." dbt Labs' 2026 State of Analytics Engineering report (third-party; 363 responses, Dec 5, 2025 to Feb 1, 2026) adds: "While 72% prioritize AI-assisted coding, only 24% prioritize AI-assisted pipeline management, including testing, observability, and quality controls." Writing pipelines got faster. Running them did not.
The same week with Data Workers
Data Workers runs the operations around your pipelines. It reads your orchestrator, warehouse, sources, dbt project and quality tools, keeps that context in one governed graph with provenance, and hands you proposals with evidence.

What it takes off your week:
- •Incident triage.
monitor_metricsandget_anomalieswatch the pipelines;diagnose_incidentandget_root_causetrace a failure to its cause, andget_incident_historyshows whether it happened before.blast_radius_analysislists every model, SLA and dashboard downstream.remediateapplies a known playbook at the level you set, with a dry run first if you ask; a novel incident goes to a person. - •Backfills. The playbook identifies the gap, prepares the backfill and starts it as a task in your orchestrator, then re-checks the data against your quality assertions. Its undo step, removing the backfilled records, is recorded before the run. If the checks stay red, the incident comes back to you with the way back ready. How to roll back an AI agent change walks one through.
- •Schema changes. A change is caught in the dbt manifest diff, in pull request review, or when
run_quality_checkfails where it lands;check_compatibilityandassess_impacton the schema agent say whether it breaks consumers.generate_migrationwrites the migration with its rollback SQL, androllback_migrationkeeps the way back ready. - •Fixes and reruns. A load-mapping change or a staging-model fix comes back as a diff for you to merge. Data Workers opens a pull request itself only for dbt write-back, when your team turns on the GitHub pull-request target. After you approve, it queues the rerun through Airflow, Dagster, Prefect or ADF and reads the run state back; dbt Cloud jobs rerun on your scheduler, with the approval on the receipt.
- •Freshness, migrations and cost.
run_quality_checkchecks nulls, keys and row counts, and the lateness baselines your team records withmonitor_metricswork against the SLAs you set withset_sla. The migration agent plans each wave with its parity checks and holds the completion gate for your sign-off. The cost agent attributes Snowflake credits to the query and the dbt model behind them through query tags, and drafts the fix for the model's owner. - •Tickets.
explain_tableanswers "what is this table, where does it come from, who owns it" with sources.create_jira_sm_ticketopens the follow-up for the source team, linking the receipt in Spellbook.
What you still own and decide:
- •The pipelines. Their design, code, schedules and every merge. Nothing executes a schema migration for you, and ingestion syncs stay in Fivetran, Airbyte or your own code.
- •The approvals. They go to a named person; an unanswered request expires and escalates, never auto-grants, and no agent can promote its own work. See how approvals work.
- •The autonomy settings. Each domain sits on L0 manual, L1 observe, L2 propose, L3 act reversibly or L4 autonomous; Data Workers ships observe-only (autonomy levels, explained). The org-wide stop halts all autonomous dispatch.
- •The receipt. Every change records the diff, who approved it, when, the blast radius and the rollback path, in a hash-chained audit log you can check with
verify_global_hash_chain.
How Data Workers fits the tools you already use
Your coding agent. Data Workers agents are MCP servers, so they sit inside Claude Code, Cursor, Codex or Copilot. From the session where you are fixing the DAG, ask "why did load_payments fail and what else reads raw.payments?" and the answer comes from the context graph, with sources. Setup follows the client setup docs: clone the open-source core, install it, and add a start-agent.sh entry per agent to your client's MCP config.
# Example: register the incident, schema and quality agents in Claude Code
git clone https://github.com/DataWorkersProject/dataworkers-claw-community.git
cd dataworkers-claw-community && npm install --ignore-optional
claude mcp add --scope user dw-incidents -- "$(pwd)/start-agent.sh" dw-incidents
claude mcp add --scope user dw-schema -- "$(pwd)/start-agent.sh" dw-schema
claude mcp add --scope user dw-quality -- "$(pwd)/start-agent.sh" dw-qualityThe core runs on a built-in sample estate; in a Data Workers deployment the same agents point at your estate and carry the approvals and receipts. List the tools with your client's own command, such as /mcp. Data Workers with your coding agents covers each client.
Orchestration. Airflow, Dagster, Prefect, ADF, Step Functions, dbt Cloud and Managed Service for Apache Airflow (formerly Cloud Composer) connect natively; Data Workers queues reruns through Airflow, Dagster, Prefect and ADF, and dbt Cloud jobs rerun on your scheduler after approval. See you're on Airflow and you're on Dagster.
Warehouses, sources and streams. Snowflake, Databricks, BigQuery, PostgreSQL, Kafka, Kafka Connect and Schema Registry connect natively, among 50+ connectors (Data Workers on Snowflake). Iceberg REST catalogs, Glue, Fivetran, Airbyte, dlt, Debezium and Redshift connect over their APIs or MCP servers today; Data Workers watches the schemas they land and leaves the syncs to them (you're on Fivetran, integrations).
Quality and paging. Monte Carlo, Great Expectations, Soda and Datadog feed signals in; PagerDuty carries the page with the diagnosis attached (you're on PagerDuty). Spellbook Data Catalog (in preview) holds the proposals, approvals and receipts, and lets analysts check freshness and lineage before they open a ticket.
The metrics you are judged on
| Metric | How Data Workers moves it | Where you see it |
|---|---|---|
| Pipeline uptime and freshness | Source changes flagged with their blast radius; freshness checked against your SLAs | SLA breaches per pipeline over time |
| MTTR | Root cause, blast radius and a fix waiting when the page arrives | Time from page to approved fix, from the receipts |
| Number of incidents | Repeats matched to history; the cause filed with the source team | Incident history per pipeline |
| Delivery of new sources | Fewer interrupts; generate_pipeline and validate_pipeline draft and check the scaffolding | Sources shipped per quarter |
| Cost | Snowflake credits attributed to each dbt model; fixes drafted for the owner | Spend per dbt model; the design target is 25 to 40% lower warehouse spend |
Take a baseline before the pilot. How to measure AI data agents sets out the method, and the ROI calculator turns your incident and ticket numbers into a range.
A worked example: a source migration at 23:40 and a 07:00 revenue SLA
This is an illustration, not a customer case. A payments service writes to PostgreSQL; an Airflow DAG, load_payments, loads it nightly into Snowflake; dbt models it; finance reads daily revenue in Tableau at 07:00. The payments domain runs at L2 propose.
| Time | System | What happened |
|---|---|---|
| 23:40 | PostgreSQL | The payments team ships a migration: amount moves from integer cents to numeric, currency becomes currency_code |
| 01:14 | Snowflake | The nightly load_payments merge into raw.payments fails on a type mismatch |
| 01:15 | Data Workers | The migration's merged pull request shows the change; get_root_cause points to the 23:40 migration |
| 01:17 | Data Workers | Blast radius: four dbt models, the 07:00 revenue SLA and one Tableau workbook |
| 01:19 | Data Workers | Proposes a migration for raw.payments with rollback SQL, a load-mapping diff and a backfill plan for the failed window |
| 01:20 | PagerDuty | The on-call engineer is paged with the diagnosis and the diffs |
| 01:34 | Data engineer | Reviews in Claude Code, applies the migration, merges the mapping diff, approves the backfill in Spellbook |
| 01:38 | Data Workers, Airflow | Records the undo (remove the backfilled rows); queues the backfill and the rerun through Airflow |
| 02:05 | Data Workers | Re-checks raw.payments against its quality assertions: all pass |
| 02:07 | Jira SM | Opens a ticket for the payments team, linking the receipt |
| 07:00 | Tableau | Daily revenue reads complete, on time |

Without Data Workers, the on-call engineer starts with a stack trace and works back to a migration in another team's repo. With it, the engineer starts with the cause, reviews three proposals, and is back asleep before the backfill finishes. The receipt holds the diffs, approvals, check results and the undo.
How to bring it to your team
- •Try it on the sample estate. Register the incident, schema and quality agents and ask the questions you answer on call.
- •Pick one pipeline that pages you. A nightly load with frequent upstream changes makes the before and after easy to count.
- •Start at L1 observe. Connect the orchestrator and a read-only warehouse role. For two weeks Data Workers logs what it would have flagged and proposed; you review that record and its receipts against what actually happened.
- •Move that domain to L2 propose, then to L3 act reversibly for the playbooks you trust, such as a backfill with a recorded undo.
- •Bring the record to the retro: pages with a cause attached, time to an approved fix, schema breaks caught. How to get your team to trust AI agents has the playbook.
When your lead asks about risk, send is it safe to let AI agents change production data and where your data goes: the agents run in your infrastructure and hold your credentials, your data stays in your systems, and the hosted Conductor sees workflow metadata only. Build it ourselves answers "why not Claude Code and a few MCP servers", and the ROI of agentic data operations makes the case in money. Then start a pilot (what a pilot looks like; $7,500 one-time, credited in full against the first year, on /pricing/). Your colleagues have their own pages: analytics engineers and data platform engineers. Our guides on on-call for data engineers and the cost of data engineering toil cover the wider picture.
FAQ
Will the agents hallucinate on our data? They work from the context graph: lineage, schemas, run history and approved definitions, each with provenance. At L2 propose every change reaches you as a diff before anything runs. How Data Workers avoids hallucinations on your data covers the mechanism.
Will it replace me? It takes the toil (triage, backfills, schema chases, tickets) and leaves the design, the pipelines and the approvals with you. Will Data Workers replace my data team? answers this directly.
Can it run a backfill on its own? At L2 propose it waits for your approval. At L3 act reversibly the backfill playbook can run within the domain you allow, because its undo is recorded before it starts and failed checks send the incident back to you.
Does it run DDL in production? No. Schema migrations come with rollback SQL for you to apply. Fixes come as diffs you merge, and reruns are queued through your orchestrator.
I already use Claude Code or Cursor. Why add this? You keep them. Data Workers runs as MCP servers inside them, adding the context, root cause, playbooks, approvals and receipts a coding agent does not hold on its own.
Sources
- •Stack Overflow, 2025 Developer Survey, AI section (third-party; 49,009 responses from 177 countries, published July 2025; "51% of professional developers use AI tools daily", 46% distrust vs 33% trust accuracy, 87% concerned about agent accuracy): https://survey.stackoverflow.co/2025/ai (checked Oct 2, 2026)
- •dbt Labs, 2026 State of Analytics Engineering Report (third-party; 363 responses, Dec 5, 2025 to Feb 1, 2026; 72% AI-assisted coding vs 24% AI-assisted pipeline management): https://www.getdbt.com/resources/state-of-analytics-engineering-2026 (checked Oct 2, 2026)
- •Google Cloud, Cloud Composer release notes (Cloud Composer "is evolving to become Managed Service for Apache Airflow"): https://cloud.google.com/composer/docs/release-notes (checked Oct 2, 2026)
- •Data Workers client setup documentation: https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)
- •Data Workers public repository tools (
diagnose_incident,get_root_cause,remediate,monitor_metrics,get_incident_history,detect_schema_change,check_compatibility,assess_impact,generate_migration,rollback_migration,blast_radius_analysis,explain_table,check_freshness,run_quality_check,get_anomalies,set_sla,generate_pipeline,validate_pipeline,create_jira_sm_ticket,verify_global_hash_chain): https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026) - •Data Workers product repository,
data-workers-agent-swarmmain @ 0c2491e3: agents/dw-incidents remediation playbook registry (backfill playbook steps, rollback step "remove backfilled records") and remediate tool (operation recorded for rollback up front, post-fix re-validation against quality assertions, escalation to a person), agents/dw-migration (wave planning, completion gate), agents/dw-cost (attribution, unused data with dependency check), agents/dw-connectors (Jira SM ticket creation), core/enterprise (approvals, promotion guard, audit log) (checked Oct 2, 2026) - •Data Workers product pages: https://dataworkers.io/product/autonomous-data-conductor/ , https://dataworkers.io/product/data-agents-swarm/ , https://dataworkers.io/product/spellbook-data-catalog/ (checked Oct 2, 2026)
- •Data Workers security and pricing: https://dataworkers.io/security/ , https://dataworkers.io/pricing/ (checked Oct 2, 2026)