You're on Datafold: Prove Every Change Before Merge, and Let Data Workers Run Production After It
Already on Datafold? It's where your team proves a change is safe before merge. Data Workers does the operations work in production after merge, behind approvals, and cites Datafold's diff as the evidence.
Your team doesn't merge a dbt change on faith. Every pull request gets a data diff: Datafold compares the PR's models value by value against production and posts the result on the PR, with column-level lineage to the Looker or Power BI assets downstream. AI Code Reviews, an optional add-on to Datafold CI, comment on the SQL. Metric monitors watch row counts and freshness with ML anomaly detection, and data diff monitors reconcile your source database against the warehouse. The Datafold Migration Agent translates legacy code and proves parity, and Datafold's MCP server lets Claude Code or Cursor run a diff from the editor.
Datafold is where your team proves a change is safe before merge. Data Workers does the operations work in production after merge, behind approvals. Production keeps asking new questions after the merge, and when the fix needs proof, Datafold's diff is the evidence the approval cites.
Key takeaways
- •Datafold keeps its job. Diffs in CI, AI Code Reviews, monitors, lineage and the Migration Agent stay put.
- •Data Workers picks up after the merge. When a Datafold monitor fails, the on-call brings the result to Data Workers, with Datafold's MCP server in the same client; Data Workers traces the cause past the dbt project and opens one incident with the blast radius.
- •Datafold's diff becomes the evidence. Data Workers proposes the fix as a diff for the owner to merge; the Datafold diff on that pull request shows exactly what it changes.
- •Production work runs behind approvals. After a named owner approves and reruns the dbt platform job, Data Workers re-checks the model and keeps a receipt.
- •Migrations get an operator too. Datafold proves parity; Data Workers plans the waves with their parity checks, tracks them and holds the completion gate for the owner's sign-off.
- •Start with a pilot. One dbt domain already under Datafold CI, read-only first, on the ladder from L0 manual to L4 autonomous.
Datafold is the proof before merge. Data Workers is the operator after it.
Datafold holds the evidence about a change before merge. Data Workers holds what happens once it is live: one context across the source database, the replication task, the lakehouse, dbt and the dashboards, one approval flow and one audit trail, with Spellbook Data Catalog (in preview) as the place your team reviews it. Data Workers connects to Datafold over its API or MCP server today, and natively to PostgreSQL (Aurora included), Databricks with Unity Catalog and the dbt platform.
A Monday night at an online marketplace: Aurora PostgreSQL for orders, AWS DMS replicating change data to S3, Auto Loader into Databricks, dbt platform jobs, Datafold CI on GitHub and Omni for finance's dashboards. Times are UTC. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| Fri 16:40 | GitHub + Datafold | PR #1182 changes int_orders_latest to keep each order's newest version by updated_at. Datafold's CI diff shows 214 expected status corrections. The PR merges |
| Mon 21:40 | AWS DMS | The platform team restarts the orders replication task with new endpoint settings. updated_at now lands in whole seconds |
| 21:41 on | Aurora PostgreSQL | Checkout writes pending then paid within the same second for most new orders. In the lakehouse, both versions now share one updated_at |
| Tue 02:00 | dbt platform on Databricks | The nightly job builds fct_orders. With ties, the dedupe keeps either version. Every dbt test passes |
| 03:05 | Datafold | A metric monitor flags Monday's paid orders 19% below range and alerts #data-alerts in Slack |
| 03:08 | Data Workers | The on-call's assistant reads the monitor run over Datafold's MCP server and hands it to Data Workers, which opens an incident |
| 03:21 | Data Workers | The on-call's assistant hands it the latest run of the Datafold data diff monitor that reconciles Aurora orders against fct_orders: 1,876 orders are paid in Aurora and pending in the lakehouse. The query Data Workers proposes for the bronze table, run by the owner, finds two versions tied on updated_at for 61% of orders after 21:40, none before |
| 03:30 | Data Workers | It lists the blast radius from lineage (fct_orders, fct_daily_gmv, the Omni "Daily GMV" dashboard the team recorded in the context graph) and proposes a tiebreaker diff on the DMS change sequence, a rerun of prod_orders_nightly after merge, and a note to the DMS task's owner. The request reaches the model's owner in Slack |
| 07:40 | GitHub + Datafold | The owner opens a PR from the proposed diff. Datafold's CI diff: 1,876 orders change from pending to paid, nothing else |
| 07:52 | Spellbook | The owner approves the plan with the Datafold diff linked as evidence, then merges |
| 07:55 | dbt platform | The owner reruns prod_orders_nightly after approval; it succeeds at 08:31 |
| 08:36 | Datafold | The owner re-runs the data diff monitor: Aurora and fct_orders match at 9,870 paid orders |
| 08:40 | Data Workers | It re-runs its quality checks on fct_orders and writes the receipt |
| 09:00 | Omni | Finance opens Daily GMV on Monday's real number. The DMS owner restores microsecond precision that afternoon |

Datafold did its part exactly. The diff on PR #1182 was right for the data that existed then. The monitor caught the drop hours before anyone read the dashboard, and Datafold's source-to-lakehouse diff measured it to the order. The cause sat in a replication setting changed three days after the merge, in a system Datafold doesn't run. Without the trace, finance would have opened on a 19% GMV drop and a debate about reverting a PR that was never wrong.
| Job | What Datafold does | What Data Workers does |
|---|---|---|
| Before merge | Diffs PR #1182 against production | Keeps the PR's lineage and blast radius in its context |
| The signal | Metric monitor flags paid orders 19% low | Picks up the run and opens one incident |
| The cause | Data diff monitor shows 1,876 orders out of step with Aurora | Traces it past dbt to the DMS restart and the tied versions in bronze |
| The fix | Waits for a pull request to diff | Writes the tiebreaker diff, rerun plan and note to the DMS owner |
| The proof | Diffs the fix PR, then confirms parity with Aurora | Cites both diffs in the approval, re-checks the model after the owner's rerun |
| The record | Keeps the diffs and monitor history | Keeps a receipt with cause, diff, approver, run and checks |
Why doesn't Datafold just do this itself?
Because Datafold is built to produce proof, and it scopes its agents to that job with care: specialized agents for migrations, optimization and code reviews, plus a context layer for coding agents. The Migration Agent translates code and validates parity, supervised by Datafold's engineers and delivered with a guaranteed price and timeline. The MCP server lets an agent run diffs, query connected sources, read monitor results and provision monitors from YAML, with each tool shown only when the API key's owner holds that permission. Every action lands on a Datafold object.
The incident needed different actions: read a replication task Datafold doesn't run, write a code fix, rerun a production job on a schedule Datafold doesn't control, and keep a record an auditor can read. Taking on liability for changes in your replication, orchestrator and lakehouse would pull a validation vendor away from the job it does well, as the neutral referee. Writing to production across systems, with blast-radius scoping, approvals, an undo path and receipts, is its own product category. That's the product Data Workers is.
Every tool owns a slice. Data Workers covers the whole lifecycle
Datafold owns two slices: data quality proof through diffs, and migration parity. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the testing tool you already run.

| Stage | Data Workers | Datafold | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 5 | Column-level lineage reaches BI tools, and the Data Knowledge Graph (private beta) serves context over MCP. Data Workers keeps one governed context across every platform. |
| Analytics & Insights | 8 | 3 | The Datafold Assistant answers questions about sources, diffs and monitors in Slack. Data Workers answers from approved definitions and checks the number behind the answer. |
| Data Quality | 8 | 9 | Datafold's home stage: value-level diffs in and across databases, on every pull request. Data Workers finds the cause of a production break and re-checks after the fix. |
| Observability & Incidents | 8.5 | 6 | Metric, data diff and schema change monitors raise the alert. Data Workers traces the cause across systems and runs the repair behind approval. |
| Pipelines & Ingestion | 8.5 | 4 | Datafold tests pipeline changes in CI and leaves running them to the orchestrator. Data Workers queues approved runs through the orchestrators and owns the outcome. |
| Schema & Migration | 8 | 9 | Datafold's second home stage: the Migration Agent translates code and proves parity with diffs. Data Workers detects source schema changes and plans migrations in approved waves. |
| Governance & Access | 8.5 | 3 | Roles, custom groups and service accounts govern Datafold and its MCP tools. Data Workers dry-runs proposed grants and routes them to a named approver. |
| Security & Privacy | 8 | 4 | VPC deployment and governed LLM inference keep data in your perimeter. Data Workers flags new sensitive column names in pull request review and proposes masking for the owner to apply. |
| Cost / FinOps | 8 | 4 | Slim Diff keeps CI cost down, and Datafold names an optimization agent. Data Workers attributes Snowflake credits to the dbt model behind them and drafts the fix. |
| MLOps & Models | 7.5 | 2 | Datafold diffs the tables models read and leaves models to MLOps tools. Data Workers keeps the data under the models healthy. |
How Datafold and Data Workers work together
People keep working in Claude Code or Cursor, GitHub, Datafold and Omni. Spellbook is where the data team reviews each proposed repair with its cause, blast radius, the Datafold diff it cites, the approver and the undo path; approval requests reach the approver in Slack or email. Underneath, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm does the work with 20+ specialist agents, and the Autonomous Data-Conductor runs each repair: detect, diagnose, fix, review, verify, remember.

What comes in. Datafold connects over its REST API or MCP server today: your team's assistant reads monitor runs and diff results side by side with Data Workers, whose own checks rest on the warehouse and the dbt manifest. Context Wizard joins them to Aurora tables, the dbt manifest, through native PostgreSQL and dbt connectors. AWS DMS and Omni connect over their APIs today. Your agent can call explain_table, trace_cross_platform_lineage and blast_radius_analysis for context and reach, get_quality_score and run_quality_check for the data, diagnose_incident and get_incident_history for the incident, and assess_impact when the cause is a source change.
What goes back. Nothing writes into Datafold from Data Workers; diffs and monitors stay your team's. Model fixes are proposed as a diff for the owner to merge, or opened as a pull request when your team turns on the GitHub pull-request target, and Datafold CI diffs that PR like any other. After approval, the owner reruns the dbt platform job and Data Workers re-checks the model. Changes to systems it doesn't run, such as the DMS task, go to their owner as a written proposal.
Setup over MCP today. Keep the Datafold MCP server and add Data Workers beside it. For Datafold, create an API key (a service account in a custom group narrows the tools). For Data Workers, clone the open-source repository and add start-agent.sh entries, as the client setup docs show.
// Example: .mcp.json for Claude Code
{
"mcpServers": {
"datafold": {
"type": "http",
"url": "https://app.datafold.com/mcp/",
"headers": { "Authorization": "Key YOUR_DATAFOLD_API_KEY" }
},
"dw-context-catalog": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-context-catalog"]
},
"dw-incidents": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-incidents"]
},
"dw-quality": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-quality"]
}
}
}List the tools with /mcp. Then "why did paid orders drop overnight?" gets Datafold's diffs and Data Workers' cause and blast radius in one answer.
Migrations, side by side. If the Datafold Migration Agent is moving you off a legacy warehouse, keep it on translation and parity proof; its final report links the diffs for every dataset. Data Workers never runs those data diffs itself. It works the operating side: maps dependencies, plans each wave with its parity checks, tracks which have passed and holds the completion gate until the owner signs off, with Datafold's diffs as the evidence on file.
One request, L0 to L4. The autonomy ladder is set per domain.

- •L0 manual. Connected, not acting. On-call works the Datafold alert by hand.
- •L1 observe. Every alert arrives with cause, blast radius and a plan; nothing changes, and the log records what each action would have needed.
- •L2 propose. Data Workers drafts the diff, rerun plan and notes; a named owner approves in Spellbook first, with the Datafold diff attached.
- •L3 act reversibly. For proven classes, such as re-checking a model after an approved fix and rerun, Data Workers acts on its own, re-checks and keeps the receipt.
- •L4 autonomous. For a scoped, trusted class in one domain, it repairs and re-checks on its own and posts the receipt.
An unanswered request expires and escalates; it never auto-grants. No agent can promote its own work, and an org-wide stop halts all autonomous dispatch. More in autonomy levels L0 to L4, is it safe to let AI agents change production data, how approvals work and who owns the agents. The agents run in your infrastructure and hold your warehouse credentials and model key; your data stays in your systems, and the hosted Conductor sees workflow metadata only (where does our data go).
Teams on dbt, SQLMesh and Elementary, Data Workers integrations, the dbt integration guide, plus Data Workers vs Datafold, the best data diff tools for pull requests and how to review data changes in dbt pull requests.
What changes for your team

Teams on Datafold already spend less time on bad merges. The time goes to the 3 a.m. alert, the hunt through replication logs, the rerun and the message to finance. With Data Workers on top, those jobs run on autopilot at the level you set: incidents arrive traced and planned, production breaks become fixes proven by a Datafold diff, Snowflake credits are tied to the dbt model behind them, grant requests get a dry run first, every repair carries its audit record, and migration waves are tracked to the owner's sign-off. The overnight archaeology becomes a review of a plan that is already written, with the proof attached.
Keep Datafold, or consolidate?
Keep Datafold if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
Most teams keep Datafold: value-level diffs on every pull request are a habit worth keeping, and a migration under contract is no time to change referees. What teams consolidate is the tooling around it: the runbook for when a monitor fires, the second alert channel, the spreadsheet of who owns which model. Thinking about building the after-merge layer yourself on the Datafold MCP server and a coding agent? Read build it ourselves with Claude Code and MCP servers: the connection is the easy part; cross-system context, approvals and rollback are the work.
The case for your CFO
The outcome: the testing investment the company already made keeps paying off after the merge. Datafold stops bad changes before they ship; Data Workers repairs what production breaks, the same night, so the numbers finance reads each morning are complete and right. Above, that is the difference between a standup opening on a 19% GMV drop and one opening on the real number.
The risk story is plain. Every repair shows its blast radius, goes to a named owner, carries the Datafold diff as evidence, runs through the dbt platform like any other job and leaves a receipt: trigger, cause, diff, approver, checks and undo path. Autonomy is set per domain from L0 manual to L4 autonomous. Zero migration: Datafold, dbt, Databricks and your source databases stay where they are.
Why now: every tool in the stack ships an agent that acts on its own objects, and costly incidents cross tools. One record of who approved what is cheaper to set up before agents change production on their own. The first win is one domain where every Datafold alert arrives explained and planned. What stays the same: your CI, diffs, monitors and Datafold contract. For the numbers, see the ROI of agentic data operations. Start with a pilot (pricing); the pilot is credited in full against the first year.
The sentence to repeat upstairs: "Datafold proves every change before we ship it; Data Workers fixes what production breaks after, with an owner approving every fix and Datafold's diff as the proof."
Getting started
Start with a pilot. Pick one dbt domain already under Datafold CI with a monitor on its key tables, connect Data Workers beside your Datafold MCP server with a read-only key, and let Data Workers explain every alert before you turn on the first approved rerun. Plans are on the pricing page, and the pilot is credited in full against the first year.
FAQ
Does Data Workers have a native Datafold connector? Datafold connects over its REST API or MCP server today: your team's assistant reads monitor runs and diff results with a scoped key, side by side with Data Workers. Its native connectors cover the systems around it: PostgreSQL, Databricks and Unity Catalog, Snowflake, BigQuery, dbt and the dbt platform, Airflow, Dagster and Prefect.
Does Data Workers replace Datafold's data diff? No. Datafold's value-level diff is the proof, and Data Workers cites it in approvals. Data Workers runs the work after merge: tracing the cause, proposing the fix, recording the approved rerun and re-checking the model.
If Datafold's CI diff was green, how did the incident happen? A diff proves a change against the data that exists when it runs. Production keeps changing after the merge; Datafold's monitors catch many of those changes, and Data Workers takes the alert to a verified fix.
Can Data Workers run Datafold diffs or change our monitors? No, by design. Engineers run diffs and provision monitors through the Datafold MCP server under their own permissions, and Datafold CI diffs the pull request your owner opens from a Data Workers fix. Data Workers reads the results and cites them.
How do Data Workers and the Datafold Migration Agent split a migration? The Migration Agent translates the code and proves parity with diffs. Data Workers plans each wave with its parity checks, tracks them and holds the completion gate until the owner signs off.
We use Datafold AI Code Reviews. Does Data Workers review PRs too? Yes. Data Workers reviews pull requests with a blast radius across every connected system and can block risky changes. Teams keep Datafold's code review and value-level diff, and add the cross-system reach.
Sources
- •Datafold, homepage ("Automate Data Engineering"; specialized agents for migrations, optimization and code reviews; Data Knowledge Graph, Beta; anomaly detection; VPC deployment; governed LLM inference), https://www.datafold.com/ (checked Oct 3, 2026)
- •Datafold docs, Welcome (data engineering automation platform; Data Knowledge Graph in private beta), https://docs.datafold.com/welcome.md (checked Oct 3, 2026)
- •Datafold docs, MCP (remote server at app.datafold.com/mcp/, API key, permission-scoped tools), https://docs.datafold.com/datafold-mcp.md (checked Oct 3, 2026)
- •Datafold docs, MCP Tool Permissions (diff, query, monitor and knowledge graph tools; provision_monitors; trigger_monitor_run), https://docs.datafold.com/security/mcp-tool-permissions.md (checked Oct 3, 2026)
- •Datafold docs, How Datafold in CI Works (diff results posted on the PR), https://docs.datafold.com/deployment-testing/how-it-works.md (checked Oct 3, 2026)
- •Datafold docs, AI Code Reviews (optional CI add-on, enabled through Datafold support; review supervisor; PR summary and inline comments), https://docs.datafold.com/deployment-testing/ai-code-reviews.md (checked Oct 3, 2026)
- •Datafold blog, Introducing AI in CI (January 29, 2025), https://www.datafold.com/blog/introducing-ai-in-ci (checked Oct 3, 2026)
- •Datafold docs, Monitor Types (data diff, metric, data test, schema change), https://docs.datafold.com/data-monitoring/monitor-types.md (checked Oct 3, 2026)
- •Datafold docs, Lineage (column-level and tabular lineage), https://docs.datafold.com/data-explorer/lineage.md, and Looker and Power BI integrations, https://docs.datafold.com/integrations/bi-data-apps/looker.md, https://docs.datafold.com/integrations/bi-data-apps/power-bi.md (checked Oct 3, 2026)
- •Datafold docs, Datafold Migration Agent (translation, parity diffs, supervised delivery), https://docs.datafold.com/data-migration-automation/datafold-migration-agent.md (checked Oct 3, 2026)
- •Datafold blog, The Datafold Migration Agent 2.0 (September 29, 2025), https://www.datafold.com/blog/the-datafold-migration-agent-2-0 (checked Oct 3, 2026)
- •Datafold blog, Datafold joins Databricks delivery provider program (April 13, 2026), https://www.datafold.com/blog/datafold-joins-databricks-delivery-provider-program (checked Oct 3, 2026)
- •Datafold docs, Slack Bot (Datafold Assistant over MCP tools), https://docs.datafold.com/integrations/agents/slack-bot.md (checked Oct 3, 2026)
- •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)
- •Data Workers open-source repository, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)