You're on dbt: dbt Defines and Tests Your Models. Data Workers Runs the Operations Around Them
dbt is where your team defines and tests transformations. Data Workers diagnoses the incidents around them, proposes fixes as diffs and verifies, behind approvals.
Your analytics engineers think in models, refs, tests and exposures. Every transformation is SQL in a pull request, every model has a unique and not_null test, contracts guard the marts other teams depend on, and a job on the dbt platform (formerly dbt Cloud) builds production every night. Since September 16, 2026, the engine is simply "dbt": the Fusion engine graduated to general availability as dbt v2, and dbt Core v2 is now dbt OSS, the Apache 2.0 distribution. dbt Wizard (public preview) builds models and investigates failed runs inside the project, and Fivetran and dbt Labs have been one company since June 1, 2026. dbt is where your team defines and tests the transformations. Data Workers runs the operations around them: it diagnoses the incident, proposes the fix as a diff, queues the approved run and verifies the numbers. That matters most on the mornings when every test passed and the number is still wrong, because the cause was a late load the project never saw.
Key takeaways
- •dbt keeps its job. Models, tests, contracts, the job scheduler, Wizard and Copilot stay where your analytics engineers use them.
- •Data Workers takes the operations around the models. It watches volume on sources and models, tracks lateness against baselines your team records, diagnoses incidents across systems, and catches the runs that pass with wrong data.
- •Exactly what it touches in dbt. It reads the manifest and catalog, and proposes the exact job runs for the owner to approve and run. Model changes arrive as a diff for the owner to merge; pull requests, for docs changes, only when your team turns on the dbt write-back GitHub target.
- •One approval, one receipt. A named owner approves in Spellbook; the receipt records the cause, the diff, the run, the verification and the undo.
dbt is the transformation layer your team reviews. Data Workers is the operator around it.
dbt runs the code your team reviewed, correctly. Data Workers owns whether the data it produced is complete and right. Here is a weekend at a food-delivery marketplace, an illustration rather than a customer case. Courier app events stream through Amazon Data Firehose into Snowflake. A dbt platform job builds fct_deliveries as a microbatch incremental model (event_time='delivered_at', batch_size='day', the default lookback of 1, on the v1 Latest release track), fct_courier_payouts builds on it, and a dbt exposure, also recorded in the context graph, points at the Sigma workbook finance operations uses to approve the weekly courier payout file on Monday at 10:00.
| Time | System | What happens |
|---|---|---|
| Sat 13:50 | Snowflake, Firehose | A scheduled key rotation replaces the key on the Firehose service user; deliveries start failing, and Firehose writes the records it cannot deliver to its S3 error bucket for manual backfill |
| Sun 02:00 | dbt platform | The nightly job builds Saturday's batch from the rows that arrived before 13:50. Every test passes, and source freshness (warn after 24 hours) passes too |
| Sun 02:30 | Data Workers | monitor_metrics flags raw.courier_events as stalled since 13:50 and Saturday's volume at 58% of the trailing four Saturdays, points at the Firehose stream as the table's only writer, and alerts the platform on-call in Slack |
| Mon 01:15 | S3, Snowflake | The platform engineer updates the key and reloads 36 hours of events from the S3 error bucket with COPY INTO. The raw table is complete |
| Mon 02:00 | dbt platform | The job runs. With lookback of 1, microbatch rebuilds Sunday's and Monday's batches; Saturday's batch is not reprocessed. Tests pass and the job is green |
| Mon 02:40 | Data Workers | Compares daily volumes: raw Saturday is now at 100% of the trailing Saturdays, fct_deliveries Saturday is still at 58%. It traces the gap to the lookback window and maps the blast radius through the manifest and the context graph: fct_courier_payouts and the Sigma workbook, 21,340 deliveries and 3,118 couriers short |
| Mon 02:55 | Spellbook | Data Workers proposes the repair to the payouts model owner: a backfill of Saturday for fct_deliveries+ with --event-time-start and --event-time-end, a diff raising lookback to 3, and the Time Travel version of both tables as the undo |
| Mon 07:20 | Spellbook | The owner reviews the evidence and approves, and sets the dates on the team's backfill job |
| Mon 07:25 | dbt platform | The owner runs the approved backfill job |
| Mon 08:30 | Snowflake | Data Workers verifies the model's daily volumes for the last seven days against the raw table and writes the receipt |
| Mon 09:10 | GitHub | The analytics engineer merges the lookback change after CI passes |
| Mon 10:00 | Sigma | Finance operations approves a payout file with every Saturday delivery in it |

Nothing in dbt misbehaved. Microbatch does what its docs say: a standard run processes batches by the current timestamp and the lookback, and older late data waits for an explicit backfill with --event-time-start and --event-time-end. The tests checked the rows that were there. Data Workers compared the source with the model after the upstream fix and handed a named owner a precise backfill with its undo before the payout file was approved.
| Job | What dbt does | What Data Workers does |
|---|---|---|
| Defining the logic | Models, refs, macros and contracts, reviewed in pull requests | Reads the manifest as part of one context graph that also covers ingestion, the warehouse and BI |
| Testing | Data tests, unit tests and source freshness on every build | Adds volume checks with run_quality_check and freshness baselines with monitor_metrics on sources and models |
| Running | The job scheduler builds production on schedule, with dbt State skipping unchanged work | Proposes the exact runs and backfills; the owner runs them after approval, and Data Workers verifies |
| Investigating | Wizard investigates job and run failures inside the project | Diagnoses incidents across systems, including green runs that built wrong data, with diagnose_incident and trace_cross_platform_lineage |
| Impact | Exposures and lineage show who reads a model | Maps the blast radius from the source to every model and recorded reader with blast_radius_analysis; pull-request review's lineage diff reads exposures |
| Fixing | Your engineers write and merge the change | Proposes the change as a diff for the owner to merge, with the run that applies it |
| The record | Run history and the Git log | A receipt: cause, diff, approver, run, verification and undo, in Spellbook and the audit trail |
Why doesn't dbt just do this itself?
Because dbt is designed to be a compiler and build system for transformation logic, and that focus is why teams trust it. A dbt run builds the code and inputs it is given, deterministically. It does not guess that a raw table received rows for a day it already built; guessing would make builds slower, costlier and less predictable. Microbatch makes the trade explicit: the lookback is a cost setting your team chooses, and late data outside it is a backfill decision left to a person. That is the right design for a build tool.
dbt's AI follows the same scope. dbt Wizard builds and refactors models, writes YAML and investigates job and run failures, with "Ask for approval" as its default mode: "You approve each file change before it is persisted." dbt Copilot generates SQL, docs, tests and semantic models. The self-hosted dbt MCP server can run dbt commands, and its README warns that doing so "could modify your data models, sources, and warehouse objects"; its Admin API tools trigger, cancel and retry job runs. Each acts on dbt's own objects, and each is good at it. On Monday at 02:40 there was no failed run to investigate.
The weekend's fix crossed four systems and needed a backfill scoped to one day, an undo ready first, a named approver and a record finance could read. That is a different product, with liability for changes beyond the project. Data Workers is built for it, on top of dbt.
Every tool owns a slice. Data Workers covers the whole lifecycle
dbt owns transformation, and nobody does it better. Each point tool around it adds another console, another contract and another handoff: a warehouse monitor, an ingestion alert, a wiki of backfill scripts. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and keeps dbt at the center of it.

| Stage | Data Workers | dbt | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 6 | dbt Catalog, docs and model-level lineage describe the project well. Data Workers keeps one governed context graph across dbt, the warehouse, ingestion and BI, with owners and usage. |
| Analytics & Insights | 8 | 5 | The Semantic Layer and dbt Charts (public beta) serve governed metrics. Data Workers answers data questions from governed definitions with lineage behind every number. |
| Data Quality | 8 | 7 | Tests, unit tests and contracts run on every build, and Copilot and Wizard write more of them. Data Workers adds volume checks and recorded lateness baselines that span sources, models and consumers. |
| Observability & Incidents | 8.5 | 4 | Job notifications and run history tell you a run failed. Data Workers diagnoses the incident across systems, including runs that passed with wrong data. |
| Pipelines & Ingestion | 8.5 | 9 | dbt's home stage: models, incremental strategies, dbt State and the job scheduler. Data Workers plans approved runs and backfills for the owner and proposes fixes as diffs. |
| Schema & Migration | 8 | 6 | Contracts, model versions and CI catch breaking changes inside the project. Data Workers catches schema changes at the source first and plans migrations in waves. |
| Governance & Access | 8.5 | 4 | Model access and groups govern who can ref a model. Data Workers drafts each warehouse access request as a scoped grant with an expiry for a named owner to approve. |
| Security & Privacy | 8 | 2 | dbt does not classify warehouse data. Data Workers' pull request review flags new columns whose names or annotations look sensitive, and leaves a receipt on every data change. |
| Cost / FinOps | 8 | 4 | dbt State skips unneeded builds and Cost Insights (GA on Snowflake, BigQuery and Databricks) estimates cost per model. Data Workers attributes Snowflake credits to the dbt model behind them and drafts the fix. |
| MLOps & Models | 7.5 | 1 | dbt builds feature tables but does not watch models. Data Workers keeps the data under your models fresh and correct. |
For the category view, read what is an agentic data platform and how Data Workers differs from data observability.
How dbt and Data Workers work together
dbt stays on top, where engineers build and review models with Wizard and Copilot. Spellbook Data Catalog (in preview) is where the data team reviews, approves and rolls back. Data Context Wizard joins the dbt manifest to the warehouse, ingestion and BI in one governed context graph, the Data-Agents Swarm does the work with 20+ specialist agents, and the Autonomous Data-Conductor carries each fix from detection to verification.

Exactly what Data Workers reads, runs and writes in dbt. dbt and the dbt platform are native connections.
| What it covers | How | |
|---|---|---|
| Reads | manifest.json (models, sources, tests, exposures), catalog.json, Semantic Layer definitions | From CI artifacts or target/; list_dbt_models, get_dbt_model_lineage |
| Runs | Job runs and backfills on jobs your team names | Proposed with the exact job and dates; the owner runs them after a named approval, or your orchestrator queues them |
| Writes | Model SQL and config changes, such as the lookback above | A diff for the owner to merge; nothing is pushed to your branches |
| Pull requests | Documentation changes (docs blocks and schema.yml descriptions) | Only when your team turns on the dbt write-back GitHub target; it is off by default and writes local files otherwise |
Data Workers also reviews your team's dbt pull requests with a blast-radius comment and an impact diff, and can block a risky change. It reads Snowflake natively and connects to Firehose and Sigma over their APIs today. The Data Workers + dbt integration guide covers the full wiring; this page is about the operations on top.
Two MCP servers in one client. Every Data Workers agent is an MCP server, set up by the documented path in the client setup guide: clone the open-source repo and add one start-agent.sh entry per agent. The dbt MCP server sits beside it, answering questions about the project, while job runs stay on the Data Workers side behind an approval; dbt documents DISABLE_DBT_CLI and DISABLE_TOOLS for that. Example for Claude Desktop or Cursor:
{
"mcpServers": {
"dbt": {
"command": "uvx",
"args": ["dbt-mcp"],
"env": {
"DBT_HOST": "cloud.getdbt.com",
"DBT_TOKEN": "<read-scoped service token>",
"DBT_PROD_ENV_ID": "<prod environment id>",
"DISABLE_DBT_CLI": "true",
"DISABLE_TOOLS": "trigger_job_run,cancel_job_run,retry_job_run"
}
},
"dw-context-catalog": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-context-catalog"] },
"dw-incidents": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-incidents"] },
"dw-quality": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-quality"] },
"dw-connectors": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-connectors"] }
}
}Ask "why are Saturday's payouts low?" and the client calls both: dbt returns the model's config, lineage and last runs; Data Workers returns the volume gap, the reload behind it, the workbook downstream and a proposed backfill awaiting its approver. For a shared endpoint, the Data Workers remote server serves /mcp with an API key (bearer) or OAuth tokens from your identity provider, verified through JWKS.
One dbt incident, L0 to L4. The same weekend at each level of the ladder, set per domain.

- •L0 manual. Finance spots short payouts; an engineer finds the gap and writes the backfill by hand.
- •L1 observe. Data Workers posts the diagnosis and every model and recorded reader downstream. Nothing changes.
- •L2 propose. Data Workers proposes the Saturday backfill, the
lookbackdiff and the undo. Nothing runs until the payouts owner approves. - •L3 act reversibly. For change classes with a proven record, such as a one-day backfill with Time Travel as the undo, Data Workers queues the run through your orchestrator, verifies it and records the receipt.
- •L4 autonomous. A scoped domain with a long clean record runs the repeatable steps end to end; model code changes still go to the owner as a diff.
For the safety model, read is it safe to let AI agents change production data and how approvals work. On where data lives: the agents run in your infrastructure and hold the warehouse and dbt credentials, your data stays in your systems, and the hosted Conductor sees workflow metadata only.
What changes for your team
dbt gave your analytics engineers a tested codebase. Data Workers handles what happens to the data between commits.

- •Incidents. A wrong number in a model gets a cross-system diagnosis, a blast radius across models and recorded readers, and a fix proposal.
- •Data quality. Volume and freshness checks run beside dbt tests, so a green run that built short or duplicated data is caught.
- •Cloud spend. On Snowflake, credits are traced to the dbt model behind them through query tags, and each fix goes to its owner drafted.
- •Access. A request for warehouse access arrives as a scoped grant with an expiry, drafted for the data owner to approve.
- •Audits. Every backfill and model change carries a receipt: the cause, the diff, the approver, the run, the verification and the undo.
- •Migrations. A warehouse move runs in approved waves, with parity checks planned and tracked per wave and a completion gate the owner signs.
The analytics engineering on-call changes most: backfill decisions stop living in someone's head and arrive as one proposal with its evidence. For incremental strategies in depth, see our dbt incremental models guide and AI agents for root cause analysis on dbt test failures.
Keep dbt, or consolidate?
Keep dbt if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
For dbt, keeping it is the norm: your models, tests and contracts are the reviewed definition of the business, and Data Workers builds on them. What teams consolidate is the tooling around dbt: a monitor that only sees the warehouse, alert rules across three consoles, backfill scripts in a wiki and a hand-run check before every payout or close. The same pattern runs through you're on Elementary, you're on Airflow and you're on Datafold; for metric definitions, see you're on the dbt Semantic Layer. Building the operator yourself on the dbt MCP server and a coding agent? Read build it ourselves with Claude Code and MCP servers: the MCP calls are the easy part; the context graph, approvals and rollback are the work.
The case for your CFO
The outcome: the numbers your company pays and reports on come out of dbt models. Data Workers makes sure those models are complete and right before anyone acts on them, and turns each gap into an approved fix instead of a correction after the money moved.
The risk story: at L1 Data Workers only reads; at L2 it changes nothing until a named person approves; at L3 it acts only on changes it can undo, such as a one-day backfill with the previous table version kept. Model code changes arrive as a diff for the owner to merge. An unanswered request expires and escalates, never auto-grants, and no agent can promote its own work. Every change carries a receipt, and an org-wide stop halts all autonomous dispatch. Zero migration: dbt, the dbt platform, Snowflake, Firehose and Sigma stay where they are.
Why now: dbt's own agents now write models and investigate failed runs well, so the costly incidents are the ones with no failed run at all. The first win is read-only: every model in one domain, such as payouts or revenue, gets daily null, uniqueness and row-count checks, lateness baselines on its sources, and a diagnosis on every gap. Your models, tests, CI and job schedules stay the same. For the numbers, see the ROI of agentic data operations; for leaders, the dbt data leaders guide.
The sentence to repeat upstairs: "dbt is where we define and test our numbers; Data Workers checks they came out complete, fixes the gap with an owner's approval and shows us the receipt."
Getting started
Start with a pilot. Pick one domain where a wrong number costs real money, such as the models behind payouts or the revenue close, connect Data Workers to the dbt project, the dbt platform and the warehouse, and run at L1 so every gap between source and model gets a diagnosis and a blast radius. Then turn on the first write class at L2: approved runs and backfills on named jobs. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.
FAQ
How does Data Workers connect to dbt? Natively. It reads manifest.json and catalog.json from your CI artifacts or target/. It works with dbt platform jobs and with dbt run by an orchestrator.
Does Data Workers open pull requests on our dbt repo? Only if your team turns on the dbt write-back GitHub target, which is off by default and carries documentation changes such as schema.yml descriptions. Model SQL and config changes arrive as a diff for the owner to merge, under your CI and branch protection.
Can Data Workers run dbt jobs? It proposes runs of the jobs your team names, each behind a named approval at L2; the owner runs them on the dbt platform, or your orchestrator queues them. At L3, a narrow class with a proven record, such as one-day backfills, can run through your orchestrator with its undo recorded. It never edits job definitions or environment settings.
We already have dbt Wizard. What does Data Workers add? Wizard is the right tool for changing the project and investigating a failed run. Data Workers covers the incidents around the project: late data, reloads, source changes and green runs that built wrong data, traced across ingestion, the warehouse and BI, with an approval and a receipt. An engineer can open any Data Workers diff in the Studio IDE and refine it with Wizard.
How does Data Workers catch a green run with wrong data? It checks models against their sources: daily volumes and freshness on both sides, with each table's volume compared to the same weekday in prior weeks, so a Saturday at 58% stands out even when every test passes.
Does the Fivetran and dbt Labs merger change anything? Not for Data Workers. It reads dbt through dbt's own artifacts and API, and reads ingestion where it lands in the warehouse, whichever ingestion tool you run.
Sources
- •dbt Developer Blog, "dbt v2.0 is GA" (Sep 16, 2026), https://docs.getdbt.com/blog/dbt-v2-is-ga (checked Oct 3, 2026)
- •dbt Developer Blog, "What's the difference between dbt and dbt OSS?" (Sep 16, 2026), https://docs.getdbt.com/blog/comparing-dbt-and-dbt-oss (checked Oct 3, 2026)
- •dbt release notes (dbt Summit 2026 naming; Cost Insights GA for Snowflake, BigQuery and Databricks), https://docs.getdbt.com/docs/dbt-versions/dbt-cloud-release-notes (checked Oct 3, 2026)
- •dbt Docs, dbt platform features, https://docs.getdbt.com/docs/cloud/about-cloud/dbt-cloud-features (checked Oct 3, 2026)
- •dbt Docs, dbt Wizard in Studio IDE (preview; approval modes), https://docs.getdbt.com/docs/dbt-ai/developer-agent (checked Oct 3, 2026)
- •dbt Docs, dbt Copilot, https://docs.getdbt.com/docs/dbt-ai/copilot-overview (checked Oct 3, 2026)
- •dbt Docs, About the dbt MCP server, https://docs.getdbt.com/docs/dbt-ai/about-mcp, and MCP environment variables, https://docs.getdbt.com/docs/dbt-ai/mcp-environment-variables (checked Oct 3, 2026)
- •dbt Labs, dbt-mcp README and releases (v2.5.0, Sep 28, 2026), https://github.com/dbt-labs/dbt-mcp (checked Oct 3, 2026)
- •dbt Docs, microbatch incremental models, https://docs.getdbt.com/docs/build/incremental-microbatch (checked Oct 3, 2026)
- •Fivetran press, "Fivetran + dbt Labs Complete Merger" (Jun 1, 2026), https://www.fivetran.com/press/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents (checked Oct 3, 2026)
- •Fivetran press, dbt Summit 2026 announcements (Sep 16, 2026), https://www.fivetran.com/press/fivetran-dbt-labs-announces-new-capabilities-to-make-enterprise-data-agent-ready-at-dbt-summit-2026 (checked Oct 3, 2026)
- •AWS Docs, Firehose destination settings for Snowflake, https://docs.aws.amazon.com/firehose/latest/dev/create-destination.html, and handling data delivery failures, https://docs.aws.amazon.com/firehose/latest/dev/retry.html (checked Oct 3, 2026)
- •Data Workers open-source repository, https://github.com/DataWorkersProject/dataworkers-claw-community, and client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)