Product
Product11 min readBy The Data Workers Team

You're on dbt: dbt Defines and Tests Your Models. Data Workers Runs the Operations Around Them

dbt is where your team defines and tests transformations. Data Workers diagnoses the incidents around them, proposes fixes as diffs and verifies, behind approvals.

Your analytics engineers think in models, refs, tests and exposures. Every transformation is SQL in a pull request, every model has a unique and not_null test, contracts guard the marts other teams depend on, and a job on the dbt platform (formerly dbt Cloud) builds production every night. Since September 16, 2026, the engine is simply "dbt": the Fusion engine graduated to general availability as dbt v2, and dbt Core v2 is now dbt OSS, the Apache 2.0 distribution. dbt Wizard (public preview) builds models and investigates failed runs inside the project, and Fivetran and dbt Labs have been one company since June 1, 2026. dbt is where your team defines and tests the transformations. Data Workers runs the operations around them: it diagnoses the incident, proposes the fix as a diff, queues the approved run and verifies the numbers. That matters most on the mornings when every test passed and the number is still wrong, because the cause was a late load the project never saw.

Key takeaways

  • •dbt keeps its job. Models, tests, contracts, the job scheduler, Wizard and Copilot stay where your analytics engineers use them.
  • •Data Workers takes the operations around the models. It watches volume on sources and models, tracks lateness against baselines your team records, diagnoses incidents across systems, and catches the runs that pass with wrong data.
  • •Exactly what it touches in dbt. It reads the manifest and catalog, and proposes the exact job runs for the owner to approve and run. Model changes arrive as a diff for the owner to merge; pull requests, for docs changes, only when your team turns on the dbt write-back GitHub target.
  • •One approval, one receipt. A named owner approves in Spellbook; the receipt records the cause, the diff, the run, the verification and the undo.

dbt is the transformation layer your team reviews. Data Workers is the operator around it.

dbt runs the code your team reviewed, correctly. Data Workers owns whether the data it produced is complete and right. Here is a weekend at a food-delivery marketplace, an illustration rather than a customer case. Courier app events stream through Amazon Data Firehose into Snowflake. A dbt platform job builds fct_deliveries as a microbatch incremental model (event_time='delivered_at', batch_size='day', the default lookback of 1, on the v1 Latest release track), fct_courier_payouts builds on it, and a dbt exposure, also recorded in the context graph, points at the Sigma workbook finance operations uses to approve the weekly courier payout file on Monday at 10:00.

TimeSystemWhat happens
Sat 13:50Snowflake, FirehoseA scheduled key rotation replaces the key on the Firehose service user; deliveries start failing, and Firehose writes the records it cannot deliver to its S3 error bucket for manual backfill
Sun 02:00dbt platformThe nightly job builds Saturday's batch from the rows that arrived before 13:50. Every test passes, and source freshness (warn after 24 hours) passes too
Sun 02:30Data Workersmonitor_metrics flags raw.courier_events as stalled since 13:50 and Saturday's volume at 58% of the trailing four Saturdays, points at the Firehose stream as the table's only writer, and alerts the platform on-call in Slack
Mon 01:15S3, SnowflakeThe platform engineer updates the key and reloads 36 hours of events from the S3 error bucket with COPY INTO. The raw table is complete
Mon 02:00dbt platformThe job runs. With lookback of 1, microbatch rebuilds Sunday's and Monday's batches; Saturday's batch is not reprocessed. Tests pass and the job is green
Mon 02:40Data WorkersCompares daily volumes: raw Saturday is now at 100% of the trailing Saturdays, fct_deliveries Saturday is still at 58%. It traces the gap to the lookback window and maps the blast radius through the manifest and the context graph: fct_courier_payouts and the Sigma workbook, 21,340 deliveries and 3,118 couriers short
Mon 02:55SpellbookData Workers proposes the repair to the payouts model owner: a backfill of Saturday for fct_deliveries+ with --event-time-start and --event-time-end, a diff raising lookback to 3, and the Time Travel version of both tables as the undo
Mon 07:20SpellbookThe owner reviews the evidence and approves, and sets the dates on the team's backfill job
Mon 07:25dbt platformThe owner runs the approved backfill job
Mon 08:30SnowflakeData Workers verifies the model's daily volumes for the last seven days against the raw table and writes the receipt
Mon 09:10GitHubThe analytics engineer merges the lookback change after CI passes
Mon 10:00SigmaFinance operations approves a payout file with every Saturday delivery in it
Incident timeline across the stack: what dbt, your team and Data Workers each do, step by step

Nothing in dbt misbehaved. Microbatch does what its docs say: a standard run processes batches by the current timestamp and the lookback, and older late data waits for an explicit backfill with --event-time-start and --event-time-end. The tests checked the rows that were there. Data Workers compared the source with the model after the upstream fix and handed a named owner a precise backfill with its undo before the payout file was approved.

JobWhat dbt doesWhat Data Workers does
Defining the logicModels, refs, macros and contracts, reviewed in pull requestsReads the manifest as part of one context graph that also covers ingestion, the warehouse and BI
TestingData tests, unit tests and source freshness on every buildAdds volume checks with run_quality_check and freshness baselines with monitor_metrics on sources and models
RunningThe job scheduler builds production on schedule, with dbt State skipping unchanged workProposes the exact runs and backfills; the owner runs them after approval, and Data Workers verifies
InvestigatingWizard investigates job and run failures inside the projectDiagnoses incidents across systems, including green runs that built wrong data, with diagnose_incident and trace_cross_platform_lineage
ImpactExposures and lineage show who reads a modelMaps the blast radius from the source to every model and recorded reader with blast_radius_analysis; pull-request review's lineage diff reads exposures
FixingYour engineers write and merge the changeProposes the change as a diff for the owner to merge, with the run that applies it
The recordRun history and the Git logA receipt: cause, diff, approver, run, verification and undo, in Spellbook and the audit trail

Why doesn't dbt just do this itself?

Because dbt is designed to be a compiler and build system for transformation logic, and that focus is why teams trust it. A dbt run builds the code and inputs it is given, deterministically. It does not guess that a raw table received rows for a day it already built; guessing would make builds slower, costlier and less predictable. Microbatch makes the trade explicit: the lookback is a cost setting your team chooses, and late data outside it is a backfill decision left to a person. That is the right design for a build tool.

dbt's AI follows the same scope. dbt Wizard builds and refactors models, writes YAML and investigates job and run failures, with "Ask for approval" as its default mode: "You approve each file change before it is persisted." dbt Copilot generates SQL, docs, tests and semantic models. The self-hosted dbt MCP server can run dbt commands, and its README warns that doing so "could modify your data models, sources, and warehouse objects"; its Admin API tools trigger, cancel and retry job runs. Each acts on dbt's own objects, and each is good at it. On Monday at 02:40 there was no failed run to investigate.

The weekend's fix crossed four systems and needed a backfill scoped to one day, an undo ready first, a named approver and a record finance could read. That is a different product, with liability for changes beyond the project. Data Workers is built for it, on top of dbt.

Every tool owns a slice. Data Workers covers the whole lifecycle

dbt owns transformation, and nobody does it better. Each point tool around it adds another console, another contract and another handoff: a warehouse monitor, an ingestion alert, a wiki of backfill scripts. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and keeps dbt at the center of it.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, dbt goes deep on its own area
StageData WorkersdbtWhy we scored it this way
Catalog & Context96dbt Catalog, docs and model-level lineage describe the project well. Data Workers keeps one governed context graph across dbt, the warehouse, ingestion and BI, with owners and usage.
Analytics & Insights85The Semantic Layer and dbt Charts (public beta) serve governed metrics. Data Workers answers data questions from governed definitions with lineage behind every number.
Data Quality87Tests, unit tests and contracts run on every build, and Copilot and Wizard write more of them. Data Workers adds volume checks and recorded lateness baselines that span sources, models and consumers.
Observability & Incidents8.54Job notifications and run history tell you a run failed. Data Workers diagnoses the incident across systems, including runs that passed with wrong data.
Pipelines & Ingestion8.59dbt's home stage: models, incremental strategies, dbt State and the job scheduler. Data Workers plans approved runs and backfills for the owner and proposes fixes as diffs.
Schema & Migration86Contracts, model versions and CI catch breaking changes inside the project. Data Workers catches schema changes at the source first and plans migrations in waves.
Governance & Access8.54Model access and groups govern who can ref a model. Data Workers drafts each warehouse access request as a scoped grant with an expiry for a named owner to approve.
Security & Privacy82dbt does not classify warehouse data. Data Workers' pull request review flags new columns whose names or annotations look sensitive, and leaves a receipt on every data change.
Cost / FinOps84dbt State skips unneeded builds and Cost Insights (GA on Snowflake, BigQuery and Databricks) estimates cost per model. Data Workers attributes Snowflake credits to the dbt model behind them and drafts the fix.
MLOps & Models7.51dbt builds feature tables but does not watch models. Data Workers keeps the data under your models fresh and correct.

How dbt and Data Workers work together

dbt stays on top, where engineers build and review models with Wizard and Copilot. Spellbook Data Catalog (in preview) is where the data team reviews, approves and rolls back. Data Context Wizard joins the dbt manifest to the warehouse, ingestion and BI in one governed context graph, the Data-Agents Swarm does the work with 20+ specialist agents, and the Autonomous Data-Conductor carries each fix from detection to verification.

How Data Workers fits with dbt: your coding agent on top, Data Workers in the middle, your estate underneath

Exactly what Data Workers reads, runs and writes in dbt. dbt and the dbt platform are native connections.

What it coversHow
Readsmanifest.json (models, sources, tests, exposures), catalog.json, Semantic Layer definitionsFrom CI artifacts or target/; list_dbt_models, get_dbt_model_lineage
RunsJob runs and backfills on jobs your team namesProposed with the exact job and dates; the owner runs them after a named approval, or your orchestrator queues them
WritesModel SQL and config changes, such as the lookback aboveA diff for the owner to merge; nothing is pushed to your branches
Pull requestsDocumentation changes (docs blocks and schema.yml descriptions)Only when your team turns on the dbt write-back GitHub target; it is off by default and writes local files otherwise

Data Workers also reviews your team's dbt pull requests with a blast-radius comment and an impact diff, and can block a risky change. It reads Snowflake natively and connects to Firehose and Sigma over their APIs today. The Data Workers + dbt integration guide covers the full wiring; this page is about the operations on top.

Two MCP servers in one client. Every Data Workers agent is an MCP server, set up by the documented path in the client setup guide: clone the open-source repo and add one start-agent.sh entry per agent. The dbt MCP server sits beside it, answering questions about the project, while job runs stay on the Data Workers side behind an approval; dbt documents DISABLE_DBT_CLI and DISABLE_TOOLS for that. Example for Claude Desktop or Cursor:

{
  "mcpServers": {
    "dbt": {
      "command": "uvx",
      "args": ["dbt-mcp"],
      "env": {
        "DBT_HOST": "cloud.getdbt.com",
        "DBT_TOKEN": "<read-scoped service token>",
        "DBT_PROD_ENV_ID": "<prod environment id>",
        "DISABLE_DBT_CLI": "true",
        "DISABLE_TOOLS": "trigger_job_run,cancel_job_run,retry_job_run"
      }
    },
    "dw-context-catalog": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-context-catalog"] },
    "dw-incidents": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-incidents"] },
    "dw-quality": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-quality"] },
    "dw-connectors": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-connectors"] }
  }
}

Ask "why are Saturday's payouts low?" and the client calls both: dbt returns the model's config, lineage and last runs; Data Workers returns the volume gap, the reload behind it, the workbook downstream and a proposed backfill awaiting its approver. For a shared endpoint, the Data Workers remote server serves /mcp with an API key (bearer) or OAuth tokens from your identity provider, verified through JWKS.

One dbt incident, L0 to L4. The same weekend at each level of the ladder, set per domain.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. Finance spots short payouts; an engineer finds the gap and writes the backfill by hand.
  • •L1 observe. Data Workers posts the diagnosis and every model and recorded reader downstream. Nothing changes.
  • •L2 propose. Data Workers proposes the Saturday backfill, the lookback diff and the undo. Nothing runs until the payouts owner approves.
  • •L3 act reversibly. For change classes with a proven record, such as a one-day backfill with Time Travel as the undo, Data Workers queues the run through your orchestrator, verifies it and records the receipt.
  • •L4 autonomous. A scoped domain with a long clean record runs the repeatable steps end to end; model code changes still go to the owner as a diff.

For the safety model, read is it safe to let AI agents change production data and how approvals work. On where data lives: the agents run in your infrastructure and hold the warehouse and dbt credentials, your data stays in your systems, and the hosted Conductor sees workflow metadata only.

What changes for your team

dbt gave your analytics engineers a tested codebase. Data Workers handles what happens to the data between commits.

Six jobs that run on autopilot with Data Workers next to dbt, with a concrete example of each
  • •Incidents. A wrong number in a model gets a cross-system diagnosis, a blast radius across models and recorded readers, and a fix proposal.
  • •Data quality. Volume and freshness checks run beside dbt tests, so a green run that built short or duplicated data is caught.
  • •Cloud spend. On Snowflake, credits are traced to the dbt model behind them through query tags, and each fix goes to its owner drafted.
  • •Access. A request for warehouse access arrives as a scoped grant with an expiry, drafted for the data owner to approve.
  • •Audits. Every backfill and model change carries a receipt: the cause, the diff, the approver, the run, the verification and the undo.
  • •Migrations. A warehouse move runs in approved waves, with parity checks planned and tracked per wave and a completion gate the owner signs.

The analytics engineering on-call changes most: backfill decisions stop living in someone's head and arrive as one proposal with its evidence. For incremental strategies in depth, see our dbt incremental models guide and AI agents for root cause analysis on dbt test failures.

Keep dbt, or consolidate?

Keep dbt if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

For dbt, keeping it is the norm: your models, tests and contracts are the reviewed definition of the business, and Data Workers builds on them. What teams consolidate is the tooling around dbt: a monitor that only sees the warehouse, alert rules across three consoles, backfill scripts in a wiki and a hand-run check before every payout or close. The same pattern runs through you're on Elementary, you're on Airflow and you're on Datafold; for metric definitions, see you're on the dbt Semantic Layer. Building the operator yourself on the dbt MCP server and a coding agent? Read build it ourselves with Claude Code and MCP servers: the MCP calls are the easy part; the context graph, approvals and rollback are the work.

The case for your CFO

The outcome: the numbers your company pays and reports on come out of dbt models. Data Workers makes sure those models are complete and right before anyone acts on them, and turns each gap into an approved fix instead of a correction after the money moved.

The risk story: at L1 Data Workers only reads; at L2 it changes nothing until a named person approves; at L3 it acts only on changes it can undo, such as a one-day backfill with the previous table version kept. Model code changes arrive as a diff for the owner to merge. An unanswered request expires and escalates, never auto-grants, and no agent can promote its own work. Every change carries a receipt, and an org-wide stop halts all autonomous dispatch. Zero migration: dbt, the dbt platform, Snowflake, Firehose and Sigma stay where they are.

Why now: dbt's own agents now write models and investigate failed runs well, so the costly incidents are the ones with no failed run at all. The first win is read-only: every model in one domain, such as payouts or revenue, gets daily null, uniqueness and row-count checks, lateness baselines on its sources, and a diagnosis on every gap. Your models, tests, CI and job schedules stay the same. For the numbers, see the ROI of agentic data operations; for leaders, the dbt data leaders guide.

The sentence to repeat upstairs: "dbt is where we define and test our numbers; Data Workers checks they came out complete, fixes the gap with an owner's approval and shows us the receipt."

Getting started

Start with a pilot. Pick one domain where a wrong number costs real money, such as the models behind payouts or the revenue close, connect Data Workers to the dbt project, the dbt platform and the warehouse, and run at L1 so every gap between source and model gets a diagnosis and a blast radius. Then turn on the first write class at L2: approved runs and backfills on named jobs. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.

FAQ

How does Data Workers connect to dbt? Natively. It reads manifest.json and catalog.json from your CI artifacts or target/. It works with dbt platform jobs and with dbt run by an orchestrator.

Does Data Workers open pull requests on our dbt repo? Only if your team turns on the dbt write-back GitHub target, which is off by default and carries documentation changes such as schema.yml descriptions. Model SQL and config changes arrive as a diff for the owner to merge, under your CI and branch protection.

Can Data Workers run dbt jobs? It proposes runs of the jobs your team names, each behind a named approval at L2; the owner runs them on the dbt platform, or your orchestrator queues them. At L3, a narrow class with a proven record, such as one-day backfills, can run through your orchestrator with its undo recorded. It never edits job definitions or environment settings.

We already have dbt Wizard. What does Data Workers add? Wizard is the right tool for changing the project and investigating a failed run. Data Workers covers the incidents around the project: late data, reloads, source changes and green runs that built wrong data, traced across ingestion, the warehouse and BI, with an approval and a receipt. An engineer can open any Data Workers diff in the Studio IDE and refine it with Wizard.

How does Data Workers catch a green run with wrong data? It checks models against their sources: daily volumes and freshness on both sides, with each table's volume compared to the same weekday in prior weeks, so a Saturday at 58% stands out even when every test passes.

Does the Fivetran and dbt Labs merger change anything? Not for Data Workers. It reads dbt through dbt's own artifacts and API, and reads ingestion where it lands in the warehouse, whichever ingestion tool you run.

Sources

  • •dbt Developer Blog, "dbt v2.0 is GA" (Sep 16, 2026), https://docs.getdbt.com/blog/dbt-v2-is-ga (checked Oct 3, 2026)
  • •dbt Developer Blog, "What's the difference between dbt and dbt OSS?" (Sep 16, 2026), https://docs.getdbt.com/blog/comparing-dbt-and-dbt-oss (checked Oct 3, 2026)
  • •dbt release notes (dbt Summit 2026 naming; Cost Insights GA for Snowflake, BigQuery and Databricks), https://docs.getdbt.com/docs/dbt-versions/dbt-cloud-release-notes (checked Oct 3, 2026)
  • •dbt Docs, dbt platform features, https://docs.getdbt.com/docs/cloud/about-cloud/dbt-cloud-features (checked Oct 3, 2026)
  • •dbt Docs, dbt Wizard in Studio IDE (preview; approval modes), https://docs.getdbt.com/docs/dbt-ai/developer-agent (checked Oct 3, 2026)
  • •dbt Docs, dbt Copilot, https://docs.getdbt.com/docs/dbt-ai/copilot-overview (checked Oct 3, 2026)
  • •dbt Docs, About the dbt MCP server, https://docs.getdbt.com/docs/dbt-ai/about-mcp, and MCP environment variables, https://docs.getdbt.com/docs/dbt-ai/mcp-environment-variables (checked Oct 3, 2026)
  • •dbt Labs, dbt-mcp README and releases (v2.5.0, Sep 28, 2026), https://github.com/dbt-labs/dbt-mcp (checked Oct 3, 2026)
  • •dbt Docs, microbatch incremental models, https://docs.getdbt.com/docs/build/incremental-microbatch (checked Oct 3, 2026)
  • •Fivetran press, "Fivetran + dbt Labs Complete Merger" (Jun 1, 2026), https://www.fivetran.com/press/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents (checked Oct 3, 2026)
  • •Fivetran press, dbt Summit 2026 announcements (Sep 16, 2026), https://www.fivetran.com/press/fivetran-dbt-labs-announces-new-capabilities-to-make-enterprise-data-agent-ready-at-dbt-summit-2026 (checked Oct 3, 2026)
  • •AWS Docs, Firehose destination settings for Snowflake, https://docs.aws.amazon.com/firehose/latest/dev/create-destination.html, and handling data delivery failures, https://docs.aws.amazon.com/firehose/latest/dev/retry.html (checked Oct 3, 2026)
  • •Data Workers open-source repository, https://github.com/DataWorkersProject/dataworkers-claw-community, and client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)