Industry
Industry9 min readBy The Data Workers Team

Your Team Standardized on dbt. Now Let Data Workers Run the Work Around the Project: A Guide for Data Leaders

A guide for VPs and heads of data whose team standardizes on dbt: what dbt, Wizard, the MCP server, the Semantic Layer and Catalog cover, and how Data Workers runs the operations work around the project, with model owners approving.

dbt builds your models. Data Workers keeps what they produce true: it diagnoses the runs that fail or warn, proposes the fix to the model's owner, queues the rerun and checks the number afterwards. This guide is for the VP or head of data whose team standardized on dbt, on the dbt platform or the open source line.

Your analytics engineering team already lives in dbt. Hundreds of models sit in Git with tests, contracts and exposures to the dashboards that read them. Jobs build them in dependency order on the dbt platform or under Airflow. The Semantic Layer, powered by MetricFlow, defines revenue once so every tool reads the same number. dbt Catalog (formerly dbt Explorer) shows each model's lineage and, on Enterprise plans, who queries it. Since September 16, dbt v2, the Rust engine previously called Fusion, is generally available, and dbt Wizard sits in the Studio IDE to build and refactor models with your engineers.

Then the project grows. Four hundred models, six teams, and an owner field filled in on some of them. A nightly job warns, and the message lands in a channel where it waits for someone to claim it. The cause is usually outside the project: a renamed source column, a late sync, a retyped field. Someone fixes the model, someone else remembers to rerun the job, and nobody checks the board-deck number against the source. The same people field the warehouse bill and the audit request.

Your team has become the glue between excellent tools. dbt made building models cheap. The expensive part is operating them: the triage, the fixing, the rerunning, the checking and the record of what changed.

What dbt solved, and what it didn't

dbt solved building. It turned transformation into reviewed code, with tests and contracts that say what correct means and a review culture that is the safest place for any change to land. Its AI is aimed at that project and does it well. dbt Wizard, in preview, answers project-aware questions, builds and refactors models, generates YAML for tests, docs and metrics, and investigates job failures with dbt Agent Skills. By default it asks you to approve each file change before it is saved. dbt Copilot generates SQL, docs, tests and semantic models in the Studio IDE, Canvas and Insights. dbt Cost Insights, generally available on Enterprise plans since July, shows what each model and job costs to run. The dbt MCP server brings the Semantic Layer, Discovery metadata, SQL and the Admin API into Claude, Cursor and other clients, including tools that trigger, retry and cancel job runs; the remote server is built for consumption, and the self-hosted server adds dbt CLI commands such as build for development.

What dbt doesn't own is the work that starts and ends outside the project. A source system renames a column, the model warns, a dashboard shows the wrong number. Tracing that across the source, the warehouse, the orchestrator and BI, getting the right owner to approve a fix, rerunning the right job, checking the number against the source and keeping a receipt an auditor can read is not a job dbt was built for. That is the right focus for a framework whose strength is the project.

dbt builds your models. Data Workers is the agentic data platform that keeps what they produce true, and runs the rest of your data back office.

Or, in one sentence for your team: dbt is where we build and review the models; Data Workers runs the incidents, reruns and checks around them, under the approvals our model owners set.

What goes on autopilot

Each of these is a queue your team works by hand today, behind signals dbt already produces in its run results, tests and manifest.

Eight back-office jobs next to dbt, today versus with Data Workers

Nothing is replaced. Data Workers connects to dbt natively: it reads the manifest, run results and catalog, the dbt platform API and the Semantic Layer definitions, and works inside the warehouse and orchestrator you already run. Model fixes are proposed as diffs the model's owner merges, through your normal review and CI. When your team turns on the opt-in dbt write-back to GitHub, Data Workers opens those changes, such as docs and descriptions, as pull requests on a branch. Your repository stays the source of truth for the code, by design.

What changes for your organization

Here's what that looks like in a normal week. Access and migrations follow the same pattern: the change is drafted for its owner, approved and recorded.

Six jobs that run on autopilot with Data Workers next to dbt, with a concrete example of each
  • •Ownership stops being a channel. A warned run arrives at the model's owner already diagnosed, with every downstream model and recorded dashboard it reaches. Nobody has to claim it first.
  • •Fixes land where they belong. When the cause is upstream, Data Workers finds it there and proposes the change in the right place, so the project stops collecting defensive patches.
  • •Reruns become deliberate. After the owner approves, Data Workers queues the rerun of the affected job through the dbt job API and checks the result, instead of someone kicking off everything to be safe.
  • •The bill gets a cause. dbt Cost Insights shows what each model costs. On Snowflake, Data Workers traces credits through query tags to the dbt model and run behind them, ties them to the incident or rerun that caused them and drafts the fix for the owner.

One month on a 400-model project

This is an illustration, not a customer case.

  • •A 400-model dbt project on Snowflake serves finance, product and marketing, and six teams ship to it. The owner field is filled in on about half the models.
  • •In September the nightly job warns or fails on 14 nights. Most causes start upstream: a renamed Salesforce field, a late Fivetran sync, a column retyped in Postgres.
  • •Each warning waits in the data channel for someone to claim it, and each fix ends with a manual rerun of the whole job.
  • •At the quarterly review, the CFO asks why the warehouse bill went up and which board numbers were wrong for a morning. The answer takes a week of reading Git history and Slack threads.

With Data Workers, each warned run arrives traced to its cause in the source or the warehouse, with every model and recorded dashboard downstream of it. The fix goes to the model's owner as a diff, or, where no owner is set, to the approver your team named for that domain, and is merged through your normal review. After approval, Data Workers queues the rerun of the affected job and re-runs the agreed checks on the result. Snowflake credits are traced to the models and runs that consumed them, reruns included. At the quarterly review, the receipts answer both questions: which models broke and why, who approved each fix, and what the reruns cost.

You choose how far and how fast

You don't have to jump to full autonomy. You choose the altitude, one domain at a time.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous

L0 manual is the work your team does by hand today. Every domain starts at L1 observe: each failed or warned run gets a diagnosis and a blast radius, and nothing changes. At L2 propose, agents draft the fix and the rerun, and a named owner approves. At L3 act reversibly, change classes with a proven record, such as rerunning a job after an approved upstream fix, run on their own with the undo written down first and a receipt left behind. At L4 autonomous, a domain that has earned it runs end to end, and any domain can be dialled back at any time. An unanswered approval request expires and escalates, never auto-grants, and no agent can promote its own work.

The ladder in detail is in autonomy levels L0 to L4 explained, and how approvals work covers who signs off on what.

Why doesn't dbt just do this itself?

Focus and risk. dbt earns trust by making the project the governed home of transformation logic, and its AI follows that line: Wizard works inside the project with per-file approvals, and the MCP server exposes the project's context, commands and jobs to other tools. The incident that matters to a leader usually starts outside the project and ends outside it, in a source system and a board deck. Fixing it means scoping what else a change touches across systems dbt doesn't run, getting the right owner to approve, rerunning the right step, proving the number against the source and owning the undo. That is a different product with a different liability, and the product we built. Data Workers reads dbt the same way on the dbt platform or the open source line.

How it fits with dbt

  • •dbt keeps its job. Models, tests, contracts, the Semantic Layer, jobs, Catalog, Wizard and Copilot stay where they are.
  • •Your review stays the gate. Every model change is a diff your reviewers merge, with your CI and branch protection in front of it. Data Workers adds a blast-radius check to dbt pull requests, whoever wrote them.
  • •Your engineers keep their client. Data Workers runs beside the dbt MCP server in Claude Code, Cursor or Codex: dbt answers what the project says, and Data Workers works out what broke across the estate and how to repair it.
  • •Your platforms' permissions stay the lock. Data Workers works through the roles you grant in the dbt platform, GitHub, your warehouse and your orchestrator, and holds no more access than you give it.
  • •Your data stays in your systems. The agents run in your infrastructure and hold the warehouse credentials. Nothing is migrated, and the hosted orchestration service sees workflow metadata only. Our security and deployment guide has the detail.

When you don't need Data Workers

  • •One team owns the whole project and incidents rarely leave dbt.
  • •Failed and warned runs are claimed and fixed the same morning, every time.
  • •Nobody upstairs is asking what the warehouse spend buys or which numbers were wrong last quarter.

If all three are true, keep your budget. If any of them isn't, the rest of this guide is about you.

The case for your CFO

The outcome. The company already pays for a dbt team to build trusted models. Data Workers makes that investment hold up in production: breaks that start upstream are traced and fixed at the source before the business reads the number, each rerun is deliberate and checked, and the warehouse bill is traced to the models behind it. Analytics engineers' hours go to new models instead of morning triage.

The risk story. Autonomy is set per domain on the ladder above, and nothing changes at L2 until a named person approves. Model changes go through your reviewers and CI like anyone else's. Every change carries a receipt: the cause, the diff, who approved it, what it touched downstream, how it was verified and how to undo it. An org-wide stop halts all autonomous dispatch. Our safety guide goes deeper, and who owns the agents covers accountability.

Why now. dbt v2 is generally available, Wizard can edit files in the project, and the dbt MCP server lets every coding agent your engineers use read the project and trigger its jobs. Agents are coming to your dbt project either way. The choice is one approval flow and one record across all of them, or each engineer's own wiring: see build it ourselves with Claude Code and MCP servers.

The first win. Read-only: every failed or warned run in one domain, such as the finance models, arrives with a diagnosis, its owner and its blast radius.

What stays the same. dbt, your repository, reviewers, CI, jobs and warehouse permissions. The path is a pilot, run on your own project before you commit, and the ROI guide shows how to size it.

The sentence to repeat upstairs: dbt is where we build the models; Data Workers keeps what they produce true, under our owners' approvals, with a receipt for every change.

Where to start

Pick one domain whose models break the same way every time: the finance marts behind the board deck, the models fed by one source system, or the most expensive models on the Snowflake bill.

Start with a pilot. A forward-deployed engineer connects Data Workers to your dbt project, the dbt platform, your warehouse and your orchestrator, and runs the first domain alongside your team, read-only first. See pricing for how the pilot works.

For the detail, send your analytics engineers you're on dbt, the practitioner version of this guide with a worked incident, and Data Workers + dbt for the wiring: what Data Workers reads from the project and what it sends back. Data Workers for analytics engineers covers the role, you're on the dbt Semantic Layer covers metrics, and the Airflow leaders guide makes the same move for the orchestrator. The practitioner write-up on root cause analysis for dbt test failures goes deeper on triage. They cover the four products: the Autonomous Data-Conductor picks up dbt's run results and runs each fix from detection to a verified result, the Data-Agents Swarm does the work, Data Context Wizard keeps one governed graph of definitions, lineage and owners across every platform with your dbt manifest and Semantic Layer as first-class sources, and Spellbook Data Catalog (in preview) is where your team approves, audits and rolls back. The thinking behind them is in our thesis.

Sources

  • •dbt, homepage (Fusion engine; dbt Catalog and lineage; "Fivetran and dbt are one company"), checked Oct 3, 2026: https://www.getdbt.com/
  • •dbt Docs blog, dbt v2 is GA (Fusion engine GA under the name dbt; Sep 16, 2026), checked Oct 3, 2026: https://docs.getdbt.com/blog/dbt-v2-is-ga
  • •dbt Docs, dbt Wizard in Studio IDE (preview; approval modes; job-failure investigation with dbt Agent Skills), checked Oct 3, 2026: https://docs.getdbt.com/docs/dbt-ai/developer-agent
  • •dbt Docs, dbt Copilot (SQL, docs, tests and semantic models in Studio IDE, Canvas and Insights), checked Oct 3, 2026: https://docs.getdbt.com/docs/cloud/dbt-copilot
  • •dbt Docs, About dbt MCP (self-hosted vs remote servers; remote for data consumption), checked Oct 3, 2026: https://docs.getdbt.com/docs/dbt-ai/about-mcp
  • •dbt Docs, dbt MCP available tools (Admin API trigger_job_run, retry_job_run, cancel_job_run; CLI tools can modify models and warehouse objects), checked Oct 3, 2026: https://docs.getdbt.com/docs/dbt-ai/mcp-available-tools
  • •dbt Docs, Cost Insights (Enterprise, Enterprise+; cost and compute time per model and job), checked Oct 3, 2026: https://docs.getdbt.com/docs/explore/cost-insights
  • •dbt Docs, dbt platform release notes (Cost Insights GA for Snowflake, BigQuery and Databricks, July 2026; dbt Summit naming: dbt v2, dbt OSS, dbt v1), checked Oct 3, 2026: https://docs.getdbt.com/docs/dbt-versions/dbt-cloud-release-notes
  • •dbt Docs, dbt Semantic Layer (powered by MetricFlow; Starter, Enterprise, Enterprise+), checked Oct 3, 2026: https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl
  • •dbt Docs, Discover data with dbt Catalog (formerly dbt Explorer; Enterprise features such as model query history), checked Oct 3, 2026: https://docs.getdbt.com/docs/explore/explore-projects
  • •dbt Labs, Fivetran and dbt Labs complete merger (Jun 1, 2026), checked Oct 2, 2026: https://www.getdbt.com/blog/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents