Product
Product12 min readBy The Data Workers Team

Data Workers + dbt: A dbt Integration Where AI Agents Work Through Your Project

A dbt integration for AI agents: Data Workers reads your manifest, runs and metric definitions, traces breaks outside dbt, and proposes fixes as diffs, or as pull requests once you turn on the GitHub target.

dbt owns your models, tests, contracts and jobs. Data Workers owns the investigation and the paperwork around a change, and hands every change to dbt's own review and CI. That is the whole shape of this dbt integration: Data Workers reads your project as context, traces breaks that start outside it, proposes the fix as a diff (a pull request, once your team turns on the GitHub pull-request target) that your reviewers handle like any other, and checks the number afterwards with a receipt on every change.

Your estate probably looks like this. Fivetran lands source data in Snowflake, BigQuery or Databricks. dbt turns it into models, with tests, contracts and a Semantic Layer on top. dbt platform jobs or Airflow run the builds. A BI tool reads the marts through exposures you've declared. When a model goes wrong, the cause is often a step or two outside the project: a renamed source column, a sync that changed shape, a DAG run that started late. Your analytics engineers find out from a test warning or from a finance question at 08:15.

Data Workers is the agentic data platform that runs the whole data lifecycle around that project. Here is how it wires into dbt, what it reads, what it sends back, and what stays exactly where it is.

Key takeaways

  • •dbt stays the home of your transformation logic. Models, tests, contracts, the Semantic Layer and jobs stay in dbt, with your Git history and your reviewers.
  • •Data Workers reads dbt as context. It reads the manifest, catalog and Semantic Layer definitions into one governed context graph, through dbt's artifacts.
  • •Every change goes back through your review. Data Workers proposes model, test and contract changes as a diff, and opens the pull request when your team turns on the GitHub pull-request target; your dbt CI runs the tests and a named engineer merges.
  • •The loop runs past dbt. Data Workers traces a break into the source, the warehouse, Airflow and BI, hands the approved rerun to your dbt job schedule, and compares the metric with the source before it closes the incident.
  • •dbt's own AI keeps its job. dbt Wizard and dbt Copilot help engineers author inside the project. Data Workers works alongside them across the estate.
  • •Start read-only. Connect the manifest and a read token first; turn on the GitHub pull-request target when the proposals earn it.

What dbt does, and why teams keep it

dbt is where analytics engineering happens. SQL models live in Git, tests and contracts say what "correct" means, the Semantic Layer and MetricFlow define metrics once, and jobs build everything in dependency order. Teams keep dbt because it turned transformation into reviewed code, and that review culture is the safest place for any agent to send a change.

2026 has been a big year for the project. Fivetran and dbt Labs completed their merger on June 1. On September 16, dbt v2 (the Rust engine previously called Fusion) reached general availability, and the open source line is now called dbt OSS. dbt's own AI story grew too: dbt Wizard in the Studio IDE, a Wizard CLI, and a remote dbt MCP server that any agent can call.

AreaWhat dbt shipsStatus, October 2026
CompanyFivetran + dbt Labs mergerCompleted June 1, 2026
dbt v2Rust engine, formerly Fusion; pip install dbt installs itGA since Sept 16, 2026
dbt v2 adaptersSnowflake, BigQuery, Redshift, Databricks (local and platform); DuckDB (local)GA, Sept 2026
dbt OSSFormerly dbt Core 2.0; Apache 2.0; dbt v1 stays open sourceBeta (Aug 2026); renamed Sept 2026
dbt MCP server (local)Apache 2.0; SQL, Semantic Layer, Discovery, Admin API, dbt CLI, Codegen, column lineagev2.5.0, Sept 28, 2026
Remote dbt MCP serverSemantic Layer, SQL, Discovery and Admin API toolsets; no CLI or CodegenGA since Oct 14, 2025; OAuth in public beta
dbt WizardAgent in the Studio IDE: builds and refactors models, writes YAML, investigates job failures; Explore modePublic preview
Wizard CLI, Wizard DesktopWizard outside the IDEPublic beta; private beta
dbt CopilotGenerates SQL, docs, tests and semantic models in the IDEGA
MetricFlowThe Semantic Layer's engineApache 2.0 since Oct 14, 2025
Semantic Layer YAML specSemantic models embedded in model YAMLNew Jan 2026; live in dbt v2 Feb 2026
Apache OssieOpen Semantic Interchange, renamed on entering the Apache IncubatorIncubating since July 10, 2026
Model contractsEnforce column names and types; not_null enforced; Databricks also enforces check; keys are metadataAvailable
Advanced CI (compare changes)Row, column and key diffs posted to the PREnterprise and Enterprise+
dbt StateReuses unchanged models locally, on the platform and with external orchestratorsGA, Sept 2026

Contracts, tests, CI and compare changes give any agent a governed way to prove a change is safe. Data Workers routes every change through them.

What Data Workers reads from dbt, and what it sends back

This is the core of the integration. Data Workers reads dbt through dbt's own surfaces: the artifacts your runs produce, the Semantic Layer definitions and your Git repository. It sends changes back the same way a person would.

Data Workers reads from dbtData Workers sends back through dbt
manifest.json: models, sources, tests, metrics, exposures and lineageModel, test and contract changes, proposed as a diff; a pull request once the GitHub target is on
Test outcomes your team reports on the incident: what passed, warned and failedA blast-radius comment on every dbt pull request, whether a person or an agent wrote it
catalog.json: columns and types as builtDocs blocks and descriptions in schema YAML, through the approvals queue
Semantic Layer and MetricFlow definitions: metrics, dimensions, grainThe rerun plan, for your job schedule to run after a named approval
Your Git history and open PRs (base and head manifests)A receipt linked on the PR and in Spellbook
What Data Workers reads from dbt and what it writes back through dbt

A few agents do most of the dbt work. The Data Context & Catalog agent reads the manifest and MetricFlow definitions into Data Context Wizard, so every agent knows what a model means and what reads it. The Incident Debugging agent starts from a failed or warning run your team reports, or a failed quality check, and follows lineage upstream. The Schema Evolution agent diffs each new manifest against the last, alongside the quality checks on landed tables, and marks every dbt model that depends on a changed column. The Data Change Review agent diffs the base and head manifests of a PR and posts the column-level blast radius as a comment. The Autonomous Data-Conductor sequences them, holds each change at the autonomy level you set, and closes the incident only after verification.

Every fact carries its source and the time it was observed. When the Semantic Layer definition of plan_revenue and a dashboard's calculated field disagree, the context graph records both and says which one it trusts.

One incident, through dbt

Here is a scenario most dbt teams will recognize. It's an illustration, not a customer case.

A product team renames plan_tier to plan_code in the Postgres billing database at 01:50. Fivetran's PostgreSQL connector handles a rename by adding the new column and leaving the old one in place, so plan_tier stops filling for new rows. At 02:30 the dbt platform job builds stg_billing__plans and fct_subscriptions. The not_null test on plan_tier is configured with severity: warn, so the job succeeds with a warning. At 08:15 finance would ask why revenue by plan is off.

StepWhere it runsWhat happensWho decides
1. DetectSnowflakeAfter the 02:30 run, run_quality_check on the landed table finds plan_tier null on new rows; the on-call engineer confirms the new plan_code column in the landed table.Data Workers, read-only
2. Diagnosedbt manifest, SnowflakeData Workers follows lineage from the source to stg_billing__plans, fct_subscriptions, the plan_revenue metric and the dashboard the team recorded in the context graph, and confirms in the landed table that the nulls start with the 02:05 Fivetran sync.Data Workers, read-only
3. ProposeGitHubThis team has turned on the GitHub pull-request target, so Data Workers opens the fix as a pull request: the staging model maps plan_code, the source YAML and the contract are updated, and a test is added for the new column. The PR carries the blast radius and the rollback path.Data Workers proposes
4. Reviewdbt CIYour CI job runs the tests and compare changes on the PR, exactly as it does for a person's PR.dbt CI
5. ApproveGitHub or SpellbookThe on-call analytics engineer reviews and merges. Branch protection still applies.A named engineer
6. Rerundbt platformAfter approval, the engineer reruns the job so the fix lands before the morning, instead of waiting for the next schedule. Data Workers records the approval and the run.Approved in step 5
7. Verifydbt platform, Snowflake, PostgresThe engineer confirms the run's tests passed. Data Workers re-runs the agreed checks and compares plan revenue in the mart with the source. It records a tamper-evident receipt.Data Workers, read-only
Incident timeline across the stack: what dbt, your team and Data Workers each do, step by step

Every change ran through dbt: its repository, its tests, its CI and its jobs. Data Workers supplied the parts no single dbt surface owns: the diagnosis across Postgres and Snowflake, through the Fivetran sync, the pull request with its blast radius, the sequencing, and the check that the metric matches the source again.

Why doesn't dbt just do this itself?

dbt built a great product for one job: turning transformation into tested, reviewed code. Its design follows from that job, and it's the right design.

dbt Wizard is scoped to the dbt project. Its docs describe it building and refactoring models, generating YAML for tests, docs and semantic models, and investigating job and run failures. Its approvals are per file inside the session ("you approve each file change before it is persisted"). That is exactly what an engineer at the keyboard wants. The remote dbt MCP server is framed for consumption, and the local server's README warns that letting a client run dbt commands "could modify your data models, sources, and warehouse objects", leaving that trust decision with the team.

Writing to production data across systems is a different product category. It needs blast-radius scoping across tools dbt doesn't run, approvals that cover a warehouse change and a dbt model in the same incident, rollback for each step, receipts an auditor can read, context about every other system in the estate, and someone accountable for changes in tools dbt doesn't own. dbt's focus on the project is why it's good at the project. That cross-system layer is the product Data Workers is.

Why a neutral operating layer matters now. With the merger closed, ingestion and transformation come from one company, Fivetran + dbt Labs, and the combined roadmap includes a Fivetran Context Layer in private beta. That's a sensible move for them. Most estates still run a warehouse, an orchestrator and a BI tool from other vendors, and some run Airbyte or custom ingestion beside Fivetran. An operating layer that reads each system through its own API, keeps one approval flow and one audit trail across all of them, and stays useful whatever you choose for each slice protects the choices you've already made. Data Workers reads dbt through its own artifacts and API, the same way it reads Airflow, and catches Fivetran's changes where they land in the warehouse.

dbt's own AI and Data Workers, side by side

dbt Copilot (GA) and dbt Wizard (public preview in the Studio IDE, CLI in public beta, Desktop in private beta) are authoring aids. They help an engineer write a model, generate tests and docs, and understand a failing job inside the project. Data Workers works alongside them.

The split is simple. An engineer writing a new mart uses Wizard or Copilot. An incident that starts in a source system, a schema change that lands overnight, or a check that the CFO's number matches the source runs through Data Workers, with nobody required at the keyboard until the approval. When Data Workers proposes a fix, an engineer can open the change in the Studio IDE and refine it with Wizard; the receipt records the final diff either way.

If you're weighing tools that act on a dbt project with AI, two sibling pages cover the choice directly: Altimate and Data Workers for dbt-aware AI agents, and Recce and Data Workers for data diffs on dbt pull requests. This page is about the wiring.

dbt MCP and dbt artifacts: how to connect today

Data Workers is MCP-native: each agent is an MCP server, so your coding agent (Claude Code, Codex or Cursor) can ask for dbt context, hand off a fix or look up a receipt from the same session. dbt's own MCP server sits next to ours in the same client; they're separate servers with separate jobs. If you use the dbt MCP server for agent access, disable toolsets you don't want exposed with its x-dbt-disable-toolsets header.

Week one: read only.

  • •Point Data Workers at your dbt artifacts: the manifest and catalog, from your CI artifacts or a local target/ directory.
  • •Add a dbt platform service token with Semantic Layer read access for metric definitions.
  • •Give Data Workers read access to the Git repository that holds the project.
  • •Give the warehouse connection a read-only role for profiling and verification queries.
  • •Every agent starts observe-only. The first thing you see is what each recent dbt failure's cause and fix would have been, with its blast radius.

Week two onward: pull request comments, then pull requests.

  • •Turn on blast-radius comments on dbt PRs. Nothing changes in the repo; each PR gets a column-level impact report.
  • •Turn on the GitHub pull-request target so Data Workers opens pull requests on branches, with no permission to merge or push to the main branch.
  • •Route job reruns for specific jobs through the approvals queue; each approved rerun runs on your dbt schedule.
  • •Keep your dbt CI and merge jobs exactly as they are. They test and build every agent PR the same way they build a person's.

The scopes above are an illustration; use the permission sets your dbt plan offers.

Guardrails: approvals, and what dbt owns

What dbt stays responsible for. The transformation code, its Git history and review. Tests, unit tests, contracts and Advanced CI. The Semantic Layer and MetricFlow definitions. Job scheduling, dbt State and builds. Wizard and Copilot for authoring. Data Workers never bypasses any of them.

What Data Workers enforces on top.

  • •Read-only start. New deployments are observe-only. You extend autonomy one domain at a time as the receipts earn trust.
  • •Autonomy per domain, L0 to L4. Documentation updates can move faster while model changes stay at "propose".
  • •Approvals where they belong. Model changes are diffs your reviewers merge (pull requests under branch protection once the GitHub target is on). Anything irreversible needs a named human.
  • •No self-approval. No agent can promote its own work, and every model change waits for a named reviewer.
  • •Receipts. Every change records the diff, the approver, the blast radius, the checks run, the before and after values, and the rollback path in a tamper-evident, hash-chained log.
  • •Rollback. A merged fix reverts like any commit, and the receipt records how.
  • •Least privilege. Data Workers acts with the grants you give it, through dbt's, GitHub's and the warehouse's own permission systems.
The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous

How it fits together

How Data Workers works through dbt: your coding agent on top, Data Workers in the middle, dbt and the rest of your estate underneath

Your team works in its coding agent and reviews in Spellbook Data Catalog, which is in preview. The Data-Agents Swarm does the work, and the Autonomous Data-Conductor runs each fix end to end. dbt stays where the code, tests and jobs live. Nothing migrates, and Data Workers stores metadata and scrubbed facts, not copies of your tables. It connects to the rest of the estate through 50+ connectors.

What changes for your team

Analytics engineers stop starting the morning by reconstructing what broke overnight. They open a proposed change that already names the cause, the blast radius and the rollback path, and they review it the way they review a colleague's. The on-call rotation changes from "find it" to "approve it". Data platform leads get one record of every change across dbt, the warehouse and the orchestrator, which makes audits and post-incident reviews a matter of reading receipts. Finance gets the plan-revenue tile right before they open it. For the step-by-step path from alerts to agents that fix and verify, read the autonomous data platform playbook.

The case for your CFO

The outcome. Your dbt team spends less of its week re-fixing breaks that start upstream, and every number finance relies on comes with a record of what changed and who approved it.

The risk story. At the observe level, agents read your manifest, runs and warehouse and report; they change nothing. At the propose level, they propose diffs (pull requests, once the GitHub target is on) and a named engineer merges. Acting levels are reversible and only for domains you choose. Nothing merges without your reviewers, no agent can promote its own work, and every change has a rollback path. Each receipt holds the diff, the approver, the blast radius, the checks run and the before and after values. There is no migration: dbt, your repo and your CI stay as they are.

Why now. dbt v2 is GA and dbt's own agents are entering the repo through the IDE, the CLI and MCP. Agents are coming to your dbt project either way, so you want one approval flow and one audit trail across all of them.

The first win. Upstream schema changes. Put Data Workers in propose mode for one source system; each change arrives as a dbt diff with its blast radius, and as a draft PR ready for your CI once the GitHub pull-request target is on.

What stays the same. dbt, your team's tools, your reviewers, and the coding agent your engineers already use as the way in.

The pilot path. Start with a pilot on one source system and one project, read-only first, then PR comments, then the GitHub pull-request target. The pilot is credited in full against the first year.

One sentence for upstairs: "dbt stays where our transformation logic, tests and reviews live; Data Workers traces breaks that start outside it and routes every fix back through our normal dbt review with a receipt."

When dbt on its own is enough

If one engineer owns the whole project and incidents almost never leave dbt, dbt's tests, CI and Wizard can carry you. Once breaks start in Fivetran, Airflow or the warehouse, or more than one team ships to the same project, an operating layer that sees the whole estate and verifies every fix starts paying for itself.

FAQ

How do AI agents integrate with dbt? Through dbt's own surfaces. Data Workers reads the manifest, catalog and Semantic Layer definitions through dbt artifacts, and sends changes back as diffs, or as pull requests once you turn on the GitHub target, that your dbt CI tests and your reviewers merge.

Does Data Workers replace dbt? No. Data Workers works through dbt. dbt holds your transformation logic, tests and semantics, and Data Workers changes the project only through changes your team approves and merges: diffs, or pull requests once the GitHub target is on.

Can I use dbt Wizard and Data Workers together? Yes. Wizard and Copilot help engineers author inside the project. Data Workers handles incidents, schema changes and verification across the estate, with one approval flow and one audit trail. An engineer can refine a Data Workers change with Wizard before merging.

Does Data Workers work with the dbt MCP server? Yes. Both run as separate MCP servers in the same client, so a coding agent can call dbt's server for project commands and Data Workers for cross-system context, fixes and receipts.

Who runs our dbt tests? Your dbt CI and jobs, exactly as they do today. Your team reports the results on the incident, and after approval the rerun goes to your dbt job schedule with the approval on record.

Can an agent merge into our dbt repo? No. Data Workers proposes on a branch; your branch protection, CI and named reviewers decide, and no agent can promote its own work.

Does it work with dbt v2, dbt OSS and dbt v1? Data Workers reads dbt artifacts and Git rather than the engine itself, so it follows the project whichever line you run. Confirm your manifest version during the pilot.

Sources

Sources for dbt capabilities and statuses, current as of October 2, 2026: the Fivetran + dbt Labs merger completion, dbt v2 is GA, the dbt release notes (product naming updates, adapters, dbt State, Wizard, Semantic Layer spec, Ossie), the dbt Summit 2026 announcements, dbt Wizard, dbt Copilot, the dbt MCP server, local and remote MCP, remote MCP setup, the remote dbt MCP server launch, dbt Agent Skills, Advanced CI, model contracts, MetricFlow open-sourcing, Apache Ossie and the Fivetran PostgreSQL connector (column renames). Product names and statuses change quickly; if we've got something wrong, tell us and we'll fix it.