Industry
Industry9 min readBy The Data Workers Team

Data Workers for analytics engineers

What Data Workers changes in an analytics engineer's week: blast radius before a dbt model change, upstream breaks traced and fixed as dbt diffs, metric definitions with owners and docs kept current, while you keep the models and the merge.

Data Workers takes the investigation and the paperwork out of an analytics engineer's week: it shows the blast radius of a dbt model change before you merge, traces an upstream break to its source and proposes the fix as a dbt diff, and keeps metric definitions and docs current with a named owner on each. You keep owning the models, the reviews and every merge: each agent action runs at the autonomy level your team sets, waits for a named approver wherever that level asks for one, and leaves a receipt.

Data Workers is the agentic data platform. For an analytics engineer, it works from the coding agent you already use, from the pull request you are already reviewing, and from Spellbook, its catalog and approvals app.

Key takeaways

  • •Blast radius before the merge. Ask your coding agent over MCP, or read the review comment Data Workers posts on the pull request: every model, metric, Explore and dashboard a change reaches.
  • •Upstream breaks arrive diagnosed. A renamed source column is caught by a quality check in the warehouse, traced to its source, and handed to you as a dbt diff with dry-run counts, not as a 02:00 test failure.
  • •Definitions have owners. Each metric has one approved definition and a named owner, and a conflicting copy goes to that owner instead of to a Slack argument.
  • •Docs stay current. Data Workers drafts model and column docs into your dbt project as a proposal you approve.
  • •You keep the job that matters. Models, tests, reviews and merges stay yours; at L2 propose the agents propose and you decide, and you set the level per domain.

Your week today

dbt Labs' 2026 State of Analytics Engineering report (third-party; 363 responses collected Dec 5, 2025 to Feb 1, 2026) puts it plainly: "A majority of respondents report spending most of their time maintaining or organizing datasets." It also found that "while 72% prioritize AI-assisted coding, only 24% prioritize AI-assisted pipeline management, including testing, observability, and quality controls," and that "71% are concerned about hallucinated or incorrect data reaching stakeholders."

You write and refactor dbt models, often with Claude Code or Cursor, and review pull requests from teammates and analysts. In between, the interrupts arrive:

  • •A source changes under you. A Salesforce admin renames a field, Fivetran lands the new column in the warehouse, and a staging model quietly fills with nulls.
  • •"Why doesn't this number match?" Finance and sales see different pipeline totals, and you spend an hour in SQL finding which definition each dashboard used.
  • •Metric debates. Net revenue has three definitions and the dbt Semantic Layer holds one. The same report found 41% cite ambiguous data ownership as an ongoing challenge.
  • •Tests, docs, reruns. Written when there is time; kicked off by hand once someone notices.

The coding agent made writing models faster. It did not take on the running of them.

The same week with Data Workers

Data Workers runs the operations around your dbt project. It reads the project as context (manifest, run results, catalog, Semantic Layer definitions), watches the warehouse and sources underneath, and hands you proposals with evidence; Data Workers + dbt covers the wiring.

Comparison matrix of Your analytics stack and Data Workers on the outcomes a data leader buys

What it takes off your week:

  • •Impact analysis. blast_radius_analysis and assess_impact find every downstream model, metric, Explore and dashboard a change reaches. On a pull request, Data Workers reviews the change with a blast radius and an impact diff in one review comment it keeps current, and can block a risky change from merging.
  • •Upstream diagnosis. A breaking change in a source table is caught in the dbt manifest diff, in pull request review, or when run_quality_check fails on the table it lands in; diagnose_incident and get_root_cause trace a failing or drifting model back to the change that caused it. The fix comes back as a diff to your staging model, with the test that would have caught it.
  • •Definition lookups. explain_table returns a table's definition, lineage, documentation and trust score in one answer.
  • •Docs. generate_documentation drafts model and column docs with provenance, written into your dbt project as docs blocks and schema.yml descriptions through the approvals queue. A description a person wrote stays the authoritative one.
  • •Reruns. After you approve a fix, Data Workers queues the rerun through Airflow, Dagster or dbt's job API and checks the result.

What you still own and decide:

  • •The models and the merge. Data Workers proposes each model change as a diff for you to merge. For the docs it writes into your dbt project, it opens the pull request itself when your team turns on the GitHub pull-request target; local files are the default.
  • •The approval level. Each domain sits on the ladder L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous. Start your dbt domains at L1 or L2; autonomy levels L0 to L4, explained walks through each step.
  • •Definitions. You or the metric's named owner sign a definition with define_business_rule and mark_authoritative. No agent can promote its own work.
  • •The receipt. Every change records the diff, who approved it, when, the blast radius and the rollback path. How approvals work and how to roll back an agent change cover both sides.

How Data Workers fits the tools you already use

Your coding agent. Data Workers agents are MCP servers, so they sit inside Claude Code, Cursor, Codex or Copilot next to the dbt MCP server. Ask "what breaks if I change the grain of fct_orders?" and the answer comes from the context graph, with sources. Setup follows the client setup docs: clone the open-source core, install it, and add a start-agent.sh entry per agent to your client's MCP config.

# Example: register the catalog, schema and quality agents in Claude Code
git clone https://github.com/DataWorkersProject/dataworkers-claw-community.git
cd dataworkers-claw-community && npm install --ignore-optional
claude mcp add --scope user dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
claude mcp add --scope user dw-schema -- "$(pwd)/start-agent.sh" dw-schema
claude mcp add --scope user dw-quality -- "$(pwd)/start-agent.sh" dw-quality

The core runs on a built-in sample estate; in a Data Workers deployment the same agents point at your estate and carry the approvals and receipts. List the tools with your client's own command, such as /mcp. Data Workers with your coding agents covers each client, and our guide to using Claude Code with dbt covers the authoring side.

dbt and the Semantic Layer. dbt stays where your models and metrics live. Data Workers imports the semantic manifest as context, checks each metric against the warehouse, and flags a Looker or Tableau calculation that has drifted from the dbt definition; you're on the dbt Semantic Layer walks that loop.

Warehouse, orchestration and tests. Snowflake, BigQuery, Databricks, Airflow, Dagster and Great Expectations connect natively; Fivetran, Elementary, Hex and Sigma connect over their APIs or MCP servers today. Looker and Tableau are read, by design: BI stays your BI team's system.

Spellbook. Spellbook Data Catalog (in preview) holds the proposals, approvals and receipts, and lets stakeholders look up a definition before they ask you.

The metrics you are judged on

MetricHow Data Workers moves itWhere you see it
Test pass rate and freshness SLAsSource changes caught before the build; run_quality_check and get_anomalies on the models that matterIncident history and receipts per model
Stakeholder trust in numbersOne approved definition per metric; drifted copies routed to the ownerOpen contradictions and their resolution time
Cycle time for requests"Which table, which definition" answered with sources, in your editor or in SpellbookRequests closed without a Slack thread
Documentation coverageDocs drafted with provenance into your dbt project for approvalDocumented models and columns over time
Warehouse cost of modelsSnowflake credits attributed to each dbt model through query tags; fixes drafted for the ownerSpend per dbt model; the design target is 25 to 40% lower warehouse spend

Measure against a baseline you take before the pilot. How to measure AI data agents sets out the method, and the ROI calculator turns your numbers into a range.

A worked example: a renamed Salesforce field before the Monday pipeline review

This is an illustration, not a customer case. Fivetran syncs Salesforce into Snowflake, dbt models it on an Airflow schedule, and sales reviews pipeline by region in Looker every Monday at 09:00. The sales domain runs at L2 propose; the analytics engineer who owns the staging models approves.

TimeSystemWhat happened
06:10SalesforceAn admin renames Region__c to Sales_Region__c
06:40FivetranThe sync lands the new column in Snowflake; the old one stops filling
06:52Data Workersrun_quality_check fails in Snowflake: a new Sales_Region__c column, and region is null in new rows
06:58Data WorkersBlast radius: stg_salesforce__opportunities, fct_pipeline, the pipeline_by_region metric and the Pipeline Explore
07:00Airflow, dbtThe build passes; the not_null test on region is set to warn
07:06Data WorkersProposes a staging-model diff that reads the new column, plus a test, with dry-run counts
07:07SlackThe model owner gets the diff and the blast radius
08:15Analytics engineerReviews the diff in Claude Code, opens the pull request; CI passes; merges
08:22Data WorkersQueues the rerun through Airflow
08:41Data WorkersRegion nulls back to zero; receipt filed
09:00LookerThe Monday review reads correct regions
Incident timeline across the stack: what Your analytics stack, your team and Data Workers each do, step by step

Without Data Workers, the first signal would have been a sales leader asking why half the pipeline had no region. With it, the engineer reviewed one diff before the meeting.

How to bring it to your team

You do not need a platform project to start:

  • •Try it on the sample estate. Register two or three agents in your coding agent and ask the questions you answer every week.
  • •Pick one domain and one pain. Upstream source breaks in your sales or finance models are a good first choice: frequent, visible and easy to measure.
  • •Start at L1 observe. Connect the dbt manifest and a read-only warehouse role. For two weeks, Data Workers logs what it would have flagged and proposed, and you review that record against what actually happened.
  • •Move to L2 propose for that domain. Fixes arrive as diffs you merge. Turn on the GitHub pull-request target for docs when the proposals have earned it.
  • •Show the receipts. Bring the record to your team retro: breaks caught before the build, time to merge a fix, definitions settled. How to get your team to trust AI agents has the playbook.

When your lead asks about risk, send is it safe to let AI agents change production data and where your data goes: the agents run in your infrastructure, your data stays in your systems, and the hosted Conductor sees workflow metadata only. Build it ourselves answers "why not Claude Code and a few MCP servers", and the ROI of agentic data operations makes the case in money. The next step is to start with a pilot (what a pilot looks like; the pilot is $7,500 one-time, credited in full against the first year, on /pricing/). Your colleagues have their own pages: Data Workers for data engineers and Data Workers for data analysts.

FAQ

Will it write plausible SQL with the wrong grain? Data Workers works from the context graph: approved definitions with named owners, lineage, and the tests on each model. At L2 propose, every change reaches you as a diff with dry-run counts before anything runs. How Data Workers avoids hallucinations on your data covers the mechanism.

Will it push to main? No. It proposes model changes as diffs for you to merge; for docs it opens a pull request only when your team turns on the GitHub target. Your branch protection and CI still apply.

I already use Claude Code or Cursor. Why add this? You keep them. Data Workers runs as MCP servers inside them, adding the blast radius, definitions, upstream diagnosis and approvals a coding agent does not hold on its own.

Does it replace dbt's own AI features or the dbt MCP server? No. dbt's tools help you author inside the project. Data Workers works across the estate around it: sources, warehouse, orchestrator and BI.

Will it replace analytics engineers? It takes the toil and leaves the design, the definitions and the decisions with you. Will Data Workers replace my data team? and who owns the agents answer this directly.

Do I need the dbt platform, or does dbt Core work? Both. Data Workers reads dbt's manifest and catalog from your project, whichever runs your builds.

Sources

  • •dbt Labs, 2026 State of Analytics Engineering Report (third-party survey; 363 responses, Dec 5, 2025 to Feb 1, 2026): https://www.getdbt.com/resources/state-of-analytics-engineering-2026 (checked Oct 2, 2026)
  • •Data Workers client setup documentation: https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)
  • •Data Workers public repository tools (blast_radius_analysis, assess_impact, detect_schema_change, diagnose_incident, get_root_cause, explain_table, generate_documentation, define_business_rule, mark_authoritative, run_quality_check, get_anomalies, check_freshness): https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)
  • •Data Workers product repository, data-workers-agent-swarm main @ 0c2491e3: agents/dw-review (pull request review and review comment), agents/dw-context-catalog dbt doc write-back (governed proposal; local files by default, GitHub pull-request target when turned on) (checked Oct 2, 2026)
  • •Data Workers product pages: https://dataworkers.io/product/autonomous-data-conductor/ , https://dataworkers.io/product/data-context-wizard/ , https://dataworkers.io/product/spellbook-data-catalog/ (checked Oct 2, 2026)
  • •Data Workers security page: https://dataworkers.io/security/ (checked Oct 2, 2026)
  • •Data Workers pricing: https://dataworkers.io/pricing/ (checked Oct 2, 2026)