Can't we just build this ourselves with Claude Code and a few MCP servers?
You can build a data agent demo with Claude Code and vendor MCP servers in a week. The year goes into lineage, approvals, rollback, receipts and evaluation. Build on Data Workers' Apache 2.0 core instead.
Yes, you can build the demo yourselves in about a week: Claude Code, the Snowflake, Databricks and dbt MCP servers, and a good prompt will answer real questions about your warehouse. What takes a year is everything around the agent that makes production changes safe and the context true, and Data Workers is that layer, with an Apache 2.0 core you build on instead of starting from scratch while Claude Code, Cursor and Codex keep calling it over MCP.
Building is a real option, and we respect the teams who choose it. The open core exists so you can.
Key takeaways
- •The demo is the easy part. A coding agent plus vendor MCP servers answers questions and writes code within days. Every vendor has made that path good.
- •The year goes into the layer around the agent. Cross-engine lineage, blast radius, approvals and autonomy per domain, rollback, verification, receipts, connector upkeep and evaluation.
- •Vendor MCP servers are governed for their own platform, by design. Writing across systems with approvals and rollback is a different product.
- •Data Workers is that product, with an Apache 2.0 core. Every agent in the core is the same agent the platform runs. Fork it, extend it, and let the platform carry the production layer.
- •Your coding agents stay. Claude Code, Cursor and Codex call Data Workers over MCP, the same way they call anything else.
What the week-one demo gets you
Start where your engineers would. An engineer runs claude mcp add for the dbt MCP server, connects the Snowflake-managed MCP server through Snowflake OAuth, and adds the Databricks SQL server for the lakehouse. By Friday the agent explains a model, lists a table's parents in dbt, runs a query and drafts a pull request. It's impressive, and it's real.
That demo works because the vendors made careful choices. Snowflake's managed MCP server (generally available) exposes Cortex Agents, Cortex Analyst, Cortex Search, SQL execution and stored procedures under role-based access control, and Snowflake recommends a Cortex Agent as the only client-facing tool for governed business questions, with direct SQL on a separate server under a least-privileged role. Databricks managed MCP servers (Public Preview) cover Genie One, Genie Agents, AI Search, Databricks SQL and Unity Catalog functions, and Unity Catalog enforces permissions so agents reach only what you grant. The remote dbt MCP server is built for data consumption; the self-hosted server can run dbt build and dbt run, and dbt's README says plainly that those commands "could modify your data models, sources, and warehouse objects."
Why don't Snowflake, Databricks and dbt just do this?
Focus and risk. Each vendor built its MCP server for its own platform, governed by that platform's grants and steered toward governed answers. That is the right design for a vendor. A change that spans Postgres, Fivetran, dbt, Snowflake and Looker needs a plan, a blast radius, an owner's approval, a rollback and a record across tools the vendor doesn't own. Taking on that liability is a different product category, and in a DIY build it lands on your team.
The MCP specification leaves the same work to the application. The 2025-11-25 spec says there "SHOULD always be a human in the loop with the ability to deny tool invocations" and that clients "MUST consider tool annotations to be untrusted unless they come from trusted servers." The protocol defines how an agent calls a tool. Who approves a write, how it rolls back and what record it leaves belong to whoever builds the application.
What the year goes into
None of this is exotic. All of it has to exist before anyone lets an agent write.
| Capability | What you build yourselves | What Data Workers ships |
|---|---|---|
| Cross-engine lineage and context | Stitch each vendor's lineage into one graph and keep it fresh | One governed context graph across engines in Data Context Wizard |
| Blast radius | Walk that graph for every proposed change | Column-level impact through blast_radius_analysis before any change |
| Approvals and autonomy | Route each write to the right owner, per domain | Approval workflows with timeouts and escalation, autonomy set per domain |
| Rollback | Write an undo for each change type | Operation history and undo for schema changes, pipeline deploys and data promotions |
| Verification | Check the fix held downstream | The Conductor's verify step, with before and after values |
| Receipts | Record who approved what, and why | A receipt on every change: diff, approver, checks, rollback path |
| Connectors | Track every vendor API and MCP release | A maintained connector library across warehouses, catalogs and ops tools |
| Evaluation | Build golden sets and regression tests per model upgrade | An evaluation harness with golden datasets, plus an observe-only stage where the ladder logs what each action would have needed |

The DIY path is strong on the top three rows, and Data Workers matches them because it runs through the coding agent and model you already chose. The difference is everything below them.

The 12 months are an estimate, and these are its assumptions: two to three engineers working on it part-time, Snowflake or Databricks with dbt, Airflow and one BI tool, and production writes in scope. A larger estate stretches the months; the list of work stays the same. Upkeep never ends: vendors ship new MCP tools, rename products and change APIs, and someone owns each break. Build-vs-buy spreadsheets miss that cost most often. The ROI page and the ROI calculator put numbers on your own estate.
How Data Workers does it
Data Workers is the agentic data platform: it runs the whole data lifecycle and builds on the tools you already run. Four parts carry that layer.
Data Context Wizard holds the context. One governed graph of lineage, owners, freshness and usage across Snowflake, Databricks, BigQuery, dbt, orchestrators and BI. Public tools such as trace_cross_platform_lineage, explain_table and search_across_platforms answer from it.
The Data-Agents Swarm does the work. More than 20 specialist agents for incidents, quality, schema, pipelines, cost, governance and migration, each an MCP server. The incident agent runs diagnose_incident and get_root_cause; the schema agent runs assess_impact and generate_migration; the quality agent runs run_quality_check.
The Autonomous Data-Conductor runs the loop. Detect, diagnose, fix, review, verify and remember, with a receipt on every change. Approvals go to a named owner per domain, with timeouts and escalation. Credentials are scoped per agent. A global kill switch halts every agent at once. At L1 observe the permission ladder logs what approval each action would have needed, and your team reviews the agents' records against what it did, before any write is turned on.
Spellbook Data Catalog is where people look. Asset pages, lineage, approvals, incidents and the audit trail in the browser (in preview).
Autonomy is set per domain on one ladder, and it only moves up when the receipts show the agents were right.

Your team can run access requests at L3 while schema changes stay at L2, and lower any domain at any time. The safety page covers what an agent can and can't do at each level, and how approvals work shows a request end to end.
Build on it, not from scratch
The open core is Apache 2.0 and lives on GitHub. Every agent in it is the same agent the platform runs. Fork it, patch it, add a tool for your internal system, and run it with no account and no license key. It ships with a sample estate, so an engineer can drive the full tool surface from Claude Code right away:
# Example: add the incident agent to Claude Code (from /opensource-docs/client-setup/)
git clone https://github.com/DataWorkersProject/dataworkers-claw-community
claude mcp add dw-incidents -- /path/to/dataworkers-claw-community/start-agent.sh dw-incidentsThe platform is built on that core and adds your estate in place of the sample, governed writes inside approval workflows, your context graph, the Conductor and Spellbook, in your infrastructure on every tier. Your engineers build what is specific to your business; the production layer comes with the platform. See the open-source docs and what the platform adds.
Your coding agents don't change. The integration guide has the config for Claude Code, Cursor and Codex, and You're on Claude covers Claude Team and Enterprise seats.
One incident, two ways
An illustration, not a customer case. At 01:50 a product team renames plan_tier to plan_code in Postgres. Fivetran syncs it to Snowflake at 02:05. The dbt run at 02:30 succeeds and writes null plans into fct_subscriptions. Airflow marks the DAG run green. At 08:10 the revenue tile in Looker drops.
The DIY path. An engineer opens Claude Code at 08:40, asks the dbt MCP server for the model's parents, queries Snowflake through its MCP server, and finds the rename by 09:30. The agent drafts the dbt fix. The engineer checks by hand which Looker dashboards read the column, gets a review, merges, reruns and spot-checks the tile. It works, because a good engineer did the parts the agent couldn't see, and the record lives in a chat session and a pull request.
The Data Workers path. The Conductor detects the null spike at 02:40, traces it through Fivetran to the Postgres rename, and lists the blast radius: two dbt models, one Airflow DAG, three Looker tiles. It proposes the dbt change as a diff and routes it to the revenue domain owner, who approves at 07:30 from Spellbook or from Claude Code. The backfill runs, the verify step re-runs the checks on the rebuilt tables and records before and after results, and the receipt records the diff, the approver, the checks and the rollback path. Finance opens a correct tile at 08:10.
Same engineers, same coding agent, same dbt project. The difference is the layer around the agent.
Where the other paths fit
- •Build on a coding agent and vendor MCP servers. Best for teams with platform engineers to spare and a narrow, read-heavy scope (our warehouse MCP server tutorial shows the pattern). Building on the open core gets you further on the same budget.
- •Platform-native agents. Genie, Cortex agents and the Databricks and Snowflake coding agents are excellent inside their own platform, under its governance. Data Workers works across them; see what an agentic data platform is.
- •Catalogs and observability. Catalogs hold the inventory and observability tools raise the alarm. Data Workers reads both and closes the loop with a fix.
- •Coding agents on their own. Claude Code, Cursor and Codex are where engineers work. Data Workers is what they call when the work touches production data.
The case for your CFO
The outcome is a data team that ships fixes to production data with an audit trail, without first spending a year building the machinery to make that safe. Building it in-house means two to three engineers for most of a year, by our estimate, then permanent upkeep, and those engineers aren't working on the data products the business asked for.
The risk story has to be answered either way. Agents propose before they act. A named owner approves anything irreversible. Every change carries a receipt with its blast radius, checks and rollback path. Autonomy is set per domain and rises only when the record earns it. Data Workers ships that layer, runs in your infrastructure, and needs zero migration: your warehouse, dbt project, orchestrator and coding agents stay as they are.
Why now: your engineers already use coding agents and the vendors already ship MCP servers, so agents will touch your data. The question is whether they do it with approvals and receipts.
The first win is cross-system incident triage, approved by the domain owner, inside the pilot. Start with a pilot: $7,500 one-time with a named engineer, and the pilot is credited in full against the first year. After it, Scale starts from $1,000/month and Enterprise from $3,000/month, billed annually, with unlimited seats, no usage meter, no markup on model spend, and your own model. See pricing and the ROI calculator.
The sentence to repeat upstairs: "We'll build on an open-source core instead of spending a year building the safety layer ourselves, and our engineers keep the coding agents they already use."
FAQ
Can we use the open core and never buy anything? Yes. The core is Apache 2.0, with no account and no license key. Teams that need governed writes against their own estate, the context graph and the Conductor add the platform.
We already run the Snowflake, Databricks and dbt MCP servers in Claude Code. Do we give anything up? No. Data Workers is one more MCP server in the same client. Vendor servers answer on their own platform; Data Workers works across platforms and carries changes through approval, verification and rollback.
What happens when a vendor changes its API or MCP server? In a DIY build, your team finds the break and fixes it. With Data Workers, the connector library is maintained as part of the platform.
How do we know the agents are right before they write? At L1 observe the ladder logs what each action would have needed and your team reviews the agents' records against its own fixes, and an evaluation harness with golden datasets tests the agents. Autonomy rises only when the receipts support it. See how to measure AI data agents.
Where does our data go? Data Workers runs in your infrastructure on every tier, and you bring your own model key. The security and deployment page covers the details.
What if we start building and change our minds later? Nothing is wasted. Tools you write against the open core run on the platform, because the platform is built on that core. If you already rolled out AI assistants, see the assistants hub and bring your own context.
Sources
- •Snowflake, Snowflake-managed MCP server: generally available; tool types (Cortex Agent, Cortex Analyst, Cortex Search, SQL execution, UDFs and procedures), RBAC, Snowflake OAuth or External OAuth, and the recommendation to expose a Cortex Agent as the only client-facing tool for governed questions with direct SQL on a separate least-privileged server. Checked October 2, 2026.
- •Databricks, Databricks managed MCP servers: Public Preview; Genie One, Genie Agent, AI Search, Databricks SQL and Unity Catalog functions servers; Unity Catalog enforces permissions; Claude Code and Cursor named as clients. Page last updated September 21, 2026; checked October 2, 2026.
- •dbt Labs, dbt MCP server: self-hosted and remote servers; remote is for data consumption and runs no CLI commands. Checked October 2, 2026.
- •dbt Labs, dbt-mcp on GitHub: CLI tools (
build,run,test,show,clone) and the warning that they could modify data models, sources and warehouse objects. Checked October 2, 2026. - •Model Context Protocol, Specification 2025-11-25: Tools: human in the loop and untrusted tool annotations. Checked October 2, 2026.
- •Anthropic, Claude Code MCP docs:
claude mcp add, scopes and project-server approvals. Checked October 2, 2026. - •Cursor, MCP docs:
mcp.jsonconfiguration and approval before MCP tool use by default. Checked October 2, 2026. - •OpenAI, Codex MCP docs: Codex CLI and IDE extension support MCP servers via
codex mcp addorconfig.toml. Checked October 2, 2026. - •Data Workers, open-source core (Apache 2.0, public tool registrations), client setup, what the platform adds, pricing and licensing, pricing and Autonomous Data-Conductor. Checked October 2, 2026.