What if Data Workers goes away? What you own, what you subscribe to, and how to leave
Your data operation keeps running. The core is Apache 2.0, the agents run in your infrastructure on your model account, and your context graph, receipts and changes live in systems and open formats you own.
You keep running. The core of Data Workers is Apache 2.0 open source that you can run, modify and fork, the agents run in your infrastructure on your own model account, and everything they build for you lives in systems you own: your context graph and receipts export in open formats, and every change it made already sits as normal SQL, dbt changes and config diffs in your own repos and warehouse.
That is true whether Data Workers the company went away or you simply decided to leave. Nothing in your data path runs through us. The one piece we host, the Autonomous Data-Conductor, holds workflow metadata (goals, routines and run records) and never queries your warehouse or holds your rows, credentials or model keys. Your warehouse, pipelines, catalog and BI keep doing their jobs, because Data Workers works on top of them with zero migration. This page shows exactly what is yours, what the subscription adds, and the exit checklist a risk reviewer can test during the pilot.
Key takeaways
- •The core is Apache 2.0. Eleven specialist agents, each an MCP server, in a public repo with an Apache 2.0 license. No licence-key check and nothing that expires.
- •The agents run where you run them. The agents run in your infrastructure on every tier, and Enterprise adds a dedicated VPC, your own cloud or on-premise. The Conductor is hosted by us and holds workflow metadata, never rows.
- •Your context exports in open formats. The context graph exports to JSON-LD with its audit chain, or to an OKF markdown bundle any tool can read.
- •Your receipts are verifiable on their own. The audit log is a SHA-256 hash chain you verify with
verify_global_hash_chainand keep in your own append-only storage. - •Your models are your accounts. Anthropic, OpenAI, Bedrock, Vertex, Azure OpenAI or a local model, on your key and your invoice.
- •Your changes are already landed. Agents propose SQL, dbt changes and config diffs the way a colleague would, and approved changes live in your repos and warehouse.
What you own, and what you subscribe to
Data Workers is the agentic data platform, and it is designed so that the value it creates lands in your systems while the work happens. The subscription pays for the managed platform that keeps that work going. The split is published on our pricing page: if you leave, you keep the open-source core and everything you built on it, your context graph, your configuration and the eleven open-source agents, running on your own infrastructure. What stops is the managed platform: the proprietary agents, the hosted Conductor and the control plane. Precisely: the continuous loop stops, so nothing detects, routes or proposes new work on its own, and governed writes across the estate stop. What keeps running is your whole stack, every change already applied, and the eleven open-source agents, which keep reading, analysing and recommending whenever your engineers call them from their MCP clients.

The Apache 2.0 core. The open-source core is a set of MCP servers for cataloguing, quality, schema, pipelines, incidents, governance, usage, observability, orchestration, connectors and ML. You clone it, build it and run it as local processes over stdio, launched by the MCP client you already use, as the client setup guide shows. The pricing page states it plainly: the core is whole, with no enterprise-only files hidden in the source tree, no licence-key check and nothing that stops working after a trial. Apache 2.0 lets you modify it, fork it and run your fork for as long as you like, with a patent grant from every contributor.
Your context graph. The graph is the shape of your estate: tables, columns, lineage across engines, owners, definitions, quality and freshness signals, and the facts your team approved. Data Workers exports it in two open formats. The JSON-LD export carries entities, edges and the audit chain, plus a manifest that separates what a fresh crawl of your systems would rebuild from what only your people could have decided: approvals, supersessions, reconciled definitions and edit history. That second list is the institutional memory, and it is the part worth protecting. The export checks itself for completeness: every entity the manifest lists must be in the file. The OKF export writes the same graph as a markdown bundle, one file per concept with YAML frontmatter and cross-links as plain links, so another catalog or a person with a text editor can read it. More on keeping context portable is in Bring your own context.
Your receipts and audit log. Every applied change carries a receipt: the diff, the approver, the blast radius and the rollback path. Each action is written to an audit log where each entry's SHA-256 hash covers the previous entry, so the chain proves itself without us. get_audit_trail returns the entries, verify_global_hash_chain checks the chain end to end, get_usage_activity_log shows who called which tool, and generate_audit_report assembles the record for a review. The security page covers where it lives.
Your changes. Agents do not keep a private copy of your logic. They propose changes in the same form a colleague would open for review: the SQL, the dbt model change, the config diff, with the reasoning. generate_migration writes forward SQL with its rollback SQL. Generated dbt documentation goes back into your own dbt project, either into the project directly or as a pull request. Once approved, a change is a commit in your repo or a statement your warehouse ran, and it stays there.
Your model account. You bring the model: Anthropic, OpenAI, Amazon Bedrock, Google Vertex AI, Azure OpenAI, or a local model through Ollama or vLLM. Your provider bills you directly at your rate, so leaving changes nothing about that contract. The cost side is in Which model does Data Workers use, and what does it cost to run?.
What the subscription adds. The platform agents that make governed writes across the estate, the hosted Autonomous Data-Conductor that runs the detect, diagnose, fix, verify loop continuously, Context Wizard authoring in the browser, the connectors we maintain as vendor APIs change, a named engineer alongside your team, and on Enterprise the Spellbook Data Catalog (in preview) where people review, approve and audit. We earn those every year; your assets do not depend on them.
Why the design works this way
Every architectural choice that makes Data Workers safe to adopt also makes it safe to leave. The agents run in your infrastructure and the hosted Conductor never queries your warehouse, so there is no hosted copy of your data to retrieve; the receipt and audit entry for every change already sit in your deployment. It reads through grants you issue, so leaving is a revoke. It uses your identity provider through JWKS and issues no tokens of its own, so your directory never depended on us. It writes through your approval flow into your repos, so there is no shadow state. And because pricing has no usage meter, there is no stock of prepaid credits to burn down or forfeit; Why is there no usage meter? explains that choice.

The autonomy dial is part of the exit plan too. Autonomy is set per domain, from L0 manual through L1 observe, L2 propose, L3 act reversibly to L4 autonomous. The first step of any exit is to set every domain to L1 observe, which stops all writes cleanly while the record is exported. After that, the open-source core keeps reading, analysing and recommending in your engineers' coding agents, and your team applies what it recommends. How individual changes are reversed is covered in How do we roll back an agent's change?, and the write-safety model in Is it safe to let AI agents change production data?.
A worked example: one week to leave
This is an illustration, not a customer event. The Data Workers agents run in the customer's AWS account, the Autonomous Data-Conductor is hosted by Data Workers, the model is Claude on Amazon Bedrock in the same account, the warehouse is Snowflake, transformations are a dbt project on GitHub, and Okta is the identity provider.

On day one the data platform lead sets every domain to L1 observe. Data Workers runs verify_global_hash_chain and generate_audit_report, then exports the context graph to JSON-LD with the audit chain and the manifest of human decisions. On day two the export and the audit log land in an S3 bucket with Object Lock, the team's own append-only store. There is nothing to pull out of GitHub: every approved dbt change is already a merged commit. On day three the team revokes the Data Workers role in Snowflake. The masking policies, grants and tables approved through Data Workers stay in place, because they were always the customer's objects. Okta unassigns the Data Workers app; the directory is untouched. On day four nothing happens at Bedrock, because the model account was the customer's all along. On day five the subscription ends and the hosted Conductor stops; it held goals, routines and run records, and the receipt for each run is already in the archived audit log. The same day an engineer clones the Apache 2.0 core and adds the agents to Claude Code, and they read the same estate over stdio.
The exit checklist
Run this during the pilot, not after you decide to leave. A rehearsed exit is the strongest answer to a vendor-continuity review.
| Step | What you do | What you keep |
|---|---|---|
| 1. Pause writes | Set every domain to L1 observe | A clean stopping point, recorded in the audit log |
| 2. Verify the record | Run verify_global_hash_chain and generate_audit_report | Proof the chain is intact at exit |
| 3. Export context | JSON-LD export with the audit chain, plus the OKF bundle | Graph, lineage, approved facts and the human-decision manifest |
| 4. Archive receipts | Copy the audit log and receipts to your append-only storage | Every diff, approver, blast radius and rollback path |
| 5. Check your repos | Confirm approved changes are merged and migrations carry rollback SQL | All changes, as normal code and SQL |
| 6. Revoke access | Drop the agents' warehouse roles and unassign the IdP app | Your grants and directory, unchanged |
| 7. Keep the core | Clone or fork the Apache 2.0 repo and add it to your MCP client | Eleven agents that keep reading and recommending |
| 8. Keep the model | Nothing to do | Your provider account, rate and terms |
The alternatives, and how each one handles leaving
Each option has made sensible portability choices; ask each the same question.
Build it yourself. You own all of it, and you also maintain all of it: lineage across engines, approvals, rollback, receipts and connectors as vendor APIs change. Can't we just build this ourselves? maps that work. With Data Workers you get the same ownership of the core under Apache 2.0, and the platform layer is maintained for you.
Catalogs. Open-source catalogs such as DataHub and OpenMetadata are Apache 2.0 (GitHub, checked Oct 2, 2026). Data Workers works with the catalog you already run and writes approved facts back into it, so your catalog stays the system of record.
Observability. Elementary's open-source package is Apache 2.0 (GitHub, checked Oct 2, 2026); hosted observability services keep monitor history in their service, a sensible design for a hosted product. Data Workers keeps its incident receipts in your audit log, in your deployment.
Platform-native agents. Databricks Genie (Genie One, Genie Agents, Genie Code) grounds every answer in your data, governed through Unity Catalog (Databricks docs, updated Sep 18, 2026). Snowflake lets you read a semantic view back out as YAML or as OSI YAML (Snowflake docs, checked Oct 2, 2026). Strong portability for one platform; Data Workers carries context and receipts across all of them in one export.
Coding agents. Claude Code keeps a project's MCP configuration in a .mcp.json file you commit to the repo (Claude Code docs, checked Oct 2, 2026), and OpenAI's Codex CLI is Apache 2.0 (GitHub, checked Oct 2, 2026). Because the Data Workers core is a set of MCP servers, it moves with whichever client you choose.
The case for your CFO
The outcome: the data team gets agents that do real work across the estate, and the company carries no stranded asset if the relationship ends. Everything of lasting value, the context graph, the receipts, the changes and the model contract, sits in systems the company already owns.
The risk story is short. The agents run in your infrastructure, the hosted Conductor holds workflow metadata and never your rows, Data Workers reads through grants you can revoke, writes only through approvals into your repos and warehouse, and keeps a hash-chained audit trail you can verify without us. The core is Apache 2.0, so the agents can keep running under your own name. There is no data migration to unwind, because there was no migration.
Why now: agents are arriving across the stack, and the exit plan is easiest to settle before they hold years of institutional memory. A JSON-LD export with a manifest of every human decision is the version of that memory you can take anywhere.
The first win is the rehearsal itself: run the exit checklist during the pilot, show the reviewer the verified chain and the export, and close the vendor-continuity question with evidence. What stays the same: your warehouse, catalog, identity provider, permission system and model provider.
The path is a pilot at $7,500 one time, and the pilot is credited in full against the first year. Scale starts from $1,000 a month and Enterprise from $3,000 a month, billed annually, with unlimited seats, no usage meter, no markup on model spend, bring your own model and an Apache 2.0 core. See /pricing/, model the return in the ROI calculator, and read the ROI of agentic data operations.
The sentence to repeat upstairs: "Everything Data Workers builds lands in systems we own, and we rehearsed leaving before we signed."
FAQ
What keeps working the day the subscription ends? Your whole data stack, because Data Workers is not in its path, and the Apache 2.0 core, which keeps reading, analysing and recommending in your engineers' MCP clients. The hosted Conductor, the platform agents' governed writes and the browser control plane stop, so the core recommends and your team applies.
Can we fork the core and change it? Yes. The repo is licensed under Apache 2.0, which permits modification and redistribution, including inside your company, as long as you keep the license and NOTICE files.
Can another tool read our context graph export? Yes. The JSON-LD export is open JSON with entities, edges, the audit chain and the manifest, which any tool can parse, and the OKF bundle is plain markdown with YAML frontmatter, one file per concept. The OKF bundle also imports back into Data Workers, through the same governed write gate as any other source, if you return.
Do we lose the audit trail? No. The log is a SHA-256 hash chain, so it proves its own integrity after export. Archive it in your own append-only storage and any auditor can check it.
What do we have to unwind in the warehouse? Only the agents' access. Policies, grants, tables and models the agents applied are your objects and stay in place, each with its receipt.
Will our model costs or contracts change? No. Your key, your account and your rate stay exactly as they are.
Can we test this before we sign? Yes, and we recommend it. Run the exit checklist during the pilot: export the graph, verify the chain, revoke a role, and run the core in Claude Code. For deployment details see Can Data Workers run in our VPC or air-gapped?, and for regulatory mapping How does it help with SOX, HIPAA, GDPR and the EU AI Act?.
Sources
- •Data Workers open-source repo (LICENSE and NOTICE, Apache 2.0; agents; tool registrations for
get_audit_trail,verify_global_hash_chain,get_usage_activity_log,generate_audit_report,generate_migration), checked Oct 2, 2026: https://github.com/DataWorkersProject/dataworkers-claw-community - •Data Workers, Documentation (open-source core), checked Oct 2, 2026: https://dataworkers.io/opensource-docs/
- •Data Workers, Deployment (open-source core), checked Oct 2, 2026: https://dataworkers.io/opensource-docs/deployment/
- •Data Workers, What the platform adds, checked Oct 2, 2026: https://dataworkers.io/opensource-docs/platform/
- •Data Workers, Pricing (rate card Sep 10, 2026), checked Oct 2, 2026: https://dataworkers.io/pricing/
- •Data Workers, Security and trust (last updated Sep 10, 2026), checked Oct 2, 2026: https://dataworkers.io/security/
- •DataHub repository (license Apache-2.0), checked Oct 2, 2026: https://github.com/datahub-project/datahub
- •OpenMetadata repository (license Apache-2.0), checked Oct 2, 2026: https://github.com/open-metadata/OpenMetadata
- •Elementary repository (license Apache-2.0), checked Oct 2, 2026: https://github.com/elementary-data/elementary
- •Databricks, Genie (last updated Sep 18, 2026), checked Oct 2, 2026: https://docs.databricks.com/aws/en/genie/
- •Snowflake, SYSTEM$READ_YAML_FROM_SEMANTIC_VIEW, checked Oct 2, 2026: https://docs.snowflake.com/en/sql-reference/functions/system_read_yaml_from_semantic_view
- •Anthropic, Claude Code MCP documentation (project scope,
.mcp.json), checked Oct 2, 2026: https://code.claude.com/docs/en/mcp - •OpenAI Codex CLI repository (license Apache-2.0), checked Oct 2, 2026: https://github.com/openai/codex