How is Data Workers different from a data catalog?
A data catalog records what your data is and serves it to AI over MCP. Data Workers acts on that context: it finds what's wrong, fixes it through approvals, verifies the fix and leaves a receipt. Keep your catalog.
A data catalog records what your data is (inventory, glossary, lineage, ownership and policies) and increasingly serves that context to AI over MCP. Data Workers acts on it: it reads that context, including from your catalog, finds what's wrong, fixes it through approvals, verifies the fix and leaves a receipt, and it ships its own catalog (Spellbook Data Catalog) and context layer (Data Context Wizard) when you want one.
Keep the catalog you love. Data Workers works with it from day one.
Key takeaways
- •A catalog is the inventory. Data Workers is the control plane. The catalog answers "what is this table, who owns it, what feeds it". Data Workers answers "what broke, what's the fix, who approved it, and did it hold".
- •Catalogs now speak MCP. Collibra, Atlan, Alation, DataHub, OpenMetadata and Google Knowledge Catalog expose catalog context to AI tools over MCP, and Databricks and Snowflake expose governed data through managed MCP servers. That makes them excellent sources for Data Workers' agents.
- •The difference is the write path. Data Workers detects, proposes a fix in the system that owns the cause, routes it to a named owner, checks the result downstream and records a tamper-evident receipt with an undo path.
- •Your catalog stays, and gets more accurate. Approved fixes and notes flow back through dbt and your catalog's next ingestion, so the lineage it shows matches what actually runs.
- •No catalog yet, or one per platform? Spellbook Data Catalog (in preview) and Data Context Wizard give you one governed view across every engine.
What a data catalog does, and does well
A catalog's job is to describe the estate. It harvests metadata from warehouses, lakehouses, dbt, orchestrators and BI, and turns it into assets with owners, domains, glossary terms, classifications and lineage. Stewards curate it, workflows approve new terms, and analysts search it to find the right table. That job is hard, and the best catalogs do it very well.
In 2026, catalogs also became context servers for AI. Each vendor's own pages, checked this month:
- •Collibra calls its MCP Server production-ready. It connects Claude, ChatGPT, Databricks, Snowflake Cortex, Microsoft Copilot, GitHub Copilot and Cursor, with read tools for assets, lineage and glossary and write tools for enriching entries and proposing terms, all inside Collibra's permission model. On September 23, 2026 it added Maestro, no-code agents for governance teams (Maestro Studio and Maestro Assistant in public preview).
- •Atlan runs a hosted MCP server, enabled for all tenants, to search context, traverse lineage, read governed definitions and run SQL. Its write tools preview each change and wait for approval.
- •Alation expanded its Alation Intelligence Operating System (AIOS) with six products on September 17, 2026 (AI Governance and Semantic Model Mastering available now, four more in early access), and offers a remote MCP server and an AI Agent SDK.
- •DataHub hosts an MCP server on DataHub Cloud and ships an open-source one for Core. Mutation tools for tags, terms, owners and descriptions are opt-in, and each is marked so MCP clients can ask for confirmation.
- •OpenMetadata 2.0 ships MCP "installed and enabled by default", with OAuth through your SSO.
- •Unity Catalog offers managed MCP servers in Public Preview (the Genie One MCP server went GA on September 25, 2026), plus automatic column-level lineage and GA data classification for tables (view scanning in Beta since September 15, 2026).
- •Snowflake Horizon Catalog governs data inside and outside Snowflake and writes table and column descriptions with Cortex; the Snowflake-managed MCP server (GA) exposes Cortex Agents, Cortex Analyst, Cortex Search, SQL and custom tools, each granted separately.
- •Google Knowledge Catalog (renamed from Dataplex Universal Catalog in April 2026) serves
search_entries,lookup_contextandlookup_entryover MCP, with data-product write tools in a separate toolset. - •Microsoft Purview Unified Catalog organizes assets into governance domains and data products, attaches policies to glossary terms, scores data quality (standalone and incremental scans are GA) and adds copilot search.
That is a strong foundation. Every one of these catalogs makes an AI agent better informed. Data Workers is the agent platform that takes that information and runs the estate with it.
What Data Workers does with catalog context
Data Workers is the agentic data platform. It does the operational work of a data estate, across every engine, under governance. Four products split the job:
Data Context Wizard reads your catalog as a first-class source. It builds one governed graph across Snowflake, Databricks, BigQuery, dbt, orchestration and BI, and brings in your existing catalogs, glossaries and semantic layers with provenance on every fact. DataHub, OpenMetadata, Microsoft Purview, Google Knowledge Catalog, AWS Glue, Iceberg catalogs, Collibra, Atlan and Alation connect over their MCP servers or APIs today, and each stays the inventory its owners check. Your catalog's glossary and owners sit next to live run results, query usage, grants and test outcomes, which is what an agent needs before it touches anything. Public tools such as explain_table, search_across_platforms and trace_cross_platform_lineage answer from that graph. See bring your own context for how outside context stays yours.
The Data-Agents Swarm finds what's wrong. More than 20 specialist agents watch quality, schema, usage and lateness against recorded baselines. The quality agent runs run_quality_check and get_quality_score, the schema agent runs assess_impact, and the incident agent runs get_root_cause over lineage that spans systems.
The Autonomous Data-Conductor runs the fix. It computes the blast radius with blast_radius_analysis, drafts the change in the system that owns the cause (a dbt diff for the owner to merge, a SQL change, an orchestrator rerun), routes it by the domain's autonomy level, checks the outcome downstream and records what it learned. When a model is renamed or moved, the graph's lineage is updated and checked for consistency, so later answers point at the right asset.
Spellbook Data Catalog is where people look. It's the agentic catalog and human control plane: one inbox to approve, steer, send back or roll back, with the full trace of what each agent saw and why it acted. No agent can promote its own work to authoritative or canonical. Approved descriptions go back as dbt docs changes, so your catalog picks them up on its next ingestion.

The matrix is the short answer. A catalog leads on inventory, glossary and lineage display, its home job. Data Workers leads on the jobs that come after the catalog shows you something: detect, fix, approve, verify and receipt.
One lineage break: what the catalog shows, what Data Workers fixes
An illustration, not a customer case. The estate runs Snowflake, dbt, Airflow, DataHub and Tableau. A dbt refactor moves fct_revenue from the analytics schema to finance_marts. The old table stays in place but stops updating, and Tableau's Revenue Daily workbook still reads it.
| Time | System | What happens |
|---|---|---|
| Mon 16:20 | dbt | Refactor merged; fct_revenue now builds in finance_marts |
| 18:00 | Airflow | The nightly DAG run builds the new table; analytics.fct_revenue stops updating |
| 18:40 | DataHub | Ingestion shows the old table with no upstream, and Tableau still downstream of it |
| Tue 05:30 | Data Workers | Flags analytics.fct_revenue as 11 hours stale while it is still being queried |
| 05:33 | Data Workers | Traces the cause from the dbt manifest diff and Snowflake query history |
| 05:35 | Data Workers | Blast radius: Revenue Daily, two extracts and the finance export job |
| 05:41 | Data Workers | Proposes a dbt diff adding a compatibility view over the new model, with a deprecation note |
| 05:42 | Data Workers | Revenue runs at L2 propose, so the change goes to the domain's named owner |
| 08:35 | Spellbook | The analytics engineer reviews the diff, blast radius and rollback path, approves it and merges it |
| 08:47 | Airflow | The DAG run builds the view; dbt tests pass |
| 08:55 | Data Workers | Verifies revenue totals match the new model and the extracts refreshed |
| 08:56 | Spellbook | Receipt: trigger, diff, approver, checks run with before and after values, undo path |
| 09:20 | DataHub | The next ingestion shows lineage reconnected end to end |
| 09:30 | Tableau | Finance opens Revenue Daily on yesterday's numbers |

DataHub did its job well: at 18:40 its lineage graph showed exactly where the chain broke. What turned that picture into a fixed dashboard was a different job. Someone had to notice, trace the cause across four systems, write the fix, get it approved, check it and record it. Data Workers did that overnight, and a person made the one decision that mattered.
Why doesn't a catalog just do this itself?
Focus and risk. A catalog is built to be the trusted record of what exists and what it means, and its design choices follow from that. Writes stay inside the catalog's own metadata: terms, tags, owners, descriptions, data products. That is a sensible boundary, and several vendors add confirmation steps even for those writes (Atlan previews each change, DataHub's mutation tools are opt-in).
Fixing production data is a different product with a different liability. It means changing dbt models, warehouse objects and orchestrator runs that the catalog doesn't own, computing blast radius across them, holding scoped credentials per agent, routing approvals to a named owner, keeping a rollback path and proving the fix held downstream. Data Workers is built for exactly that, and it treats your catalog as one of the best sources it reads. The safety page covers what an agent can and can't do at each level.
How autonomy works on catalog-driven work
Each domain sits on one ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous. Start catalog-adjacent work, such as descriptions and ownership gaps, at L2, and keep revenue models at L2 until the receipts show a record. Reversible, low-risk fixes like a stale-table refresh can move to L3. Anything irreversible goes to a named owner; an unanswered request expires and escalates, and never grants itself.

Keep your catalog, or use Spellbook?
Keep your catalog if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too, because Spellbook Data Catalog gives business users search by meaning and one approvals inbox across every platform, and Data Context Wizard keeps the graph current as a byproduct of the agents' work.
Three situations come up often. If you run Collibra, Alation, DataHub or OpenMetadata, start with the catalog comparison; on Atlan, read Atlan vs Spellbook. If your catalog is your platform's own, see Unity Catalog vs Spellbook, Horizon Catalog vs Spellbook and Knowledge Catalog vs Spellbook. Unity Catalog, Horizon and Knowledge Catalog stay the permission systems on their platforms, by design; Data Workers applies approved grants through them. If you have no catalog, Spellbook and the Context Wizard give you one from the start. Our resource guides on the agentic data catalog and AI data catalogs cover the category, and what is an agentic data platform places catalogs among the other neighbours, including data observability.
The case for your CFO
The outcome: the catalog investment starts paying off in fixed problems. Today a catalog shows a broken lineage chain or a stale table, and someone still has to find the cause, write the fix and chase an approval. With Data Workers, that work runs overnight and the team approves a reviewed change in the morning. Dashboards open on correct numbers, and engineers spend their hours on the data products the business asked for.
The risk story: agents propose before they act, a named owner approves anything irreversible, every change carries its blast radius, checks and rollback path, and autonomy rises per domain only when the receipts show the agents were right. Data Workers runs in your infrastructure on every tier, and your data stays in your systems; the hosted Conductor sees workflow metadata only. The security page has the detail.
Why now: catalogs now hand context to every AI tool over MCP, so agents will act on your data either way. The question is whether they act through one governed write path.
What stays the same: your catalog, its stewards and its workflows, plus the warehouse, dbt and orchestrator. Zero migration.
The first win is overnight triage of stale or broken lineage at L2, approved by the domain owner, inside the pilot. Start with a pilot: $7,500 one-time, and the pilot is credited in full against the first year. After it, Scale starts from $1,000/month and Enterprise from $3,000/month, billed annually, with unlimited seats, no usage meter, no markup on model spend, bring your own model and an Apache 2.0 core. See pricing, the ROI calculator and the ROI page. Weighing a build instead? Read build it ourselves with Claude Code and MCP servers.
The sentence to repeat upstairs: "Our catalog tells us what our data is; Data Workers keeps it right, under our approvals, with a receipt on every change."
FAQ
Is Data Workers a data catalog? It includes one. Spellbook Data Catalog (in preview) is the agentic catalog and control plane, and Data Context Wizard is the context graph underneath. The first step is connecting the catalog you already run.
Our catalog has an MCP server now. Isn't that enough for agents? It gives agents good context, which is valuable. Acting on production data also needs blast radius, scoped credentials, named-owner approvals, rollback, downstream verification and receipts across systems the catalog doesn't own. That is what Data Workers adds, and your team's assistant can use your catalog's MCP server side by side with it.
Will Data Workers overwrite our stewards' definitions? No. Definitions from your catalog come in with their source. Agent proposals go to a named owner, and no agent can promote its own work to authoritative. Data Workers doesn't write into your catalog: approved facts land in the Context Wizard graph, and approved notes flow back through dbt docs for your catalog to ingest. A steward can apply them in the catalog directly, or your team's MCP client can write them where the catalog's MCP server allows.
Which catalogs does it connect to? Unity Catalog and Snowflake through their platforms. DataHub, OpenMetadata, Microsoft Purview, Google Knowledge Catalog, AWS Glue, Iceberg REST catalogs, Collibra, Atlan and Alation connect over their MCP servers or APIs today. The integrations page lists the rest.
Does Data Workers replace our access controls? No. Unity Catalog, Snowflake and your other platforms keep enforcing grants and policies. Data Workers runs the request queue, applies approved grants through Unity Catalog and drafts the grant for every other platform's owner to apply.
How do we prove it to auditors? Every change leaves a tamper-evident receipt: trigger, diff, blast radius, checks with before and after values, approver and undo path. See how approvals work.
Sources
- •Collibra, MCP Server: read and write tools within Collibra's permission model; connects Claude, ChatGPT, Databricks, Snowflake Cortex, Microsoft Copilot, GitHub Copilot, Cursor and VS Code; production-ready. Maestro announcement (September 23, 2026; Maestro Studio and Assistant in public preview). Checked October 2, 2026.
- •Atlan, remote MCP overview (hosted, enabled for all tenants) and MCP tools reference (39 tools in nine categories; tools that change anything "return a preview of the change and wait for your approval before writing"). Checked October 2, 2026.
- •Alation, AIOS expansion with six products (September 17, 2026) and AI Agent SDK and remote MCP server. Checked October 2, 2026.
- •DataHub, MCP server guide: hosted on DataHub Cloud, open source for Core, opt-in mutation tools marked for client confirmation. Checked October 2, 2026.
- •OpenMetadata, MCP setup and authentication: installed and enabled by default; OAuth 2.0 through SSO. Checked October 2, 2026.
- •Databricks, managed MCP servers and September 2026 release notes (Genie One MCP GA September 25, 2026; view classification Beta September 15, 2026); data classification and lineage. Checked October 2, 2026.
- •Snowflake, Snowflake-managed MCP server (GA; "Access to the MCP Server does not give access to the tools") and Horizon Catalog (open catalog for data inside and outside Snowflake; Cortex-generated descriptions). Checked October 2, 2026.
- •Google Cloud, Knowledge Catalog MCP reference (last updated September 21, 2026; renamed from Dataplex Universal Catalog April 10, 2026). Checked October 2, 2026.
- •Microsoft, Purview Unified Catalog and what's new in Microsoft Purview (updated September 21, 2026; standalone and incremental data quality scans and configurable thresholds GA in May 2026). Checked October 2, 2026.
- •Data Workers, Spellbook Data Catalog, Data Context Wizard, Autonomous Data-Conductor, pricing and the open-source core (Apache 2.0, public tool registrations). Checked October 2, 2026.