You're on Coalesce Catalog (formerly CastorDoc): Business Users Trust What It Shows. Data Workers Keeps It True
Coalesce Catalog is where business users find and trust data. Data Workers does the operations work that keeps it true across your stack, behind approvals.
Your business users open Coalesce Catalog, the catalog formerly known as CastorDoc, when they need a table they can trust. They ask AI Search a question in Slack, Microsoft Teams or Google Chat, filter by Certified and PII, read a description that Catalog Scribe wrote and a steward reviewed, and follow lineage from warehouse columns to the Sigma workbook they already use. Analysts write SQL with SQL Copilot; managers ask Dashboard Q&A how a report is built. Since Coalesce 2.0 (Sep 15, 2026), Catalog runs on one context layer with Coalesce Transform and Coalesce Quality, and assistants read it over the Catalog MCP, which signs each user in with OAuth and only reads. Data Workers does the operations work that keeps that trust earned: when something changes outside Coalesce, it diagnoses the break, proposes each fix to the person who owns it, and verifies the result behind approvals, with a receipt for every change.
Key takeaways
- •Coalesce Catalog keeps its job. Search, Certified tables, owners, Knowledge pages, lineage and PII tags stay where your users already look.
- •The catalog records; Data Workers acts. When an app release or a replication change breaks a Certified table, Data Workers traces it to the cause and proposes the repair in the system that owns it.
- •Access requests get a dry-run, not a ticket. Data Workers shows which sensitive columns a proposed Snowflake grant would reach after role inheritance, and proposes a scoped, time-bound alternative for the owner.
- •Catalog changes stay with the steward. Corrected descriptions and tags are proposed for the Catalog owner to apply; approved facts land in the Data Context Wizard graph.
- •Read-only MCP on one side, approvals on the other. The Catalog MCP gives assistants ten search and metadata tools; Data Workers adds per-domain autonomy, from L0 manual to L4 autonomous, for changes to the data itself.
Coalesce Catalog is the storefront. Data Workers restocks and relabels the shelves behind it.
A storefront earns trust by showing the right thing on the shelf; someone behind it has to pull what is wrong before a customer picks it up. Coalesce Catalog does the showing exceptionally well. Data Workers does the back-room work: one context across the systems Catalog describes, one approval flow and one audit trail, with Spellbook Data Catalog (in preview) as the place owners review and approve.
A Monday and Tuesday at an online learning company. The product runs on Postgres on Amazon RDS, AWS DMS replicates the app schema into Snowflake, Coalesce Transform builds the ANALYTICS.DIM_LEARNER table that Coalesce Catalog shows as Certified, and the customer success team runs renewals from a Sigma workbook built on it. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| Mon 16:40 | Postgres on Amazon RDS | An app release moves email, phone and date_of_birth from learners into a new learner_contacts table. New signups stop writing learners.email |
| Mon 16:55 | AWS DMS | The replication task, mapped to the whole app schema, creates RAW_APP.LEARNER_CONTACTS in Snowflake. RAW_APP.LEARNERS.EMAIL arrives null for every new row |
| Tue 02:00 | Coalesce Transform | The scheduled job builds DIM_LEARNER from LEARNERS alone and succeeds. 2,310 learners who signed up since Monday have no email |
| 05:30 | Coalesce Catalog | The next Snowflake sync adds LEARNER_CONTACTS, and PII tagging suggestions flag its three columns from metadata. DIM_LEARNER keeps its Certified badge and its description, "one row per learner with primary contact email" |
| 06:10 | Data Workers | run_quality_check fails the null check on RAW_APP.LEARNERS.EMAIL in Snowflake, and the team's monitor_metrics baseline flags the null rate on new rows. Pull request review had recorded the Monday migration that moved the contact columns into learner_contacts, so Data Workers traces the cause to that release and maps the blast radius: DIM_LEARNER, the Sigma "Learner Renewals" workbook and the renewal audience export |
| 06:25 | Data Workers | The steward confirms Catalog's PII suggestions for email, phone and date_of_birth. No masking policy covers the raw table. Data Workers drafts one as a proposal for the Snowflake admin |
| 06:40 | Data Workers | It proposes the repair in Spellbook: a DIM_LEARNER node diff that takes the latest contact per learner from LEARNER_CONTACTS for the Coalesce Transform owner, the masking proposal, and a corrected description for the Catalog steward |
| 08:40 | Coalesce Catalog | A lifecycle marketing analyst finds LEARNER_CONTACTS through AI Search and clicks Request access. The request is recorded as a comment on the table |
| 08:50 | Data Workers | The table owner hands the request to Data Workers. The dry-run shows the analyst role would inherit schema-wide reads and reach date_of_birth and phone unmasked, a conflict with the PII policy. provision_access returns review and grants nothing; Data Workers proposes SELECT on DIM_LEARNER, with email masked, for 30 days |
| 09:30 | Spellbook | Three named owners approve their parts: the Transform owner merges and deploys the node change, the Snowflake admin applies the masking policy, and the data owner approves the scoped grant, which the admin applies |
| 10:05 | Coalesce Transform | The Transform owner reruns the DIM_LEARNER job |
| 10:20 | Snowflake | Data Workers checks the email null rate on new learners against its baseline, row counts against Monday and the masking policy on the raw table, records the grant and its expiry, and writes the receipt |
| 10:40 | Coalesce Catalog | The steward applies the corrected description and keeps the Certified badge. The Sigma workbook lists all 2,310 new learners with contact emails before the noon renewal send |

Coalesce Catalog did its job well: it picked up the new table, suggested the right PII tags without reading a row of data, and recorded the analyst's request where the owner would see it. The break sat in a Postgres release and a DMS mapping, systems the catalog describes and does not run. Without the trace, the renewal send would have skipped 2,310 new learners and the request would have opened raw birth dates to a marketing role.
| Job | What Coalesce Catalog does | What Data Workers does |
|---|---|---|
| The record | Holds the table, its owner, Certified badge, description, tags and Knowledge pages | Brings them in as sourced context and checks the data against them |
| Discovery | AI Search, SQL Copilot and Dashboard Q&A help people find and use data | Answers from governed definitions and flags when a trusted table has drifted from its description |
| Lineage | Column-level lineage in the warehouse and asset-level lineage to Sigma workbooks, served over the Catalog MCP | Joins it to schema changes from pull request review and the dbt manifest diff, run history and null rates to find the cause and the blast radius |
| Classification | PII tagging suggestions from metadata | Flags likely PII column names in pull request review and proposes masking for the owner to apply |
| Access | Records the request as a table comment for the owner | Dry-runs the proposed grant, names the sensitive columns it would reach and proposes a scoped, time-bound grant |
| The fix | Shows what is downstream | Proposes the node diff, the masking policy and the description change, each to its owner, under one approval flow |
| The proof | History tab shows who changed the metadata | A receipt: the cause, the diffs, who approved what, how it was verified and how to undo it |
Why doesn't Coalesce Catalog just do this itself?
Because a catalog's value rests on being the record people trust, and Coalesce designed its write surface around that. The Catalog MCP exposes ten tools for search and metadata reads: tables, dashboards, columns, glossary terms, asset and field lineage, tags, users and teams. Sign-in is OAuth, so an assistant sees only what that user may see; static API-token auth for MCP is deprecated. The Public API reads and updates metadata with a Read & Write token. Everything it changes is catalog metadata: descriptions, owners, tags, lineage and quality results. Even a data access request ends as a comment for the table owner, because the grant belongs to the warehouse.
Coalesce 2.0 extends that care into building: a central policy layer checks testing, PII masking, ownership and documentation rules before a Coalesce node reaches production, and Scout, "an always-on data SRE, triages problems and proposes fixes in code." Inside the Coalesce estate that is a strong design. Our Tuesday table arrived through DMS from a Postgres release, before any Coalesce node touched it.
Acting outside the catalog is a different product with a different liability: tracing a change into an app database, scoping what a grant reaches after role inheritance, routing each change to its owner and proving the result downstream. That is the product Data Workers is. Our comparison with data catalogs covers the category view.
Every tool owns a slice. Data Workers covers the whole lifecycle
Coalesce Catalog goes deepest on catalog and context. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail.

| Stage | Data Workers | Coalesce Catalog | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 9.5 | Coalesce Catalog's home stage: AI documentation, column-level lineage from source to BI, owners, Certified badges, Knowledge pages and glossary, served to assistants over the Catalog MCP. Data Workers reads it as a first-class source. |
| Analytics & Insights | 8 | 6 | AI Search, SQL Copilot and Dashboard Q&A help business users find data and write queries. Data Workers answers data questions from governed definitions with lineage behind every number. |
| Data Quality | 8 | 5 | The Data Quality dashboard gathers test results from Coalesce Quality, dbt, Soda and the API next to each table. Data Workers writes, runs and repairs checks across every platform. |
| Observability & Incidents | 8.5 | 3 | Owners get notified on comments and description changes, and incidents link to assets through Coalesce Quality. Data Workers traces the cause across systems, proposes the fix and verifies it. |
| Pipelines & Ingestion | 8.5 | 2 | Catalog documents pipelines; building them is Coalesce Transform's job. Data Workers proposes pipeline changes for their owners and queues reruns through your orchestrator. |
| Schema & Migration | 8 | 4 | Column lineage shows what a schema change reaches downstream. Data Workers catches the change in pull request review and the dbt manifest diff, scores its blast radius and generates migrations with rollback SQL for the owner. |
| Governance & Access | 8.5 | 7 | Role-based asset access, audit trails and access requests recorded as table comments. Data Workers dry-runs each proposed grant and proposes a scoped, time-bound grant for the owner's approval. |
| Security & Privacy | 8 | 6 | PII tagging suggestions flag sensitive columns from metadata. Data Workers' pull request review flags new columns whose names look personal and proposes masking for the owner to apply. |
| Cost / FinOps | 8 | 2 | Popularity counts read queries; spend is outside a catalog's job. Data Workers attributes Snowflake spend down to the query and drafts the fix for its owner. |
| MLOps & Models | 7.5 | 3 | Lineage to service accounts shows data feeding ML apps. Data Workers keeps the data under your models fresh and correct. |
Coalesce Catalog leads where a catalog should. For the category, see the agentic data catalog, AI data catalog and data catalog vs context layer guides, and how Atlan teams run the same pattern.
How Coalesce Catalog and Data Workers work together
Coalesce Catalog stays where people search, read and request. Spellbook Data Catalog is where owners review, approve and roll back each proposed change. Between them, Data Context Wizard keeps one governed context graph across the estate, the Data-Agents Swarm does the work with 20+ specialist agents, the Autonomous Data-Conductor runs each fix end to end, and per-domain guardrails hold approvals, receipts and rollback.

Coalesce Catalog to Data Workers. Coalesce Catalog connects over its API or MCP server today: your team's assistant reads lineage, owners, tags, glossary terms and descriptions side by side with Data Workers and hands what matters to Context Wizard as sourced proposals. Postgres and Snowflake are native; AWS DMS and Sigma connect over their APIs, and DMS syncs stay with their owner (Data Workers does not trigger them). Pull request review and the dbt manifest diff catch a schema change, run_quality_check catches the nulls it causes, trace_cross_platform_lineage follows it through Snowflake and Coalesce to Sigma, blast_radius_analysis maps what it reaches, and explain_table pulls each table's definition, documentation and lineage. Pull request review flags new columns whose names look personal; provision_access dry-runs a grant and request_governance_review opens a trackable review when the verdict is review. generate_documentation drafts the steward's description, run_quality_check and monitor_metrics verify, and every step lands in get_audit_trail.
What Data Workers changes, and who applies it. Coalesce node changes go to the Transform owner as a diff to merge. Snowflake masking policies and grants are proposed for the owner to apply; Data Workers records each grant and its expiry. Catalog descriptions and tags are proposed for the Catalog steward; approved facts land in the Context Wizard graph, so the next agent question starts from them.
Side by side in one client. Coalesce documents the Catalog MCP for Claude Desktop, Cursor and Dust, with OAuth sign-in and per-user access control. Add Data Workers beside it: every agent is an MCP server, and the client setup guide documents the path (clone the open-source repo, add one start-agent.sh entry per agent). Example for Cursor's ~/.cursor/mcp.json, with the Catalog's US endpoint:
{
"mcpServers": {
"Catalog": { "url": "https://api.us.castordoc.com/mcp/server" },
"dw-context-catalog": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-context-catalog"] },
"dw-schema": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-schema"] },
"dw-governance": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-governance"] }
}
}Ask "can marketing use DIM_LEARNER for the renewal send, and is it complete?" and the client calls both: Coalesce Catalog returns the owner, the Certified badge, the description and the lineage to Sigma; Data Workers returns the Monday release, the null emails, the sensitive columns in play and the proposed repair. The Catalog decides what each user may see; Data Workers decides whether a data change may run in that domain, with a named approver. A shared Data Workers endpoint takes an API key or OAuth tokens from your identity provider, verified through JWKS.
One trusted table, L0 to L4. The same broken Certified table at each level, set per domain.

- •L0 manual. Customer success notices blank emails at the noon send and files a ticket.
- •L1 observe. Data Workers posts the diagnosis: the release, the null spike, the raw PII table and everything downstream. Nothing changes.
- •L2 propose. Data Workers proposes the node diff, the masking policy, the scoped grant and the description. Nothing runs until each owner approves in Spellbook.
- •L3 act reversibly. For proven change classes, such as re-running the verification checks after an owner's fix and recording approved facts in the context graph, Data Workers runs the step, verifies it and records the receipt. The Coalesce Transform job stays the Transform owner's to run.
- •L4 autonomous. In a scoped domain with a long clean record, Data Workers runs the repeatable steps end to end; code, grants and masking still go to their owners.
On safety, see is it safe to let AI agents change production data, how approvals work, autonomy levels and who owns the agents. On where data lives: the agents run in your infrastructure, your data stays in your systems, and the hosted Conductor sees workflow metadata only.
What changes for your team
Coalesce Catalog taught the business where to look. Data Workers keeps what they find there right.

- •Trusted tables. A silently broken Certified table gets a diagnosis, a proposed fix and a corrected description.
- •Access requests. A catalog request becomes a dry-run and a scoped grant proposal.
- •Classification. New PII columns are classified as they land, with masking proposed for the owner.
- •Incidents. Breaks upstream of the catalog are traced to the source change and closed with a receipt.
- •Audits. Catalog keeps the metadata history; Data Workers records what changed in the data and who approved it.
- •Migrations. A legacy warehouse move runs in approved waves, with parity checks planned and tracked.
The steward gains most: one proposal to review at 9:30 a.m., not a Certified table found wrong at renewal time.
Keep Coalesce Catalog, or consolidate?
Keep Coalesce Catalog if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
Most Coalesce teams keep it: business users already search there, and Coalesce 2.0 ties the catalog to Transform and Quality. What teams consolidate is the work around it: access tickets that sit for days, tables nobody re-checked after a release, and the Slack thread that is the only record of why a badge stayed on. Building that layer yourself on the Catalog MCP? Read build it ourselves with Claude Code and MCP servers: the reads are easy; the context graph, approvals and rollback are the work. For the rest of this stack, see you're on Coalesce Quality, you're on Coalesce, you're on AWS DMS, you're on Sigma and Data Workers integrations.
The case for your CFO
The outcome: the tables your business trusts in Coalesce Catalog stay right after the changes nobody announces. Data Workers turns each silent break into a diagnosis on arrival and a complete fix approved by the people who own each part, before the number reaches a customer or a board.
The risk story is plain. Autonomy is set per domain: at L1 Data Workers only reads; at L2 it changes nothing until a named person approves; at L3 it applies changes it can undo. Grants and masking are proposed for the owner to apply, and every grant is dry-run first, so you see which sensitive columns it would reach. An unanswered approval request expires and escalates, never auto-grants. No agent can promote its own work. Every change carries a receipt with the cause, the approver and how to undo it. An org-wide stop halts all autonomous dispatch. Zero migration: Postgres, DMS, Snowflake, Coalesce and Sigma stay where they are.
Why now: assistants already read your catalog over MCP and repeat what it says, so a stale Certified badge misleads software as well as people. The first win is read-only: every schema change upstream of a Certified table that comes through pull request review or the dbt manifest gets a diagnosis and a blast radius. Your catalog, stewards, Snowflake roles and schedules stay the same. For the numbers, see the ROI of agentic data operations.
The sentence to repeat upstairs: "Coalesce Catalog shows us which data to trust; Data Workers keeps it worth trusting, with an approval on every change and proof it held."
Getting started
Start with a pilot. Pick the Certified tables one team relies on, connect Data Workers to Coalesce Catalog, Postgres and Snowflake, and run at L1 so every upstream change gets a diagnosis and a blast radius. Then turn on the first write class at L2, such as access requests dry-run for the data owner. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.
FAQ
Is CastorDoc the same product as Coalesce Catalog? Yes. Coalesce integrated its enterprise data catalog, formerly CastorDoc, into its platform; since Coalesce 2.0 (Sep 15, 2026) Catalog, Transform and Quality run on one context layer. The Catalog MCP and Public API still use castordoc.com hosts.
How does Data Workers connect to Coalesce Catalog? Over Coalesce Catalog's API or MCP server today, reading lineage, owners, tags, glossary terms and descriptions. Postgres and Snowflake are native. You can also run both MCP servers in one client.
Does Data Workers write to Coalesce Catalog? Corrected descriptions and tags are proposed for the Catalog steward to apply, so the catalog stays the record your stewards control. Approved facts land in the Data Context Wizard graph right away.
Coalesce has Scout. Why add Data Workers? Scout triages problems and proposes fixes in Coalesce code, and does it well. Many breaks start outside that code: an app release, a replication mapping, a grant request. Data Workers traces those, routes each change to its owner and verifies the result.
What happens when someone requests access in Coalesce Catalog? The request is recorded on the table for its owner. With Data Workers, the owner gets a dry-run of the proposed Snowflake grant: effective privileges after role inheritance, the sensitive columns it would reach, any policy conflict and a least-privilege recommendation with an expiry. The owner approves and applies the grant.
Does Data Workers change our Coalesce Transform nodes? It proposes node changes as a diff for the owner to merge through your normal review, then verifies the rebuilt table in Snowflake.
Sources
- •Coalesce, Coalesce Catalog product page, https://coalesce.io/product/catalog/ (checked Oct 3, 2026)
- •Coalesce, "Coalesce 2.0: The Agentic Data Engineering Platform" (Sep 15, 2026), https://coalesce.io/product-technology/coalesce-2-0-the-agentic-data-engineering-platform/ (checked Oct 3, 2026)
- •Coalesce homepage (Transform, Catalog, Quality; Scout), https://coalesce.io/ (checked Oct 3, 2026)
- •Coalesce Docs, Catalog MCP integration (endpoints, OAuth, token auth deprecated, ten read tools), https://docs.coalesce.io/docs/catalog/developer/mcp-integration.md (checked Oct 3, 2026)
- •Coalesce Docs, Catalog MCP glossary entry, https://docs.coalesce.io/docs/reference/glossary/catalog-mcp.md (checked Oct 3, 2026)
- •Coalesce Docs, Introduction to the Catalog AI Assistant, https://docs.coalesce.io/docs/catalog/ai-assistant/introduction.md (checked Oct 3, 2026)
- •Coalesce Docs, Catalog Public API, https://docs.coalesce.io/docs/catalog/developer/catalog-apis/public-api.md (checked Oct 3, 2026)
- •Coalesce Docs, Tables (Certified, PII, owners, popularity), https://docs.coalesce.io/docs/catalog/assets/tables.md (checked Oct 3, 2026)
- •Coalesce Docs, PII Tagging Suggestions, https://docs.coalesce.io/docs/catalog/document-your-data/pii-tagging-suggestions.md (checked Oct 3, 2026)
- •Coalesce Docs, Request Data Access, https://docs.coalesce.io/docs/catalog/collaborate/request-data-access.md (checked Oct 3, 2026)
- •Coalesce Docs, Column Description Propagation, https://docs.coalesce.io/docs/catalog/document-your-data/catalog-scribe/column-description-propagation.md (checked Oct 3, 2026)
- •Coalesce Docs, Snowflake integration for Catalog, https://docs.coalesce.io/docs/catalog/integrations/data-warehouses/snowflake.md (checked Oct 3, 2026)
- •Coalesce Docs, Sigma integration for Catalog (lineage is asset-level), https://docs.coalesce.io/docs/catalog/integrations/data-viz/sigma.md (checked Oct 3, 2026)
- •Data Workers open-source repository, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
- •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)