Product
Product12 min readBy The Data Workers Team

You're on Coalesce Catalog (formerly CastorDoc): Business Users Trust What It Shows. Data Workers Keeps It True

Coalesce Catalog is where business users find and trust data. Data Workers does the operations work that keeps it true across your stack, behind approvals.

Your business users open Coalesce Catalog, the catalog formerly known as CastorDoc, when they need a table they can trust. They ask AI Search a question in Slack, Microsoft Teams or Google Chat, filter by Certified and PII, read a description that Catalog Scribe wrote and a steward reviewed, and follow lineage from warehouse columns to the Sigma workbook they already use. Analysts write SQL with SQL Copilot; managers ask Dashboard Q&A how a report is built. Since Coalesce 2.0 (Sep 15, 2026), Catalog runs on one context layer with Coalesce Transform and Coalesce Quality, and assistants read it over the Catalog MCP, which signs each user in with OAuth and only reads. Data Workers does the operations work that keeps that trust earned: when something changes outside Coalesce, it diagnoses the break, proposes each fix to the person who owns it, and verifies the result behind approvals, with a receipt for every change.

Key takeaways

  • •Coalesce Catalog keeps its job. Search, Certified tables, owners, Knowledge pages, lineage and PII tags stay where your users already look.
  • •The catalog records; Data Workers acts. When an app release or a replication change breaks a Certified table, Data Workers traces it to the cause and proposes the repair in the system that owns it.
  • •Access requests get a dry-run, not a ticket. Data Workers shows which sensitive columns a proposed Snowflake grant would reach after role inheritance, and proposes a scoped, time-bound alternative for the owner.
  • •Catalog changes stay with the steward. Corrected descriptions and tags are proposed for the Catalog owner to apply; approved facts land in the Data Context Wizard graph.
  • •Read-only MCP on one side, approvals on the other. The Catalog MCP gives assistants ten search and metadata tools; Data Workers adds per-domain autonomy, from L0 manual to L4 autonomous, for changes to the data itself.

Coalesce Catalog is the storefront. Data Workers restocks and relabels the shelves behind it.

A storefront earns trust by showing the right thing on the shelf; someone behind it has to pull what is wrong before a customer picks it up. Coalesce Catalog does the showing exceptionally well. Data Workers does the back-room work: one context across the systems Catalog describes, one approval flow and one audit trail, with Spellbook Data Catalog (in preview) as the place owners review and approve.

A Monday and Tuesday at an online learning company. The product runs on Postgres on Amazon RDS, AWS DMS replicates the app schema into Snowflake, Coalesce Transform builds the ANALYTICS.DIM_LEARNER table that Coalesce Catalog shows as Certified, and the customer success team runs renewals from a Sigma workbook built on it. This is an illustration, not a customer case.

TimeSystemWhat happens
Mon 16:40Postgres on Amazon RDSAn app release moves email, phone and date_of_birth from learners into a new learner_contacts table. New signups stop writing learners.email
Mon 16:55AWS DMSThe replication task, mapped to the whole app schema, creates RAW_APP.LEARNER_CONTACTS in Snowflake. RAW_APP.LEARNERS.EMAIL arrives null for every new row
Tue 02:00Coalesce TransformThe scheduled job builds DIM_LEARNER from LEARNERS alone and succeeds. 2,310 learners who signed up since Monday have no email
05:30Coalesce CatalogThe next Snowflake sync adds LEARNER_CONTACTS, and PII tagging suggestions flag its three columns from metadata. DIM_LEARNER keeps its Certified badge and its description, "one row per learner with primary contact email"
06:10Data Workersrun_quality_check fails the null check on RAW_APP.LEARNERS.EMAIL in Snowflake, and the team's monitor_metrics baseline flags the null rate on new rows. Pull request review had recorded the Monday migration that moved the contact columns into learner_contacts, so Data Workers traces the cause to that release and maps the blast radius: DIM_LEARNER, the Sigma "Learner Renewals" workbook and the renewal audience export
06:25Data WorkersThe steward confirms Catalog's PII suggestions for email, phone and date_of_birth. No masking policy covers the raw table. Data Workers drafts one as a proposal for the Snowflake admin
06:40Data WorkersIt proposes the repair in Spellbook: a DIM_LEARNER node diff that takes the latest contact per learner from LEARNER_CONTACTS for the Coalesce Transform owner, the masking proposal, and a corrected description for the Catalog steward
08:40Coalesce CatalogA lifecycle marketing analyst finds LEARNER_CONTACTS through AI Search and clicks Request access. The request is recorded as a comment on the table
08:50Data WorkersThe table owner hands the request to Data Workers. The dry-run shows the analyst role would inherit schema-wide reads and reach date_of_birth and phone unmasked, a conflict with the PII policy. provision_access returns review and grants nothing; Data Workers proposes SELECT on DIM_LEARNER, with email masked, for 30 days
09:30SpellbookThree named owners approve their parts: the Transform owner merges and deploys the node change, the Snowflake admin applies the masking policy, and the data owner approves the scoped grant, which the admin applies
10:05Coalesce TransformThe Transform owner reruns the DIM_LEARNER job
10:20SnowflakeData Workers checks the email null rate on new learners against its baseline, row counts against Monday and the masking policy on the raw table, records the grant and its expiry, and writes the receipt
10:40Coalesce CatalogThe steward applies the corrected description and keeps the Certified badge. The Sigma workbook lists all 2,310 new learners with contact emails before the noon renewal send
Incident timeline across the stack: what Coalesce Catalog, your team and Data Workers each do, step by step

Coalesce Catalog did its job well: it picked up the new table, suggested the right PII tags without reading a row of data, and recorded the analyst's request where the owner would see it. The break sat in a Postgres release and a DMS mapping, systems the catalog describes and does not run. Without the trace, the renewal send would have skipped 2,310 new learners and the request would have opened raw birth dates to a marketing role.

JobWhat Coalesce Catalog doesWhat Data Workers does
The recordHolds the table, its owner, Certified badge, description, tags and Knowledge pagesBrings them in as sourced context and checks the data against them
DiscoveryAI Search, SQL Copilot and Dashboard Q&A help people find and use dataAnswers from governed definitions and flags when a trusted table has drifted from its description
LineageColumn-level lineage in the warehouse and asset-level lineage to Sigma workbooks, served over the Catalog MCPJoins it to schema changes from pull request review and the dbt manifest diff, run history and null rates to find the cause and the blast radius
ClassificationPII tagging suggestions from metadataFlags likely PII column names in pull request review and proposes masking for the owner to apply
AccessRecords the request as a table comment for the ownerDry-runs the proposed grant, names the sensitive columns it would reach and proposes a scoped, time-bound grant
The fixShows what is downstreamProposes the node diff, the masking policy and the description change, each to its owner, under one approval flow
The proofHistory tab shows who changed the metadataA receipt: the cause, the diffs, who approved what, how it was verified and how to undo it

Why doesn't Coalesce Catalog just do this itself?

Because a catalog's value rests on being the record people trust, and Coalesce designed its write surface around that. The Catalog MCP exposes ten tools for search and metadata reads: tables, dashboards, columns, glossary terms, asset and field lineage, tags, users and teams. Sign-in is OAuth, so an assistant sees only what that user may see; static API-token auth for MCP is deprecated. The Public API reads and updates metadata with a Read & Write token. Everything it changes is catalog metadata: descriptions, owners, tags, lineage and quality results. Even a data access request ends as a comment for the table owner, because the grant belongs to the warehouse.

Coalesce 2.0 extends that care into building: a central policy layer checks testing, PII masking, ownership and documentation rules before a Coalesce node reaches production, and Scout, "an always-on data SRE, triages problems and proposes fixes in code." Inside the Coalesce estate that is a strong design. Our Tuesday table arrived through DMS from a Postgres release, before any Coalesce node touched it.

Acting outside the catalog is a different product with a different liability: tracing a change into an app database, scoping what a grant reaches after role inheritance, routing each change to its owner and proving the result downstream. That is the product Data Workers is. Our comparison with data catalogs covers the category view.

Every tool owns a slice. Data Workers covers the whole lifecycle

Coalesce Catalog goes deepest on catalog and context. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Coalesce Catalog goes deep on its own area
StageData WorkersCoalesce CatalogWhy we scored it this way
Catalog & Context99.5Coalesce Catalog's home stage: AI documentation, column-level lineage from source to BI, owners, Certified badges, Knowledge pages and glossary, served to assistants over the Catalog MCP. Data Workers reads it as a first-class source.
Analytics & Insights86AI Search, SQL Copilot and Dashboard Q&A help business users find data and write queries. Data Workers answers data questions from governed definitions with lineage behind every number.
Data Quality85The Data Quality dashboard gathers test results from Coalesce Quality, dbt, Soda and the API next to each table. Data Workers writes, runs and repairs checks across every platform.
Observability & Incidents8.53Owners get notified on comments and description changes, and incidents link to assets through Coalesce Quality. Data Workers traces the cause across systems, proposes the fix and verifies it.
Pipelines & Ingestion8.52Catalog documents pipelines; building them is Coalesce Transform's job. Data Workers proposes pipeline changes for their owners and queues reruns through your orchestrator.
Schema & Migration84Column lineage shows what a schema change reaches downstream. Data Workers catches the change in pull request review and the dbt manifest diff, scores its blast radius and generates migrations with rollback SQL for the owner.
Governance & Access8.57Role-based asset access, audit trails and access requests recorded as table comments. Data Workers dry-runs each proposed grant and proposes a scoped, time-bound grant for the owner's approval.
Security & Privacy86PII tagging suggestions flag sensitive columns from metadata. Data Workers' pull request review flags new columns whose names look personal and proposes masking for the owner to apply.
Cost / FinOps82Popularity counts read queries; spend is outside a catalog's job. Data Workers attributes Snowflake spend down to the query and drafts the fix for its owner.
MLOps & Models7.53Lineage to service accounts shows data feeding ML apps. Data Workers keeps the data under your models fresh and correct.

Coalesce Catalog leads where a catalog should. For the category, see the agentic data catalog, AI data catalog and data catalog vs context layer guides, and how Atlan teams run the same pattern.

How Coalesce Catalog and Data Workers work together

Coalesce Catalog stays where people search, read and request. Spellbook Data Catalog is where owners review, approve and roll back each proposed change. Between them, Data Context Wizard keeps one governed context graph across the estate, the Data-Agents Swarm does the work with 20+ specialist agents, the Autonomous Data-Conductor runs each fix end to end, and per-domain guardrails hold approvals, receipts and rollback.

How Data Workers fits with Coalesce Catalog: your coding agent on top, Data Workers in the middle, your estate underneath

Coalesce Catalog to Data Workers. Coalesce Catalog connects over its API or MCP server today: your team's assistant reads lineage, owners, tags, glossary terms and descriptions side by side with Data Workers and hands what matters to Context Wizard as sourced proposals. Postgres and Snowflake are native; AWS DMS and Sigma connect over their APIs, and DMS syncs stay with their owner (Data Workers does not trigger them). Pull request review and the dbt manifest diff catch a schema change, run_quality_check catches the nulls it causes, trace_cross_platform_lineage follows it through Snowflake and Coalesce to Sigma, blast_radius_analysis maps what it reaches, and explain_table pulls each table's definition, documentation and lineage. Pull request review flags new columns whose names look personal; provision_access dry-runs a grant and request_governance_review opens a trackable review when the verdict is review. generate_documentation drafts the steward's description, run_quality_check and monitor_metrics verify, and every step lands in get_audit_trail.

What Data Workers changes, and who applies it. Coalesce node changes go to the Transform owner as a diff to merge. Snowflake masking policies and grants are proposed for the owner to apply; Data Workers records each grant and its expiry. Catalog descriptions and tags are proposed for the Catalog steward; approved facts land in the Context Wizard graph, so the next agent question starts from them.

Side by side in one client. Coalesce documents the Catalog MCP for Claude Desktop, Cursor and Dust, with OAuth sign-in and per-user access control. Add Data Workers beside it: every agent is an MCP server, and the client setup guide documents the path (clone the open-source repo, add one start-agent.sh entry per agent). Example for Cursor's ~/.cursor/mcp.json, with the Catalog's US endpoint:

{
  "mcpServers": {
    "Catalog": { "url": "https://api.us.castordoc.com/mcp/server" },
    "dw-context-catalog": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-context-catalog"] },
    "dw-schema": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-schema"] },
    "dw-governance": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-governance"] }
  }
}

Ask "can marketing use DIM_LEARNER for the renewal send, and is it complete?" and the client calls both: Coalesce Catalog returns the owner, the Certified badge, the description and the lineage to Sigma; Data Workers returns the Monday release, the null emails, the sensitive columns in play and the proposed repair. The Catalog decides what each user may see; Data Workers decides whether a data change may run in that domain, with a named approver. A shared Data Workers endpoint takes an API key or OAuth tokens from your identity provider, verified through JWKS.

One trusted table, L0 to L4. The same broken Certified table at each level, set per domain.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. Customer success notices blank emails at the noon send and files a ticket.
  • •L1 observe. Data Workers posts the diagnosis: the release, the null spike, the raw PII table and everything downstream. Nothing changes.
  • •L2 propose. Data Workers proposes the node diff, the masking policy, the scoped grant and the description. Nothing runs until each owner approves in Spellbook.
  • •L3 act reversibly. For proven change classes, such as re-running the verification checks after an owner's fix and recording approved facts in the context graph, Data Workers runs the step, verifies it and records the receipt. The Coalesce Transform job stays the Transform owner's to run.
  • •L4 autonomous. In a scoped domain with a long clean record, Data Workers runs the repeatable steps end to end; code, grants and masking still go to their owners.

On safety, see is it safe to let AI agents change production data, how approvals work, autonomy levels and who owns the agents. On where data lives: the agents run in your infrastructure, your data stays in your systems, and the hosted Conductor sees workflow metadata only.

What changes for your team

Coalesce Catalog taught the business where to look. Data Workers keeps what they find there right.

Six jobs that run on autopilot with Data Workers next to Coalesce Catalog, with a concrete example of each
  • •Trusted tables. A silently broken Certified table gets a diagnosis, a proposed fix and a corrected description.
  • •Access requests. A catalog request becomes a dry-run and a scoped grant proposal.
  • •Classification. New PII columns are classified as they land, with masking proposed for the owner.
  • •Incidents. Breaks upstream of the catalog are traced to the source change and closed with a receipt.
  • •Audits. Catalog keeps the metadata history; Data Workers records what changed in the data and who approved it.
  • •Migrations. A legacy warehouse move runs in approved waves, with parity checks planned and tracked.

The steward gains most: one proposal to review at 9:30 a.m., not a Certified table found wrong at renewal time.

Keep Coalesce Catalog, or consolidate?

Keep Coalesce Catalog if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

Most Coalesce teams keep it: business users already search there, and Coalesce 2.0 ties the catalog to Transform and Quality. What teams consolidate is the work around it: access tickets that sit for days, tables nobody re-checked after a release, and the Slack thread that is the only record of why a badge stayed on. Building that layer yourself on the Catalog MCP? Read build it ourselves with Claude Code and MCP servers: the reads are easy; the context graph, approvals and rollback are the work. For the rest of this stack, see you're on Coalesce Quality, you're on Coalesce, you're on AWS DMS, you're on Sigma and Data Workers integrations.

The case for your CFO

The outcome: the tables your business trusts in Coalesce Catalog stay right after the changes nobody announces. Data Workers turns each silent break into a diagnosis on arrival and a complete fix approved by the people who own each part, before the number reaches a customer or a board.

The risk story is plain. Autonomy is set per domain: at L1 Data Workers only reads; at L2 it changes nothing until a named person approves; at L3 it applies changes it can undo. Grants and masking are proposed for the owner to apply, and every grant is dry-run first, so you see which sensitive columns it would reach. An unanswered approval request expires and escalates, never auto-grants. No agent can promote its own work. Every change carries a receipt with the cause, the approver and how to undo it. An org-wide stop halts all autonomous dispatch. Zero migration: Postgres, DMS, Snowflake, Coalesce and Sigma stay where they are.

Why now: assistants already read your catalog over MCP and repeat what it says, so a stale Certified badge misleads software as well as people. The first win is read-only: every schema change upstream of a Certified table that comes through pull request review or the dbt manifest gets a diagnosis and a blast radius. Your catalog, stewards, Snowflake roles and schedules stay the same. For the numbers, see the ROI of agentic data operations.

The sentence to repeat upstairs: "Coalesce Catalog shows us which data to trust; Data Workers keeps it worth trusting, with an approval on every change and proof it held."

Getting started

Start with a pilot. Pick the Certified tables one team relies on, connect Data Workers to Coalesce Catalog, Postgres and Snowflake, and run at L1 so every upstream change gets a diagnosis and a blast radius. Then turn on the first write class at L2, such as access requests dry-run for the data owner. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.

FAQ

Is CastorDoc the same product as Coalesce Catalog? Yes. Coalesce integrated its enterprise data catalog, formerly CastorDoc, into its platform; since Coalesce 2.0 (Sep 15, 2026) Catalog, Transform and Quality run on one context layer. The Catalog MCP and Public API still use castordoc.com hosts.

How does Data Workers connect to Coalesce Catalog? Over Coalesce Catalog's API or MCP server today, reading lineage, owners, tags, glossary terms and descriptions. Postgres and Snowflake are native. You can also run both MCP servers in one client.

Does Data Workers write to Coalesce Catalog? Corrected descriptions and tags are proposed for the Catalog steward to apply, so the catalog stays the record your stewards control. Approved facts land in the Data Context Wizard graph right away.

Coalesce has Scout. Why add Data Workers? Scout triages problems and proposes fixes in Coalesce code, and does it well. Many breaks start outside that code: an app release, a replication mapping, a grant request. Data Workers traces those, routes each change to its owner and verifies the result.

What happens when someone requests access in Coalesce Catalog? The request is recorded on the table for its owner. With Data Workers, the owner gets a dry-run of the proposed Snowflake grant: effective privileges after role inheritance, the sensitive columns it would reach, any policy conflict and a least-privilege recommendation with an expiry. The owner approves and applies the grant.

Does Data Workers change our Coalesce Transform nodes? It proposes node changes as a diff for the owner to merge through your normal review, then verifies the rebuilt table in Snowflake.

Sources

  • •Coalesce, Coalesce Catalog product page, https://coalesce.io/product/catalog/ (checked Oct 3, 2026)
  • •Coalesce, "Coalesce 2.0: The Agentic Data Engineering Platform" (Sep 15, 2026), https://coalesce.io/product-technology/coalesce-2-0-the-agentic-data-engineering-platform/ (checked Oct 3, 2026)
  • •Coalesce homepage (Transform, Catalog, Quality; Scout), https://coalesce.io/ (checked Oct 3, 2026)
  • •Coalesce Docs, Catalog MCP integration (endpoints, OAuth, token auth deprecated, ten read tools), https://docs.coalesce.io/docs/catalog/developer/mcp-integration.md (checked Oct 3, 2026)
  • •Coalesce Docs, Catalog MCP glossary entry, https://docs.coalesce.io/docs/reference/glossary/catalog-mcp.md (checked Oct 3, 2026)
  • •Coalesce Docs, Introduction to the Catalog AI Assistant, https://docs.coalesce.io/docs/catalog/ai-assistant/introduction.md (checked Oct 3, 2026)
  • •Coalesce Docs, Catalog Public API, https://docs.coalesce.io/docs/catalog/developer/catalog-apis/public-api.md (checked Oct 3, 2026)
  • •Coalesce Docs, Tables (Certified, PII, owners, popularity), https://docs.coalesce.io/docs/catalog/assets/tables.md (checked Oct 3, 2026)
  • •Coalesce Docs, PII Tagging Suggestions, https://docs.coalesce.io/docs/catalog/document-your-data/pii-tagging-suggestions.md (checked Oct 3, 2026)
  • •Coalesce Docs, Request Data Access, https://docs.coalesce.io/docs/catalog/collaborate/request-data-access.md (checked Oct 3, 2026)
  • •Coalesce Docs, Column Description Propagation, https://docs.coalesce.io/docs/catalog/document-your-data/catalog-scribe/column-description-propagation.md (checked Oct 3, 2026)
  • •Coalesce Docs, Snowflake integration for Catalog, https://docs.coalesce.io/docs/catalog/integrations/data-warehouses/snowflake.md (checked Oct 3, 2026)
  • •Coalesce Docs, Sigma integration for Catalog (lineage is asset-level), https://docs.coalesce.io/docs/catalog/integrations/data-viz/sigma.md (checked Oct 3, 2026)
  • •Data Workers open-source repository, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
  • •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)