You're on Collibra: Your Governance Office Sets the Policy. Data Workers Keeps the Estate True to It
On Collibra? Keep it as the governance office of record. Data Workers reads its policies, classifications and AI use cases and keeps the running estate true to them, behind approvals.
Your governance office runs on Collibra. Assets live in domains inside communities, every business term has a data steward, and workflows route each change to the person who answers for it. Data classes and classification matches mark which columns hold personal data. AI Command Center keeps the registry of every AI use case, model and agent, with the assessment that says what each may touch. Collibra Data Access turns classifications into masking and filtering in Snowflake, Databricks and BigQuery. Since September 23, Maestro agents help the governance team itself, and the Collibra MCP Server hands that context to Claude, ChatGPT, Databricks, Cursor and the rest of your AI estate. Collibra is the governance office of record. Data Workers is the operations crew that keeps the governed estate true: it catches drift from what your stewards decided, proposes the fix where the cause lives, routes the Collibra-side change to the steward and leaves a receipt.
Policies are written once; the estate changes every night. A new column lands and an AI agent reads it before the next classification scan. That gap is this guide's job.
Key takeaways
- •Collibra keeps its job. Glossary, stewardship, classification, policies, AI Command Center and Data Access stay where they are. Data Workers works next to them from day one, over Collibra's MCP server or REST API.
- •Collibra's AI agents serve the governance team. Maestro (public preview since September 23, 2026) builds no-code agents for stewards. Guardian agents with agent contracts are announced for October 2026 to supervise how AI agents behave. Data Workers does the operations work in the data systems underneath.
- •Drift is caught in the estate, not at the next scan. Data Workers flags new columns that are unclassified or whose names look sensitive and checks them against the policies and AI use cases Collibra holds.
- •Fixes land where each owner works. The dbt change goes to the model owner as a diff. Classifications and notes go to the Collibra steward as proposals. Collibra Data Access does the masking.
- •Every change leaves a receipt: trigger, diff, blast radius, checks before and after, approver and undo path, ready for the steward to link in Collibra.
- •Start with a pilot on one governed domain, read-only first, on the ladder from L0 manual to L4 autonomous.
Collibra is the governance office of record. Data Workers is the operations crew that keeps the estate true.
Collibra is where an enterprise writes down what its data means, who answers for it and what each use may touch. The estate keeps moving: SaaS tools add fields, ingestion lands them, dbt passes them along, AI agents read them. Data Workers covers that side, measured against what Collibra allows.
Here is a night at a subscription software company. Intercom holds support conversations, Fivetran lands them in Databricks, dbt Cloud builds the models, Unity Catalog holds the tables, and a support copilot built on Databricks answers from a context table. The copilot is registered in Collibra AI Command Center, and its approved assessment says its context data holds no direct identifiers. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| Tue 15:20 | Intercom | Support ops adds a free-text conversation attribute, "Issue summary". Some agents paste the customer's email address and callback number into it |
| Wed 01:00 | Fivetran | The Intercom connector syncs on its own schedule; a new column issue_summary lands in raw_intercom.conversation in Databricks |
| 02:00 | dbt Cloud | The nightly job builds stg_intercom__conversations and support_copilot_context. The context model selects with dbt_utils.star and an except list, so the new column passes straight through. Every test passes |
| 02:15 | Data Workers | dbt's refreshed catalog shows a new column on support_copilot_context. Its name, issue_summary, gives no PII hint and Data Workers does not read values, so it flags an unclassified free-text column feeding the copilot |
| 02:18 | Collibra | The on-call's assistant reads the table's asset over the Collibra MCP server and hands it to Data Workers: it is linked to the "Support copilot" AI use case, whose assessment allows no direct identifiers, and the new column has no classification match. blast_radius_analysis shows the copilot's index (next refresh 12:00) and one support QA dashboard downstream |
| 02:20 | Spellbook and Slack | Data Workers opens a governance review with request_governance_review and drafts one change set: a dbt diff that adds issue_summary to the except list, a request for the steward to classify the staging and raw columns, and a finding note for the AI use case. Approval requests reach the analytics engineer and the steward in Slack |
| 08:35 | Spellbook and GitHub | The analytics engineer reviews the diff and blast radius, approves and merges it |
| 08:50 | Collibra | The steward checks the column, finds email addresses and phone numbers, and applies Email Address and Phone Number classification matches. The Collibra Data Access policy for personal data now masks issue_summary in Databricks outside the support QA group |
| 09:00 | dbt Cloud | The owner reruns support_copilot_context in dbt Cloud after approval; it succeeds |
| 09:20 | Databricks | Data Workers verifies in the rerun's dbt catalog that the column is gone from the context table; the steward confirms the column mask on the staging column. It writes the receipt |
| 09:30 | Collibra | The steward links the receipt to the Support copilot assessment |
| 12:00 | Databricks | The copilot's index refreshes from a context table with no direct identifiers |

Collibra did its job: the policy existed, the use case was assessed, and Data Access was ready to mask anything classified as personal data. Nobody had classified a column that did not exist the day before. The fix lived in dbt, which Collibra doesn't run, and in a classification only the steward should apply. Both went through a named person, and the identifiers never reached the copilot.
| Job | What Collibra does | What Data Workers does |
|---|---|---|
| The policy | Holds policies, data classes, AI use cases and assessments, with stewards and workflows | Reads them as governed context and measures the running estate against them |
| The inventory | Harvests assets and lineage; the MCP server serves them to AI clients | Joins Collibra's assets to live schema changes, dbt runs, usage and quality results |
| The drift | Classifies what its scans and stewards have seen | Catches a new column the night it lands and scans its values for personal data |
| The fix | Records the decision and routes Collibra changes through workflows | Proposes the change in the system that owns the cause, such as a dbt diff, and queues the rerun |
| The Collibra-side change | The steward applies classifications, notes and status | Drafts them as proposals for the steward, with the evidence attached |
| The enforcement | Data Access masks and filters by classification in Snowflake, Databricks and BigQuery | Checks the mask is attached after the steward classifies |
| The proof | Keeps the governance record and the assessment | Verifies downstream and writes a receipt the steward links to the assessment |
Why doesn't Collibra just do this itself?
Because Collibra built the product an enterprise trusts to decide, and that trust depends on staying the neutral record for every system around it. Its write surface follows from that job. The Collibra MCP Server, which Collibra calls production-ready, writes Collibra objects: it creates and edits assets, adds or removes classification matches, pushes data contract manifests and proposes glossary terms. "Every action uses Collibra's existing permission model, so AI acts within your governance framework, not around it." Its data quality tools stay off until you turn them on, and their writes preview until called again with confirm=true. Maestro builds agents for governance teams, and guardian agents, announced for October 2026 in AI Command Center, are designed to supervise AI agents against agent contracts. All of it serves governance, the right design for a governance office.
Rewriting a dbt model, changing a Databricks table or rerunning a job is a different product with a different liability: blast radius across every system, scoped credentials in each, a named owner's approval, a rollback path and downstream proof. Those systems belong to other vendors and your data team; a governance platform that edited them would be grading its own homework. Data Workers is the product on the other side of that line, and it hands every Collibra change back to the steward who owns it.
Every tool owns a slice. Data Workers covers the whole lifecycle
Collibra owns its slice outright: the catalog of record and the governance on it. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the governance you already run in Collibra.

| Stage | Data Workers | Collibra | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 9.5 | Collibra's home stage: assets in domains and communities, a business glossary with stewards, lineage and semantic context, served to AI clients over its MCP server. Data Workers brings that record in as context with provenance, next to schema changes, runs, usage and quality. |
| Analytics & Insights | 8 | 4 | Collibra helps people and agents find and trust data rather than analysing it. Data Workers' Insights agent answers questions from governed context and checks the numbers behind them. |
| Data Quality | 8 | 7 | Collibra Data Quality & Observability runs monitors and rules, and its MCP tools preview changes before applying them. Data Workers repairs what fails, in the pipeline or model that caused it, and proves the fix. |
| Observability & Incidents | 8.5 | 5 | Collibra surfaces quality results and lineage for impact. Data Workers detects, diagnoses, fixes and verifies across systems, with a receipt for every incident. |
| Pipelines & Ingestion | 8.5 | 2 | Collibra harvests metadata from pipelines; it does not run them. Data Workers queues reruns and backfills through your orchestrator and catches source changes before they spread. |
| Schema & Migration | 8 | 3 | Collibra records schemas as assets and shows lineage. Data Workers detects schema changes the night they land, assesses their impact and plans migrations in approved waves. |
| Governance & Access | 8.5 | 9 | Collibra's other home stage: policies, stewardship workflows, assessments, AI Command Center's registry with EU AI Act, NIST AI RMF and AIUC-1 templates, and Data Access policies. Data Workers runs the request queue, dry-runs grants and keeps the estate matched to the policy. |
| Security & Privacy | 8 | 7 | Data Access masks and filters by classification in Snowflake, Databricks and BigQuery. Data Workers flags new sensitive column names in pull request review and asks the steward to classify them. |
| Cost / FinOps | 8 | 2 | Collibra does not manage warehouse spend. Data Workers traces Snowflake credits to the dbt model behind them and drafts the fix for its owner. |
| MLOps & Models | 7.5 | 7 | AI Command Center registers every AI use case, model and agent and links them to their data and assessments. Data Workers keeps the data under those models healthy and true to what the assessment allows. |
How Collibra and Data Workers work together
Your people stay where they are: Claude, ChatGPT or Cursor for questions and code, Collibra for stewardship and policy. Spellbook Data Catalog (in preview) is where the data team looks: each finding, proposed change, blast radius, approver and rollback, with approval requests reaching people in Slack or email. Underneath, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm does the work with more than 20 specialist agents, the Autonomous Data-Conductor runs each fix end to end (detect, diagnose, fix, review, verify, remember), and per-domain guardrails hold approvals, receipts and rollback.

What Context Wizard does with Collibra. It reads the assets, terms, classifications, contracts and AI use cases you point it at, with provenance (domain, steward, when observed), and joins them to the systems Data Workers reaches natively, including Databricks with Unity Catalog, Snowflake, BigQuery, dbt and Airflow, so a Collibra asset traces to the dbt model that builds it and the agent that reads it. Agents get one governed view through explain_table, trace_cross_platform_lineage and blast_radius_analysis. No agent can promote its own work to authoritative; that takes a named person, the way your stewards already work.
Setup over MCP today. Data Workers connects to Collibra over the Collibra MCP Server or its REST API today. The open-source server, collibra/chip, is a single binary that reads its Collibra URL and credentials from a config file or environment variables, and every tool call runs under that account's Collibra permissions. Give the account Data Workers uses read scopes: in this design, changes to Collibra are proposals the steward applies. Data Workers' agents come from the open-source repository: clone it and add start-agent.sh entries to your client config, as the client setup docs show.
// Example: .mcp.json for Claude Code
{
"mcpServers": {
"collibra": {
"type": "stdio",
"command": "/usr/local/bin/chip",
"env": { "COLLIBRA_MCP_API_URL": "https://<your-instance>.collibra.com" }
},
"dw-context-catalog": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-context-catalog"]
},
"dw-governance": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-governance"]
},
"dw-schema": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-schema"]
}
}
}List the tools with your client's own command (/mcp in Claude Code). A governance lead can then ask "does anything feeding the Support copilot hold personal data its assessment doesn't allow?" and get the use case from Collibra, the lineage and the steward's classifications in one answer.
Where writes go. The dbt fix goes to the model owner as a diff to merge. Reruns are queued through your orchestrator, or by the owner in dbt Cloud. Ingestion stays with Fivetran: Data Workers reads what lands in the warehouse and never triggers a sync. Classification matches, notes and status changes go to the Collibra steward as proposals, and approved facts land in the Context Wizard graph. Masking stays with Collibra Data Access, by design. Where no policy covers a column, Data Workers drafts the masking change as a dry run for the owner to apply; it never applies a mask itself.
One request, L0 to L4. The autonomy ladder is set per domain.

- •L0 manual. Connected, not acting. Stewards check the estate by hand.
- •L1 observe. Data Workers checks new columns against Collibra's policies and AI use cases, reports every gap with lineage and owners and logs what each action would have needed. Nothing changes.
- •L2 propose. Data Workers drafts the dbt diff and the Collibra classification proposals with their blast radius. The owner and the steward approve before anything runs.
- •L3 act reversibly. For proven change classes, such as queuing the rerun of an affected model and re-scanning it after a fix merges, Data Workers acts on its own, verifies and records it. Code changes still arrive as diffs for the owner to merge, or as pull requests when your team turns on the GitHub pull-request target.
- •L4 autonomous. For a scoped, trusted class in one domain, Data Workers runs the loop end to end and posts the receipt for review. Collibra changes still go to the steward.
Read more on safety, approvals and where your data goes. To choose a catalog, see our Collibra, Alation, DataHub and OpenMetadata comparison; for monitoring, the Collibra Data Quality guide. The same pattern holds for Atlan and Alation, and the data catalog question ties them together. Earlier pages: Data Workers vs Collibra, Collibra alternatives with AI agents and an MCP server for Collibra metadata.
What changes for your team

Governance offices spend much of their week after the decision: chasing which systems implement a policy, classifying new columns and explaining changes to auditors. With Data Workers, those jobs run on autopilot at the level you set.
- •Incidents. A source change that breaks a governed model is caught the night it lands, traced, fixed through its owner and recorded.
- •Data quality. Every failing check on a governed asset gets a cause, a fix in the system that owns it and a downstream check that it held.
- •Cloud spend. Snowflake credits are traced to the dbt model behind them, and the fix is drafted for that model's owner.
- •Access. A request for data Collibra classifies as sensitive arrives dry-run: effective privileges after role inheritance, the sensitive columns it reaches and a least-privilege grant with an expiry for the owner.
- •Audits. Each fix links the policy or assessment it serves, the approver, the diff and the rollback path, ready for the steward to attach in Collibra.
- •Migrations. When a platform moves, each wave is planned with its parity checks and Collibra links, and the owner signs off before it closes.
Stewards get their week back for the decisions only people can make: what a term means, what a use case may touch, who owns a domain.
Keep Collibra, or consolidate?
Keep Collibra if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
For most regulated enterprises, the glossary, stewardship, assessments, AI Command Center and Data Access belong in Collibra. What teams consolidate is the tooling around it: scripts that diff a Collibra export against dbt YAML, a separate column scanner, spreadsheets mapping AI use cases to tables, tickets asking an engineer to change a model. If you are weighing building this layer yourself on the Collibra MCP server, read build it ourselves with Claude Code and MCP servers: reading Collibra is the easy part; cross-system context, approvals and rollback are the work.
The case for your CFO
The outcome: what the governance office approves is what the data estate does. Personal data stays out of places its assessment forbids, and audit evidence exists before anyone asks. The Collibra investment shows up as fewer exposures and faster fixes.
The risk story is plain. Data Workers reads Collibra under its own permission model and proposes every Collibra change to the steward who owns it. Every change in the estate shows its blast radius, goes to a named approver, lands through your existing dbt review, is verified downstream and leaves a receipt: trigger, diff, approver, checks before and after, and how to undo it. Unanswered requests expire and escalate; they never auto-grant. Autonomy is set per domain from L0 manual to L4 autonomous, and an org-wide stop halts all autonomous dispatch. Zero migration.
Why now: AI agents read whatever the estate holds the moment it lands, and a policy enforced at the next scan arrives after the agent did. The first win is one governed domain, such as the data behind one registered AI use case, read-only, with a report of every gap between Collibra's policies and the running tables. What stays the same: your Collibra licence, stewards, workflows and policies. Start with a pilot; the path is on the pricing page, and the pilot is credited in full against the first year. For the numbers, see the ROI of agentic data operations.
The sentence to repeat upstairs: "Collibra is where we decide the rules; Data Workers makes sure every system follows them, and every fix comes with an approval and a receipt."
Getting started
Start with a pilot. Pick the domain your governance office cares about most, such as the data behind a registered AI use case or a regulated report, connect the Collibra MCP server read-only next to Data Workers, and let it report gaps between Collibra's policies and the running estate for a few weeks before turning on the first fix class. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.
FAQ
Does Data Workers write to Collibra? Data Workers proposes Collibra changes, such as classification matches, notes or status changes, and the steward applies them in Collibra. The account Data Workers uses can stay read-only. Collibra's own MCP server can write, under Collibra's permission model, for the people and agents you allow.
How is this different from Collibra Maestro and guardian agents? Maestro gives governance teams no-code agents for governance work, and guardian agents, announced for October 2026, are designed to supervise how AI agents behave against agent contracts. Data Workers works in the data systems underneath: dbt, Databricks, Snowflake and your orchestrator. Its own agents act only through approvals and leave receipts your governance office can review.
Does Data Workers replace Collibra Data Access masking? No. Data Workers proposes the classification, the steward applies it, and Collibra Data Access applies the masking policy in Snowflake, Databricks or BigQuery. Data Workers then checks that the mask is attached and records it in the receipt.
Which connections are native, and which go through Collibra? Data Workers connects natively to Databricks with Unity Catalog, Snowflake, BigQuery, dbt, dbt Cloud, Airflow and Slack, among 50+ connectors. It connects to Collibra over the Collibra MCP Server or REST API today, and to Fivetran and Intercom over their APIs.
Where does our data go? The agents run in your infrastructure on every tier and hold the warehouse credentials and model key. Your data stays in your systems, and the hosted Conductor sees workflow metadata only: goals, signals, table names, proposals with diffs, run records and approval handles.
Sources
- •Collibra, Collibra MCP Server product page, https://www.collibra.com/products/mcp-server (checked Oct 3, 2026)
- •Collibra, "Collibra launches new capabilities to reduce the hallucination tax on enterprise AI" (press release, Sep 23, 2026), https://www.collibra.com/company/newsroom/press-releases/collibra-launches-new-capabilities-to-reduce-the-hallucination-tax-on-enterprise-ai (checked Oct 3, 2026)
- •Collibra, AI Command Center (AI governance) product page, https://www.collibra.com/products/ai-governance (checked Oct 3, 2026)
- •Collibra, Collibra Data Access product page, https://www.collibra.com/products/protect (checked Oct 3, 2026)
- •Collibra,
chipopen-source MCP server README and tool list (release v0.0.47, Aug 7, 2026), https://github.com/collibra/chip (checked Oct 3, 2026) - •dbt Labs, dbt-utils
starmacro, https://github.com/dbt-labs/dbt-utils (checked Oct 3, 2026) - •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)
- •Data Workers open-source repository, agent tools and start-agent.sh, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
- •Data Workers, Spellbook Data Catalog product page, https://dataworkers.io/product/spellbook-data-catalog/ (checked Oct 3, 2026)