Product
Product11 min readBy The Data Workers Team

You're on Collibra Data Quality: It Scores Quality Inside Your Governance Platform. Data Workers Fixes the Cause and Proves It

Collibra Data Quality & Observability scores and observes quality inside the governance platform. Data Workers diagnoses each breaking rule, fixes it behind approvals and verifies it.

Your governance office runs on Collibra. Every business term has a steward, critical data elements are flagged in their domains, and Collibra Data Quality & Observability watches the tables behind them. DQ jobs push down into your warehouse's compute or run on a dedicated Spark engine. Adaptive rules learn each table's row count, nulls, empties and uniqueness, then flag a run as breaking when it leaves the band. Custom rules, written by hand or drafted by Collibra's AI, cover the business logic. Scores land on assets in the data marketplace and count as compliance evidence against your policies. When a rule breaks, lineage shows the root cause upstream and the owners downstream, and an issue workflow assigns the steward and logs every action. Collibra Data Quality scores and observes quality inside the governance platform. Data Workers diagnoses the cause, fixes it behind approvals and verifies the fix.

Put simply, Collibra Data Quality is the smoke alarm, wired into your governance office. Data Workers is the crew. When a rule on a CDE breaks, the cause usually sits in a system Collibra describes but does not run: a dbt seed, a Databricks job, an SAP configuration change.

Key takeaways

  • •Collibra keeps its job. Rules, scores, CDEs, policies, issue workflows and the catalog stay with your governance office.
  • •A breaking rule arrives with a diagnosis. The breaking rule reaches Data Workers through the on-call, whose assistant can use Collibra's MCP server side by side with Data Workers; Data Workers traces the cause across SAP S/4HANA, Databricks, dbt and Prefect, and maps everything downstream.
  • •One change set, one named approver. The dbt change, the rerun and a new test go to the owner in Slack or email; nothing changes until they approve in Spellbook.
  • •Verified before the steward closes the issue. Data Workers queues the rerun, checks every table the fix touched and closes the page; the steward resolves the Collibra issue with the receipt link.
  • •Autonomy per domain. Finance, risk and customer data each climb from L0 manual to L4 autonomous at the pace your governance office sets.

Collibra Data Quality is the scorekeeper. Data Workers is the crew.

Month-end close week at a manufacturer that bought a UK distributor on October 1. The estate: SAP S/4HANA, an SAP Datasphere replication flow to Amazon S3, Databricks on AWS with Unity Catalog, dbt on Databricks scheduled by Prefect, Sigma for the close workbooks, Collibra Data Quality & Observability with a pushdown job on the general ledger, and Opsgenie for the finance data on-call. Times are US Eastern. This is an illustration, not a customer case.

TimeSystemWhat happens
Tue 18:00SAP S/4HANAFinance activates company code 4100 for the new UK entity, and its first journal entries post
Wed 01:00SAP Datasphere and DatabricksThe replication flow writes the ACDOCA delta to S3; Auto Loader lands 38,600 new journal lines for 4100 in finance.bronze.acdoca
02:00Prefect and dbtThe gl_nightly flow runs dbt on Databricks. stg_sap__acdoca inner-joins the company_codes seed, which has no row for 4100, so every 4100 line drops out of finance.gold.fct_gl_journal. Every test passes
03:00Collibra DQThe pushdown job on fct_gl_journal runs. The adaptive row-count rule is breaking at 11% below its learned band and the job score falls from 98 to 71. The score on the CDE "Net revenue" drops in the data marketplace, and the issue workflow assigns the finance data steward
03:01OpsgenieThe job's notification reaches Opsgenie, which pages the finance data on-call
03:06Data WorkersReads the job run and rule results from Collibra. Profiles bronze and gold in Databricks: every missing line is company code 4100. Lineage runs from fct_gl_journal back through stg_sap__acdoca to the seed join; The dbt manifest diff rules out a schema change. Blast radius: fct_gl_journal, fct_revenue_by_entity, the FP&A consolidation snapshot and the Sigma "Close: revenue by entity" workbook
03:15Data WorkersReads the Collibra business term "Net revenue": consolidated across every company code from the acquisition date. Proposes one change set: a seed row mapping 4100 to entity UK01 in GBP from October 1, a left join with a dbt relationships test so an unmapped company code fails the build instead of disappearing, and a rerun of gl_nightly. The approval request goes to the finance data owner
07:30SpellbookThe finance data owner reviews the diff, the definition it follows and the blast radius, and approves; the analytics engineer merges the dbt change
07:40PrefectData Workers queues the gl_nightly flow run; dbt builds, and the new test passes
08:20Databricks and OpsgenieThe monitor_metrics baselines show line counts per company code from bronze to gold and the UK01 revenue total back in range, the owner reports the Collibra score recovered, and Data Workers writes the receipt and closes the Opsgenie alert
08:30Collibra DQThe steward reruns the job; the adaptive rule passes, the score returns to 98, and the steward resolves the Collibra issue with the receipt link
09:00SigmaThe close workbook queries Databricks live and shows UK01 revenue for the 10:00 close review
Incident timeline across the stack: what Collibra Data Quality, your team and Data Workers each do, step by step

Collibra did what a governance platform should: the adaptive rule caught a drop nobody had written a rule for, and the issue reached the right steward. What changed is everything after the alert. The on-call woke up to a diagnosis that reached a missing seed row, a blast radius that included the close workbook, and one change set waiting for one approval. The steward got a receipt, the evidence an auditor asks for later.

JobWhat Collibra Data Quality doesWhat Data Workers does
Defining good dataAdaptive rules learn each table; custom rules and templates encode the business logic; AI drafts rule SQL from plain languageReads those rules and the glossary as context and runs its own checks on the tables it repairs
DetectionPushdown or Spark jobs score every run and flag breaking rulesWatches freshness, volume and schema itself, so many breaks are caught before a rule runs
Scoring and evidenceScores sit on assets, data products and CDEs, linked to policiesJoins the breaking rule to lineage, pipeline runs, schema changes and incident history to find the cause
RoutingIssue workflows assign the steward, notify owners and log every action, in Collibra or in Jira and ServiceNowRoutes the proposed change to the named owner of the system that broke
The fix inside CollibraRules, jobs, templates and assets are created and tuned, including from an AI client over the MCP serverLeaves rules, jobs and stewardship to Collibra
The fix outside CollibraThe steward takes the issue to the owning teamProposes the change to the model, job or table with blast radius and queues reruns after approval
VerificationThe next job run updates the scoreChecks freshness, volume, schema and quality on every downstream table the fix touched
The recordThe issue history and the score over timeA receipt: the cause, the change, who approved it, what it touched, how it was verified and how to undo it

Why doesn't Collibra just do this itself?

Because Collibra drew the right line for a governance platform. Every write its MCP server offers changes a Collibra object: an asset, a classification, a data contract, a DQ job or a rule, and "every action uses Collibra's existing permission model." On the open-source server, the data quality tools are off by default and opt-in as a group with one flag. Collibra's own pages differ on their status: the Product Resource Center's MCP tools reference (Aug 28, 2026) lists them as experimental and available on the local server only, while the server's README (updated Oct 2, 2026) calls them generally available. Either way, you turn them on deliberately, and most of their writes return a preview until called again with confirm=true. Its agent work points the same way: Maestro, a no-code agent builder for governance teams, entered public preview on September 23, 2026, with guardian agents due through AI Command Center in October. None of it edits your dbt project or reruns your pipelines.

Stewards and auditors trust a quality score because the platform that produces it does not also change the data it scores. Fixing a dbt seed, rerunning a Prefect flow and checking the tables downstream is a different product with a different liability: each change scoped by blast radius, approved by its owner, reversible, verified and recorded, in systems Collibra does not own. Data Workers is built for that job.

Every tool owns a slice. Data Workers covers the whole lifecycle

Collibra Data Quality owns data quality inside the governance platform: rules that learn, scores tied to CDEs and policies. Each point tool adds another console, another contract and another handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Collibra Data Quality goes deep on its own area
StageData WorkersCollibra Data QualityWhy we scored it this way
Catalog & Context97.5Quality scores sit on catalog assets, data products and CDEs in the Collibra Platform. Data Workers keeps one governed context graph of definitions, owners, lineage, quality and usage across every platform.
Analytics & Insights83Executive dashboards roll up quality scores. Data Workers answers data questions from governed definitions with lineage behind every number.
Data Quality88.5Collibra's home stage: profiling, adaptive rules that learn each table, custom rules, AI rule authoring, pushdown jobs and scores tied to policies and CDEs. Data Workers runs checks too and fixes the cause when a rule breaks.
Observability & Incidents8.56Lineage shows the root cause upstream and the owners downstream, and issue workflows assign and log the work. Data Workers diagnoses across systems, fixes with approval and verifies the fix.
Pipelines & Ingestion8.52Jobs run inside the warehouse or on Spark, and they watch the pipeline's output. Data Workers queues reruns and backfills through your orchestrator, with approvals.
Schema & Migration83Profiling notices column and type changes on monitored tables. Data Workers scores a schema change's blast radius and drafts the migration with rollback SQL.
Governance & Access8.57.5Quality links to policies, CDEs and stewards, so the score counts as compliance evidence. Data Workers dry-runs each warehouse access request and proposes a time-bound grant for the owner.
Security & Privacy84Role-gated scorecards, exports and sample data controls protect what the jobs read. Data Workers' pull request review flags new columns whose names or annotations look sensitive, and leaves a receipt on every data change.
Cost / FinOps81Pushdown spends warehouse compute on checks; it does not manage spend. Data Workers reads warehouse spend next to each incident and drafts setting changes for their owner.
MLOps & Models7.54Quality signals connect to AI models in the platform, and AI Command Center governs them. Data Workers keeps the data under your models fresh and correct.

These scores cover the Data Quality & Observability product. Collibra's catalog and governance, scored as a whole platform, are in the data catalog comparison and in you're on Collibra; for the category view, read Data Workers vs data observability and beyond data observability: autonomous resolution.

How Collibra Data Quality and Data Workers work together

Spellbook Data Catalog (in preview) is where data owners review each proposed change, see who approved it and roll it back. Data Context Wizard keeps one governed context graph, the Data-Agents Swarm does the work with 20+ specialist agents, and the Autonomous Data-Conductor runs each fix end to end: detect, diagnose, fix, review, verify, remember.

How Data Workers fits with Collibra Data Quality: your coding agent on top, Data Workers in the middle, your estate underneath

Collibra to Data Workers. Data Workers connects to Collibra over its API or MCP server today. With the data quality tools turned on, Collibra's MCP server returns DQ jobs, runs, scores and per-rule results, plus business terms, column semantics, lineage and data contracts. Collibra's public API also returns DQ job and monitor results (generally available in release 2026.07, Oct 2, 2026). Databricks and Unity Catalog, dbt, Prefect, Opsgenie and Slack are native Data Workers connectors. SAP S/4HANA, SAP Datasphere and Sigma connect over their APIs; Sigma is read by design.

diagnose_incident and get_root_cause work the evidence, trace_cross_platform_lineage follows the data back through the models, blast_radius_analysis maps every table, snapshot and workbook downstream, the dbt manifest diff rules schema in or out, and get_incident_history checks for repeats. remediate proposes the change set; the dbt change arrives as a diff for the owner to merge. After approval, Data Workers queues the rerun through Prefect, the monitor_metrics baseline and the Collibra rules verify the result, and every step lands in get_audit_trail, a hash-chained log. Data Workers closes the Opsgenie alert it verified; rules, jobs and issues stay with your Collibra administrators and stewards, and the steward resolves the issue with the receipt from Spellbook.

Side by side in one client. Every Data Workers agent is an MCP server; the client setup guide documents the path: clone the open-source repo and add one start-agent.sh entry per agent. Example .mcp.json for Claude Code, with Collibra's open-source MCP server over stdio, its credentials in ~/.config/collibra/mcp.yaml as Collibra's README describes, and the data quality tools turned on:

{
  "mcpServers": {
    "collibra": {
      "type": "stdio",
      "command": "/path/to/chip",
      "args": ["--data-quality"]
    },
    "dw-context-catalog": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-context-catalog"] },
    "dw-incidents": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-incidents"] },
    "dw-quality": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-quality"] },
    "dw-schema": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-schema"] }
  }
}

Ask "why did the row-count rule on fct_gl_journal break, and what does the fix touch?" and the client calls both: Collibra returns the run, the breaking rule and the CDE; Data Workers returns the cause upstream, everything downstream and a proposed repair. For a shared endpoint, the Data Workers remote server takes an API key (bearer) or OAuth tokens from your identity provider, such as Okta or Entra ID, verified through JWKS.

One breaking Collibra rule, L0 to L4, set per domain:

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. An engineer traces the issue by hand.
  • •L1 observe. Data Workers posts the diagnosis and blast radius. Nothing changes.
  • •L2 propose. The dbt change, rerun and test wait for the data owner's approval.
  • •L3 act reversibly. Proven classes, such as a rerun after an approved fix, are queued and verified.
  • •L4 autonomous. In a scoped domain with a clean record, repeatable steps run end to end; model changes still go to their owners.

For the safety model, read is it safe to let AI agents change production data and how approvals work. On where data lives: the agents run in your infrastructure and hold the credentials, your data stays in your systems, and the hosted Conductor sees workflow metadata only. Connectors are listed in Data Workers integrations.

What changes for your team

Collibra gave your governance office a score everyone trusts and a steward for every CDE. Data Workers gives those stewards a crew.

Six jobs that run on autopilot with Data Workers next to Collibra Data Quality, with a concrete example of each
  • •Incidents. A breaking Collibra rule arrives with a diagnosis, a blast radius and a proposed fix, and closes with a receipt.
  • •Data quality. A falling CDE score becomes a fixed cause upstream, verified before the next job run, with a new test where the break entered.
  • •Cloud spend. Warehouse spend sits next to each incident, and setting changes go to their owner drafted.
  • •Access. A warehouse access request is dry-run and becomes a scoped, time-boxed grant proposed for the data owner.
  • •Audits. Collibra records the score and the steward's work; Data Workers records what changed, who approved it and how to undo it.
  • •Migrations. A move off legacy systems runs in approved waves, with parity checks planned and tracked for each wave.

The stewards change most: their issues arrive with the cause and a fix awaiting approval, so their time goes to rules, definitions and policies. Read Data Workers for data governance leads.

Keep Collibra Data Quality, or consolidate?

Keep Collibra Data Quality if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

Most governance offices keep Collibra: CDE scores the auditors accept are hard-won. What they consolidate is the work around it: a second monitoring tool on the same tables, hand-built remediation runbooks and spreadsheets of open quality issues. If you are weighing a fix layer of your own on Collibra's MCP server and a coding agent, read build it ourselves with Claude Code and MCP servers: the read calls are the easy part; the context graph, approvals, rollback and receipts are the work. See also our earlier page on Data Workers vs Collibra, and across the category, you're on Ataccama, you're on Informatica Data Quality and you're on DQLabs.

The case for your CFO

The outcome: Data Workers turns each breaking rule from a morning of cross-team tracing into a diagnosis on arrival and a fix behind one approval, so the close, the board pack and regulatory filings run on corrected numbers, on schedule.

The risk story is plain. Autonomy is set per domain: L1 only reads; L2 changes nothing until a named person approves; L3 applies only changes it can undo. An unanswered approval request expires and escalates, never auto-grants. No agent can promote its own work. Every change carries a receipt, and an org-wide stop halts all autonomous dispatch. Zero migration: Collibra, Databricks, dbt and every other system stay where they are.

Why now: adaptive rules and AI-drafted checks find problems faster than ever, so the slow part is the fix across teams, and close windows are short. The first win is read-only: every breaking rule in one domain gets a diagnosis and a blast radius, while rules, stewardship and change control stay the same. For the numbers, see the ROI of agentic data operations.

The sentence to repeat upstairs: "Collibra tells us which data we can trust and who owns it; Data Workers fixes what breaks across our systems, with one approval, and hands the steward the proof."

Getting started

Start with a pilot. Pick one governed domain such as finance close, connect Collibra, your orchestrator, warehouse and on-call tool, and run at L1 so every breaking rule gets a diagnosis and a blast radius. Then turn on a first write class at L2, such as reruns after an approved fix. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.

FAQ

How does Data Workers connect to Collibra Data Quality? Over Collibra's API or MCP server today. With the data quality tools turned on, Collibra's MCP server exposes DQ jobs, runs, scores and rule results along with business terms, lineage and data contracts; Collibra's public API also returns DQ job and monitor results.

Are Collibra's MCP data quality tools generally available? Collibra's pages differ. The server's README (Oct 2, 2026) says generally available; the MCP tools reference (Aug 28, 2026) says experimental and local server only. Both agree they are off by default and opt-in, so check with Collibra for your deployment.

Collibra's AI already writes rules, and its MCP server can create jobs. Why add Data Workers? Collibra's AI and MCP tools work on Collibra's objects: rules, jobs, templates and assets. Data Workers works on the systems that produced the bad data, from diagnosis to approved change, rerun, verification and receipt.

Does Data Workers change our Collibra rules, jobs or issues? No. Rules, jobs, templates, scores and issue workflows stay with your Collibra administrators and stewards. Data Workers reads them, and its receipt lives in Spellbook and the audit trail for the steward to attach.

Can Data Workers fix bad records or our dbt models directly? It proposes. Model changes arrive as a diff for the owner to merge, and data cleanups go to the owner to approve and apply. After approval, Data Workers queues the reruns through your orchestrator and verifies the result.

Our auditors need proof a CDE was fixed and stayed fixed. What do they get? Collibra's score history shows the CDE recovered. Data Workers adds a receipt per change (cause, change, approver, what it touched, checks passed, how to undo it) in a hash-chained audit trail.

Sources

  • •Collibra, Data Quality & Observability product page (monitors, pushdown or Spark, scores, CDEs, issue workflows), https://www.collibra.com/products/data-quality-and-observability (checked Oct 3, 2026)
  • •Collibra, MCP Server product page ("production-ready", supported clients, "Every action uses Collibra's existing permission model"), https://www.collibra.com/products/mcp-server (checked Oct 3, 2026)
  • •Collibra, open-source MCP server collibra/chip README (data quality tools off by default behind --data-quality, described as generally available, confirm=true checkpoints; updated Oct 2, 2026), https://github.com/collibra/chip (checked Oct 3, 2026)
  • •Collibra Product Resource Center, MCP tools reference (Aug 28, 2026: data quality tools listed as experimental, opt-in, local server only), https://productresources.collibra.com/docs/collibra/latest/Content/ModelContextProtocol/ref_mcp-tools.htm (checked Oct 3, 2026)
  • •Collibra Product Resource Center, Release 2026.07 (Oct 2, 2026: adaptive rule reset, DQ job and monitor results via public API, DQ services in the core platform), https://productresources.collibra.com/docs/collibra/latest/Content/ReleaseNotes/Archive/ref_release-202607.htm (checked Oct 3, 2026)
  • •Collibra Product Resource Center, Data Quality & Observability (self-hosted) Release 2026.04 (May 1, 2026: Databricks Pushdown, SQL Assistant for Data Quality), https://productresources.collibra.com/docs/collibra/dqc/latest/Content/DataQuality/DQReleaseNotesArchive/ref_dq-release-202604.htm (checked Oct 3, 2026)
  • •Collibra Product Resource Center, Create a Pullup job (default adaptive monitors, data lookback and learning phase), https://productresources.collibra.com/docs/collibra/latest/Content/UnifiedDataQuality/ta_create-pullup-job.htm (checked Oct 3, 2026)
  • •Collibra, press release, AI Command Center launch (May 6, 2026), https://www.collibra.com/company/newsroom/press-releases/collibra-launches-ai-command-center-to-scale-agentic-ai (checked Oct 3, 2026)
  • •Collibra, press release, Maestro (public preview), guardian agents and agent contracts (October via AI Command Center) (Sep 23, 2026), https://www.collibra.com/company/newsroom/press-releases/collibra-launches-new-capabilities-to-reduce-the-hallucination-tax-on-enterprise-ai (checked Oct 3, 2026)
  • •Data Workers open-source repository, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
  • •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)