Industry
Industry8 min readBy The Data Workers Team

Data Workers for data governance leads

What Data Workers changes for a data governance lead: privacy checks on changes, access requests and definition conflicts handled by agents, controls proposed for a named approver, and audit evidence from a hash-chained log. Your catalog stays.

For a data governance lead, Data Workers turns policy documents into work that gets done: its agents flag sensitive columns in pull requests, dry-run and draft scoped access grants, capture definitions and catch conflicting ones, and propose every control for a named person to approve. Your week moves from chasing owners, tags and screenshots to setting the rules, approving the exceptions and handing auditors evidence the platform already wrote.

Key takeaways

  • •The agents do the stewardship legwork. pull request review flags new columns that look sensitive, provision_access drafts column-level grants with an expiry, and define_business_rule records definitions with author and source.
  • •Every control is a proposal a person approves. Masking is always a dry-run proposal a person applies, and catalog write-back is proposed for the steward. No agent can promote its own work to authoritative.
  • •Evidence is a by-product. Every agent action lands in a SHA-256 hash-chained audit log; generate_audit_report exports access, PII and violation reports for any period.
  • •Your catalog stays. Collibra, Atlan, Purview or Unity Catalog remains the system people browse; Data Workers does the operational work around it.
  • •You set the autonomy, per domain. Each domain runs from L0 manual to L4 autonomous, and regulated domains can stay at L2 propose for as long as you choose.

A governance lead's week today

The work is keeping rules true in an estate that changes daily. A dbt change adds a column nobody classified. Two teams define "active customer" differently and both numbers reach a deck. An owner field sits blank on a table of customer data. Access requests pile up in ServiceNow, each needing someone to check what the table holds. Internal audit asks for evidence of the PII controls, and the answer is a week of screenshots. Now the AI council asks how you will govern the agents too.

The tools: a catalog and glossary (Collibra, Alation, Atlan, Informatica CDGC, Microsoft Purview or Unity Catalog), privacy tools (BigID, OneTrust, Immuta), a ticket queue and policy documents in a shared drive.

dbt Labs' 2026 State of Analytics Engineering report found that 41% of respondents cite ambiguous data ownership as an ongoing challenge, effectively unchanged year over year, while the share who say increasing trust in data is important rose from 66% to 83%.

The same week with Data Workers

The agents take the legwork; you and your stewards keep the decisions.

Comparison matrix of Your governance program and Data Workers on the outcomes a data leader buys

What the agents take off.

  • •Classification. Data Workers' pull request review reads new column names and annotations, not values, and flags the ones that look personal; your classification tool and data owners stay the source of truth for what is in a table. Findings arrive as proposed tags and a dry-run masking policy for the platform's own engine. How Data Workers handles PII covers each step.
  • •Access requests. check_policy tests the request against your rules, and Data Workers dry-runs the proposed grant: the effective privileges after role inheritance, the PII, PCI or PHI columns it would reach, any policy conflicts, and a least-privilege recommendation with an expiry. provision_access then drafts a column-level grant, 90 days by default. On a review verdict it grants nothing and request_governance_review opens a trackable review for the owner. Unity Catalog grants are applied after approval; on other platforms the owner applies the drafted grant. Every grant and its expiry is recorded in a ledger your own access reviews can start from.
  • •Definitions. define_business_rule records a definition, calculation or constraint for a table or column with its author and source, and import_tribal_knowledge brings in what lives in runbooks and wikis. When catalogs disagree on a metric's definition, Data Workers flags the contradiction and routes it to the review inbox for a named person; it never picks a winner.
  • •Lineage for audits. trace_cross_platform_lineage follows a number back to its sources across engines.

What you still own and decide.

  • •The rules. You write the policies check_policy evaluates, such as data minimisation for a domain.
  • •Who owns what. Approvals go to a named person. An unanswered request expires and escalates; it never auto-grants. Promotion to authoritative needs a named human through mark_authoritative, and no agent can promote its own work. Who owns the agents sets out the split.
  • •How far each domain goes. Each domain runs at L0 manual, L1 observe, L2 propose, L3 act reversibly or L4 autonomous, and Data Workers ships observe-only. Your team reviews the agents' receipts against the bar you set and moves a domain up when it clears it. Autonomy levels L0 to L4 explains each rung, and how approvals work shows the approval flow.

How Data Workers fits your catalog and tools

Your catalog stays the place people browse and the system of record for your glossary. Catalogs do that job well, and Collibra and Atlan now ship MCP servers that bring their governance context to AI assistants. Unity Catalog connects to Data Workers natively for grants, and Microsoft Purview, DataHub, OpenMetadata, Collibra, Alation, Atlan and Informatica CDGC connect over their APIs or MCP servers today. Catalog write-back is proposed for the steward: a new PII tag or description arrives as a change to approve and apply, and approved facts land in the Data Context Wizard graph with provenance. Bring your own context shows how the Wizard plugs your glossary in as a first-class source.

Spellbook Data Catalog (in preview) is where you and your stewards look: one inbox to approve, send back or roll back the agents' work, with the receipt behind every item. From Claude or ChatGPT Enterprise, the same agents answer over MCP, with sign-in through your identity provider (Okta or Entra ID) and Data Workers verifying the tokens via JWKS. Requests that belong in a queue go there through create_servicenow_ticket or create_jira_sm_ticket.

The agents run in your infrastructure on every tier, holding your warehouse credentials and model key. Your data stays in your systems; the hosted Autonomous Data-Conductor sees workflow metadata only, never rows, credentials or model keys. Personal data is scrubbed from the facts agents record, and a right-to-be-forgotten request purges a subject from the Data Workers context graph with a tombstone in the hash chain. Erasure in your own warehouse stays with your process.

The metrics you are judged on

Governance leads are measured on ownership and classification coverage, exceptions closed, audit findings, time to answer a regulator or data subject, and quality scores.

MetricHow Data Workers moves itWhere the number comes from
Classification coveragePrivacy flags on pull requests in the domains you choose, tags proposed for approvalThe receipt for each flag and tag decision
Ownership coverageDefinitions recorded with an author; authority set by a named ownerContext Wizard records, mark_authoritative history
Access request timeGrants dry-run, policy-checked and drafted, ready to approveRequest-to-approval time in the receipts
Audit findingsEvidence exported from a log you can verify end to endget_audit_trail, verify_global_hash_chain
Data quality scoresChecks and scores per domain, fixes behind approvalget_quality_score

Record the baseline in the pilot and set targets per domain. How to measure AI data agents sets out the scorecard, and the ROI calculator puts hours on it.

A worked example: a new PII column, an access request and an audit ask

This is an illustration, not a customer case. A retailer runs Snowflake and dbt, catalogs in Collibra and takes requests in ServiceNow. The customer domain runs at L2 propose.

TimeWhoWhat happened
Mon 08:10Snowflake, dbtA dbt change adds date_of_birth to crm.customers
Mon 08:25Data WorkersPull request review flags the new column name as a likely date of birth
Mon 08:30Data WorkersDrafts a dry-run masking policy and a PII tag for the Collibra asset; routes both to the customer data steward
Mon 10:00ServiceNowA marketing analyst requests crm.customers for a churn study
Mon 10:10Data Workerscheck_policy runs; the grant's dry-run shows it would reach date_of_birth and email; provision_access drafts a 90-day, column-level grant without them
Mon 11:30Customer data stewardApproves the masking policy and the grant in Spellbook; the platform team applies both in Snowflake
Tue 09:00CollibraThe steward applies the proposed PII tag to the asset
Thu 14:00Internal auditAsks for evidence of PII controls on customer data this quarter
Thu 14:20Data Workersgenerate_audit_report exports the access report and the receipts for the masking and tag decisions; verify_global_hash_chain confirms the log is intact
Thu 15:00Customer data stewardReviews the evidence pack and sends it to internal audit
Incident timeline across the stack: what Your governance program, your team and Data Workers each do, step by step

The column was classified the morning it appeared, the analyst had scoped access by lunch, and nobody built the evidence by hand. SOX, HIPAA, GDPR and the EU AI Act maps these controls to the rules' own text; your auditors and counsel decide what satisfies them. This is not legal advice.

How your role differs from your CDO and CISO

Your CDO sets strategy, budget and which domains go first (Data Workers for CDOs). Your CISO owns security risk, identity and vendor review (Data Workers for CISOs). You own the rules in between: ownership, meaning, sensitivity, access, and the evidence that they held. Each of you reads the same receipts for a different question.

The case for your CFO

The outcome is governance that keeps up with the estate without adding headcount. Today stewards spend their weeks on classification, access tickets and evidence, so the program covers only the domains they have time for. Data Workers does that work behind approvals, so the same team covers more domains and answers auditors faster.

The risk story: autonomy is set per domain, starting observe-only, with regulated domains held at L2 propose. Masking and catalog changes are always proposals a person applies, and every change has a named approver and a receipt in a tamper-evident log. Nothing migrates: your catalog, privacy tools, warehouses and ticket queue stay, and the core is Apache 2.0.

Why now: AI assistants and coding agents already touch your data, adding columns to classify and requests to review. The first win is one domain, usually customer data, with privacy checks on pull requests and access requests at L2 for a quarter, and the receipts as the scorecard.

Start with a pilot ($7,500 one-time, credited in full against the first year). Scale is from $1,000 a month and Enterprise from $3,000 a month, billed annually, with unlimited seats, no usage meter, no markup on model spend, and you bring your own model. See pricing, and the ROI of agentic data operations for the arithmetic. The sentence for upstairs: "Agents do our stewardship legwork under rules we set, and every control they propose has a named approver and a receipt."

FAQ

Will agents change our definitions without sign-off? No. define_business_rule records who wrote a definition, conflicts go to a person, and only a named human can mark a source authoritative. No agent can promote its own work.

Can an agent apply masking on its own? No. Masking is always a dry-run proposal that a person applies or approves for their own change process, and it can never be promoted to run on its own. Is it safe to let AI agents change production data? walks through the controls.

Do we have to replace Collibra or Purview? No. Your catalog stays the system people browse, connected natively or over its API or MCP server. Write-back arrives as proposals for your stewards.

Will it keep evidence our auditors accept? Every agent action is in a hash-chained log that verify_global_hash_chain checks end to end, and generate_audit_report exports the reports. Your auditors decide what satisfies them.

Does our data leave our environment? Your data stays in your systems; the hosted Conductor sees workflow metadata only, and sovereign mode blocks external model calls. Where does our data go? covers it in full.

Could our engineers build this with Claude Code and MCP servers? They could build agents. Approvals, receipts, blast radius and per-domain autonomy are the larger job. Build it ourselves prices that path.

For more on stewardship, see the AI data stewardship guide and data steward vs data owner.

Sources

  • •dbt Labs, 2026 State of Analytics Engineering Report: https://www.getdbt.com/resources/state-of-analytics-engineering-2026 (checked Oct 2, 2026)
  • •Data Workers public repository tools (scan_pii, check_policy, provision_access, request_governance_review, generate_audit_report, define_business_rule, import_tribal_knowledge, mark_authoritative, trace_cross_platform_lineage, detect_schema_change, get_audit_trail, verify_global_hash_chain, get_quality_score, create_servicenow_ticket, create_jira_sm_ticket): https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)
  • •Data Workers product repository, data-workers-agent-swarm main @ 0c2491e3: approvals and promotion guard, grant dry-run (simulate-access.ts: effective privileges, sensitive columns, policy conflicts), contradiction routing to the review inbox (contradiction-sentinel), hash-chained audit log, PII middleware, right-to-be-forgotten purge (checked Oct 2, 2026)
  • •Collibra, Model Context Protocol (MCP) Server: https://www.collibra.com/products/mcp-server (checked Oct 2, 2026)
  • •Atlan, Atlan MCP documentation: https://docs.atlan.com/product/capabilities/atlan-ai/how-tos/atlan-mcp-overview (checked Oct 2, 2026)
  • •Data Workers product pages: https://dataworkers.io/product/data-context-wizard/ , https://dataworkers.io/product/spellbook-data-catalog/ , https://dataworkers.io/product/autonomous-data-conductor/ (checked Oct 2, 2026)
  • •Data Workers security, pricing and ROI calculator: https://dataworkers.io/security/ , https://dataworkers.io/pricing/ , https://dataworkers.io/roi-calculator/ (checked Oct 2, 2026)