How does Data Workers handle PII?
Data Workers flags new PII columns in pull request review, flags PII in tool responses, redacts it before anything is stored, routes PII changes to a person and erases a subject with a tamper-evident tombstone. Step by step, with a worked example.
Data Workers flags PII where it changes, with agents that run in your infrastructure: its pull request review reads new column names and annotations, never values, and flags the ones that look personal, while your classification tool and the data owner decide what a table holds. PII is flagged in tool responses and redacted before anything is written to the context graph or the audit records, masking changes and access grants go to a named person, your own engine enforces the mask, and an erasure request deletes the subject and leaves a tamper-evident tombstone.
Data Workers is the agentic data platform, and PII handling is built into the path every agent takes. Below is each step, what the code does at that step, and one worked example from a new email column to an approved mask.
Key takeaways
- •Flagged in review, classified by your tools. Data Workers' pull request review flags new columns whose names or annotations look personal, reading names, not values. Your classification tool and the data owner decide what a table holds.
- •Every tool response is scanned. Findings are logged against the tool call with a count, so a PII leak into an agent's output is visible and countable.
- •Redacted before it is stored. Writes to the context graph are redacted with typed placeholders such as
[REDACTED_EMAIL], raw samples are dropped, and audit records redact PII-named arguments and secrets. - •A person decides on PII. Policies route PII writes to review, a review verdict never grants access, and PII masking can never be promoted to auto-resolve, so every masking change waits for the data owner.
- •Your warehouse enforces the mask. Snowflake masking policies, Unity Catalog column masks and BigQuery data policies stay the enforcement engine, by design.
- •Erasure is provable. A subject's data leaves the graph, and a tombstone in the hash chain records that it happened, never what it was.
How it works, step by step

1. Flag it where it changes. When a pull request adds or renames a column, Data Workers' review reads the column names and annotations in the change for hints such as email, phone, ssn and dob, and flags the ones that look personal before merge. It reads names, never values. What a table already holds is classified by your own tool (Snowflake classification, Databricks data classification, Sensitive Data Protection or your catalog) and confirmed by the data owner, and the classifications the owner confirms land in the context graph, where every agent works from them.
2. Flag it in every response. Every agent tool runs inside the same middleware, and each block of every response is scanned with the shared detector (emails, phone numbers, US Social Security numbers, card numbers with a Luhn check, IP addresses, titled names, IBANs, credential formats). Findings are logged against the call with a count; the response itself is passed on unchanged. A Luhn check drops card-shaped numbers that aren't valid cards, an IBAN checksum drops look-alikes, and context checks keep credential detectors from firing on ordinary IDs.
3. Redact it before anything is stored. Facts that agents contribute to the context graph, and the learnings they record from tool output, go through one scrubber. Free text and node names are redacted with typed placeholders, sample_values, raw_data and connection_string are dropped outright, and a write carrying a credential or a prompt-injection string is blocked. The shared detector also supports hash or remove for your own redaction paths, and scan_pii checks what Data Workers' own store holds, so you can confirm no raw value was kept. PII classification tags land in the graph, so the next agent knows contact_email is PII without ever seeing a value. Audit records follow the same rule: PII-named arguments become [PII_REDACTED], secrets become [REDACTED], and Article 12 logs redact content and hash user identifiers.
4. Keep it in your systems. The agents run in your infrastructure on every tier and query the warehouse in place with your credentials. Your data stays in your systems; the hosted Autonomous Data-Conductor sees workflow metadata only: goals, signals, table and column names, proposals with diffs, run records and approval handles. Where does our data go? maps every boundary.
5. Put a person on every PII change. check_policy evaluates each action against your rules, and a rule such as "review PII writes" returns a review verdict. A review is never a yes: provision_access treats a review verdict as a no and grants nothing, and the grants it does record expire after 90 days by default. request_governance_review opens a trackable review for the owner. Each domain sits on the autonomy ladder you choose:

PII masking sits outside the climb by design: the tool that promotes a fix class to standing auto-resolve permanently refuses PII masking and column renames. The agent drafts the mask or the grant, and a named person from your identity provider approves it. An unanswered request expires and escalates, never auto-grants. How approvals work for AI data agents walks through the queue.
6. Propose the mask, let your engine enforce it. For a classified column, Data Workers drafts the masking change for the platform's own engine (a Snowflake masking policy, a Unity Catalog column mask, a BigQuery data policy), with the blast radius from blast_radius_analysis attached. Renames get the same care: a pull request that renames a column shows the old and new names side by side in review, the privacy check flags a new name that looks personal, and the owner confirms the classification so the mask follows the data, not the name. Every masking change is a dry-run proposal that always requires approval: the data owner or platform admin applies it in the warehouse, and Data Workers never writes a mask on its own. Each receipt carries the diff, the approver, the blast radius and the rollback path. How to roll back an AI agent change shows the undo.
7. Erase on request, and prove it. On an erasure request, Data Workers deletes the subject's nodes and edges from the context graph and appends a tombstone to the SHA-256 hash-chained audit log: subject, counts, requester and time, never the erased content. The chain is re-verified after the append. verify_global_hash_chain checks it end to end and generate_audit_report assembles the evidence.
A worked example: a new email column lands
This is an illustration, not a customer incident. The agents run in the customer's Azure subscription, the model is Azure OpenAI in the same tenant, people sign in through Microsoft Entra ID, and the Conductor is hosted by Data Workers. The CRM domain runs at L2 propose.

At 10:14 a pull request adds contact_email and contact_name to the dbt staging model over main.crm.contacts in Databricks. At 10:16 Data Workers' review reads the change, the column names and the model's annotations, never the values, and flags both columns as likely PII on the pull request. At 10:17 the privacy owner confirms both and adds the free-text support_notes column, which Databricks data classification had already tagged as holding email addresses. The Conductor records the signal: table, column names and source, with no values.
At 10:19 trace_cross_platform_lineage shows that the dbt model dim_customer reads all three columns and feeds the Customer 360 report in Power BI. At 10:21 the proposal lands with the privacy owner: Unity Catalog column masks on the three columns, with the blast radius (one dbt model, one report) and the rollback path. The owner signs in through Entra ID and approves at 10:48. The platform admin applies the approved masks in Unity Catalog and confirms that a read as a non-privileged user returns masked values; Data Workers writes the receipt, the audit entry and the classification. No values left the workspace.
Where the warehouse-native features fit
Each platform ships strong PII controls, and Data Workers drives them.
| Platform | What it provides | Source, date |
|---|---|---|
| Snowflake | Masking policies: schema-level objects applied to a column at query time; sensitive data classification assigns semantic and privacy categories as system tags, with an optional AI mode in preview; both need Enterprise Edition or higher | Snowflake docs, checked Oct 2, 2026 |
| Databricks | Unity Catalog column masks: a SQL UDF applied at query time; ABAC policies apply masks across tables from governed tags | Databricks docs, updated Sep 24, 2026 |
| Databricks | Data classification: an agentic AI system classifies and tags sensitive data with class.* governed tags; view scanning in Beta | Databricks docs, updated Sep 17, 2026 |
| BigQuery | Sensitive Data Protection discovery builds table and column data profiles with infoTypes and sensitivity levels; dynamic data masking applies rules by role through policy tags and data policies | Google Cloud docs, updated Sep 30, 2026; checked Oct 2, 2026 |
Keep them. Each classifies and masks inside its own platform, the right design for a warehouse. A data team's PII crosses platforms: it lands through Fivetran, flows through dbt, shows up in a BI extract and gets asked about by an assistant. Data Workers runs the loop across all of them with one context, one approval flow and one audit trail: flag the column, trace where it flows, propose the mask in each engine, get the owner's approval, record the owner's check and keep the receipt.
The other options buyers weigh. A catalog records the classification and leaves the change to a person (Data Workers vs a data catalog). A coding agent such as Claude Code writes a masking policy on its user's credentials; connected to Data Workers over MCP, the same request goes through policy check, review and receipt. Building the privacy check, scrubber, approval gate and erasure chain yourself is priced in Can't we just build this ourselves?. Writing PII changes to production across systems, with approvals, rollback and receipts, is a different product, and it is the one Data Workers is. Is it safe to let AI agents change production data? covers write safety in full.
The case for your CFO
The outcome: new sensitive columns are flagged in review and protected across every platform, and the evidence is already written when an auditor or a customer asks.
The risk story: agents start read-only, PII masking can never be set to auto-resolve, and every masking policy or access grant waits for a named person from your directory. Data Workers stores classifications and redacted text, never raw samples, and every change carries a receipt in a hash-chained log. Your data stays in your systems, model calls go to your own provider account or a local model, and the hosted Conductor sees workflow metadata only.
Why now: assistants and agents already reach your warehouse, so PII exposure is a question of when, not whether. One governed path is cheaper than reviewing each tool on its own.
The first win is one warehouse's PII loop: privacy checks on every pull request, your existing classifications into the context graph, masking proposals for owners to approve, receipts for the auditor. What stays the same: your warehouse, its masking engine, your catalog, your identity provider and your model provider. Nothing migrates.
Start with a pilot: $7,500 one-time, and the pilot is credited in full against the first year. Scale is from $1,000 a month and Enterprise from $3,000 a month, billed annually, with unlimited seats, no usage meter, no markup on model spend, bring your own model and an Apache 2.0 core. See /pricing/, the ROI calculator and the ROI of agentic data operations.
The sentence for upstairs: "Every new PII column is flagged in review, masked with an owner's approval and logged, across all our platforms, without our data leaving our systems."
FAQ
Does any PII ever leave our environment? No values. Data Workers' privacy check reads column names and annotations, not values, and model calls go to the model account you hold the key for. Turn on sovereign mode and inference stays on your own Ollama or vLLM endpoint. The hosted Conductor receives table and column names and proposals, never values.
Will Data Workers mask a column without asking? No. PII masking is a change class that can never be promoted to standing auto-resolve, so a named owner approves every masking change, and access grants on a review verdict wait for a person too.
Does it replace Snowflake classification, Databricks data classification or Sensitive Data Protection? No; keep them. Their classifiers serve their own platforms well, and their masking engines enforce what owners approve. Data Workers adds the loop across platforms with approvals and receipts.
How does erasure work with an append-only audit log? The subject's data is deleted from the context graph and a tombstone recording the erasure, never the content, is appended to the hash chain. The chain still verifies. SOX, HIPAA, GDPR and the EU AI Act maps this to Article 17.
Can we run all of this with no internet access? Yes, on an Enterprise on-premise deployment with sovereign mode on and outbound traffic blocked; see Can Data Workers run in our VPC or air-gapped?. For one assistant, read How to give Claude access to Snowflake without exposing PII; background in automated PII detection in data warehouses, column masking policy automation and GDPR DSAR automation.
Sources
- •Data Workers product repository,
data-workers-agent-swarmmain @ 0c2491e3: agents/dw-governance/src (pii-scanner, tools/scan-pii, policy-store, engine/access-provisioning), core/enterprise/src (pii-middleware, agent-middleware, graph-pii-scrubber, audit/tool-call-audit-middleware, audit/tamper-evident-log, eu-ai-act/article12-pii-redactor, eu-ai-act/pii-approval-gate, rtbf), agents/dw-migration/src/tools (auto-resolve-policy), agents/dw-context-catalog/src/tools (auto-tag-dataset, graph-query-tools, purge-subject), checked Oct 2, 2026 - •Data Workers public repository tools (
scan_pii,check_policy,request_governance_review,provision_access,trace_cross_platform_lineage,blast_radius_analysis,verify_global_hash_chain,generate_audit_report), checked Oct 2, 2026: https://github.com/DataWorkersProject/dataworkers-claw-community - •Data Workers, Pricing (agents run in your infrastructure; the Autonomous Data-Conductor hosted by us), checked Oct 2, 2026: https://dataworkers.io/pricing/
- •Data Workers, Security and trust, checked Oct 2, 2026: https://dataworkers.io/security/
- •Snowflake, Understanding Dynamic Data Masking, checked Oct 2, 2026: https://docs.snowflake.com/en/user-guide/security-column-ddm-intro
- •Snowflake, Sensitive data classification, checked Oct 2, 2026: https://docs.snowflake.com/en/user-guide/classify-intro
- •Databricks, Row filters and column masks (updated Sep 24, 2026), checked Oct 2, 2026: https://docs.databricks.com/aws/en/tables/row-and-column-filters
- •Databricks, Data classification (updated Sep 17, 2026), checked Oct 2, 2026: https://docs.databricks.com/aws/en/data-governance/unity-catalog/data-classification
- •Google Cloud, Sensitive Data Protection data profiles (updated Sep 30, 2026), checked Oct 2, 2026: https://docs.cloud.google.com/sensitive-data-protection/docs/data-profiles
- •Google Cloud, Introduction to BigQuery column data masking, checked Oct 2, 2026: https://docs.cloud.google.com/bigquery/docs/column-data-masking-intro