How does Data Workers help with SOX, HIPAA, GDPR and the EU AI Act?
Data Workers doesn't make you compliant. It gives you the controls and evidence auditors ask for when agents touch data: named approvals, scoped access, PII handling, tamper-evident receipts and human oversight, mapped to each regulation's text.
Data Workers doesn't make you compliant; your controls, your people and your auditors do that. What Data Workers gives you is the controls and the evidence auditors ask for when agents touch data: change approvals with a named approver, scoped least-privilege access, privacy checks on pull requests and redaction, tamper-evident receipts in an audit trail, human oversight set per domain, and agents that run in your own environment.
An auditor asks about agents what they ask about any engineer: who changed what, who approved it, what could it reach, and can you prove the record wasn't edited. Data Workers, the agentic data platform, answers those questions by design. This page maps each control to the regulation's own words. It is not legal advice.
Key takeaways
- •Controls and evidence, not a compliance claim. You and your auditors decide how the evidence fits your control framework.
- •SOX: change management for agents. Production changes at a gated level wait for a named person, never the agent that proposed them, and leave a receipt with the diff, the approver, the blast radius and the rollback path.
- •HIPAA and GDPR: least privilege and minimisation. Agents act through scoped roles, access grants expire, pull request review flags new columns whose names look sensitive, and the facts and learnings agents write to the context graph are PII-scrubbed first.
- •EU AI Act: oversight and logging, where they apply. The Digital Omnibus on AI, Regulation (EU) 2026/1744, moved the high-risk requirements, including Articles 12, 14 and 26, to 2 December 2027 for Annex III systems. Procurement teams already borrow these articles as the standard, and Data Workers ships those controls today.
- •Your data stays in your systems. The agents run in your infrastructure on every tier with your credentials and your model key; the hosted Conductor sees workflow metadata only.
What auditors ask, and the evidence Data Workers leaves
Every control below comes from the product and the live security, Autonomous Data-Conductor and Spellbook Data Catalog pages.

Who approved this change? You set each domain on one ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous. Data Workers ships observe-only, and a domain can run in shadow mode, where nothing executes and the permission ladder logs what approval each action would have needed.

When a domain's level calls for a person, the proposal waits in the Spellbook inbox (in preview) with its diff, blast radius and rollback plan, and the approval records the approver's identity. The approval guards reject agent identities, placeholders and self-approval, so an agent never signs off its own work. An unanswered approval expires and escalates; it never turns into a yes. How approvals work for AI data agents walks through the queue.
What could the agent reach? Agents act through the warehouse roles you grant, with secrets scoped per agent, so your platform's permissions bound every change. Remote access authenticates through your identity provider, such as Okta or Entra, with Data Workers verifying tokens against your JWKS endpoint; it issues no tokens of its own. provision_access applies least privilege and column-level permissions with a 90-day expiry by default, and check_policy evaluates each action against your policies before it runs.
Was personal or health data exposed? Your classification tool and data owners say what a table holds; Data Workers' pull request review flags new columns whose names or annotations look sensitive, reading names, not values. Facts agents contribute to the context graph, and the learnings they record, are PII-scrubbed before they land. How Data Workers handles PII goes deeper.
What exactly happened, and can the record be trusted? Every tool call is written to a SHA-256 hash-chain audit log: agent, tool, tenant, user, action, resource, sanitized arguments, outcome (including denied calls) and time. An edit anywhere breaks the chain, and verify_global_hash_chain checks it end to end. Auditors read it through get_audit_trail, and generate_audit_report produces a period report with its evidence chain.
Can a person stop it or undo it? Every production change records its rollback path before it runs, and a domain can drop back to L1 observe at any time. For erasure requests, Data Workers purges a person's data from its context graph and writes a tombstone into the audit chain that records the erasure, never the content; the chain still verifies.
Where does it run? The agents run in your infrastructure on every tier, hold the warehouse credentials and model key, and write the context graph, receipts and audit log there; Enterprise also runs on-premise and air-gapped. The Autonomous Data-Conductor is hosted by us and receives workflow metadata only (goals, signals, table names, proposals with diffs, run records), never rows, credentials or model keys. In sovereign mode, external model calls are blocked, inference goes to a local Ollama or vLLM endpoint, and operations are logged with an Article 12 tag. Data Workers in your VPC or air-gapped and where does our data go? cover deployment.
The regulations, in their own words
SOX. Section 404(a) requires each annual report to "state the responsibility of management for establishing and maintaining an adequate internal control structure and procedures for financial reporting" and to assess its effectiveness (15 U.S.C. 7262). SEC Rule 13a-15(f) defines that control to include procedures giving reasonable assurance that receipts and expenditures are made "only in accordance with authorizations of management and directors". PCAOB AS 2201 tells the auditor to "understand how IT affects the company's flow of transactions". An agent change to a model that feeds revenue sits inside your IT general controls, and the named approver and the receipt are what a change-management walkthrough samples.
HIPAA. The Security Rule requires access control that allows access "only to those persons or software programs that have been granted access rights" (45 CFR 164.312(a)(1)), audit controls that "record and examine activity" in systems holding ePHI (164.312(b)), integrity controls against "improper alteration or destruction" (164.312(c)(1)). The minimum necessary standard limits PHI "to the minimum necessary to accomplish the intended purpose" (164.502(b)), and the administrative safeguards require regular review of "audit logs, access reports" (164.308(a)(1)(ii)(D)). Scoped roles, expiring grants, privacy checks on pull requests, the hash chain and the audit report map to those lines, and because the agents run in your environment, PHI stays in your systems.
GDPR. Article 5(1)(c) requires personal data to be "adequate, relevant and limited to what is necessary", and Article 5(2) makes the controller "able to demonstrate compliance". Article 25(2) limits default processing to what is necessary, including "their accessibility". Article 32(1) names confidentiality and integrity and "a process for regularly testing, assessing and evaluating" security measures. Article 17(1) gives the right to erasure "without undue delay". Scoped access, scrub-on-write to the graph, the audit trail as accountability evidence, and erasure with a provable tombstone are the Data Workers side of each.
EU AI Act. Regulation (EU) 2024/1689 entered into force on 1 August 2024 and became generally applicable on 2 August 2026. The Digital Omnibus on AI, Regulation (EU) 2026/1744 of 8 July 2026, in force since 27 July 2026, moved the high-risk requirements in Chapter III Sections 1 to 3 to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I products. Those sections hold Article 12, under which high-risk systems "shall technically allow for the automatic recording of events (logs)", Article 14, under which they can be "effectively overseen by natural persons", including the ability to stop the system so it can "come to a halt in a safe state", and Article 26(6), under which deployers keep logs "of at least six months". Per-domain autonomy levels, named approvals, shadow mode and an exportable hash-chained log line up with those articles wherever your counsel decides they apply.
One worked example: an agent fix during the quarter-end close
This is an illustration, not a customer case. On day 2 of the close, an upstream currency table lands in Snowflake with a duplicated rate date. The finance domain runs at L2 propose.
| Step | System | What happened | Evidence left |
|---|---|---|---|
| Detect | Snowflake, dbt | fct_revenue falls off its monitor_metrics revenue baseline | Audit entry for the check and the trace |
| Assess | Data Workers | blast_radius_analysis finds four dbt models, the revenue Explore in Looker and the close export | Blast radius attached to the proposal |
| Propose | dbt, Spellbook | A dedupe step in stg_fx_rates with dry-run row counts and a rollback path | Diff and rollback plan in the queue |
| Approve | Okta, Spellbook | The assistant controller signs in through the company IdP and approves | Approver identity and time; never an agent |
| Apply and verify | Data Workers, dbt | Change runs with the finance role; the baseline holds and the owner's reconciliation passes | Receipt in the hash chain |
| Test | Internal audit | Pulls the quarter's finance receipts with generate_audit_report | Evidence chain for the walkthrough sample |

Nobody had to reconstruct the change from chat logs. The auditor's sample is a query.
The alternatives buyers weigh
Platform-native audit logs. Databricks exposes an audit log system table at system.access.audit, in Public Preview (docs updated Sep 11, 2026), and Snowflake's ACCESS_HISTORY view shows object access for the last 365 days on Enterprise Edition or higher (checked Oct 2, 2026). Keep them. Data Workers adds the layer above them: which agent proposed a change across dbt, Airflow and Looker, who approved it, and how to undo it.
Compliance automation platforms. Vanta, for example, says it will "automate audit prep and evidence collection" across 400+ integrations, and lists AI governance among its products (checked Oct 2, 2026). Data Workers produces the agent-change evidence those frameworks ask for, and its audit report is something they can collect.
Coding agents with broad credentials. Claude Code, Cursor and Codex are excellent at writing changes, and their approval prompts go to the person running the session, on that person's credentials. Connected to Data Workers over MCP, the same agents act through the per-domain gate, and the approval goes to the domain's owner with a receipt. Build it ourselves with Claude Code and MCP servers prices the alternative, and is it safe to let AI agents change production data? compares platform-native agents and read-only observability tools.
Why don't these tools do this themselves? Focus. A warehouse log is right to record its own platform, and a compliance tracker is right to collect evidence rather than change data. Writing to production across systems, with approvals, rollback and receipts, is a different product, and it is the one Data Workers is.
The case for your CFO
The business outcome is a shorter, calmer audit. When agents fix data, the evidence your auditors sample is already written, chained and exportable, so nobody spends the week before fieldwork reconstructing changes from tickets and chat.
The risk story: each domain sits at the autonomy level you choose, from L0 manual to L4 autonomous. Finance can stay at L2 propose with a named approver on every change, and no agent can approve its own work. Sensitive column names are flagged in pull request review, PII is scrubbed from the context graph, the agents run in your environment with your model key, and the hosted Conductor sees workflow metadata only.
Why now: agents already change production data, and auditors are starting to ask how. The EU's high-risk dates now start in December 2027, which gives teams time to put oversight and logging in place before a regulator or a customer asks.
The first win is the finance domain at L2 propose for one quarter, with internal audit pulling receipts for its walkthrough (see Data Workers for data governance leads). What stays the same: your warehouse, dbt, orchestrator, BI, identity provider and control framework. Nothing migrates.
Start with a pilot ($7,500 one-time, credited in full against the first year); Scale is from $1,000 a month and Enterprise from $3,000 a month, billed annually, with unlimited seats and no usage meter. See pricing, the ROI calculator and the ROI of agentic data operations. The sentence for upstairs: "Agents change our data under the same change controls as our engineers, and every change has a receipt our auditors can test."
FAQ
Does Data Workers make us SOX, HIPAA or GDPR compliant? No product can; compliance is judged by your auditors and regulators. Data Workers supplies controls and evidence for the part of those programs where agents touch data. This page is not legal advice.
Can an agent approve its own change? No. The approval guards reject agent identities, placeholders and self-approval, and the receipt records the approving person's identity. Which people may approve a domain follows the roles you set.
Do we need a business associate agreement with Data Workers for PHI? The agents run in your environment and read through grants you give them, so PHI stays in your systems, model calls go to a provider account you hold or a local model, and the hosted Conductor receives workflow metadata only. Your counsel decides what agreements your setup needs, and we answer the questions they bring.
How long are audit records kept? The hash-chained log lives in your deployment, so retention follows your policy. Set it to cover your audit cycle and, where the EU AI Act applies to a system, the six-month minimum in Article 26(6).
How does erasure work with an append-only audit log? Data Workers purges the person's data from its context graph and writes a tombstone into the chain that records the erasure, never the content. The chain still verifies afterwards.
Is Data Workers a high-risk AI system under the EU AI Act? Classification depends on the use case, and data operations agents are not an Annex III use case on their face. Your counsel decides; either way, Data Workers gives you human oversight per domain and exportable logs.
Sources
- •15 U.S.C. 7262 (Sarbanes-Oxley Act Section 404), U.S. Code 2024 edition via GPO govinfo: https://www.govinfo.gov/link/uscode/15/7262?type=usc&year=mostrecent&link-type=html (checked Oct 2, 2026; uscode.house.gov was under maintenance)
- •SEC Rule 13a-15(f), 17 CFR 240.13a-15, eCFR current as of Sep 30, 2026: https://www.ecfr.gov/current/title-17/section-240.13a-15 (checked Oct 2, 2026)
- •PCAOB AS 2201, An Audit of Internal Control Over Financial Reporting: https://pcaobus.org/oversight/standards/auditing-standards/details/AS2201 (checked Oct 2, 2026)
- •HHS HIPAA Security and Privacy Rules, 45 CFR 164.308, 164.312 and 164.502, eCFR current as of Sep 30, 2026: https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164 (checked Oct 2, 2026)
- •GDPR, Regulation (EU) 2016/679, Articles 5, 17, 25 and 32, EUR-Lex: https://eur-lex.europa.eu/eli/reg/2016/679/oj (checked Oct 2, 2026)
- •EU AI Act, Regulation (EU) 2024/1689, Articles 12, 14, 26 and 113, EUR-Lex: https://eur-lex.europa.eu/eli/reg/2024/1689/oj (checked Oct 2, 2026)
- •Digital Omnibus on AI, Regulation (EU) 2026/1744 of 8 July 2026, OJ 24.7.2026: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=OJ:L_202601744 (checked Oct 2, 2026)
- •European Commission, AI Act policy page (updated Aug 3, 2026): https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai (checked Oct 2, 2026)
- •European Commission, AI Omnibus enters into force: https://digital-strategy.ec.europa.eu/en/news/ai-omnibus-enters-force (checked Oct 2, 2026)
- •Databricks, Audit log system table reference (updated Sep 11, 2026): https://docs.databricks.com/aws/en/admin/system-tables/audit-logs (checked Oct 2, 2026)
- •Snowflake, ACCESS_HISTORY view: https://docs.snowflake.com/en/sql-reference/account-usage/access_history (checked Oct 2, 2026)
- •Vanta, automated compliance: https://www.vanta.com/products/automated-compliance (checked Oct 2, 2026)
- •Data Workers product repository,
data-workers-agent-swarmmain @ 0c2491e3: core/enterprise/src (approval-workflow, promotion-guard, autonomy-controller, credential-scoper, audit/tamper-evident-log, rtbf, pii-middleware, sovereign-config), agents/dw-governance/src/sovereign-mode.ts (checked Oct 2, 2026) - •Data Workers public repository tools (
scan_pii,provision_access,check_policy,get_audit_trail,verify_global_hash_chain,generate_audit_report,blast_radius_analysis): https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026) - •Data Workers security page: https://dataworkers.io/security/ (checked Oct 2, 2026)
- •Data Workers, where does our data go? (what runs where, the pull request privacy check): https:///blog/where-does-our-data-go-security-privacy-deployment/ (checked Oct 2, 2026)
- •Data Workers pricing: https://dataworkers.io/pricing/ (checked Oct 2, 2026)