Where does our data go? Security, privacy and deployment with Data Workers
The Data Workers agents run in your infrastructure with your credentials and your model key. The hosted Conductor sees workflow metadata, never your rows. What runs where, what crosses each boundary, and the receipts it leaves.
Your table contents, credentials and model keys stay in your environment: the Data Workers agents run in your infrastructure on every tier, query your warehouse in place through grants you give them, and send model calls only to a model account you hold the key for or to a local model you run. The Autonomous Data-Conductor, the orchestrator that keeps the loop turning, is hosted by us on every tier, and what it receives is workflow metadata: goals, detected signals, the names of the tables involved, proposals and run records.
Below is each component, what crosses each boundary, and the control on it, so a security reviewer can check every line against the product.
Key takeaways
- •The agents run where you run them. On every tier the agents and the context graph run in your infrastructure, with your credentials. Enterprise adds a dedicated VPC, your own cloud or on-premise.
- •The hosted Conductor orchestrates from metadata. It has no warehouse connector. Rows, credentials and model keys never reach it.
- •Your model, your key. Prompts go to the provider account you control or a local model, and Data Workers does not train on your data.
- •Read-only by default. Writes switch on per domain, from L0 manual to L4 autonomous, and anything irreversible needs a named approver.
- •Your identity provider stays in charge. Tokens are verified against your IdP's JWKS, or you issue API keys.
- •Every change leaves a receipt, chained into a SHA-256 audit log you can verify and export.
What runs where
Data Workers is the agentic data platform, built so the parts that touch your data sit inside your boundary. Per the rate card on /pricing/:
| Component | Where it runs | What it holds or sees |
|---|---|---|
| Data-Agents Swarm (the specialist agents) | Your infrastructure, on every tier; inside your coding agent | Warehouse credentials, connector secrets, your model key; queries data in place |
| Context graph and Data Context Wizard | Your deployment | Schemas, lineage, owners, quality signals, approved facts, PII-scrubbed |
| Receipts and the audit log | Your deployment, exportable | Diff, approver, blast radius, rollback path, hash-chained entries |
| Autonomous Data-Conductor | Hosted by us, on every tier | Goals, detected signals, table names, proposals (description and diff), run records, approval handles |
| Your model | Your provider account, or Ollama or vLLM on your hardware | Prompts and responses, under your agreement with that provider |
| Open-source core | Your laptop or server, launched by your MCP client over stdio | Your credentials and model key; nothing reaches us |
Why the Conductor can be hosted safely. It is the orchestration layer: it turns a goal such as "every table is documented" into routines, dispatches work to the domain agents, and records each run. It carries no warehouse connector and never queries the warehouse; warehouse access stays with the agents in your infrastructure. So the Conductor knows that raw.crm_contacts changed and that a masking proposal is waiting for approval; it never sees a row of raw.crm_contacts. Your no-touch list and autonomy ceilings live there too, and they only narrow what agents may do.
Air-gapped, on-premise. For estates with no internet path, Enterprise deploys on-premise from an offline bundle, scoped with you at contract: you verify the checksums, load the images into your private registry (the guide uses Harbor), and install with Helm onto your Kubernetes cluster (the guide uses RKE2). Postgres holds metadata, MinIO holds artifacts, and Ollama or vLLM serves the model on your own nodes, with GPUs optional. The guide's network table has four internal flows and zero outbound internet access, confirmed by a sovereign self-test and your network policy. Kubernetes and Helm are the platform's path; the open-source core runs locally over stdio. More in Can Data Workers run in our VPC or air-gapped?.
What crosses which boundary
Metadata, results and identity come in to the agents. Proposals, receipts and model calls go out to systems you own; workflow metadata goes to the hosted Conductor.

| What | Where it lives | Who can see it |
|---|---|---|
| Table contents | Your warehouse and lake, where they are now | Your users, and agents through the grants you give |
| Credentials and connector secrets | Where the agents run: your cloud or your machine | Your operators |
| Model keys | Your environment, on your provider account | Your provider bills you directly |
| Metadata, lineage, quality results | The context graph, in your deployment | Agents and people with the right role |
| Goals, signals, proposals, run records | The hosted Conductor | Your team, through approvals; Data Workers operates it |
| Receipts and audit log | Your deployment, exportable | Approvers, auditors, your security team |
What it never copies. Data Workers does not migrate your data into a new store. Rows stay in Snowflake, Databricks, BigQuery or Postgres, and agents query them in place when a task needs it. Credentials and connection strings never go into a prompt. Facts agents contribute to the context graph, and the learnings they record, go through one governed gate that scrubs PII first and fails closed, checks the ontology, and checks who has authority to promote a fact; fields such as sample_values, raw_data and connection_string are dropped outright.
Erasure. When a person exercises a right to be forgotten, the erasure tool deletes that subject's nodes from the graph and writes an erasure tombstone into the hash-chained audit: the fact of the deletion stays provable, the deleted content does not.
How agents reach your systems
Agents use the grants you issue, scoped per domain, and start read-only.

At L0 manual and L1 observe, agents only read. At L2 propose, an agent drafts the SQL, dbt change or policy and a named person approves it. At L3 act reversibly, it applies changes that carry a recorded rollback. L4 autonomous is reserved for change classes you have decided are routine. check_policy evaluates each action against your rules, request_governance_review routes anything sensitive to its owner, and provision_access turns approved requests into grants in your own engine. Unity Catalog, Snowflake roles and your IAM stay the permission system, by design. The full answer on write safety is in Is it safe to let AI agents change production data?.
How models are used
You bring the model. Calls go through one provider interface that supports Anthropic, OpenAI, Amazon Bedrock, Google Vertex AI, Azure OpenAI, and local models through Ollama or vLLM. With Bedrock, Vertex or Azure OpenAI, calls stay in your own cloud account; with Ollama or vLLM, on your hardware.
Most of what an agent sends is structure: table and column names, SQL, your question and error text, and only when local parsing and rules are not enough. When a task needs values, such as a query result or a few samples to classify a column, they go to the model account you control, under your agreement with that provider. Data Workers does not use customer data to train or fine-tune models. More in Which model does Data Workers use, and what does it cost to run?.
Sovereign mode. With it on, every call to an external model API is blocked in code with a SOVEREIGN_MODE_BLOCKED error, inference routes to your Ollama or vLLM endpoint, external telemetry is blocked, and operations are logged with the actor, resource, local model endpoint and time. If no local model is configured, inference stops; it never falls back to an external provider.
PII handling
Data Workers does not sample cell values for PII. Its pull request review reads new column names and annotations, never values, and flags the ones that look sensitive; your classification tool and the data owner decide what a table holds.
The same detector covers emails, phone numbers, US Social Security numbers, card numbers (with a Luhn check to cut false positives), IP addresses and titled names. Every tool response is scanned, and findings are logged with a count. Facts agents contribute to the context graph, and the learnings they record, are redacted with typed placeholders such as [REDACTED_EMAIL]; the shared detector also supports hash or remove, and you add patterns for your own identifiers, such as member or policy numbers. Classifications land in the context graph for the next agent. More in How does Data Workers handle PII?, and a walkthrough for one stack in How to give Claude access to Snowflake without exposing PII.
Identity and the audit trail
Remote endpoints accept an API key you issue or an OAuth token from your own IdP, such as Okta or Microsoft Entra ID. Data Workers verifies tokens against your issuer's published keys (JWKS), caches them and follows rotation, and maps claims to a tenant and a role. It issues no tokens and registers no clients, so people, roles and offboarding stay in your directory.
Every applied change carries a receipt: the diff, who approved it and when, the assessed blast radius, and the rollback path. Receipts are keyed to the approval that authorised them. Each action is also written to an audit log where every entry's SHA-256 hash covers its content plus the previous entry's hash; edit or delete one entry and the chain stops verifying. get_audit_trail returns entries, verify_global_hash_chain checks the chains end to end, get_usage_activity_log shows who called which tool with what outcome, and generate_audit_report assembles the record for a review.
A worked example
This is an illustration, not a customer incident. The agents run in the customer's AWS VPC, the model is Claude on Amazon Bedrock in the same AWS account, and the Conductor is hosted by Data Workers.

At 09:02 a pull request adds contact_email from raw.crm_contacts to the dbt staging model in Snowflake. Data Workers' review reads the column names and annotations, no values, and flags contact_email as likely PII; the data owner confirms it and adds the free-text notes column, which the team's classification tool had already tagged. The Conductor records the signal: table, columns, source. trace_cross_platform_lineage shows the dbt model dim_customer feeds the customers Explore in Looker. The model call that drafts the fix sees schema, lineage and policy text, with no row values.
At 09:07 the proposal lands with the governance owner: a Snowflake masking policy on both columns, with the blast radius (one dbt model, one Explore, the dashboards on it). The owner signs in through Okta and approves at 09:20. The owner applies the approved policy in Snowflake, a read as a non-privileged user returns masked values, and Data Workers writes the receipt and the audit entry. No row, credential or key left the AWS account; the Conductor holds the table names, the proposal and the run record.
How the alternatives handle your data
Platform assistants and coding agents have each made sensible choices about data, and Data Workers works alongside all of them.
| Option | Where the model runs | What the vendor commits to | Source, date |
|---|---|---|---|
| Databricks AI assistive features | Partner models (Azure OpenAI, OpenAI on Databricks, Anthropic on Databricks) or Databricks-hosted open models, set by a workspace setting | Not used to train foundation models offered to third parties; partner endpoints are zero data retention | Databricks docs, updated Sep 11, 2026 |
| Snowflake Cortex AI Functions | LLMs deployed within the Snowflake service perimeter; cross-region inference can move the prompt transiently to another region | Customer data stays stored in your account region and is not persisted in the processing region | Snowflake docs, checked Oct 2, 2026 |
| Gemini for Google Cloud, including BigQuery | Google Cloud | Prompts and responses are not used to train models | Google Cloud docs, updated Sep 30, 2026 |
| Coding agents such as Claude Code | Runs locally; prompts go to Anthropic or to Amazon Bedrock, Google Cloud's Agent Platform or Microsoft Foundry | Commercial terms: no training on code or prompts; 30-day standard retention | Claude Code docs, checked Oct 2, 2026 |
Each answer covers one platform or one tool. A data team's estate spans several, and the hard questions sit at the seams: which agent read what across Snowflake and dbt, who approved a write that reached Looker, what the rollback was. Data Workers answers those once, for the whole lifecycle, with one context, one approval flow and one audit trail, and it builds on the assistant or coding agent you already run. Weighing a build of your own? Can't we just build this ourselves? maps the security layer it must own; the AI assistants hub shows how each assistant connects.
What a security reviewer will ask
| The reviewer asks | The control in Data Workers |
|---|---|
| Does our data leave our environment? | Rows, credentials and keys stay with the agents in your infrastructure; the hosted Conductor receives workflow metadata; on-premise air-gapped for zero outbound |
| Which third party sees prompts? | Only the model account you hold the key for, or none with a local model |
| Will it train on our data? | No; Data Workers does not train or fine-tune on customer data |
| What can an agent change? | Read-only by default; per-domain autonomy; named approval for anything irreversible |
| How do we find and protect PII? | Privacy check on pull requests (names, not values), response scanning, scrub-on-write to the graph, masking proposals for the owner |
| How do people authenticate? | Your IdP via JWKS, or API keys you issue |
| How do we prove what happened? | Receipts plus a SHA-256 hash-chained, exportable audit log |
| How do we honour erasure requests? | An erasure tool that deletes the subject and writes an audit tombstone that keeps the chain valid |
Data Workers provides the controls; your auditor maps them to your obligations, with the receipts as evidence. We complete security questionnaires, sign DPAs and NDAs on request, and walk your reviewer through the deployment on your stack.
The case for your CFO
The outcome: agents that do real work on production data, and a short security review, because the answer to "where does our data go" is precise. Agents run in our infrastructure on our credentials, model calls go to the provider contract we already signed, rows stay where they are, and the hosted Conductor coordinates from metadata alone.
The risk story: agents start read-only, writes switch on per domain, anything irreversible waits for a named approver from your directory, and every change carries a receipt chained into a tamper-evident log. Nothing is migrated, so there is nothing to unwind.
Why now: assistants and coding agents already reach data their own ways, and one governed layer is cheaper to review once than five times.
The first win is usually a PII loop: privacy checks on every pull request, your existing classifications into the context graph, masking proposals for owners to approve, receipts for the auditor. What stays the same: your warehouse, catalog, IdP, permission system and model provider.
The path is a pilot: $7,500 one time, and the pilot is credited in full against the first year. Scale starts from $1,000 a month and Enterprise from $3,000 a month, billed annually, with unlimited seats, no usage meter, no markup on model spend and an Apache 2.0 core. See /pricing/, and model the return with the ROI calculator or ROI of agentic data operations.
The sentence to repeat upstairs: "Data Workers' agents run in our infrastructure, on our model, with our approvals, and every change they make leaves a receipt."
FAQ
Does any of our data reach Data Workers the company? Workflow metadata does: the hosted Conductor holds goals, detected signals, table and column names, proposals with their diffs, and run records. Table contents, credentials and model keys do not; they stay with the agents in your infrastructure and your model account.
Can we run it with no internet access at all? Yes, as an Enterprise on-premise deployment: an offline bundle with checksums, your private registry, Helm on your Kubernetes cluster, Postgres and MinIO in the cluster, Ollama or vLLM for inference, sovereign mode on, and outbound traffic blocked by network policy.
Can an agent write to production without a person? Only for change classes you have set to L4 autonomous in that domain. Everything else is propose-first, and anything irreversible needs a named approver.
What happens to the audit log if someone edits it? The SHA-256 chain stops verifying at that entry, and verify_global_hash_chain reports it. Export the log to your own append-only storage for retention.
How does this fit with EU or residency requirements? Sovereign mode keeps inference local and blocks external model APIs, inference stops rather than fall back to an external provider, and operations are logged. Your counsel maps those controls to your obligations; SOX, HIPAA, GDPR and the EU AI Act goes further.
Sources
- •Data Workers, Pricing (rate card Sep 10, 2026; "The Autonomous Data-Conductor, hosted by us" on every tier; "Runs in: Your infrastructure"), checked Oct 2, 2026: https://dataworkers.io/pricing/
- •Data Workers, Security and trust (last updated Sep 10, 2026), checked Oct 2, 2026: https://dataworkers.io/security/
- •Data Workers, What the platform adds, checked Oct 2, 2026: https://dataworkers.io/opensource-docs/platform/
- •Data Workers, Deployment (open-source core), checked Oct 2, 2026: https://dataworkers.io/opensource-docs/deployment/
- •Data Workers, Autonomous Data-Conductor, checked Oct 2, 2026: https://dataworkers.io/product/autonomous-data-conductor/
- •Data Workers open-source repo (tool registrations for
scan_pii,check_policy,request_governance_review,provision_access,get_audit_trail,verify_global_hash_chain,get_usage_activity_log,generate_audit_report,trace_cross_platform_lineage), checked Oct 2, 2026: https://github.com/DataWorkersProject/dataworkers-claw-community - •Databricks, Databricks AI assistive features trust and safety (updated Sep 11, 2026), checked Oct 2, 2026: https://docs.databricks.com/aws/en/databricks-ai/databricks-ai-trust
- •Snowflake, Cortex AI Functions, checked Oct 2, 2026: https://docs.snowflake.com/en/user-guide/snowflake-cortex/aisql
- •Snowflake, Cross-region inference, checked Oct 2, 2026: https://docs.snowflake.com/en/user-guide/snowflake-cortex/cross-region-inference
- •Google Cloud, How Gemini for Google Cloud uses your data (updated Sep 30, 2026), checked Oct 2, 2026: https://docs.cloud.google.com/gemini/docs/discover/data-governance
- •Anthropic, Claude Code data usage, checked Oct 2, 2026: https://code.claude.com/docs/en/data-usage