Product
Product7 min readBy The Data Workers Team

Can Data Workers run in our VPC or air-gapped?

Yes. Data Workers agents run in your infrastructure on every tier, Enterprise adds your VPC, your own cloud or on-premise, and the air-gapped build runs on local models with sovereign mode and zero outbound access.

Yes. The Data Workers agents run in your infrastructure on every tier, Enterprise adds a dedicated VPC, your own cloud or on-premise, and an air-gapped deployment runs on local models through Ollama or vLLM with sovereign mode, which blocks every external model call and never falls back to a cloud provider.

The part we host is the Autonomous Data-Conductor, and it sees workflow metadata only. Below is what runs where in each of the three deployment shapes, how each one installs, and what an air-gapped install looks like end to end.

Key takeaways

  • •Every tier runs in your infrastructure. The agents hold your warehouse credentials and model key and query data in place. Enterprise adds a dedicated VPC, your own cloud or on-premise, with SSO, single-tenant and private networking.
  • •The hosted Conductor orchestrates from metadata. It has no warehouse connector and receives goals, signals, table names, proposals and run records, never rows, credentials or model keys.
  • •Air-gapped means zero outbound. An offline bundle with checksums, your private registry, Helm on RKE2, Postgres and MinIO in the cluster, and Ollama or vLLM on your nodes.
  • •Sovereign mode is enforced in code. External model calls are blocked, there is no fallback to a cloud provider, and a network guard logs every outbound request it blocks.
  • •Your identity provider stays in charge. Tokens are verified against your IdP's JWKS, on your own network if it has to be, or you issue API keys.

Three deployment shapes

Data Workers is the agentic data platform, built so the parts that touch your data sit inside your boundary. Per the rate card on /pricing/, Start and Scale run in your infrastructure, inside the coding agent your engineers already use, and Enterprise runs in your VPC, your own cloud or on-premise, scoped with you at contract.

How Data Workers fits with Your estate: your coding agent on top, Data Workers in the middle, your estate underneath

Your VPC or your own cloud. The platform installs with Docker Compose on a single node or with Helm on your Kubernetes cluster. The Helm chart ships a NetworkPolicy that denies all egress except DNS, traffic between the release's own pods, and an allowlist you write: your model endpoint, your IdP, your SIEM. Model calls can stay inside your cloud account with Amazon Bedrock, Azure OpenAI or Google Cloud's Gemini Enterprise Agent Platform (formerly Vertex AI). Your credentials never leave your network.

On-premise. The same Helm chart on your own Kubernetes, behind your firewall, with private networking. The license is a signed key validated locally, and with telemetry off the install needs no runtime call to dataworkers.io. Audit events can stream to your SIEM over an HMAC-signed webhook.

Air-gapped. No internet path at all. On a connected machine you build an offline bundle that holds the container images, the packaged Helm chart and SHA-256 checksums, carry it across, verify the checksums, push the images into your private registry (the guide uses Harbor), and install with Helm onto RKE2. The guide's network table has four flows, all internal: the platform to Postgres for metadata, to MinIO for artifacts, to Ollama or vLLM for inference, and your ingress to the platform's API. Its stated requirement is zero outbound internet access, confirmed by the sovereign self-test in the startup log and by kubectl get networkpolicy.

What runs where in each shape

Comparison matrix of Your estate and Data Workers on the outcomes a data leader buys

The hosted Conductor, precisely. The Autonomous Data-Conductor turns a goal such as "every table is documented" into routines, dispatches work to the domain agents, and records each run. It carries no warehouse connector and never queries the warehouse. What it holds is workflow metadata: goals, detected signals, the names of the tables involved, proposals with their diffs, run records and approval handles. In a VPC or on-premise deployment with an outbound path, the Conductor knows that claims.policy_events changed and that a fix is waiting for approval; it never sees a row of it. An air-gapped network has no path out, so nothing reaches the hosted Conductor, workflow metadata included, and orchestration for that estate is scoped with you at contract as part of the Enterprise on-premise deployment. The component-by-component answer is in Where does our data go?.

The open-source core. The Apache 2.0 agents run on a laptop or server as local processes over stdio, launched by your MCP client, against your credentials and model key: the quickest way to test on a stack that can't call out.

Models, offline

You bring the model in every shape. With outbound access, calls go through one provider interface that supports Anthropic, OpenAI, Amazon Bedrock, Google Cloud, Azure OpenAI, Ollama and vLLM. Air-gapped, the Helm chart defaults to Ollama serving Mistral Small, and vLLM serving Mistral 7B Instruct is the guide's choice for production throughput. Weights are pulled on a connected machine and carried across. A GPU node is optional: Ollama runs CPU-only once you remove the GPU requests from the chart values.

Sovereign mode. The air-gap install turns on sovereign mode. With sovereign mode on, calls to external model APIs are blocked in code and inference routes to your local endpoint. If no local model is configured, inference stops; it never falls back to an external provider. A network guard checks outbound HTTP requests against your allow-list (localhost and in-cluster service names always pass), logs what it blocks, and external telemetry is blocked. Data Workers does not sample column values for PII: its pull request review reads column names and annotations only. Model choice and cost are in Which model does Data Workers use, and what does it cost to run?, and PII handling in How does Data Workers handle PII?.

The same controls in every shape

Where it runs changes nothing about how the agents behave. They start read-only, writes switch on per domain, and anything irreversible waits for a named approver.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous

Remote endpoints accept an API key you issue or an OAuth token from your own IdP. Data Workers verifies tokens against the issuer's published keys (JWKS), at the issuer's default location or a JWKS URL you set, which matters air-gapped because your IdP lives on your network too. It caches keys, follows rotation, and maps claims to a tenant and a role. It issues no tokens and registers no clients. Every applied change carries a receipt (diff, approver, blast radius, rollback path), chained into a SHA-256 audit log that verify_global_hash_chain checks end to end. The write-safety detail is in Is it safe to let AI agents change production data?.

A worked example, air-gapped

This is an illustration. An EU insurer runs Data Workers on a three-node RKE2 cluster with no internet route. Images come from Harbor, vLLM serves Mistral 7B Instruct on a GPU node, people sign in through Keycloak, and the claims data lives in PostgreSQL, modelled with dbt Core and scheduled by a self-managed Apache Airflow.

TimeSystemWhat happens
06:00RKE2The platform pod starts with sovereign mode on, and the sovereign self-test passes
07:12PostgreSQLA source team renames policy_ref to policy_id in claims.policy_events
07:15Data Workersrun_quality_check fails on the table: policy_ref is gone and policy_id is new; assess_impact and trace_cross_platform_lineage show three dbt models and the nightly reserving DAG depend on it
07:16vLLMThe model drafts the dbt change from schema, lineage and the failing SQL, on the GPU node inside the cluster
07:18Data WorkersProposal routed to the claims data owner with the diff, blast radius and rollback path
08:05KeycloakThe owner signs in; Data Workers verifies the token against Keycloak's JWKS on the internal network
08:06dbt CoreApproved change applied; the three models rebuild and their tests pass
08:20Apache AirflowThe reserving DAG run completes on schedule
08:21Postgres, MinIOReceipt and audit entry written; the network guard has logged zero blocked calls

Every byte in that sequence stayed inside the cluster. In a VPC deployment the same run would also send the hosted Conductor its workflow metadata: the table name, the signal, the proposal and the run record.

How the alternatives deploy

Each platform's own agents run inside that platform's cloud, a sound design for agents that serve one platform. Data Workers works alongside all of them.

OptionWhere the agent runsNetwork and residency controlsSource, date
Snowflake Cortex Agents"Within Snowflake's governed environment"; Snowflake's LLMs are "deployed within the Snowflake Service perimeter"Snowflake privileges; model availability by cross-region routing scope, an account parameterSnowflake docs, checked Oct 2, 2026
Databricks agents (Agent Bricks, Model Serving, Databricks Apps)Serverless compute plane "in your Databricks account", not your cloud accountNetwork boundary per workspace; Geos for residencyDatabricks docs, updated Sep 11 and Sep 30, 2026
BigQuery Data Engineering Agent (Gemini in BigQuery)Google Cloud, under your project's IAM; it generates pipelines you review and runA VPC Service Controls perimeter around Dataform, BigQuery and the Conversational Analytics APIGoogle Cloud docs, updated Sep 30, 2026
Microsoft Fabric data agentFabric capacity (F2 or higher, or P1), on Azure OpenAI Assistant APIs; read-onlyTenant settings for cross-geo processing; source and agent capacity in the same regionMicrosoft Learn, updated May 11, 2026

Building it yourself on open-source agents and MCP servers gets you self-hosting, and leaves you to own the approval gate, receipts, sovereign enforcement and offline packaging; Can't we just build this ourselves? maps that work. Data Workers runs one governed layer across Snowflake, Databricks, BigQuery, Fabric and the on-premise systems none of them reach, inside your network, with one context, one approval flow and one audit trail.

The case for your CFO

The outcome: the data team gets agents that fix schema breaks, quality failures and access requests, and the security review is short because the deployment answer is exact. The agents run in our network on our credentials and our model; the air-gapped build has no outbound path at all.

The risk story: agents start read-only, writes switch on per domain, irreversible changes wait for a named approver from our own directory, and every change leaves a receipt in a tamper-evident log stored on our side. Nothing is migrated, so there is nothing to unwind.

Why now: data residency and sovereignty rules keep tightening, and platform agents run in their vendors' clouds. A layer that runs inside our boundary lets the regulated parts of the estate use agents too.

The first win is usually a schema-change catch like the one above: caught, traced, fixed and approved before the nightly run, with a receipt for the auditor. What stays the same: our warehouse, catalog, scheduler, IdP, permission system and model provider.

The path is a pilot: $7,500 one time, and the pilot is credited in full against the first year. Scale starts from $1,000 a month and Enterprise, where the VPC, own-cloud and on-premise deployments sit, from $3,000 a month, billed annually, with unlimited seats, no usage meter, no markup on model spend and an Apache 2.0 core. See /pricing/, and model the return with the ROI calculator or ROI of agentic data operations.

The sentence to repeat upstairs: "Data Workers runs inside our network on our own model, and in the air-gapped build nothing leaves at all."

FAQ

Does the hosted Conductor ever see our data? It sees workflow metadata: goals, detected signals, table and column names, proposals with their diffs, run records and approval handles. It has no warehouse connector, so table contents never reach it, and credentials and model keys stay with the agents. An air-gapped network sends it nothing.

Which tier do we need for a VPC or air-gapped deployment? Every tier runs in your infrastructure. A dedicated VPC, your own cloud, on-premise and air-gapped installs are Enterprise deployments, scoped with you at contract, alongside SSO, single-tenant and private networking.

What happens if the local model goes down in sovereign mode? Inference stops. Sovereign mode never falls back to an external provider.

How do we prove the air gap to our security team? The sovereign self-test result is in the startup log, the network guard logs every blocked request, and the NetworkPolicy is visible with kubectl get networkpolicy.

How do licensing and upgrades work offline? The license is a signed key validated locally. Upgrades arrive as a new bundle: verify the checksums, push the images to your registry, and run helm upgrade.

Does this make us compliant with residency rules? Data Workers provides the controls: local inference, blocked external calls, zero-egress networking and a hash-chained audit trail. Your counsel maps them to your obligations; SOX, HIPAA, GDPR and the EU AI Act goes further.

Sources

  • •Data Workers, Pricing (rate card Sep 10, 2026: "Runs in" per tier; Enterprise "Dedicated VPC or your own cloud, SSO, private networking, on-premise"; "Your credentials never leave your network"; "The Autonomous Data-Conductor, hosted by us"), checked Oct 2, 2026: https://dataworkers.io/pricing/
  • •Data Workers, Security and trust (last updated Sep 10, 2026), checked Oct 2, 2026: https://dataworkers.io/security/
  • •Data Workers, Deployment (open-source core), checked Oct 2, 2026: https://dataworkers.io/opensource-docs/deployment/
  • •Data Workers, What the platform adds, checked Oct 2, 2026: https://dataworkers.io/opensource-docs/platform/
  • •Data Workers, Autonomous Data-Conductor, checked Oct 2, 2026: https://dataworkers.io/product/autonomous-data-conductor/
  • •Data Workers open-source repo (tool registrations for run_quality_check, assess_impact, trace_cross_platform_lineage, scan_pii, verify_global_hash_chain), checked Oct 2, 2026: https://github.com/DataWorkersProject/dataworkers-claw-community
  • •Snowflake, Cortex Agents, checked Oct 2, 2026: https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-agents
  • •Snowflake, Cortex AI Functions, checked Oct 2, 2026: https://docs.snowflake.com/en/user-guide/snowflake-cortex/aisql
  • •Databricks, High-level architecture (updated Sep 11, 2026), checked Oct 2, 2026: https://docs.databricks.com/aws/en/getting-started/high-level-architecture
  • •Databricks, Use agents on Databricks (updated Sep 30, 2026), checked Oct 2, 2026: https://docs.databricks.com/aws/en/generative-ai/agent-bricks/
  • •Databricks, AI assistive features trust and safety (updated Sep 11, 2026), checked Oct 2, 2026: https://docs.databricks.com/aws/en/databricks-ai/databricks-ai-trust
  • •Google Cloud, Data Engineering Agent for BigQuery pipelines (updated Sep 30, 2026; "cannot execute pipelines"; VPC Service Controls perimeter for Dataform, BigQuery and the Conversational Analytics API), checked Oct 2, 2026: https://docs.cloud.google.com/bigquery/docs/data-engineering-agent-pipelines
  • •Microsoft Learn, Fabric data agent concepts (updated May 11, 2026), checked Oct 2, 2026: https://learn.microsoft.com/en-us/fabric/data-science/concept-data-agent