Industry
Industry8 min readBy The Data Workers Team

Data Workers for CIOs and CTOs

What Data Workers changes for a CIO or CTO: one governed agent layer across clouds and engines, agents in your environment, identity through your IdP, an Apache 2.0 core with a tested exit, and flat pricing with no usage meter.

For a CIO or CTO, Data Workers replaces a growing set of per-platform, per-team data agents with one governed agent layer that runs in your own environment, signs in through your identity provider, and works across the clouds and engines you run, from Snowflake, Databricks and BigQuery to Postgres and Iceberg. Your week shifts from reviewing each new agent, vendor and exit plan one at a time to setting one rule for all of them: where agents run, who approves their changes, how they stop, and what you keep if you leave.

Key takeaways

  • •One layer across your engines. The same agents, context graph, approval flow and audit trail work across Snowflake, Databricks, BigQuery, Postgres and the catalogs, orchestrators and BI tools around them, through 50+ connectors. Platform-native agents and your engineers' coding agents keep their jobs and reach Data Workers over MCP.
  • •It runs in your environment. The agents run in your infrastructure on every tier and hold your warehouse credentials and model key. Your data stays in your systems; the hosted Autonomous Data-Conductor sees workflow metadata only.
  • •Your identity provider stays in charge. Remote access signs in through Okta or Entra ID, and Data Workers verifies those tokens via JWKS. It issues no tokens of its own.
  • •You can stop it and you can leave it. An org-wide stop halts all autonomous dispatch, a two-person admin kill switch revokes the tenant's keys, tokens and sessions, and the Apache 2.0 core plus open exports turn the exit into a checklist you rehearse in the pilot.
  • •The cost is a flat line. Unlimited seats, no usage meter, no markup on model spend, and your own model account.

A CIO's or CTO's week today

The work is portfolio, architecture and risk. The architecture review board weighs another AI proposal. Security asks where a new tool's data goes and which service account it needs. Finance asks why cloud and warehouse spend outruns the budget. The board asks which agents can change production systems, and who answers for them. Modernisation runs for quarters in the background.

The agent question is new and it is multiplying. Every major data platform now ships its own agents and MCP servers: Databricks' Genie One MCP server is generally available (docs updated Sep 21, 2026), and the Snowflake-managed MCP server serves Cortex Agents, Cortex Analyst, Cortex Search and SQL execution to outside clients. Engineering runs Claude Code, Cursor or GitHub Copilot. Each is good on its own surface, and each brings its own identity model, logs and idea of an approval. A two-cloud estate can end up with four agent stacks and no single answer to "what changed production last night?"

The tools are architecture reviews, ServiceNow, FinOps reports and security dashboards. The FinOps Foundation's State of FinOps 2026 reports that "78% of teams now report to the CTO or CIO" and that many organizations are "asked to self-fund AI investments through optimization savings". dbt Labs' 2026 State of Analytics Engineering found 72% of respondents prioritize AI-assisted coding but only 24% AI-assisted pipeline management.

The same week with Data Workers

Data Workers is the agentic data platform: 20+ specialist agents (the Data-Agents Swarm), one governed context graph across every platform (the Data Context Wizard), a loop that detects, diagnoses, fixes, verifies and remembers (the Autonomous Data-Conductor), and one place where people review, approve and roll back (the Spellbook Data Catalog, in preview). The agents take the legwork; you set the rules once.

Comparison matrix of Your technology organisation and Data Workers on the outcomes a data leader buys

What the agents take off. Incidents diagnosed across engines, lineage traced with trace_cross_platform_lineage, change impact sized with blast_radius_analysis, least-privilege grants drafted by provision_access and checked by check_policy, and Snowflake spend attributed to the query and dbt model behind it, with the fix drafted for the model's owner. Migrations are planned in waves with their parity checks, tracked, and held at a completion gate for the owner's sign-off.

What you still own and decide.

  • •The architecture rule. Where the agents run (your cloud account, or a dedicated VPC or on-premise on Enterprise), which model they call, and which systems are off limits.
  • •The autonomy policy. Each domain runs at L0 manual, L1 observe, L2 propose, L3 act reversibly or L4 autonomous. Data Workers ships observe-only. Autonomy levels L0 to L4 explained covers each rung.
  • •Who approves. Approvals go to a named person. An unanswered request expires and escalates; it never auto-grants. No agent can promote its own work. See who owns the agents and how approvals work.
  • •The off switch. The org-wide stop halts all autonomous dispatch. The admin kill switch needs two people, the caller and a different approver, and revokes the tenant's API keys, install tokens and active sessions.

How Data Workers fits your architecture

Across clouds and engines. Data Workers connects natively to Snowflake, Databricks and Unity Catalog, BigQuery, PostgreSQL, Azure storage, dbt, Airflow, Dagster, Kafka and more, and to Iceberg catalogs, Glue, DataHub, OpenMetadata and Purview over their APIs or MCP servers today: 50+ connectors in all, listed on the integrations page. Nothing migrates: your warehouses, catalogs and orchestrators stay the systems of record, and Unity Catalog, Snowflake roles and your IAM stay the permission system, by design.

With platform-native agents and coding agents. Keep them. Genie and Cortex Agents answer inside their platforms; Data Workers works across all of them with one context graph, one approval flow and one audit trail. An engineer's coding agent can hold the Genie One or Snowflake-managed MCP server next to the Data Workers agents; the Genie MCP guide shows the wiring, and your company just rolled out AI assistants. Now what? maps each assistant to its setup.

A swarm, by design. Each data job gets a specialist agent with its own tools, permissions and autonomy level, so no single agent holds every permission. Why a swarm of specialist agents gives the blast-radius argument.

Where it runs. The agents run in your infrastructure on every tier, hold your warehouse credentials and model key, and query data in place; the context graph, receipts and audit log are written there. The Autonomous Data-Conductor is hosted by us and receives workflow metadata only (goals, signals, table names, proposals with diffs, run records, approval handles), never rows, credentials or model keys. Enterprise adds a dedicated VPC, your own cloud or on-premise, including air-gapped estates. You bring the model (Anthropic, OpenAI, Bedrock, Vertex AI, Azure OpenAI, or Ollama and vLLM locally), and sovereign mode blocks external model calls entirely. See where does our data go? and Data Workers in your VPC or air-gapped.

Identity. The remote endpoint takes an API key or OAuth through your identity provider: Okta or Entra ID is the authorization server, and Data Workers verifies its tokens against your IdP's JWKS. Data Workers issues no tokens and registers no clients, so your directory stays the source of truth and offboarding is the change you already make there.

The metrics you are judged on

MetricHow Data Workers moves itWhere the number comes from
AI adoption with controlOne governed layer, autonomy per domain, a receipt on every change and a named approver at L2Autonomy levels per domain, approvals in Spellbook
Risk incidentsScoped grants, blast radius checked before a change, tamper-evident logget_audit_trail, verify_global_hash_chain, get_usage_activity_log
Uptime of data servicesIncidents diagnosed across engines; reversible fixes at L3Incident history and time to resolve in the receipts
Platform costSnowflake spend attributed to the query and dbt model; fixes drafted for the ownerDesign target of 25 to 40% lower warehouse spend (Cost Savings agent page), measured in your pilot
Modernisation deliveryMigration waves planned with parity checks and a sign-off gateDesign target of 4 to 8 weeks per migration against 6 to 12 months (Data Migration agent page)

The cost and migration lines are design targets published on the Cost Savings and Data Migration agent pages; your pilot measures your own. How to measure AI data agents sets out the scorecard and the ROI calculator runs your numbers.

A worked example: "Which agents can change production, and can we stop them?"

This is an illustration, not a customer case. The company runs Snowflake on AWS and BigQuery on Google Cloud, with dbt, Airflow and ServiceNow. The agents run in its AWS account with Claude on Amazon Bedrock; Okta is the identity provider. Incidents run at L3 act reversibly, access and finance at L2 propose.

TimeWhoWhat happened
Mon 09:00Audit committeeAsks the CIO which AI agents can change production data, who approves them, and whether they can be stopped
Mon 09:30Data WorkersLists each domain's autonomy level and its named approver
Mon 10:00Data Workersget_audit_trail pulls every agent write this quarter, each with its receipt and approval record
Mon 10:05Data Workersverify_global_hash_chain confirms the log is intact end to end
Mon 11:00Data Workersget_usage_activity_log shows who called which tool over MCP, when, and the outcome
Tue 14:00CISOStaging drill: the org-wide stop halts all autonomous dispatch
Tue 14:20Data Workerscreate_servicenow_ticket opens a change record for the drill, with the receipt linked
Wed 10:00Platform leadExit rehearsal: the context graph exports to JSON-LD with its audit chain
Wed 11:00Platform leadRevokes the staging agents' roles in Snowflake and BigQuery; grants and tables people approved stay, as the company's own objects
Thu 09:00CIOSends the committee one page: agents, levels, approvers, stop and exit, each tested
Incident timeline across the stack: what Your technology organisation, your team and Data Workers each do, step by step

"Can we stop them?" became a drill with a change record, and "what if the vendor goes away?" an export the platform lead ran. The Genie and Cortex agents stay governed by Databricks and Snowflake roles. What if Data Workers goes away? has the full exit checklist.

Build it, buy it, or wait for the platforms

Build it with coding agents and MCP servers. Your engineers can wire Claude Code to vendor MCP servers. The operating layer around the agents (approvals, blast radius, rollback, receipts, autonomy per domain, connectors kept current) is the larger job, and yours to staff. Build it ourselves with Claude Code and MCP servers? prices that path.

Rely on each platform's agents. Genie and Cortex Agents are strong inside their own platform, and that focus is the right design. Writing to production across systems they don't own is a different product (context about every other engine, approvals that span them, one audit trail), and it is the product Data Workers is.

Point tools per slice. Each adds another console, contract and handoff; Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail.

The case for your CFO

The outcome is AI doing real data operations work under one set of controls, at a cost that does not climb with use. Data Workers closes that work behind approvals and attributes the spend, so the capacity data operations consume today goes back to the roadmap.

The risk story: the agents run in your environment on your model account, start observe-only, and move up per domain only when the records clear your bar. Every change has a receipt in a tamper-evident log, and every proposal at L2 waits for a named approver. One stop halts all autonomous dispatch, and two people can revoke every key and session. Nothing migrates, the core is Apache 2.0, and the exit is something you rehearse in the pilot.

Why now: your teams already write data code with AI and the platforms add agents every quarter; one control layer is cheaper before the sprawl than after. The first win is one domain, usually incidents and freshness, at L2 or L3 for a quarter, scored on its receipts.

Start with a pilot: $7,500 one-time, and the pilot is credited in full against the first year. Scale is from $1,000 a month and Enterprise from $3,000 a month, billed annually, with unlimited seats, no usage meter, no markup on model spend and your own model; why there is no usage meter explains the choice. See pricing, and the ROI of agentic data operations for the arithmetic. The sentence for upstairs: "Every data agent we run works in our environment, under our identity provider, with a receipt on every change and a named approver where we want one, and we can stop it or leave it on our terms."

FAQ

Is it safe to let agents change production data at all? Inside the controls above: scoped grants, autonomy per domain, a named approver at L2, recorded rollback at L3, a receipt on every change. Is it safe to let AI agents change production data? walks through each.

Does Data Workers replace Genie, Cortex Agents or our coding agents? No. Keep them; they reach Data Workers over MCP. What is an agentic data platform? gives the test, and the data leaders' guides for Snowflake and Databricks go deeper.

Does Data Workers hold our credentials or route warehouse traffic through its cloud? No. The agents hold the credentials in your infrastructure; the hosted Conductor sees workflow metadata only and never queries your warehouse.

How does it fit our identity and access model? Through your IdP and JWKS, or API keys. Grants stay in your engine's permission system: provision_access drafts them least-privilege with a 90-day default duration, and the data owner approves them (and, outside Unity Catalog, applies them).

What is our exit path? Set every domain to L1 observe, verify the audit log with verify_global_hash_chain, export the context graph to JSON-LD with its audit chain (or an OKF markdown bundle), and revoke the agents' roles. Approved changes already live in your repos and warehouse. The hosted Conductor and governed writes stop; the eleven Apache 2.0 agents keep reading and recommending in your engineers' MCP clients.

Sources

  • •FinOps Foundation, State of FinOps 2026: https://data.finops.org/ (checked Oct 2, 2026)
  • •dbt Labs, 2026 State of Analytics Engineering Report: https://www.getdbt.com/resources/state-of-analytics-engineering-2026 (checked Oct 2, 2026)
  • •Databricks documentation, Genie One MCP server (generally available; page updated Sep 21, 2026): https://docs.databricks.com/aws/en/agents/mcp-tools/genie-mcp (checked Oct 2, 2026)
  • •Snowflake documentation, Snowflake-managed MCP server: https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-agents-mcp (checked Oct 2, 2026)
  • •Data Workers public repository tools (trace_cross_platform_lineage, blast_radius_analysis, provision_access, check_policy, get_audit_trail, verify_global_hash_chain, get_usage_activity_log, create_servicenow_ticket): https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)
  • •Data Workers product repository, data-workers-agent-swarm main @ 0c2491e3: remote transport JWKS verification, Conductor governor org-wide stop, two-person kill switch (checked Oct 2, 2026)
  • •Data Workers product pages: https://dataworkers.io/product/data-agents-swarm/ , https://dataworkers.io/product/autonomous-data-conductor/ , https://dataworkers.io/product/spellbook-data-catalog/ (checked Oct 2, 2026)
  • •Data Workers security page: https://dataworkers.io/security/ (checked Oct 2, 2026)
  • •Data Workers pricing and ROI calculator: https://dataworkers.io/pricing/ , https://dataworkers.io/roi-calculator/ (checked Oct 2, 2026)
  • •Data Workers agent pages (design targets): https:///blog/cost-savings-cleanup-agent/ (25 to 40% spend reduction), https:///blog/data-migration-agent/ (six to twelve months toward four to eight weeks) (checked Oct 2, 2026)
  • •Data Workers, What if Data Workers goes away? (publishes this run; exit checklist): https:///blog/what-if-data-workers-goes-away/