Product category
An autonomous agentic data platform is a system whose agents run the full detect, diagnose, fix, review and verify loop on a data platform unattended, under a named human's authority, with a receipt on every change and context that spans more than one warehouse. Data Workers is one, built from four modules that map onto that definition.
Last updated: September 10, 2026 · Dhanush Shetty, founder, Data Workers
This page defines the product category. Its sister page defines the practice these products carry out: what is agentic data engineering.
Every product in this category has these four parts, whatever it calls them. The left column is the category. The right is what they are called here. Use the left column to classify any vendor, including ours.
An orchestrator that runs the loop
Autonomous Data-Conductor
Runs detect, diagnose, fix, review and verify without waiting for someone to open an editor, and routes each piece of work to the agent that owns it. An autonomy dial sets how far it may go on its own, from propose-only upward.
Specialised agents that do the work and write
Data-Agents Swarm
Twenty specialised agents covering incidents, quality, schema evolution, pipeline building, change review, governance, cost, migration and the rest. They produce the actual artifact, not advice about it.
A context layer the agents reason on
Data Context Wizard
A semantic and knowledge graph spanning clouds: schema, ownership, metrics, PII tags, column-level lineage and past corrections, every fact provenance-stamped. Without one, an agent rediscovers the estate on every run and is wrong in a new way each time.
A control plane where humans approve and audit
Spellbook Data Catalog (in preview)
The catalog the agents write and keep current, and the place a human approves a change and reads back what happened. In preview, so treat it as a direction rather than a shipped surface.
A product that fails two of these is something else, whatever it is called. Each test names the class of product that fails it, because failing a test is usually a deliberate design choice rather than a shortcoming.
Test 01
Does it run with the IDE closed?
Data Workers: Yes. The Conductor runs unattended on a schedule and on events, with an autonomy dial that governs how far it goes without a human.
IDE coding agents fail this by design. They run while an engineer is typing and stop when the editor closes. That is not a defect in them, it is their shape.
Test 02
Does it close the whole loop rather than one step?
Data Workers: Yes, across six lanes: incident resolution, data quality, schema and change review, pipeline build, catalog and documentation, and cross-cloud scope.
Observability tools stop at detect. Pull request reviewers stop at review. Record-fixers stop at one record. Each owns a step well and does not own the loop.
Test 03
Does every write carry a receipt, with an approval gate for anything irreversible?
Data Workers: Yes. Propose first, dry run where one exists, a named human approves anything irreversible, and every applied change carries the diff, the approver, the blast radius across downstream models and dashboards, and the rollback path.
A product that applies changes with no per-change audit record and no rollback is not safer for being faster. In our own index that gap is the difference between a documented L3 and an L4 claim.
Test 04
Is the context cross-cloud rather than warehouse-native?
Data Workers: Yes: one graph across Snowflake, BigQuery and Databricks, which is what makes a change that spans two of them legible to the agent.
Warehouse-native agents answer no by design, and for a single-platform estate that is the right trade. It stops being right the day a second warehouse appears.
Test 05
Is the price published, with an open-source core you can clone today?
Data Workers: Yes. The rate card is on the pricing page and the core is Apache-2.0: 11 agents and 160+ MCP tools you can run this afternoon on your own model key.
Of the 21 vendors reviewed on 10 September 2026, six publish a rate a buyer can read without contacting sales. Five have no pricing page at all.
Our own long-form answers, including the cases where we are the wrong choice: should you use Data Workers.
In the Q4 2026 edition of the Agentic Data Engineering Index, 21 vendors are graded on public evidence across six lanes. No vendor is graded above L3, apply with approval, on documented evidence. Exactly one cell across the whole market sits above L3, and it is marked as a claim. Data Workers is graded L3 on incident resolution, data quality and schema review, and L2 elsewhere.
Every cell carries a citation, we graded ourselves last and by the same rules, and there is a public route to contest a grade: the Agentic Data Engineering Index, with the dataset under CC BY 4.0.
On the pricing half of the fifth test, of the 21 vendors reviewed on 10 September 2026 six publish a rate a buyer can read without contacting sales, six name a billing unit with no figure, four are contact-sales only and five have no pricing page at all: the published pricing census.
Inside the coding tool your team already uses, or inside your own VPC. Either way it runs on your own model provider key, so prompts and query results go to your account under your agreement with that provider, not through ours. Anything irreversible waits for a named human.
Where the agents run, what they can see, and our honest compliance status, including the fact that we are not SOC 2 certified: security and trust.
What is an autonomous agentic data platform?
A system whose agents run the full detect, diagnose, fix, review and verify loop on a data platform unattended, under a named human's authority, with a receipt on every change and context that spans more than one warehouse. It is made of four parts: an orchestrator that runs the loop, specialised agents that do the work and write, a context layer they reason on, and a control plane where humans approve and audit.
How is it different from an AI coding agent?
A coding agent runs while an engineer is typing and stops when the editor closes. An autonomous agentic data platform is defined by what happens when nobody is typing. Data Workers runs inside Claude Code, Cursor, Codex and OpenCode as well, so the two are not alternatives: the coding tool is where a person works with the agents, and the platform is what keeps running afterwards.
Is any vendor fully autonomous today?
No. In the Q4 2026 Agentic Data Engineering Index, which grades 21 vendors including Data Workers on public evidence only, no vendor is graded above L3, apply with approval, on documented evidence. Exactly one cell market-wide sits above L3 and it is marked as a claim. Data Workers is not above L3 either.
Where does it run, and who sees our data?
It runs inside the coding tool your team already uses or inside your own VPC, on your own model provider key, so prompts and query results go to your provider account under your agreement with them. Data Workers is not SOC 2 certified; a Type I audit is planned and timing is to be confirmed.
Two next steps
Run the agents yourself, free: the Apache-2.0 core is 11 agents and 160+ MCP tools inside Claude Code, Cursor, Codex or OpenCode, on your own model key. Clone it.
Or see the governed-write path and the Conductor on a live stack in a 45-minute session: book a time. Pilot Program $7,500 one-time, then Scale from $1,000 per month, no usage meter, as published on 10 September 2026. Pricing.