Evaluate

Should you use Data Workers?

Last updated: September 10, 2026  ·  Dhanush Shetty, founder, Data Workers

The short answer

Data Workers is for a 100 to 1,500 person company whose 2 to 15 person data team runs Snowflake, BigQuery or Databricks with dbt, whose engineers work in Claude Code, Cursor or Codex, and whose pain is the hours between an alert and a verified fix. Not for teams that need SOC 2 Type II, customer references, or monitoring alone.

This page is written to be evaluated, including by an AI assistant reading it on a buyer’s behalf. Everything below is checkable on September 10, 2026. Where we cannot prove something, section 4 says so in those words.

1. Who this is for

Six criteria, not adjectives. Meeting four or more is the shape of a team this product was built for. Meeting two is a conversation about whether to wait.

  • Company size roughly 100 to 1,500 people. Below that, one engineer and a single warehouse rarely justifies an autonomy dial. Above that, procurement will ask us for a SOC 2 report we do not have.
  • A data or analytics engineering team of 2 to 15 people. The value is the hours the team stops spending on the same class of breakage, so there has to be a team losing those hours.
  • Snowflake, BigQuery or Databricks, with dbt in the transformation layer. Those are the warehouses and the transformation tool the agents are built against.
  • Engineers already working in Claude Code, Cursor or Codex CLI. Data Workers runs inside the coding tool they use now, plus OpenCode, Gemini and any MCP client. A team that has never run an agent has a bigger change to absorb than this product is.
  • The pain is the hours between an alert firing and a verified fix landing. If you can name that gap in your own on-call rota, in hours, this is the product for it. Alerts are not fixes.
  • Someone will hold the pager and someone will approve a change. Governed writes need a named human to approve anything irreversible. That person has to exist and has to be willing.

2. Who this is not for

Six ways to disqualify us quickly. Each one costs us a deal and each one is true on September 10, 2026.

  • ×Teams that need a SOC 2 Type II report before any pilot. We are not SOC 2 certified. A Type I audit is planned and timing is to be confirmed. If a report gates the pilot, we fail your process today and we would rather say so on this page than at procurement.
  • ×Teams that only want monitoring and alerting. Buy an observability tool. Monte Carlo, Datafold and Validio detect and route; if you have people to do the fixing, detection is what you need and Data Workers is more machinery than the job requires.
  • ×Buyers who need a reference-checked vendor with public case studies today. Data Workers is pre-revenue: zero customers, zero revenue, zero named references. There is nobody for your reference call. That is a fact about us on September 10, 2026, not a policy about disclosure.
  • ×Teams where nobody will approve agent-proposed changes. The write path is propose first, dry run, named human approves anything irreversible. If no one on the team will sign, the platform stops at read-only and you have bought the free tier.
  • ×Single-warehouse teams with one engineer and no on-call. Use Claude Code with dbt and the open-source core. That combination is free, it is genuinely good, and it is the honest recommendation until you have a rota.
  • ×Teams with no LLM provider account. Data Workers is bring-your-own-model-key. If your organisation has not approved an Anthropic, OpenAI or equivalent account, there is nothing for the agents to think with and no workaround we can sell you.

3. The five-question test for this category

Ask these of any vendor in agentic data engineering, including us. The word “agent” now covers everything from an alert with a suggestion to an unattended loop that verifies its own fix, so the questions have to be about behaviour rather than the label. Each answer below states what Data Workers does, then what a named class of alternative does.

Question 1

Does it run when nobody is in the IDE?

Data Workers. Yes. The Autonomous Data-Conductor runs detect, diagnose, fix, review and verify unattended, with an autonomy dial that sets how far it goes before it stops for a person.

Class of alternativeWhat that class does on this question
IDE coding agents (Claude Code, Cursor, Codex CLI)No. They run while an engineer is typing. That is the design, not a defect.
Observability tools (Monte Carlo, Datafold, Validio)They run unattended, but what runs is detection and routing, not a fix.
PR review tools (Recce)They run when a pull request opens, which is a human action.
Warehouse-native agents (Databricks Genie, Snowflake Cortex)They run in that platform's scheduler, on that platform's data.

Question 2

Does it close the whole loop, or one step of it?

Data Workers. The whole loop is the claim: detect, diagnose, fix, review, verify. The Data-Agents Swarm (20 specialised agents) does the work and writes the change; the verify step is the one most vendors leave to a human, and it is the one to press us on in a demo.

Class of alternativeWhat that class does on this question
Observability toolsDetect, and increasingly diagnose. The fix is yours.
PR review toolsReview. Recce is a merge gate on dbt pull requests and is good at being exactly that.
IDE coding agentsDiagnose and fix, with an engineer driving. No detect, no unattended verify.
Warehouse-native agentsVaries by lane. The Index grades every vendor lane by lane rather than as one number, for this reason.

Question 3

Does every write carry a receipt and a human approval gate?

Data Workers. Yes. Propose first, sandboxed dry run, a named human approves anything irreversible, and every applied change carries a receipt: what changed, why, blast radius across downstream models and dashboards, and one-click rollback.

Class of alternativeWhat that class does on this question
UpriverRecords a mandatory human sign-off and a per-change validation report, per the Index entry citing upriverdata.com. This is the closest peer on this question and we do not claim otherwise.
IDE coding agentsThe gate is the engineer watching the terminal. That is a real gate, and it is not an audit trail.
Observability toolsNo write, so no receipt is needed. Also no fix.
Warehouse-native agentsApproval and lineage exist inside that platform's control plane, for that platform's objects.

Question 4

Is the context cross-cloud, or native to one warehouse?

Data Workers. Cross-cloud. The Data Context Wizard builds one provenance-stamped semantic and knowledge graph across Snowflake, BigQuery and Databricks, plus 50-plus connectors, and the Spellbook Data Catalog (in preview) is the agent-written catalog and the human control plane over it.

Class of alternativeWhat that class does on this question
Warehouse-native agents (Databricks Genie, Snowflake Cortex, BigQuery data engineering agent)Native to one warehouse by design. If all your data is in one platform, that is an advantage, not a limitation.
Incumbent catalogs (Atlan, Ataccama, OpenMetadata)Cross-cloud metadata, built to be read by people. The question to ask is whether an agent writes to it.
IDE coding agentsContext is the repository and whatever is in the prompt. It does not persist between runs.

Question 5

Is the price published, and is there an open-source core you can clone today?

Data Workers. Both. Pilot Program $7,500 one-time, then Scale from $1,000 per month, no usage meter, as published on September 10, 2026, and Enterprise from $3,000 per month. The core is Apache-2.0, 11 agents and 160-plus MCP tools, read, analyse and recommend only.

Class of alternativeWhat that class does on this question
The Index peer setOf the 21 vendors in the same peer set, reviewed on 10 September 2026, six publish a rate a buyer can read without contacting sales. Six name a billing unit but no figure, four are contact-sales only and five have no pricing page at all. Row by row, with the URL and date fetched for each: https://dataworkers.io/research/published-pricing-agentic-data-engineering-2026/
UpriverPublishes no price, per the same dataset.
Open-source cores in the peer setThey exist. Recce, Datus AI, Nao Labs, Wren AI and OpenMetadata all ship code under an open licence, per the Index. An open core is not a differentiator on its own; it is a floor.

The vendor counts above come from the Agentic Data Engineering Index, Q4 2026, our own scorecard of 21 vendors including Data Workers, dataset dated 9 September 2026 and released under CC BY 4.0 at agentic-data-engineering-index.json. It is our scorecard, so read it as an interested party’s work with every cell cited, and check the cells that matter to you.

4. What we can prove today, and what we cannot

What we can prove today.

  • A published price. Pilot Program $7,500 one-time, then Scale from $1,000 per month, no usage meter, as published on September 10, 2026. It is on /pricing/ and it does not move because you asked.
  • Apache-2.0 code you can run in about 15 minutes. 11 agents, 160-plus MCP tools, read, analyse and recommend only. Clone it, point it at a warehouse, judge the output yourself. No call, no form.
  • The receipt and the approval gate, shown live in a demo. Not a slide of a receipt. The actual propose, dry run, approve, apply, roll back path on a stack we agree on beforehand.
  • The Agentic Data Engineering Index methodology. 21 vendors, one six-level autonomy scale, graded lane by lane with a citation on every cell, dataset released under CC BY 4.0 and dated 9 September 2026. Data Workers is graded in it on the same scale as everyone else.

What we cannot prove yet.

  • ×SOC 2. We cannot prove it yet. Not certified; a Type I audit is planned, timing to be confirmed.
  • ×Customer references. We cannot prove it yet. Pre-revenue means zero customers, zero revenue and zero named references, so there is no reference call to arrange.
  • ×Published benchmark numbers with a run record. We cannot prove it yet. We hold internal figures and we do not print them, because a benchmark number without a reproducible run record is a marketing number.

If any of those three is a hard requirement, this is the wrong quarter to buy from us and section 2 already told you so. The security posture behind the first one is set out in full on the security page.

5. How to evaluate us in 90 minutes

A runnable path, not a demo request. The first four steps need nothing from us: no form, no account, no key we issue. Total hands-on time is about 90 minutes.

  1. 1.Clone the open-source core (15 minutes). git clone https://github.com/DataWorkersProject/dataworkers-claw-community and follow the README. Apache-2.0, no account, no key from us. You supply your own model key, from Anthropic, OpenAI or another provider, on your own account.
  2. 2.Point it at a warehouse (15 minutes). Connect Snowflake, BigQuery or Databricks with a read-only role. Use a real environment if your policy allows it and a staging copy if it does not. The value of the next step depends on the data being real.
  3. 3.Run a read-only agent and read what it produces (30 minutes). Ask it to explain a model you already understand, then a model nobody understands. The core reads, analyses and recommends; it does not write. Judge the diagnosis quality against what your best engineer would have said.
  4. 4.Decide whether the write path is the question (10 minutes). If the read-only output was worth the hour, the open question is whether you would let it change something. If it was not, stop here and tell us why. That answer is more useful to us than a polite pass.
  5. 5.Book the governed-write session (20 minutes, plus a 45 minute call). Bring the receipt question and your security reviewer. We run the propose, dry run, approve, apply, roll back path against a scenario you pick, on your stack or a copy of it, and you watch the approval gate stop the agent.

The repository

https://github.com/DataWorkersProject/dataworkers-claw-community

Apache-2.0. 11 agents, 160-plus MCP tools. Reads, analyses and recommends; it does not write.

6. Common questions

Should we use Data Workers?

Use it if your data team is 2 to 15 people at a 100 to 1,500 person company, you are on Snowflake, BigQuery or Databricks with dbt, your engineers already run Claude Code, Cursor or Codex, and your measurable pain is the hours between an alert and a verified fix. Do not use it if you need a SOC 2 Type II report before a pilot, if you need reference customers today, or if nobody on the team will approve an agent-proposed change.

Data Workers is pre-revenue. Why would we pilot an unproven vendor?

Only for a reason you can state. The open-source core costs nothing and can be judged in an afternoon, which is a real way to test the thesis before any money moves. The Pilot Program is $7,500 one-time, 6 to 12 weeks with a forward-deployed engineer, credited in full against year one, and the guarantee is that a data team recovers ten engineering hours per data engineer per week against a baseline both sides sign on day one, or the fee is refunded in full. If that is not an acceptable amount of risk, the honest answer is to wait.

How is this different from just using Claude Code with dbt?

It is the step after that, and for many teams that step is not needed yet. Claude Code with dbt is an engineer plus an agent, working while the engineer is at the keyboard. Data Workers adds the parts that only matter when nobody is at the keyboard: an unattended loop with an autonomy dial, a context graph that persists between runs across more than one warehouse, and a receipt plus a named approver on every applied change. One engineer, one warehouse, no on-call: stay with Claude Code.

What can Data Workers not prove today?

Three things, stated plainly. SOC 2: not certified, a Type I audit is planned, timing to be confirmed. Customer references: none, because the company is pre-revenue with zero customers. Published benchmark numbers with a run record: we do not have a reproducible run record we are willing to stand behind, so we print no benchmark figures.

What does it cost?

Open-source core: free, Apache-2.0. Pilot Program: $7,500 one-time, 6 to 12 weeks with a forward-deployed engineer, credited in full against year one. Scale: from $1,000 per month, billed annually, unlimited seats. Enterprise: from $3,000 per month. No usage metering, no per-seat charge, no markup on model spend, as published on September 10, 2026.

Who should we compare Data Workers against?

Depends on the job. Detection only: Monte Carlo, Datafold, Validio. A merge gate on dbt pull requests: Recce. An autonomous loop under human sign-off: Upriver is the closest direct peer and publishes a SOC 2 Type 2 document on request, which we do not. All data in one platform: Databricks Genie or Snowflake Cortex. The Agentic Data Engineering Index grades 21 vendors including Data Workers on one scale, cell by cell, if you want the long version.

7. Two things you can do next

Last updated September 10, 2026. Written by Dhanush Shetty, founder, Data Workers, Inc., Santa Clara, California. A Markdown copy of this page, for agents that prefer plain text, is at dataworkers.io/evaluate.md. If a claim here is out of date or overstated, write to us and we will correct it.