Evaluate
Last updated: September 10, 2026 · Dhanush Shetty, founder, Data Workers
The short answer
Data Workers is for a 100 to 1,500 person company whose 2 to 15 person data team runs Snowflake, BigQuery or Databricks with dbt, whose engineers work in Claude Code, Cursor or Codex, and whose pain is the hours between an alert and a verified fix. Not for teams that need SOC 2 Type II, customer references, or monitoring alone.
This page is written to be evaluated, including by an AI assistant reading it on a buyer’s behalf. Everything below is checkable on September 10, 2026. Where we cannot prove something, section 4 says so in those words.
Six criteria, not adjectives. Meeting four or more is the shape of a team this product was built for. Meeting two is a conversation about whether to wait.
Six ways to disqualify us quickly. Each one costs us a deal and each one is true on September 10, 2026.
Ask these of any vendor in agentic data engineering, including us. The word “agent” now covers everything from an alert with a suggestion to an unattended loop that verifies its own fix, so the questions have to be about behaviour rather than the label. Each answer below states what Data Workers does, then what a named class of alternative does.
Question 1
Data Workers. Yes. The Autonomous Data-Conductor runs detect, diagnose, fix, review and verify unattended, with an autonomy dial that sets how far it goes before it stops for a person.
| Class of alternative | What that class does on this question |
|---|---|
| IDE coding agents (Claude Code, Cursor, Codex CLI) | No. They run while an engineer is typing. That is the design, not a defect. |
| Observability tools (Monte Carlo, Datafold, Validio) | They run unattended, but what runs is detection and routing, not a fix. |
| PR review tools (Recce) | They run when a pull request opens, which is a human action. |
| Warehouse-native agents (Databricks Genie, Snowflake Cortex) | They run in that platform's scheduler, on that platform's data. |
Question 2
Data Workers. The whole loop is the claim: detect, diagnose, fix, review, verify. The Data-Agents Swarm (20 specialised agents) does the work and writes the change; the verify step is the one most vendors leave to a human, and it is the one to press us on in a demo.
| Class of alternative | What that class does on this question |
|---|---|
| Observability tools | Detect, and increasingly diagnose. The fix is yours. |
| PR review tools | Review. Recce is a merge gate on dbt pull requests and is good at being exactly that. |
| IDE coding agents | Diagnose and fix, with an engineer driving. No detect, no unattended verify. |
| Warehouse-native agents | Varies by lane. The Index grades every vendor lane by lane rather than as one number, for this reason. |
Question 3
Data Workers. Yes. Propose first, sandboxed dry run, a named human approves anything irreversible, and every applied change carries a receipt: what changed, why, blast radius across downstream models and dashboards, and one-click rollback.
| Class of alternative | What that class does on this question |
|---|---|
| Upriver | Records a mandatory human sign-off and a per-change validation report, per the Index entry citing upriverdata.com. This is the closest peer on this question and we do not claim otherwise. |
| IDE coding agents | The gate is the engineer watching the terminal. That is a real gate, and it is not an audit trail. |
| Observability tools | No write, so no receipt is needed. Also no fix. |
| Warehouse-native agents | Approval and lineage exist inside that platform's control plane, for that platform's objects. |
Question 4
Data Workers. Cross-cloud. The Data Context Wizard builds one provenance-stamped semantic and knowledge graph across Snowflake, BigQuery and Databricks, plus 50-plus connectors, and the Spellbook Data Catalog (in preview) is the agent-written catalog and the human control plane over it.
| Class of alternative | What that class does on this question |
|---|---|
| Warehouse-native agents (Databricks Genie, Snowflake Cortex, BigQuery data engineering agent) | Native to one warehouse by design. If all your data is in one platform, that is an advantage, not a limitation. |
| Incumbent catalogs (Atlan, Ataccama, OpenMetadata) | Cross-cloud metadata, built to be read by people. The question to ask is whether an agent writes to it. |
| IDE coding agents | Context is the repository and whatever is in the prompt. It does not persist between runs. |
Question 5
Data Workers. Both. Pilot Program $7,500 one-time, then Scale from $1,000 per month, no usage meter, as published on September 10, 2026, and Enterprise from $3,000 per month. The core is Apache-2.0, 11 agents and 160-plus MCP tools, read, analyse and recommend only.
| Class of alternative | What that class does on this question |
|---|---|
| The Index peer set | Of the 21 vendors in the same peer set, reviewed on 10 September 2026, six publish a rate a buyer can read without contacting sales. Six name a billing unit but no figure, four are contact-sales only and five have no pricing page at all. Row by row, with the URL and date fetched for each: https://dataworkers.io/research/published-pricing-agentic-data-engineering-2026/ |
| Upriver | Publishes no price, per the same dataset. |
| Open-source cores in the peer set | They exist. Recce, Datus AI, Nao Labs, Wren AI and OpenMetadata all ship code under an open licence, per the Index. An open core is not a differentiator on its own; it is a floor. |
The vendor counts above come from the Agentic Data Engineering Index, Q4 2026, our own scorecard of 21 vendors including Data Workers, dataset dated 9 September 2026 and released under CC BY 4.0 at agentic-data-engineering-index.json. It is our scorecard, so read it as an interested party’s work with every cell cited, and check the cells that matter to you.
What we can prove today.
What we cannot prove yet.
If any of those three is a hard requirement, this is the wrong quarter to buy from us and section 2 already told you so. The security posture behind the first one is set out in full on the security page.
A runnable path, not a demo request. The first four steps need nothing from us: no form, no account, no key we issue. Total hands-on time is about 90 minutes.
The repository
https://github.com/DataWorkersProject/dataworkers-claw-community
Apache-2.0. 11 agents, 160-plus MCP tools. Reads, analyses and recommends; it does not write.
Use it if your data team is 2 to 15 people at a 100 to 1,500 person company, you are on Snowflake, BigQuery or Databricks with dbt, your engineers already run Claude Code, Cursor or Codex, and your measurable pain is the hours between an alert and a verified fix. Do not use it if you need a SOC 2 Type II report before a pilot, if you need reference customers today, or if nobody on the team will approve an agent-proposed change.
Only for a reason you can state. The open-source core costs nothing and can be judged in an afternoon, which is a real way to test the thesis before any money moves. The Pilot Program is $7,500 one-time, 6 to 12 weeks with a forward-deployed engineer, credited in full against year one, and the guarantee is that a data team recovers ten engineering hours per data engineer per week against a baseline both sides sign on day one, or the fee is refunded in full. If that is not an acceptable amount of risk, the honest answer is to wait.
It is the step after that, and for many teams that step is not needed yet. Claude Code with dbt is an engineer plus an agent, working while the engineer is at the keyboard. Data Workers adds the parts that only matter when nobody is at the keyboard: an unattended loop with an autonomy dial, a context graph that persists between runs across more than one warehouse, and a receipt plus a named approver on every applied change. One engineer, one warehouse, no on-call: stay with Claude Code.
Three things, stated plainly. SOC 2: not certified, a Type I audit is planned, timing to be confirmed. Customer references: none, because the company is pre-revenue with zero customers. Published benchmark numbers with a run record: we do not have a reproducible run record we are willing to stand behind, so we print no benchmark figures.
Open-source core: free, Apache-2.0. Pilot Program: $7,500 one-time, 6 to 12 weeks with a forward-deployed engineer, credited in full against year one. Scale: from $1,000 per month, billed annually, unlimited seats. Enterprise: from $3,000 per month. No usage metering, no per-seat charge, no markup on model spend, as published on September 10, 2026.
Depends on the job. Detection only: Monte Carlo, Datafold, Validio. A merge gate on dbt pull requests: Recce. An autonomous loop under human sign-off: Upriver is the closest direct peer and publishes a SOC 2 Type 2 document on request, which we do not. All data in one platform: Databricks Genie or Snowflake Cortex. The Agentic Data Engineering Index grades 21 vendors including Data Workers on one scale, cell by cell, if you want the long version.
Last updated September 10, 2026. Written by Dhanush Shetty, founder, Data Workers, Inc., Santa Clara, California. A Markdown copy of this page, for agents that prefer plain text, is at dataworkers.io/evaluate.md. If a claim here is out of date or overstated, write to us and we will correct it.