For Databricks

Data Workers for Databricks

Data Workers runs the detect, diagnose, fix, review and verify loop on a Databricks lakehouse when nobody is in the workspace. A failed job at 3am is diagnosed, a fix is proposed against the offending task, a named human approves anything irreversible, and every applied change carries a receipt with its blast radius and a rollback.

For a data team already running Databricks: jobs, Unity Catalog, Delta tables and dbt, with an on-call rota and more incidents than people.

Last updated: September 10, 2026 · Dhanush Shetty, founder, Data Workers

What breaks on Databricks, and what it costs you

None of these is a Databricks defect. They are the failures a busy estate produces, and the hours between one of them starting and somebody verifying a fix are what this platform is for.

A job failed at 3am and the dashboards went stale quietly

The run shows failed in the Jobs UI. Nothing downstream says so. By the time somebody opens Slack, a gold-layer table has been a day behind for nine hours and three dashboards have been read as current.

A Unity Catalog table with no owner and no comment that three teams query

It has lineage, so you can see who reads it. Nobody can tell you what the column means or who decides when it changes. That is not a documentation gap, it is the reason the next schema change will break something.

A task that succeeded on empty input

The upstream source returned nothing, the task ran clean, and the merge wrote zero rows over a partition that used to have data. Every status light is green. The number in the report is wrong.

A schema change upstream that a gold model did not survive

A column type changed in a bronze Delta table. The silver notebook coerced it. The gold aggregate now sums a string cast that silently truncates, and no test covered that path.

Alerts are not fixes. Detection tells you the first of those sentences. The rest is the work.

The four modules on Databricks

Runs the loop

Autonomous Data-Conductor

The Conductor watches job and task outcomes and starts work without waiting for somebody to open the workspace. On a failed run it reads the task, the cluster logs and the lineage around the target table, forms a diagnosis, and produces the change it believes fixes it. An autonomy dial governs how far it goes on its own: at the lowest setting it stops at a proposal and pages a human, higher up it may run the dry run and stage the change, and anything irreversible always waits for a named approver regardless of the setting. The dial is a policy you set, not a claim about what it has already done unattended in production.

Does the work and writes

Data-Agents Swarm, 20 specialised agents

On a lakehouse the incident debugging, data quality, schema evolution and pipeline building agents carry most of the load, with the cost and cleanup agent working on clusters and jobs that run for consumers who left. The agents produce the actual artifact, the notebook change, the DLT expectation, the job config, rather than advice about it.

What the agents reason on

Data Context Wizard

The Context Wizard holds what Unity Catalog knows plus what it does not: ownership that is actually acted on, metric definitions, PII tags, column-level lineage and the record of past corrections, every fact provenance-stamped so an agent can tell what it learned from where. Cross-cloud matters even to a team that considers itself all-in on Databricks, because there is almost always a Postgres, a BigQuery export or a vendor warehouse at the edge of the estate, and that edge is where the wrong numbers come from.

Where humans approve and audit

Spellbook Data Catalog, in preview

Spellbook is the catalog the agents write and keep current, and the plane where a human approves a change and reads back what happened. It sits alongside Unity Catalog rather than replacing it: Unity Catalog stays the governance and access system of record. Spellbook is in preview, so treat it as a direction rather than a shipped surface.

What each module has to do to qualify as this kind of product, stated generically so you can use it on any vendor: the autonomous agentic data platform.

Week one, without talking to us

The first three steps need no form, no account and no key we issue. The open-source core is Apache-2.0: 11 agents and 160+ MCP tools that read, analyse and recommend.

  1. Clone the Apache-2.0 core and point it at one Databricks workspace with a read-only service principal.
  2. Ask it to explain a gold table nobody on the team fully understands, then compare its account with your best engineer's.
  3. Run it against the last three failed jobs you resolved by hand, and check whether the diagnosis matches what you actually found.
  4. If the read-only output is worth the hour, book the session and bring one incident you want to see fixed end to end behind an approval gate.

What this does not do on Databricks

  • It does not replace Unity Catalog, and it does not want to. Governance, access and lineage stay where they are.
  • It does not tune your clusters for you as a performance product would. Cost work here is about finding what runs for nobody, not about rewriting your Spark.
  • It cannot approve its own irreversible changes. If nobody on the team will sign, the platform stops at read-only and you have bought the free tier.

The full list of who we are not for, including the cases where the honest answer is to buy something else: should you use Data Workers.

Common questions

Does this replace Unity Catalog?

No. Unity Catalog remains the governance, access and lineage system of record on Databricks. The Data Context Wizard reads from it and adds what a catalog does not hold: metric definitions, acted-on ownership, past corrections and cross-cloud lineage, provenance-stamped. Spellbook, which is in preview, is the human approval and audit plane for agent-made changes, not a second governance system.

Where does it run, and who sees our data?

Inside the coding tool your engineers already use, or inside your own VPC. It runs on your own model provider key, so prompts and query results go to your provider account under your agreement with them, not through ours. Data Workers is not SOC 2 certified; a Type I audit is planned and timing is to be confirmed.

Can it change production on Databricks without us?

Not for anything irreversible. The write path is propose first, dry run where one exists, and a named human approves before an irreversible change lands. Every applied change carries a receipt: the diff, the approver, the blast radius across downstream tables and dashboards, and the rollback path.

The rest of your stack

Two next steps

Run the read-only agents on Databricks yourself, free: clone the open-source core.

Or bring one real incident to a 45-minute session and watch the approval gate stop the agent: book a time. Pilot Program $7,500 one-time, then Scale from $1,000 per month, no usage meter, as published on 10 September 2026. Pricing.