For Snowflake

Data Workers for Snowflake

Data Workers runs the detect, diagnose, fix, review and verify loop on a Snowflake estate when nobody is logged in. A task that stopped days ago is found and diagnosed, a fix is proposed, a named human approves anything irreversible, and every applied change carries a receipt with its blast radius across downstream models and dashboards.

For a data team already running Snowflake: warehouses, tasks, streams and dbt, with an on-call rota and more incidents than people.

Last updated: September 10, 2026 · Dhanush Shetty, founder, Data Workers

What breaks on Snowflake, and what it costs you

None of these is a Snowflake defect. They are the failures a busy estate produces, and the hours between one of them starting and somebody verifying a fix are what this platform is for.

A task that stopped running days ago and nobody noticed

A task in the graph suspended after repeated failures. Everything downstream kept querying the table it should have been refreshing. The first person to notice was whoever read the number and knew it was wrong.

A warehouse burning credits for a consumer who left

A scheduled query has run hourly for eleven months. The dashboard it fed belonged to someone who changed teams. Nothing is broken, which is exactly why nobody has looked at it.

A stream that went stale past its retention

The consuming task fell behind, the stream's offset aged past the source table's data retention, and the stream became unreadable. Recovering it is a reload, and knowing that early is the difference between an hour and a day.

A dbt model that compiles clean and ships wrong numbers

The SQL is valid, the tests pass, and a join fans out because an upstream table quietly stopped being unique on its key. Nothing fails. The report is wrong by a factor nobody can explain in the meeting.

Alerts are not fixes. Detection tells you the first of those sentences. The rest is the work.

The four modules on Snowflake

Runs the loop

Autonomous Data-Conductor

The Conductor watches task outcomes, stream lag and freshness and starts work without waiting for somebody to log in. On a suspended task it reads the failure history, the objects the task touches and the lineage around them, forms a diagnosis, and produces the change it believes fixes it. An autonomy dial governs how far it goes on its own: at the lowest setting it stops at a proposal, higher up it may run the dry run and stage the change, and anything irreversible always waits for a named approver regardless of the setting. The dial is a policy you set, not a claim about what it has already done unattended in production.

Does the work and writes

Data-Agents Swarm, 20 specialised agents

On Snowflake the incident debugging, data quality, schema evolution and data change review agents carry most of the load. The cost and cleanup agent has an obvious job here, finding warehouses and scheduled work running for consumers who are gone, and its output is a list of specific objects with the evidence for each, not a savings figure we would have no basis to print.

What the agents reason on

Data Context Wizard

The Context Wizard holds ownership that is actually acted on, metric definitions, PII tags, column-level lineage and the record of past corrections, every fact provenance-stamped so an agent can tell what it learned from where. Cross-cloud matters even to a team that considers itself all-in on Snowflake, because the wrong numbers usually arrive from the edge of the estate: a vendor export, a Postgres, a BigQuery dataset that somebody landed once and nobody owns.

Where humans approve and audit

Spellbook Data Catalog, in preview

Spellbook is the catalog the agents write and keep current, and the plane where a human approves a change and reads back what happened. It sits alongside Snowflake Horizon rather than replacing it: Horizon stays the governance surface. Spellbook is in preview, so treat it as a direction rather than a shipped surface.

What each module has to do to qualify as this kind of product, stated generically so you can use it on any vendor: the autonomous agentic data platform.

Week one, without talking to us

The first three steps need no form, no account and no key we issue. The open-source core is Apache-2.0: 11 agents and 160+ MCP tools that read, analyse and recommend.

  1. Clone the Apache-2.0 core and point it at one Snowflake account with a read-only role.
  2. Ask it to explain a model nobody on the team fully understands, then compare its account with your best engineer's.
  3. Run it against the last three incidents you resolved by hand, and check whether the diagnosis matches what you actually found.
  4. If the read-only output is worth the hour, book the session and bring one incident you want to see fixed end to end behind an approval gate.

What this does not do on Snowflake

  • It does not replace Horizon, and it does not want to. Governance and access stay where they are.
  • It is not a query optimiser. The cost work is about finding what runs for nobody, not about rewriting your SQL for warehouse performance.
  • It cannot approve its own irreversible changes. If nobody on the team will sign, the platform stops at read-only and you have bought the free tier.

The full list of who we are not for, including the cases where the honest answer is to buy something else: should you use Data Workers.

Common questions

Does this replace Snowflake Horizon?

No. Horizon remains the governance surface on Snowflake. The Data Context Wizard reads what is there and adds what a governance catalog does not hold: metric definitions, acted-on ownership, past corrections and cross-cloud column-level lineage, provenance-stamped. Spellbook, which is in preview, is the human approval and audit plane for agent-made changes.

Will it reduce our Snowflake bill?

The cost and cleanup agent finds warehouses, tasks and scheduled queries that run for consumers who are gone, and reports each one with the evidence. Whether that reduces your bill depends on what it finds in your account, which we cannot know in advance. We are pre-revenue and publish no savings figures, because we have no customer runs to derive one from.

Can it change production on Snowflake without us?

Not for anything irreversible. The write path is propose first, dry run where one exists, and a named human approves before an irreversible change lands. Every applied change carries a receipt: the diff, the approver, the blast radius across downstream models and dashboards, and the rollback path.

The rest of your stack

Two next steps

Run the read-only agents on Snowflake yourself, free: clone the open-source core.

Or bring one real incident to a 45-minute session and watch the approval gate stop the agent: book a time. Pilot Program $7,500 one-time, then Scale from $1,000 per month, no usage meter, as published on 10 September 2026. Pricing.