For BigQuery

Data Workers for BigQuery

Data Workers runs the detect, diagnose, fix, review and verify loop on a BigQuery estate when nobody is at the console. A scheduled query that half-succeeded is diagnosed, a fix is proposed against the job that caused it, a named human approves anything irreversible, and every applied change carries a receipt with its blast radius and a rollback.

For a data team already running BigQuery on Google Cloud: datasets, scheduled queries and dbt, with an on-call rota and more incidents than people.

Last updated: September 10, 2026 · Dhanush Shetty, founder, Data Workers

What breaks on BigQuery, and what it costs you

None of these is a BigQuery defect. They are the failures a busy estate produces, and the hours between one of them starting and somebody verifying a fix are what this platform is for.

A scheduled query that half-succeeded

The job reports success. It wrote a partition from a source that was still loading, so the partition is real, present and short by a third. Nothing retries, because nothing failed.

A dataset with no descriptions that the business reads daily

Tables, columns and no documentation anywhere. The meaning lives in one engineer's head and in a Slack thread from last year, and the next person to change it will be guessing.

A query that scans a full table every hour for nobody

Partition pruning was lost when somebody wrapped the filter in a function. The bill went up, nothing broke, and the report it feeds has no readers left anyway.

A schema change in an upstream dataset that a view did not survive

A column was added upstream and a downstream view selected by position rather than by name. The view still runs. The columns it returns are now shifted by one.

Alerts are not fixes. Detection tells you the first of those sentences. The rest is the work.

The four modules on BigQuery

Runs the loop

Autonomous Data-Conductor

The Conductor watches job outcomes and freshness across datasets and starts work without waiting for somebody to open the console. On a failed or suspicious job it reads the job, the destination table and the lineage around it, forms a diagnosis, and produces the change it believes fixes it. An autonomy dial governs how far it goes on its own: at the lowest setting it stops at a proposal, higher up it may run the dry run and stage the change, and anything irreversible always waits for a named approver regardless of the setting. The dial is a policy you set, not a claim about what it has already done unattended in production.

Does the work and writes

Data-Agents Swarm, 20 specialised agents

On BigQuery the incident debugging, data quality, schema evolution and catalog agents carry most of the load, with the cost and cleanup agent working on jobs that scan far more than they need and on scheduled work whose consumers are gone. The agents produce the actual artifact, the query change, the schema fix, the description, rather than advice about it.

What the agents reason on

Data Context Wizard

The Context Wizard holds ownership that is actually acted on, metric definitions, PII tags, column-level lineage and the record of past corrections, every fact provenance-stamped so an agent can tell what it learned from where. Cross-cloud matters here in particular, because a BigQuery estate is very often the second warehouse in an organisation rather than the only one, and the disagreements between the two are where the wrong numbers live.

Where humans approve and audit

Spellbook Data Catalog, in preview

Spellbook is the catalog the agents write and keep current, and the plane where a human approves a change and reads back what happened. It sits alongside Dataplex rather than replacing it. Spellbook is in preview, so treat it as a direction rather than a shipped surface.

What each module has to do to qualify as this kind of product, stated generically so you can use it on any vendor: the autonomous agentic data platform.

Week one, without talking to us

The first three steps need no form, no account and no key we issue. The open-source core is Apache-2.0: 11 agents and 160+ MCP tools that read, analyse and recommend.

  1. Clone the Apache-2.0 core and point it at one Google Cloud project with a read-only service account.
  2. Ask it to explain a table the business depends on and nobody has documented, then compare its account with your best engineer's.
  3. Run it against the last three incidents you resolved by hand, and check whether the diagnosis matches what you actually found.
  4. If the read-only output is worth the hour, book the session and bring one incident you want to see fixed end to end behind an approval gate.

What this does not do on BigQuery

  • It does not replace Dataplex, and it does not want to. Governance and cataloguing at the Google Cloud level stay where they are.
  • It is not a query optimiser, and it will not rewrite your SQL for performance as its main job. Cost work is about finding what runs for nobody and what scans more than it needs.
  • It cannot approve its own irreversible changes. If nobody on the team will sign, the platform stops at read-only and you have bought the free tier.

The full list of who we are not for, including the cases where the honest answer is to buy something else: should you use Data Workers.

Common questions

Does this replace Dataplex?

No. Dataplex stays the Google Cloud governance and catalog surface. The Data Context Wizard reads what is there and adds what a catalog does not hold: metric definitions, acted-on ownership, past corrections and cross-cloud column-level lineage, provenance-stamped. Spellbook, which is in preview, is the human approval and audit plane for agent-made changes.

We are on BigQuery and something else. Does that help or hurt?

It helps, and it is the case the context graph was built for. The graph spans warehouses rather than being native to one, so a change that crosses from BigQuery into another warehouse is legible to the agent instead of stopping at the boundary. Warehouse-native agents answer no to that question by design.

Can it change production in BigQuery without us?

Not for anything irreversible. The write path is propose first, dry run where one exists, and a named human approves before an irreversible change lands. Every applied change carries a receipt: the diff, the approver, the blast radius across downstream tables and dashboards, and the rollback path.

The rest of your stack

Two next steps

Run the read-only agents on BigQuery yourself, free: clone the open-source core.

Or bring one real incident to a 45-minute session and watch the approval gate stop the agent: book a time. Pilot Program $7,500 one-time, then Scale from $1,000 per month, no usage meter, as published on 10 September 2026. Pricing.