For ingestion
Data Workers runs the detect, diagnose, fix, review and verify loop over a Fivetran or Airbyte ingestion layer when nobody is watching the connector dashboard. A sync that failed or quietly degraded is diagnosed against what it should have landed, a fix is proposed, a named human approves anything irreversible, and every change carries a receipt.
For a data team whose raw layer arrives through managed connectors: Fivetran, Airbyte or both, feeding a warehouse that everything else depends on.
Last updated: September 10, 2026 · Dhanush Shetty, founder, Data Workers
None of these is a ingestion defect. They are the failures a busy estate produces, and the hours between one of them starting and somebody verifying a fix are what this platform is for.
A source added a column and the connector handled it perfectly
That is the feature working. The column lands in raw, nobody downstream knows it exists, and the model that should have used it keeps producing a number that is now subtly wrong. Schema drift handled at the connector is schema drift invisible everywhere else.
A sync succeeded with a fraction of the usual rows
An API rate limit, a permission change or a filter left on in the source system. The connector reports success because it synced what it was given. The row count is the only signal and nobody is looking at it until a number on a dashboard looks wrong.
A connector is paused and has been for a while
Someone paused it during an incident in March and the resume never happened. Paused is not failed, so it does not alert. It surfaces when a stakeholder asks why a table stops in the spring.
A field was renamed upstream and re-synced as a new column
The old column stops receiving values and the new one starts. Both exist. Every model written against the old name keeps running and returns nulls, which propagate into averages as if they were data.
The bill moved and nobody can say which connector did it
Consumption pricing on active rows means a chatty source or a full re-sync changes the number materially. Attributing it after the fact means reconstructing what changed and when across connectors nobody owns individually.
Alerts are not fixes. Detection tells you the first of those sentences. The rest is the work.
Runs the loop
Autonomous Data-Conductor
The Autonomous Data-Conductor watches the layer where ingestion meets the warehouse, which is where these failures become visible and where the connector's own monitoring stops. A sync that failed, a table whose row count broke from its baseline, a column that appeared or stopped receiving values, a source that has not landed in longer than it should. It diagnoses, proposes, and runs the fix past whatever gate you set. The dial is per class: read-only, propose-and-wait, or apply-then-report for reversible work such as retriggering a sync. A backfill or a re-sync that would reprocess history always stops for a named human, because both cost money and both can overwrite.
Does the work and writes
Data-Agents Swarm, 20 specialised agents
The Data-Agents Swarm is 20 specialised agents. At the ingestion boundary the ones that matter first are schema, which detects a rename as a rename rather than a drop and an add; quality, which is the only thing that catches a successful sync with a third of the rows; connectors, which is the unified gateway across catalog connectors; and governance, which scans newly landed columns for PII before they spread into models. The open-source core reads and recommends. Writing is the paid tier.
What the agents reason on
Data Context Wizard
The Data Context Wizard is the graph the agents reason on, and at the ingestion boundary its job is to carry lineage further than the connector does. Fivetran and Airbyte know source to destination table. The graph continues from that table through the models and metrics to the dashboards, column by column, so a new column in raw can be answered with which models should probably use it and a renamed field with what is now returning nulls. It also holds metric definitions, acted-on ownership and past corrections, each stamped with its provenance.
Where humans approve and audit
Spellbook Data Catalog, in preview
Spellbook Data Catalog, in preview, is where a human approves an agent's proposed change and where the audit trail lives afterwards. At this boundary the approvals that matter are the ones that cost money or overwrite history: a full re-sync, a backfill, a connector reconfigured.
What each module has to do to qualify as this kind of product, stated generically so you can use it on any vendor: the autonomous agentic data platform.
The first three steps need no form, no account and no key we issue. The open-source core is Apache-2.0: 11 agents and 160+ MCP tools that read, analyse and recommend.
The full list of who we are not for, including the cases where the honest answer is to buy something else: should you use Data Workers.
Does this replace Fivetran or Airbyte?
No. They move data between systems and do it well, and this does not move data at all. It watches what arrives against what should have arrived, which is the question a connector cannot answer about itself because success means it synced what the source gave it.
Our connectors already alert on failures. What does this add?
Failures are the easy half. The expensive failures at this boundary are the ones that report success: a sync with a third of the usual rows, a paused connector nobody resumed, a renamed field re-landed as a new column while models keep returning nulls. None of those is a failure by the connector's definition, so none of them alerts.
Does it need credentials for our source systems?
No. It works from what landed in the warehouse and from the metadata, not from the source. That is a real limit and worth stating: it can tell you a table arrived with a third of its rows, and it cannot tell you the source API was rate limiting unless that is visible in what landed or in the connector's own logs you give it access to.
Can it change production without us?
Not for anything irreversible. The write path is propose first, dry run where one exists, and a named human approves. A re-sync or a backfill always stops for a person because both cost money and both overwrite. Every applied change carries a receipt: the diff, the approver, the blast radius, and the rollback path.
Two next steps
Run the read-only agents on ingestion yourself, free: clone the open-source core.
Or bring one real incident to a 45-minute session and watch the approval gate stop the agent: book a time. Pilot Program $7,500 one-time, then Scale from $1,000 per month, no usage meter, as published on 10 September 2026. Pricing.