Product
Product12 min readBy The Data Workers Team

You're on Estuary: It Moves Data in Real Time From Databases and SaaS. Data Workers Owns Whether What Lands Is Right

Estuary streams CDC, SaaS and batch data into your warehouse in real time. Data Workers catches the breaking change that lands on time, traces it through dbt and BI, and gets the fix approved.

Your team thinks in captures, collections and materializations. A capture reads the PostgreSQL write-ahead log or a SaaS API once, the documents land in collections with a schema, and materializations keep Snowflake, BigQuery, Databricks or Iceberg tables current, some within a second, some hourly, all from the same read. Derivations reshape data in flight, and Dekaf hands collections to Kafka consumers such as ClickHouse or Tinybird. You run it on Estuary's public cloud, a private deployment or BYOC, where the data plane sits in your own cloud account. If you've used the product a while, you may still call it Estuary Flow; today it markets itself as the "Right-Time Data Platform for CDC, Streaming and Batch ETL."

Estuary is built to keep moving when a source changes. With "Automatically keep schemas up to date" switched on, a capture picks up a new column or a changed type by itself. When a change is incompatible, each materialization does what you configured: by default it backfills the binding, and when the change needs a different table structure, the destination table is dropped, re-created and refilled. That is the right behaviour for a data mover. It also means a table can be a fifth full at 14:15 and whole again at 14:50, every task green, while a dbt job and a dashboard read it in between. Data Workers watches what lands, traces it downstream and gets the fix approved before anything builds on it.

Key takeaways

  • •Estuary keeps its job. Captures, collections, derivations, materializations and backfills stay with your team. Data Workers works on what Estuary lands.
  • •Green tasks get checked too. Row counts against a baseline and null checks on the materialized tables turn a correct-by-design refill into a diagnosed incident before the next scheduled job reads it.
  • •Connected over Estuary's API or MCP server today. Data Workers reads the materialized tables, a PostgreSQL source and dbt natively; Estuary's agent skills and flowctl sit next to its agents in the same client.
  • •Every fix goes through a named person. The owner approves the hold and the dbt diff and runs them; Estuary settings stay with the Estuary owner; each step leaves a receipt.
  • •Autonomy is set per domain. Start at L1 observe, move to L2 propose, and open L3 act reversibly for narrow classes once the record earns it.

Estuary keeps your tables fresh to the second. Data Workers owns whether they are still right.

Estuary's job is to read every change once and land it everywhere at the latency each destination needs. The job after the landing is different: notice that a fresh table says something it shouldn't, tie it to the source, size what it touches, hold what would act on it, get the fix approved, rerun and prove the number. Here is one afternoon at a live-events ticketing company. This is an illustration, not a customer case.

TimeSystemWhat happens
Tue 14:05PostgreSQLThe ticketing app ships a migration for a new arena with named sections: orders.seat_section changes from SMALLINT to VARCHAR(12) so it can hold values like "FLOOR-A". The arena's on-sale opens at 15:00
14:06EstuaryThe Postgres capture has auto-discover on and picks up the change. The collection schema moves from integer to string, which the Snowflake materialization can't apply in place, so it follows its incompatible-schema-change setting, the default backfill. Every task stays green
14:07Estuary + SnowflakeAn incompatible column type change needs a different table structure, so, as Estuary's docs describe, TICKETING.ORDERS is dropped, re-created and refilled from the collection: 240 million rows
14:15dbt Cloud + LightdashThe hourly dbt Cloud job runs. fct_sales_pacing is a table model, so it rebuilds from the 22% of rows refilled so far. The Lightdash "On-sale pacing" dashboard now shows this week's sales down by most of their volume, and the marketing lead starts drafting a paid-promotion request for three slow shows
14:17Data Workers + SnowflakeThe row-count metric the team records on ORDERS reads 78% below its baseline. Data Workers opens an incident
14:20Data Workers + PostgreSQL + dbtData Workers reads the app's merged migration in GitHub, a native connection: seat_section is now VARCHAR(12) at the source, so the change came from the app's migration and the drop is a refill in progress, not lost orders. Blast radius: fct_sales_pacing, the Lightdash dashboard (a context-graph note the team recorded), and int_seat_map, which casts seat_section to an integer and will fail on the first "FLOOR-A" order after 15:00, taking the arena out of pacing on its launch day
14:21Opsgenie + SlackData Workers raises an Opsgenie alert with the diagnosis and sends the approval request to the analytics engineering owner in Slack
14:29SpellbookThe owner reviews three proposals: pause the hourly dbt Cloud job until the refill completes; a dbt diff that keeps seat_section as text in stg_orders and joins int_seat_map on text; and a note for the Estuary owner on the materialization's incompatible-schema-change setting for orders, with what backfill, disableBinding and abort would each have meant today. She approves the first two and forwards the note
14:33dbt Cloud + GitHubShe pauses the job schedule in dbt Cloud and merges the diff. The marketing lead holds the promotion request on the incident link
14:52Estuary + SnowflakeThe refill completes. Data Workers reads the ORDERS row count back at its baseline
14:55dbt CloudThe owner queues the approved dbt Cloud job run, recorded against her approval in Spellbook; the pacing models rebuild on the full table
15:00PostgreSQL + EstuaryThe arena on-sale opens; "FLOOR-A" orders land in Snowflake within seconds
15:06Data Workers + SnowflakeData Workers verifies: the row-count check passes, the new column type is recorded in the incident, int_seat_map built with the arena's sections in it, and pacing for the slow shows is back at its pre-refill level
15:08Opsgenie + SpellbookData Workers closes the Opsgenie alert with the receipt linked: cause, approvals, runs, checks passed and the undo (revert the dbt diff). The owner restarts the hourly schedule
16:00EstuaryThe Estuary owner sets the orders binding to her chosen behaviour for the next breaking change, in the Estuary UI
Incident timeline across the stack: what Estuary, your team and Data Workers each do, step by step

Every part of Estuary did what it was set to do: the capture saw the type change, the collection schema followed it, and the materialization refilled the table rather than writing strings into a number column. Catching it takes knowledge Estuary was never meant to hold: that a table model reads ORDERS hourly, that marketing spends on what the dashboard says, and that a downstream cast breaks when the first named section sells.

JobWhat Estuary doesWhat Data Workers does
The moveCaptures CDC, SaaS and batch sources once and materializes them at each destination's latency, from 200+ connectorsReads what Estuary lands, natively in the warehouse; connects to Estuary over its API or MCP server
The schemaAuto-discovers source changes and, on an incompatible change, backfills, disables or aborts per the materialization settingTakes the change from the source migration in review, classifies it and traces every model, dashboard and job that reads the column
The signalTask logs and stats in the web app or flowctl, email or Slack alerts when a pipeline stops or fails, and an OpenMetrics API for Prometheus or DatadogChecks the landed data itself: row counts against a baseline and nulls, so a green pipeline with a half-filled table still raises an incident
The diagnosisExplains a failing task from its logs, in the UI or through a coding agent running flowctlJoins the source, the landed table and lineage into one cause, with the blast radius across dbt models, dashboards and downstream jobs
The fixApplies whatever capture and materialization specs the team publishesProposes the hold and the dbt diff to the owner, and a setting note to the Estuary owner, who approve and apply them
The proofKeeps task logs and statsRe-checks the tables and writes a receipt: what changed, who approved it, how it was checked, how to undo it

Why doesn't Estuary just do this itself?

Because Estuary is built to move data faithfully and continuously, and it made sensible choices for that job. Its schema inference "only widens schemas, never narrows them," auto-discover keeps captures current without a ticket, and the materialization setting decides what happens on a breaking change. The default, backfill, keeps the destination correct in the end; the others stop a binding or the whole task, or reject the change for a person to resolve. Each is a trade between freshness and safety that Estuary hands to you, per materialization. Estuary does not know that a dbt table model rebuilds every hour from the table being refilled, or that a cast downstream assumes the old type.

Estuary's AI features point the same way. Agent Skills (labelled New) teach Claude Code, Cursor, Codex, Copilot or Gemini CLI to run flowctl workflows: create captures and materializations, write derivations, check task health, restart connectors. Estuary's docs describe its MCP server as documentation access, paired with the skills for operations. On its homepage, Estuary says agents connected this way "can inspect, observe, edit, and even build new pipelines." All of that is about building and running Estuary pipelines well.

Owning whether landed data is right across the source, the warehouse, dbt and BI is a different product: a context graph of every table and consumer, blast-radius scoping, named approvers, a recorded undo and receipts an auditor can read. That is Data Workers. More in is it safe to let AI agents change production data.

Every tool owns a slice. Data Workers covers the whole lifecycle

Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the Estuary pipelines already there.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Estuary goes deep on its own area
StageData WorkersEstuaryWhy we scored it this way
Catalog & Context93Collections carry a schema for every document Estuary moves. Data Workers keeps one governed context graph of what each table means, who owns it and what reads it.
Analytics & Insights82Not Estuary's job: it keeps tables fresh, while the numbers live in BI. Data Workers answers data questions from governed definitions with lineage behind every number.
Data Quality84Every document is validated against its collection schema before it lands. Data Workers checks the landed tables for nulls and volume, tracks lateness against a baseline your team records, and turns a bad load into an incident.
Observability & Incidents8.54Task logs, stats and the OpenMetrics API show a failing capture or materialization. Data Workers diagnoses the data incident when every task is green, proposes the fix and verifies it.
Pipelines & Ingestion8.59Estuary's home stage: log-based CDC, SaaS and batch from 200+ connectors, derivations in SQL, TypeScript or Python, Dekaf for Kafka consumers. Data Workers plans the reruns after a load for the owner.
Schema & Migration85Auto-discover keeps collection schemas current, and each materialization backfills, disables or aborts on a breaking change, as configured. Data Workers traces the landed change through dbt, BI and the jobs that read it.
Governance & Access8.53SSO and role-based access control govern who can change a pipeline. Data Workers routes every data change to a named approver.
Security & Privacy84Private and BYOC data planes, field-level redaction and column blocking keep the move contained. Data Workers leaves a receipt on every data change.
Cost / FinOps83Estuary prices the move; warehouse spend sits in the warehouse. Data Workers traces Snowflake credits to the dbt model behind them through query tags.
MLOps & Models7.52Feeds models and agents fresh data. Data Workers keeps the data under models and agents healthy.

How Estuary and Data Workers work together

How Data Workers fits with Estuary: your coding agent on top, Data Workers in the middle, your estate underneath

Data Workers connects to Estuary over its API or MCP server today. What it needs from each pipeline sits where Estuary writes: the materialized tables in Snowflake or BigQuery, read natively, next to the source database (PostgreSQL is a native connection) and the dbt project. Captures, collections, materializations, their schema-change settings and every backfill stay with your team: Data Workers never starts, stops or backfills an Estuary task; it proposes a setting change for the Estuary owner to make. The rest of this incident runs on native connections: Snowflake, PostgreSQL, dbt Cloud, Opsgenie, Slack and GitHub, among 50+ connectors that also cover BigQuery, Databricks, Airflow and Dagster. Lightdash connects over its API or MCP server today, and enters the blast radius through the team's context-graph note.

The owner's Estuary tools and the Data Workers agents run side by side in one client. Estuary's agent skills drive flowctl with the owner's own login and its MCP server answers from Estuary's docs, so the owner inspects the task and changes the binding in the same session where Data Workers explains what the change did downstream. The Data Workers side follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry.

# Example: Estuary's skills and docs MCP plus Data Workers agents in Claude Code
# Estuary agent skills (they run flowctl; authenticate flowctl with your access token first)
/plugin marketplace add estuary/agent-skills
/plugin install estuary-operations@estuary
claude mcp add --transport http estuary-docs https://estuary.mcp.kapa.ai  # Google sign-in on first use

# Data Workers agents, from a clone of the open-source repo
claude mcp add --scope user dw-incidents -- "$(pwd)/start-agent.sh" dw-incidents
claude mcp add --scope user dw-quality -- "$(pwd)/start-agent.sh" dw-quality
claude mcp add --scope user dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
claude mcp add --scope user dw-schema -- "$(pwd)/start-agent.sh" dw-schema
claude mcp add --scope user dw-connectors -- "$(pwd)/start-agent.sh" dw-connectors

List the tools with your client's own command (/mcp in Claude Code). In this incident: monitor_metrics (dw-incidents) flags the row-count drop against the team's baseline; run_quality_check (dw-quality) runs volume and null checks on Snowflake; diagnose_incident names the cause; blast_radius_analysis and trace_cross_platform_lineage (dw-context-catalog) map what the change reaches, with get_dbt_model_lineage (dw-connectors) for the dbt side; send_opsgenie_alert and resolve_opsgenie_alert handle the alert; remediate re-checks the quality assertions after the rebuild and escalates any failure to a person.

In production the agents run in your infrastructure and hold the warehouse credentials and model key, much as Estuary's BYOC keeps its data plane in your own cloud account. Your data stays in your systems; the hosted Conductor sees workflow metadata only. More in where does our data go.

One incident, L0 to L4, set per domain:

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. The promotion request goes in at 14:40, the arena drops out of pacing at 15:00, and an analyst untangles the refill and the failed cast by evening, after the money is spent.
  • •L1 observe. Data Workers flags the row-count drop and the type change at 14:17 with the cause and blast radius. Nothing changes.
  • •L2 propose. Data Workers proposes the hold and the dbt diff, and the setting note; nothing moves until the named owner approves, and an unanswered request expires and escalates, never auto-grants.
  • •L3 act reversibly. For a class with a clean record, Data Workers queues the approved rebuild itself and sends any failed check to a person. Estuary settings and backfills stay with the Estuary owner.
  • •L4 autonomous. For a scoped domain, Data Workers checks materialized tables as changes land, so the owner's fix is ready before the next scheduled job reads them.

What changes for your team

Six jobs that run on autopilot with Data Workers next to Estuary, with a concrete example of each
  • •On-call starts from a cause. The alert carries the source change, the damage and the proposed fix.
  • •Real-time stops meaning real-time mistakes. A breaking change that lands in seconds reaches the jobs and dashboards that read it as a held run, not a wrong number.
  • •Schema-change settings become a reviewed decision. Each incident shows the Estuary owner what backfill, disableBinding or abort would have meant for a real table, so the setting fits the binding.
  • •One record for data and platform teams. Both see the same incident and receipt in Spellbook Data Catalog (in preview), linked from the Opsgenie alert.

Keep Estuary, or consolidate?

Keep Estuary if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

For most Estuary teams the answer is keep it: one read of the source serving sub-second and hourly destinations, in-flight derivations and BYOC are hard to replace. What teams consolidate is the tooling around the landed data: a separate observability tool, hand-written row-count checks and "watch the dashboard after a migration" runbooks. Many estates also run a second mover: see you're on Fivetran, you're on Airbyte and, for self-run CDC on Kafka, you're on Debezium; for the models, you're on dbt.

Weighing a build on Estuary's agent skills and a coding agent? Running flowctl from a chat is easy; the context graph, approvals, undo and receipts are the work. See build it ourselves with Claude Code and MCP servers.

The case for your CFO

The outcome. A change that lands in the real-time tables behind sales or operations is caught and corrected before anyone spends money on it, with a record of how it was checked.

The risk story. At L0 and L1, agents only read. At L2 they propose and a named person approves; an unanswered request expires and escalates, never auto-grants. At L3 they act on reversible changes inside the domains you open; L4 is a later choice per domain. Estuary settings and backfills stay with the Estuary owner. No agent can promote its own work, and an org-wide stop halts all autonomous dispatch. Every change carries a receipt: what changed, who approved it, the blast radius and the undo.

Why now. Real-time pipelines now feed apps and agents as well as dashboards, so a wrong number reaches a decision in seconds.

The first win. L1 on the materialized tables that feed money numbers: a source migration becomes a diagnosed incident before the next scheduled job reads it.

What stays the same. Nothing migrates: Estuary, your dbt project, your warehouse and your on-call rota. For the numbers, see the ROI of agentic data operations.

The sentence for upstairs: "Estuary keeps our data fresh; Data Workers makes sure it's right and gets it fixed with our approval when it isn't, before we spend or report on it."

Getting started

Start with a pilot. Pick the Estuary materializations behind the numbers leaders and customers read, give Data Workers read access to them and their sources, and run at L1 for a few weeks. Then turn on L2 for one domain, and open L3 for a narrow class once the receipts show the agents were right. The pilot path is on the pricing page, and the pilot is credited in full against the first year.

FAQ

How does Data Workers connect to Estuary? Over Estuary's API or MCP server today, with Estuary's agent skills and flowctl in the same client for the owner's side. Data Workers reads the materialized tables, the source database and dbt natively. It works the same on Estuary's public cloud, a private deployment or BYOC.

Every Estuary task is green. How can the table be wrong? Green means Estuary did what you configured. On an incompatible schema change the default is to backfill the binding, which refills the destination table (dropping and re-creating it when the table structure must change); anything that reads it mid-refill sees part of the data. A change the destination accepts in place can still break a downstream model. Data Workers checks what landed and what reads it.

Should we switch the materialization to abort or disableBinding instead? It depends on the binding. backfill keeps the table correct in the end, disableBinding stops that table updating, disableTask stops the whole materialization, and abort leaves the decision to a person. Data Workers shows, per incident, what each would have meant downstream; the Estuary owner decides.

Will Data Workers start, stop or backfill our Estuary tasks? No. Captures, materializations, settings and backfills stay with the owner, in the Estuary UI, flowctl or a coding agent running Estuary's skills. Data Workers proposes changes, queues approved downstream runs through your orchestrator and records the dbt Cloud reruns the owner queues after approval.

Should we let agents run Estuary through its skills? They are a fast way to build and debug pipelines, and they act with the flowctl login you give them. Keep the pipelines that feed money or customers behind an approval; Data Workers adds the cross-system context and the approval flow around those changes.

Sources

  • •Estuary, homepage ("Right-Time Data Platform for CDC, Streaming and Batch ETL"; "One data pipeline for your analytics, apps and agents"; 200+ connectors; Public Cloud, Private Cloud, BYOC; Agent Skills "New"; "Agents connected to Estuary through MCP can inspect, observe, edit, and even build new pipelines"; connector builder skills Beta), https://estuary.dev/ (checked Oct 3, 2026)
  • •Estuary docs, Core concepts (captures, collections, materializations, derivations), https://docs.estuary.dev/concepts/ and Derivations, https://docs.estuary.dev/concepts/derivations/ (checked Oct 3, 2026)
  • •Estuary docs, Schema evolution guide (inference "only widens schemas"; autoDiscover; evolveIncompatibleCollections; onIncompatibleSchemaChange backfill default, abort, disableBinding, disableTask; triggers only on automated actions such as AutoDiscover), https://docs.estuary.dev/guides/schema-evolution/ and Schema evolution concept ("If the schema change requires a different table structure (such as an incompatible column type change ...), the resource is dropped and recreated; otherwise it is truncated"), https://docs.estuary.dev/concepts/advanced/evolutions/ (checked Oct 3, 2026)
  • •Estuary docs, Dekaf, https://docs.estuary.dev/reference/Connectors/materialization-connectors/Dekaf/ (checked Oct 3, 2026)
  • •Estuary docs, Private and BYOC deployments, https://docs.estuary.dev/private-byoc/ (checked Oct 3, 2026)
  • •Estuary docs, OpenMetrics API, https://docs.estuary.dev/reference/openmetrics-api/ (checked Oct 3, 2026)
  • •Estuary docs, Agent Skills, https://docs.estuary.dev/guides/agent-skills/, MCP integration, https://docs.estuary.dev/features/mcp-integration/, and Using coding agents with Estuary, https://docs.estuary.dev/features/using-coding-agents/ (checked Oct 3, 2026)
  • •Data Workers open-source repository (tool registrations in dw-incidents, dw-quality, dw-context-catalog, dw-schema, dw-connectors), https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
  • •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)