You're on Airbyte: It Moves the Data In From 700+ Sources. Data Workers Owns Whether What Lands Is Right
Airbyte syncs data from 700+ sources and records nulled values in _airbyte_meta instead of failing the sync. Data Workers catches what landed wrong, traces it downstream and gets the fix approved.
Your team runs connections: a source, a destination and the streams you selected, each sync job writing to Snowflake, BigQuery or Databricks on a schedule or when Dagster or Airflow says so. The long tail of partner APIs lives in Connector Builder sources, where the AI Assistant (beta) prefills auth, pagination and incremental sync. Each connection has a schema propagation setting, every row arrives with _airbyte_raw_id, _airbyte_extracted_at and _airbyte_meta, and when something goes wrong you refresh the stream. You run it on Airbyte Cloud, on Pro with capacity-based pricing, on Enterprise Flex with data planes in your own infrastructure, or as open-source Airbyte Core.
Airbyte's homepage names the hard part itself: "Schema changes break pipelines, dashboards, and agent answers." And by design, since Destinations V2, a value that doesn't fit the declared type no longer fails the sync: Airbyte writes null, records the change in _airbyte_meta and keeps moving your data. Right for a data mover; it also means a sync can be green while a column you bill on is a third empty. Data Workers watches what lands, traces it through dbt, BI and the jobs that read it, and gets the fix approved before anything downstream builds on it.
Key takeaways
- •Airbyte keeps its job. Connections, streams, custom connectors, schedules and refreshes stay with your team. Data Workers works on what the syncs land.
- •Green syncs get checked too. Null, volume and freshness checks on landed tables, plus the
_airbyte_metachange count against a baseline, turn a green sync with wrong values into a diagnosed incident the same night. - •Connected over Airbyte's API or MCP server today. Data Workers reads what the syncs land,
_airbyte_metaincluded, natively in the warehouse; the Airbyte MCP sits next to the Data Workers agents in the same client for sync status and logs. - •Every fix goes through a named person. The owner approves the hold, the connector change, the dbt diff and the refresh plan, and runs the Airbyte side; each step leaves a receipt.
- •Autonomy is set per domain. Start at L1 observe, move to L2 propose, and open L3 act reversibly for narrow classes once the record earns it.
Airbyte is the on-ramp for 700+ sources. Data Workers owns whether what lands is right.
Airbyte's job is to get every record from the source to the destination, keep the sync running when the source changes shape, and say exactly what it had to change on the way. The job after the load is different: notice that a successful sync landed wrong values, tie them to a partner's change, size the damage, hold anything that would ship the error, get the cause fixed, refresh and rebuild in order, and prove the result. Here is one night at an industrial supplies distributor. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| Sun 21:30 | Carrier API | The freight carrier ships an API release. freight_cost on the shipments endpoint now arrives as a formatted string ("1,240.50") for amounts of 1,000 and up; smaller amounts stay numbers |
| Mon 01:00 | Dagster + Airbyte | Dagster's nightly schedule starts the carrier_shipments connection, built on the team's Connector Builder source, into Snowflake (Incremental, Append + Deduped). The stream schema declares freight_cost a number, so Airbyte writes null for 2,066 of 6,480 records and adds a NULLED change with reason DESTINATION_SERIALIZATION_ERROR to each row's _airbyte_meta. The sync job succeeds |
| 01:20 | Dagster + dbt + Snowflake | Dagster runs dbt build. fct_order_margin coalesces missing freight to zero, so the largest truckload orders now show inflated margins. Every dbt test passes |
| 01:35 | Data Workers + Snowflake | The null check on raw_carrier.shipments.freight_cost reads 31.9% against a 30-day norm of 0.3%. Data Workers opens an incident |
| 01:42 | Data Workers + Snowflake | Data Workers reads _airbyte_meta on the new rows: every new null carries the same NULLED change and the same sync_id, all on LTL and full-truckload shipments, while freight under 1,000 landed intact. The format changed at the source, not in the connector. Blast radius: fct_order_margin, the Tableau "Contribution margin by customer" workbook, and Dagster's publish_price_floors job, which writes minimum quote prices to the quoting app's PostgreSQL at 07:00 |
| 01:45 | ServiceNow + Slack | Data Workers opens a ServiceNow incident with the diagnosis and sends the approval request to the connection owner in Slack |
| 06:10 | Spellbook | The owner reviews four proposals: pause publish_price_floors; declare freight_cost a string in the custom connector's stream schema; a dbt diff that parses it in stg_carrier__shipments; and a run plan (publish the connector, refresh the shipments stream retaining records, rebuild the margin models, then publish floors). She approves all four |
| 06:20 | Dagster + GitHub | She stops the publish_price_floors schedule in Dagster and merges the dbt diff |
| 06:35 | Airbyte | In Airbyte she publishes the connector change, refreshes the source schema and starts "Refresh stream and retain records" on shipments. The refresh lands at 07:14 as a new _airbyte_generation_id, deduped on shipment_id |
| 07:22 | Data Workers + Dagster | Data Workers queues the approved dbt run through Dagster; the margin models rebuild |
| 07:40 | Data Workers + Snowflake | Data Workers verifies: freight_cost nulls are back at the norm, the new generation carries no _airbyte_meta changes, shipment counts match the night before with no duplicate shipment_id, the column change from NUMBER to VARCHAR is recorded, and the marts' load lag is back on its baseline |
| 07:45 | ServiceNow | The on-call engineer resolves the ticket, which links the Data Workers receipt: cause, approvals, runs, checks passed and the undo |
| 07:55 | Dagster + PostgreSQL | The owner restarts the schedule and runs publish_price_floors; the quoting app gets floors built on true freight costs |
| 09:00 | Tableau | The pricing review opens the margin workbook on verified numbers |

Every part of Airbyte did its job: the connector read every record, the destination connector would not write a string into a number column, and _airbyte_meta said exactly which values it nulled and why. Catching it takes knowledge Airbyte was never meant to hold: that freight_cost feeds a margin model, that zero is a plausible but wrong default, and that a job pushes prices to sales at 07:00.
| Job | What Airbyte does | What Data Workers does |
|---|---|---|
| The move | Runs connections and sync jobs from 700+ connectors and custom Connector Builder sources; dedupes, refreshes and retries | Reads what the syncs land, natively in the warehouse; connects to Airbyte over its API or MCP server |
| The signal | Sync status, the connection Timeline, notifications and per-row _airbyte_meta changes; a content problem no longer fails the sync | Checks the landed data itself: nulls, volume and freshness, plus the _airbyte_meta change count as a recorded metric, so a green sync with wrong values still raises an incident |
| The diagnosis | Explains a failed sync from its logs, in the UI or through the Airbyte MCP | Joins the landed rows and their _airbyte_meta changes, the source and lineage into one cause, with the blast radius across dbt models, dashboards and downstream jobs |
| The fix | Applies whatever connector manifest and stream schema the team publishes | Proposes the connector change, the dbt diff and the hold to the owner, who approves and applies them |
| The refresh | Refreshes a stream, retaining or removing records, with no data downtime | Proposes the refresh in order with the rebuilds after it; the owner runs the refresh in Airbyte, and Data Workers queues the approved downstream runs |
| The proof | Keeps sync history, generations and audit logs | Re-checks the tables and writes a receipt: what changed, who approved it, how it was checked, how to undo it |
Why doesn't Airbyte just do this itself?
Because Airbyte is built to move data faithfully from any source to any destination, and it made a deliberate choice about content problems. In its docs, a failing sync now means Airbyte "could not move all of your data", while bad values are recorded in _airbyte_meta so you "decide how to handle rows with errors/changes on a case-by-case basis." Schema propagation has the same scope: per connection, Airbyte propagates changes, waits for your approval or stops syncs, and pauses on its own when a cursor or primary key disappears. Airbyte does not know that a nulled freight cost becomes a zero in a margin model, or that a pricing job reads that model at 07:00.
Airbyte's AI features point the same way. The Connector Builder's AI Assistant (beta) drafts connectors under "Human Oversight Required." The Airbyte MCP (private beta) builds connections, starts syncs and reads job logs with your own Airbyte permissions. Airbyte Agents gives your own agents typed connectors, a Context Store with semantic search, and a read-only history of every session.
Owning whether landed data is right across a partner API, a warehouse, dbt, BI and a pricing job is a different product: a context graph of every table and consumer, blast-radius scoping, named approvers, a recorded undo and receipts an auditor can read. That is Data Workers. More in is it safe to let AI agents change production data.
Every tool owns a slice. Data Workers covers the whole lifecycle
Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the Airbyte connections already there.

| Stage | Data Workers | Airbyte | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 4 | The Context Store resolves entities and searches synced records for agents. Data Workers keeps one governed context graph of what each table means, who owns it and what reads it. |
| Analytics & Insights | 8 | 2 | Not Airbyte's job: it lands the data and answers agent lookups, while business numbers live in BI. Data Workers answers data questions from governed definitions with lineage behind every number. |
| Data Quality | 8 | 3 | _airbyte_meta records every value it nulled and why, so a content problem never fails the sync. Data Workers runs null and volume checks on the landed tables, tracks lateness against a baseline your team records, and turns those records into an incident. |
| Observability & Incidents | 8.5 | 4 | Connection Timeline, sync status, notifications and MCP log reading explain a failed sync. Data Workers diagnoses the data incident when the sync succeeded, proposes the fix and verifies it. |
| Pipelines & Ingestion | 8.5 | 9 | Airbyte's home stage: 700+ connectors, the Connector Builder with its AI Assistant, CDC, refreshes and 5-minute syncs on Pro. Data Workers plans the refresh and reruns for the owner. |
| Schema & Migration | 8 | 5 | Detects source schema changes before each sync and propagates, waits for approval or pauses per connection. Data Workers traces the landed change through dbt, BI and downstream jobs. |
| Governance & Access | 8.5 | 4 | RBAC, SSO and audit logs on Pro and Flex; the MCP server acts with your own Airbyte role. Data Workers routes every data change to a named approver. |
| Security & Privacy | 8 | 5 | Column hashing, external secrets and Enterprise Flex data planes in your own infrastructure protect the move. Data Workers leaves a receipt on every data change. |
| Cost / FinOps | 8 | 3 | Credits or capacity-based Data Workers units price the syncs; warehouse spend sits in the warehouse. Data Workers traces Snowflake credits to the dbt model behind them. |
| MLOps & Models | 7.5 | 3 | Airbyte Agents, agent connectors and semantic search feed AI agents. Data Workers keeps the data under models and agents healthy. |
How Airbyte and Data Workers work together

Data Workers connects to Airbyte over its API or MCP server today. What it needs from each sync sits in the warehouse Airbyte writes to, where _airbyte_meta, _airbyte_extracted_at and _airbyte_generation_id are ordinary columns it reads natively. Connections, custom connectors, schedules, schema propagation settings and refreshes stay with your team: Data Workers never starts, refreshes or pauses a sync. It proposes the connector change for the owner to publish and the refresh for the owner to run in Airbyte. The rest of this incident runs on native connections: Snowflake, dbt, Dagster, Tableau (read), PostgreSQL and ServiceNow, among 50+ connectors that also cover BigQuery, Databricks, Airflow and Prefect.
The Airbyte MCP and the Data Workers agents run side by side in one client. The hosted Airbyte MCP (private beta, Airbyte Cloud) signs you in with OAuth and acts with your Airbyte role, so the owner reads sync logs, publishes the connector change and starts the refresh in the same session where Data Workers explains what last night's syncs did to the data. The Data Workers side follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry.
# Example: the Airbyte MCP plus Data Workers agents in Claude Code
# Airbyte MCP (hosted, private beta): authenticate from /mcp in Claude Code
claude mcp add --transport http airbyte https://mcp.airbyte.com/mcp
# Data Workers agents, from a clone of the open-source repo
claude mcp add --scope user dw-incidents -- "$(pwd)/start-agent.sh" dw-incidents
claude mcp add --scope user dw-quality -- "$(pwd)/start-agent.sh" dw-quality
claude mcp add --scope user dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
claude mcp add --scope user dw-schema -- "$(pwd)/start-agent.sh" dw-schemaList the tools with your client's own command (/mcp in Claude Code). In this incident: run_quality_check (dw-quality) runs the null and volume checks; monitor_metrics (dw-incidents) tracks the _airbyte_meta change count the team records against its baseline; diagnose_incident names the cause; blast_radius_analysis and trace_cross_platform_lineage (dw-context-catalog) map what the nulls reached; remediate re-checks the quality assertions after the rebuild and escalates any failure to a person. On the Airbyte side, get_cloud_sync_logs, update_custom_source_definition and run_cloud_sync cover the owner's part.
In production the agents run in your infrastructure and hold the warehouse credentials and model key, much as Enterprise Flex keeps Airbyte's data planes in your environment. Your data stays in your systems; the hosted Conductor sees workflow metadata only. More in where does our data go.
One incident, L0 to L4, set per domain:

- •L0 manual. A sales rep asks at 10:30 why a truckload quote went out under cost; an analyst finds the nulls by early afternoon, after the floors were live all morning.
- •L1 observe. Data Workers flags the null spike at 01:35 with the cause and blast radius. Nothing changes.
- •L2 propose. Data Workers proposes the hold, the connector change, the dbt diff and the refresh plan; nothing moves until the named owner approves, and an unanswered request expires and escalates, never auto-grants.
- •L3 act reversibly. For a class with a clean record, Data Workers queues the approved rebuilds itself and sends any failed check to a person. Connector changes and refreshes stay with the owner.
- •L4 autonomous. For a scoped domain, Data Workers checks landed tables as new rows arrive, so the owner's fix is waiting before any downstream job runs.
Who sets those levels: who owns the agents, how approvals work for AI data agents and autonomy levels L0 to L4 explained.
What changes for your team

- •On-call starts from a cause. The ticket carries the nulled rows, the damage, the proposed fix and the refresh plan.
- •Custom connectors get a safety net. When a partner API changes and a Connector Builder source starts landing nulls, the owner hears it from Data Workers, not from sales.
- •Jobs that push data out get a guard. When a pricing or billing job reads a table that just went wrong, the hold is proposed before it runs.
- •One record for data and platform teams. Both see the same incident and receipt in Spellbook Data Catalog (in preview), linked from the ServiceNow ticket.
For the choosing and building angles, see Airbyte alternatives, Airbyte vs Fivetran and building Airbyte connectors with Claude Code.
Keep Airbyte, or consolidate?
Keep Airbyte if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
For most Airbyte teams the answer is keep it: open-source connectors, a builder for the long tail and deployment where you need it are hard to replace. What teams consolidate is the tooling around the landed data: a separate observability tool, hand-written _airbyte_meta queries and "refresh and eyeball the dashboard" runbooks. Many estates also run a second mover: see you're on Fivetran, you're on dlt and, if Dagster schedules your syncs, you're on Dagster. One incident record spans them.
Weighing a build on the Airbyte MCP and a coding agent? Read build it ourselves with Claude Code and MCP servers. Reading sync logs from a chat is easy; the context graph, approvals, undo and receipts are the work.
The case for your CFO
The outcome. When an Airbyte sync lands wrong values in the tables behind margin, pricing or revenue, it is caught the same night and corrected before a customer or a sales rep acts on it, with a record of how it was checked.
The risk story. At L0 and L1, agents only read. At L2 they propose and a named person approves; an unanswered request expires and escalates, never auto-grants. At L3 they act on reversible changes inside the domains you open; L4 is a later choice per domain. Connector changes and refreshes stay with the owner in Airbyte. No agent can promote its own work, and an org-wide stop halts all autonomous dispatch. Every change carries a receipt: what changed, who approved it, the blast radius and the undo.
Why now. Airbyte now feeds AI agents as well as dashboards, through its MCP server and Airbyte Agents, so a wrong value travels further before a person sees it.
The first win. L1 on the connections that feed money numbers: a format change at a partner becomes a diagnosed incident before the first downstream job runs.
What stays the same. Nothing migrates: Airbyte, your connections and custom connectors, your schedules, your dbt project, your warehouse and your on-call rota. For the numbers, see the ROI of agentic data operations.
The sentence for upstairs: "Airbyte brings our data in; Data Workers makes sure what lands is right and gets it fixed with our approval when it isn't, before anyone prices or reports on it."
Getting started
Start with a pilot. Pick the Airbyte connections that feed the numbers leaders and customers read, give Data Workers read access to the schemas they write, and run at L1 for a few weeks. Then turn on L2 for one domain, and open L3 for a narrow class once the receipts show the agents were right. The pilot path is on the pricing page, and the pilot is credited in full against the first year.
FAQ
How does Data Workers connect to Airbyte? Over Airbyte's API or MCP server today, with the Airbyte MCP in the same client for sync status and logs. Data Workers reads the landed tables, _airbyte_meta included, natively in your warehouse. It works the same with Airbyte Cloud, Pro, Enterprise Flex or Airbyte Core.
Our sync succeeded. How can the data be wrong? Since Destinations V2, Airbyte separates moving problems from content problems. If a value doesn't match the declared type, Airbyte writes null and records a NULLED change in _airbyte_meta, and the sync still succeeds. Airbyte's docs point you to query those rows; Data Workers checks every new load and traces what the nulls reached.
We set schema propagation to "Approve all changes myself". Isn't that enough? It is a good setting for the declared schema. It doesn't catch a value that changes format inside a column, as in the carrier feed above, or a correct-looking value that means something new. Those land as successful syncs. Data Workers checks what landed.
Will Data Workers start, refresh or pause our Airbyte syncs? No. Connector changes, schema propagation settings, syncs and refreshes stay with the owner, in the Airbyte UI, the API or the Airbyte MCP. Data Workers proposes them in order and queues approved downstream runs through your orchestrator.
Airbyte prices Pro and Enterprise Flex in "Data Workers". Is that related? No. In Airbyte's pricing, Data Workers are the compute units that run your syncs. Data Workers, the Autonomous Agentic Data Platform, is a separate company and product that works with Airbyte over its API and MCP server.
Should we let agents manage Airbyte through its MCP server? Airbyte designed it carefully: it acts with your own role, destructive tools only touch resources the agent created in that session, and most clients ask before a tool changes anything. Keep connections that feed money or customers behind an approval; Data Workers adds the cross-system context and the approval flow around those calls.
Sources
- •Airbyte, homepage ("Data Movement for Analytics and AI Agents"; "700+ connectors"; "Schema changes break pipelines, dashboards, and agent answers"), https://airbyte.com/ (checked Oct 3, 2026)
- •Airbyte, pricing (Standard, Plus, Pro, Enterprise Flex; capacity-based pricing on Data Workers compute units; Airbyte Core), https://airbyte.com/pricing (checked Oct 3, 2026)
- •Airbyte docs, Manage and monitor data workers, https://docs.airbyte.com/platform/cloud/managing-airbyte-cloud/manage-data-workers (checked Oct 3, 2026)
- •Airbyte release notes, Airbyte 2.0 (released Oct 14, 2025), https://docs.airbyte.com/release_notes/self-managed/v-2.0, and Airbyte 2.1 (released Apr 3, 2026), https://docs.airbyte.com/release_notes/self-managed/v-2.1 (checked Oct 3, 2026)
- •Airbyte docs, Enterprise Flex, https://docs.airbyte.com/platform/enterprise-flex (checked Oct 3, 2026)
- •Airbyte docs, Schema change management, https://docs.airbyte.com/platform/using-airbyte/schema-change-management (checked Oct 3, 2026)
- •Airbyte docs, Typing and deduping (per-row change handling in
_airbyte_meta), https://docs.airbyte.com/platform/using-airbyte/core-concepts/typing-deduping, and Direct-Load tables, https://docs.airbyte.com/platform/using-airbyte/core-concepts/direct-load-tables (checked Oct 3, 2026) - •Airbyte docs, Airbyte metadata fields, https://docs.airbyte.com/platform/understanding-airbyte/airbyte-metadata-fields (checked Oct 3, 2026)
- •Airbyte docs, Refreshing your data, https://docs.airbyte.com/platform/operator-guides/refreshes (checked Oct 3, 2026)
- •Airbyte docs, AI Assistant for the Connector Builder (Beta), https://docs.airbyte.com/platform/connector-development/connector-builder-ui/ai-assist (checked Oct 3, 2026)
- •Airbyte docs, Airbyte MCP (private beta), https://docs.airbyte.com/platform/airbyte-mcp, tools and safety, https://docs.airbyte.com/platform/airbyte-mcp/tools, install, https://docs.airbyte.com/platform/airbyte-mcp/install, and context layer, https://docs.airbyte.com/platform/context-layer (checked Oct 3, 2026)
- •Airbyte Agents docs and release notes (latest Oct 1, 2026), https://docs.airbyte.com/ai-agents and https://docs.airbyte.com/release_notes/agents (checked Oct 3, 2026)
- •PyAirbyte 0.71.1 (Sept 30, 2026), https://pypi.org/project/airbyte/, and airbyte-agent-sdk (Oct 2, 2026), https://pypi.org/project/airbyte-agent-sdk/ (checked Oct 3, 2026)
- •Data Workers open-source repository (tool registrations in dw-incidents, dw-quality, dw-context-catalog, dw-schema, dw-connectors), https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
- •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)