Product
Product11 min readBy The Data Workers Team

Data Workers + Fivetran: A Fivetran Integration That Catches Schema Changes Before They Break Your Models

A Fivetran integration for AI agents: Data Workers catches Fivetran schema changes from Fivetran's own change record and traces them through dbt to BI.

Fivetran moves your data and keeps each sync running when a source changes shape. Data Workers works out what that change means for the dbt models, metrics and dashboards downstream, and gets the fix reviewed before the next job builds on it. That is the whole shape of this Fivetran integration: Data Workers takes each schema change Fivetran records, as your team's log alert or the owner hands it over, traces every breaking change through the dbt project to the dashboards that read it, and proposes the fix to a named engineer with a receipt on every change. Fivetran itself connects over its REST API or its official MCP server today.

Your estate probably looks like this. Fivetran connections replicate Postgres, Salesforce, Stripe and a dozen SaaS apps into Snowflake, BigQuery or Databricks. Schema change handling is set per connection to Allow all, Allow columns or Block all. dbt turns the raw schemas into staging models and marts, run by dbt platform jobs or Airflow, and a BI tool reads the marts. When a source team retypes or renames a column, Fivetran does exactly what it promises: the sync succeeds and the data arrives without loss. The model that casts that column finds out on its next run, or worse, keeps running and returns a quietly wrong number.

Data Workers is the agentic data platform that runs the whole data lifecycle around that stack. Here is how it wires into Fivetran, what it reads, what it sends back, and what stays exactly where it is.

Key takeaways

  • •Fivetran keeps moving the data. Connections, syncs, schema change handling and destinations stay in Fivetran, configured by your team.
  • •Data Workers turns Fivetran schema changes into reviewed fixes. Fivetran logs every change as an alter_table event; your team's alert on that log, or the owner, hands it to Data Workers, which classifies it as breaking or safe against the dbt manifest and checks the landed table with run_quality_check.
  • •A change is traced before it breaks anything. Data Workers maps a landed change through the dbt manifest to the models, metrics and dashboards it reaches, and proposes the fix with that blast radius attached.
  • •Fivetran stays in your team's hands. When a fix needs a re-sync, Data Workers proposes it with the reason and the columns involved, and your team runs it in Fivetran. Schema config stays your team's call.
  • •Fivetran's MCP server sits beside it. Your coding agent can hold the official Fivetran MCP server, read scope by default, and Data Workers' own MCP tools in one session.
  • •Start read-only. Connect warehouse read access and the dbt manifest; move a domain to propose mode when the reports have earned it.

Going further. This page pairs with Data Workers + dbt, which covers the transformation side of the same loop, and Data Workers + Airflow, which covers reruns through the orchestrator.

What Fivetran does, and why teams keep it

Fivetran is where many data teams stopped writing extraction code. Prebuilt connections handle authentication, incremental syncs, retries and API changes on the source side, and land tables in a destination with predictable naming. Pricing follows Monthly Active Rows, so teams know what drives the bill. Teams keep Fivetran because it turned ingestion into configuration, and because it is careful with change: a new column arrives on its own if you allow it, and a type change is promoted to a type that holds both the old and new values without losing data.

2026 has been a big year for the company. Fivetran and dbt Labs completed their merger on June 1, so ingestion and the dbt project now come from one vendor. In May, Fivetran announced it would become the steward of GX Core, the open source Great Expectations project, and the repository now lives in Fivetran's GitHub organization. Fivetran also ships an official MCP server for agents.

AreaWhat Fivetran shipsStatus, October 2026
CompanyFivetran + dbt Labs mergerCompleted June 1, 2026
GX CoreFivetran becomes steward of the Great Expectations open source project; GX Core stays open sourceAnnounced May 13, 2026; repo now under Fivetran's GitHub org
Schema Change HandlingAllow all, Allow columns, Block all per connection; ALLOW_ALL, ALLOW_COLUMNS, BLOCK_ALL in the APIAvailable
Type changesColumn type promoted to the most specific type that holds old and new data losslesslyAvailable
Column renames (PostgreSQL)A renamed column arrives as a new destination column; the old column staysDocumented connector behavior
Connection Schema APIRead and update schema, table and column config; reload a schema; re-sync tablesAvailable
WebhooksSync, connection, transformation and connector lifecycle events, signed with HMAC SHA-256 if you set a secret; no schema-change eventAvailable
Log eventsalter_table, change_schema_config, schema_migration_start and more, in the Fivetran Platform Connector LOG table and external log servicesAvailable
Connector Metadata APIConnector types and their configuration schemasAvailable
Fivetran MCP serverOfficial, MIT licensed; read scope by default, writes opt-in through FIVETRAN_SCOPEv0.5.1, September 30, 2026
Fivetran Context LayerData and metadata for LLMs and agents, built on the open Agents SchemaAnnounced in private beta Sept 16, 2026; public preview in Fivetran's docs

Every row gives an agent firmer ground: documented schema behavior, an API for schema config, a log of every alter_table, and an MCP server. Data Workers builds on all of it.

What Data Workers reads around Fivetran and what it sends back

The integration runs where Fivetran's work shows up: Fivetran's record of each change and the destination warehouse it lands in. Fivetran's webhooks report syncs and connection events but not schema changes, so the signal is Fivetran's log: the alter_table event lands in the Fivetran Platform Connector LOG table or your external log service, and your team's alert on it, or the owner, hands it to Data Workers. Data Workers classifies every change against the dbt manifest and checks the landed table with run_quality_check. A widened VARCHAR or INT to BIGINT is safe. A number that became a string, a dropped column or a likely rename is breaking for whatever reads it.

Data Workers readsData Workers sends back
Null, key and row-count checks on the tables Fivetran lands in Snowflake or BigQueryA re-sync proposal with its reason, for your team to run in Fivetran
Each alter_table event your team's log alert or the owner hands over, and the dbt manifest diff for the fixProposals for schema config, as a recommendation your team applies in Fivetran or through Fivetran's MCP server
Query history on each landed table, so the blast radius counts who reads itA blast-radius comment on the dbt pull request that carries the fix
Sync state and schema config in Fivetran's own terms, from Fivetran's official MCP server in your team's client sessionA receipt for each change, linked in Spellbook
The dbt manifest, so each raw table maps to the models that read itNothing in connections, destinations or schema handling changes without your team
What Data Workers reads from Fivetran and what it writes back through Fivetran

The reads go into Data Context Wizard, next to the dbt manifest and warehouse metadata, so the graph knows which Fivetran schema lands which table and which models, metrics and dashboards depend on each column. The writes are deliberately narrow. Data Workers leaves Fivetran's configuration with your team: schema config, column blocking and hashing, and Fivetran's catalog stay where they are. A re-sync is proposed with its reason, because a historical re-sync is a real cost decision your team should make.

One incident, through Fivetran

Here is a scenario many Fivetran teams will recognize. It's an illustration, not a customer case.

At 09:40 on a Tuesday, the shop team changes orders.discount in Postgres from a numeric column to text, so they can store values like "10%". The postgres_prod connection syncs at 10:15 over logical replication. Fivetran handles the type change as documented: it promotes discount in Snowflake to a string so old and new values both fit, logs an alter_table event and reports the sync as successful. dbt's stg_shop__orders model casts discount to a number, fct_orders and fct_daily_revenue build on it, the net_revenue metric reads fct_daily_revenue, and the Looker revenue dashboard, recorded in the context graph, reads that mart. The next hourly dbt job runs at 12:00.

StepWhere it runsHandoff
1. SyncFivetranThe sync succeeds. discount is now a string in Snowflake, and the change is in Fivetran's log as alter_table. Fivetran did its job.
2. DetectFivetranAt 10:18 the team's alert on Fivetran's log hands the alter_table event for postgres_prod to Data Workers, which checks it against the dbt manifest and classifies discount NUMBER to VARCHAR as breaking.
3. Tracedbt manifestData Workers follows the column through the manifest: one staging model casts it, two marts build on it, one metric reads the result, and the context graph adds the Looker dashboard.
4. ProposeSpellbookAt 10:30 the proposal is waiting: parse the percent string in stg_shop__orders, add a test for the new format, with the blast radius and a note for the shop team on what their change touched.
5. ApproveSpellbook and GitHubAt 11:05 the analytics engineer approves the plan, and their coding agent opens the dbt pull request from it. Data Workers posts the blast radius on the PR.
6. Testdbt CIdbt CI runs the tests on the PR. The engineer merges at 11:40.
7. Builddbt platformThe 12:00 hourly job builds green on the new staging logic. No model ever failed and no wrong number reached the dashboard.
8. Verify and closeSnowflake and LookerData Workers re-runs its pinned query for net revenue, compares it with the value pinned from Postgres, and closes the incident with a receipt: the change, the cause, the PR, the approver and the checks that passed.
Incident timeline across the stack: what Fivetran, your team and Data Workers each do, step by step

Fivetran kept the data flowing and recorded exactly what it changed. Data Workers turned that record into a decision about the dbt project before anything downstream ran on it. If the source team had renamed the column instead, the story is the same with a different signal: Fivetran adds the new column and keeps the old one, which then stops filling; Fivetran's log shows the new column, run_quality_check sees the old one go null, and it flags the pair as a likely rename and traces every model still reading the old name.

Why doesn't Fivetran just do this itself?

Because Fivetran is built for one job, and its design is right for that job. Faithful replication means the sync keeps running and nothing is lost: promote the type, keep the old column, let the team decide what to allow. A replication tool that refused to land data whenever a downstream model might object would be a worse replication tool. Fivetran's schema handling answers "should this new column arrive?", which is a question about ingestion. "What does this change do to net_revenue?" is a question about every system after it.

Answering that second question, and acting on it, is a different product. It needs blast-radius scoping across every model, metric and dashboard downstream, approvals that name a person, rollback for each change, receipts an auditor can read, context about systems Fivetran doesn't run, and someone accountable for edits in the dbt repo and the warehouse. Fivetran's own MCP server reflects the same sensible boundary: writes are opt-in, and its README says write and delete confirmations are advisory rather than enforced, leaving trust to the client. Taking on liability for cross-system changes would be stepping outside the job Fivetran does well. That cross-system product is the one Data Workers is.

Fivetran schema change handling: what each setting means downstream

Fivetran schema changes come in two kinds, and the setting decides only the first. New schemas, tables and columns follow your Schema Change Handling setting. Changes to columns you already sync, such as a new type or a rename, flow through whatever the setting, because that is how Fivetran keeps data complete.

  • •Allow all. Everything new arrives on its own. Data Workers flags new tables and columns as safe, records them in the context graph and tells you which staging models probably want them. Retypes and renames are where it raises a breaking change.
  • •Allow columns. New columns arrive, new tables don't. Same signal, smaller surface.
  • •Block all. Nothing new arrives until someone enables it. Fivetran's docs note that switching from Block all to Allow columns doesn't re-enable columns blocked earlier; you enable them and run a historical re-sync. Once your team enables the columns, Data Workers proposes the re-sync with the columns and the reason attached, and your team runs it in Fivetran.

Whichever setting you choose, the retyped column and the renamed column still land. That is the gap a downstream check covers, and it's why Data Workers works from Fivetran's change log and the landed table rather than the setting.

Fivetran MCP, webhooks and the REST API: how to connect today

You don't need anything new from Fivetran. Data Workers reads your warehouse and your dbt project, and its agents expose their tools over MCP, so Claude Code, Codex or Cursor can ask what last night's schema changes touch, see the blast radius and approve a proposal from one session. Fivetran's official MCP server can sit next to it in the same client: with its default read scope, your coding agent can also read a connection's schema config and sync timing in Fivetran's own terms. The roles below are an illustration; use what your Fivetran account and warehouse offer.

Week one: read only.

  • •Add Fivetran's official MCP server to your coding agent with its default read scope, so sync state and schema config sit in the same session as Data Workers.
  • •Route Fivetran's `alter_table` log events to an alert that hands each one to Data Workers.
  • •Grant read access to the destination schemas Fivetran writes, so Data Workers can run null, key and row-count checks on every landed table.
  • •Connect the dbt manifest, so each raw table maps to the models that read it.
  • •Every domain starts at observe. The first thing you see is every schema change the alert hands over, classified, with its downstream reach.

When you're ready: proposals.

  • •Move one domain to propose mode, so each breaking change arrives in Spellbook with a staging fix and, where one is needed, a re-sync proposal.
  • •Keep schema config changes and re-syncs with your team, in the Fivetran dashboard or through Fivetran's MCP server set to the narrowest scope that works, with DISALLOWED_ACTIONS for anything you never want an agent to call.
  • •Keep dbt fixes as pull requests your engineers open in your repository, with no direct deploy rights.

Webhooks. If you already route sync_end and connection_failure webhooks to your alerting, keep them. Data Workers doesn't depend on them, since schema changes aren't a webhook event; it works from Fivetran's log and checks the landed tables directly.

One vendor for ingestion and dbt: why a neutral operating layer matters

With Fivetran and dbt Labs now one company, and Fivetran stewarding GX Core, more of the data path comes from one vendor, and that vendor is building agents for it, from dbt Wizard to the Fivetran Context Layer. That is good news for integration inside the stack. The operating layer above it has a different job: it has to see Postgres, Snowflake or Databricks, Airflow and Looker as clearly as Fivetran and dbt, apply one approval flow and one audit trail to all of them, and keep working if you add Airbyte for one source or move a warehouse. A neutral layer keeps your incident record, your approvals and your context graph yours, whatever happens to the vendor map, and it lets you judge each tool on the job it does.

Guardrails: approvals, and what Fivetran owns

  • •Read-only when you want it. Start with read keys everywhere, then open writes one domain at a time.
  • •Autonomy per domain, L0 to L4. Flagging and tracing schema changes can run at L1 from day one, while staging fixes and re-sync proposals run at L2 (propose, you approve).
  • •Approvals. Each change to production waits for a named human, and no agent can promote its own work.
  • •Receipts. Every governed change records what changed, why, who approved it, the blast radius and how to roll it back, in a tamper-evident log.
  • •Rollback. dbt fixes revert like any commit. Re-syncs run in Fivetran, by your team, and Fivetran records them like any other.
  • •What Fivetran owns stays with Fivetran. Connections, destinations, Schema Change Handling, column blocking and hashing, sync schedules and the MAR that drives your bill stay in Fivetran, and Data Workers works from what lands in your warehouse.
The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous

How it fits together

How Data Workers fits with Fivetran: your coding agent on top, Data Workers in the middle, your estate underneath

Your team works in its coding agent and reviews in Spellbook Data Catalog, which is in preview. The Data-Agents Swarm does the work, with the Schema Evolution agent watching what Fivetran logs, and the Autonomous Data-Conductor runs each incident from detection to verification. Fivetran keeps syncing. Nothing migrates: Data Workers stores metadata and scrubbed facts, not copies of your data, and reaches the rest of the estate through 50+ connectors.

What changes for your team

  • •Schema changes become a morning list, not a surprise. Every change from overnight syncs arrives classified, with the models and dashboards it reaches.
  • •Source teams hear about their change. Each proposal carries a note for the source team on what their change touched, which is how the same break stops repeating.
  • •Fewer red dbt runs from upstream. The staging fix is reviewed and merged before the job that would have failed.
  • •Fivetran admins keep their settings. Whoever owns connections and schema handling today keeps owning them.

The fastest first win: a schema change report for one connection

Pick the Fivetran connection whose changes hurt most, usually the production application database. Connect Data Workers read-only for a month. It takes each entry in that connection's change log, as your alert or the owner hands it over, and reports each landed change with its classification and dbt blast radius, so the team sees which past breaks it would have caught. Then move that domain to propose mode: each breaking change arrives in Spellbook with a proposed staging fix. When your engineers have approved those proposals as written for a few weeks, consider L3 for the narrow, reversible part, such as recording safe changes and alerting source teams on their own.

The case for your CFO

The outcome. Revenue, finance and product numbers stop going wrong because someone upstream changed a column. Source changes are caught when they land and fixed before the models that read them run, so fewer reports are pulled, rerun and explained after the fact.

The risk story. At L0 and L1, agents only read: the change events your team hands over, warehouse checks and the dbt project. At L2 they propose and a named person approves. At L3 they act on changes that can be reversed, inside the domains you open, and L4 is a choice you make per domain, later. Every change carries a receipt with what changed, why, who approved it, the blast radius and the rollback path. Fivetran keeps its own log. Nothing migrates.

Why now. With ingestion and transformation now under one vendor and agents arriving on both sides, through Fivetran's MCP server and dbt Wizard, agents will touch the data path whether you plan for it or not. You want one approval flow and one audit trail across all of them before that happens.

The first win. A schema change report on one production connection, read-only, showing which past breaks would have been caught.

What stays the same. Fivetran and its settings, your destinations, your dbt project and CI, your orchestrator, your BI tool, and the coding agent your engineers already use.

The pilot path. Start with a pilot on one connection and the dbt models it feeds (pricing); the pilot is credited in full against the first year.

The sentence for upstairs: "Fivetran keeps moving our data; Data Workers makes sure a change at the source is caught and fixed before it reaches a report."

When Fivetran on its own is enough

If your sources rarely change shape, your dbt staging layer is thin, and Block all plus a weekly review of Fivetran's log keeps you comfortable, Fivetran's own controls can carry you for now. Once source teams change schemas often, more than a few models depend on each raw table, or a quietly wrong number has reached a dashboard, an agent layer that watches what lands, traces it and gets the fix reviewed earns its place.

FAQ

What does the Data Workers Fivetran integration connect to? The destination warehouse, for the schema Fivetran lands; the dbt manifest, for what reads each table; and Fivetran itself over its REST API or official MCP server in the same client session, for sync state and schema config.

How does Data Workers detect Fivetran schema changes? Fivetran records each change as an alter_table log event, in the Fivetran Platform Connector LOG table or your external log service. Your team's alert on that log, or the owner, hands the event to Data Workers, which classifies it against the dbt manifest (added, removed, retyped or a likely rename, breaking or safe) and checks the landed table with run_quality_check. Fivetran's webhooks don't include a schema-change event, so the change log is the reliable signal.

Does Data Workers change our Schema Change Handling settings? No. Schema config, column blocking and hashing stay with your team. Data Workers can recommend a change, and your team applies it in Fivetran or through Fivetran's MCP server.

Can we use the Fivetran MCP server and Data Workers together? Yes. Fivetran's official MCP server runs with read scope by default, and writes are opt-in through FIVETRAN_SCOPE. It sits next to Data Workers' MCP tools in Claude Code, Codex or Cursor, so one session can read Fivetran's config and see Data Workers' blast radius.

Does Data Workers replace Fivetran? No. Fivetran moves the data and Data Workers doesn't. Data Workers watches what lands, works out what it means downstream and gets the fix reviewed.

Fivetran and dbt Labs are one company now. Does that change anything? Data Workers reads the dbt project natively and connects to Fivetran over its API or MCP server, as before. The merger makes a neutral operating layer more useful, since your approvals and audit trail also cover the warehouse, the orchestrator and BI. Data Workers + dbt covers the dbt side.

Can we run it read-only? Yes. Start with warehouse read access and the dbt manifest, and move a domain to propose mode when you're ready.

Sources

Sources for Fivetran capabilities and statuses: Fivetran documentation and announcements current as of October 2, 2026, including Fivetran + dbt Labs complete merger (June 1, 2026), Fivetran to become steward of GX Core (May 13, 2026) and the GX Core repository, dbt Summit 2026 announcements (Sept 16, 2026), the Fivetran MCP server, Schema Change Handling settings, the Connection Schema API, core concepts on type changes, the PostgreSQL connector, new columns not syncing after a schema change, webhooks, log events, the Connector Metadata API and pricing, all checked October 2, 2026. Product names and statuses change quickly; if we've got something wrong, tell us and we'll fix it.