Product
Product9 min readBy The Data Workers Team

What integrations does Data Workers support? Every category, what it reads and what it can change

Data Workers connects natively to your warehouses, lakehouses, pipelines, BI, catalogs, observability, ticketing and code hosts, and to everything else over its API or MCP server. Every agent is an MCP server too.

Data Workers connects to your warehouses and lakehouses, catalogs, transformation and orchestration tools, quality and observability, BI, ticketing and code hosts through native connectors, and to everything else over its API or MCP server today. And every Data Workers agent is itself an MCP server, so any assistant or coding agent your company already runs can call it.

A list alone tells you little, so this page also shows what Data Workers reads in each system, what it can change and who has to say yes first, built only from the connector code in the product and the Apache 2.0 core.

Key takeaways

  • •50+ connectors across the whole estate. Snowflake, Databricks, BigQuery, PostgreSQL, Azure Blob, dbt, Airflow, Dagster, Kafka, Monte Carlo, Datadog, Tableau, Looker, ServiceNow, PagerDuty, Okta, GitHub and more.
  • •Everything else connects over its API or MCP server today. If a tool has a REST API or ships an MCP server, Data Workers works with it from day one.
  • •Every agent is an MCP server. Claude, ChatGPT, Claude Code, Cursor and Codex can call Data Workers agents directly, through approvals.
  • •Reading is broad; acting is gated. Agents read freely within the grants you issue, propose changes with a blast radius and rollback path, and act only through approvals set per domain.
  • •Zero migration. Data Workers works on top of the tools you run. Your warehouse stays the warehouse, your catalog stays the catalog.

How Data Workers connects

There are four paths, and most estates use all of them.

Native connectors. The Data Context Wizard lists 50+ connectors, and they are real code in the product: a connector per platform with discovery, health checks and, where the platform allows it, governed writes. Context Wizard reads them into one governed context graph, so every agent works from the same picture of tables, lineage, owners, quality and cost.

APIs. Anything else connects through its own REST, GraphQL or SQL interface, under a scoped credential you issue.

Vendor MCP servers. dbt, Fivetran, Monte Carlo, DataHub, Snowflake, Databricks, Tableau and Power BI all publish MCP servers now. Data Workers connects to those servers today and puts its approvals, receipts and cross-system context around what they expose.

Open standards. The Kafka Schema Registry means Data Workers reads and writes stream schemas once for every engine that speaks it, and Data Workers reports its own runs to your lineage server as OpenLineage run events.

And in the other direction, every agent in the Data-Agents Swarm is its own MCP server. The open-source core ships as a set of them that your engineers add to their client by following the client setup guide: clone the repo and point the client at start-agent.sh.

Example: an MCP client config with two Data Workers agents

{
  "mcpServers": {
    "dw-connectors": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-connectors"]
    },
    "dw-catalog": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-context-catalog"]
    }
  }
}

From there list_all_catalogs shows every catalog the agents can see, get_connector_status reports which connectors are live, and trace_cross_platform_lineage follows a column from source to dashboard. For hosted assistants, the product's remote endpoint accepts an API key or OAuth tokens from your own identity provider, such as Okta or Microsoft Entra ID, which Data Workers verifies against the provider's JWKS. Data Workers issues no tokens of its own. The full setup for each client is in Data Workers with your coding agents.

How Data Workers fits with Your estate: your coding agent on top, Data Workers in the middle, your estate underneath

The integration list, by category

Every row below comes from connector and agent code we checked on Oct 2, 2026. "Reads" is what agents see within the grants you issue. "Proposes or changes" is what they put in front of the owner, and what they do after a named person approves it, or on their own only where you have set that domain to act.

CategoryNative connectorsWhat Data Workers readsWhat it proposes or changes (through approvals)
Warehouses and databasesSnowflake, Databricks (Unity Catalog), BigQuery, PostgreSQLDatabases, schemas, DDL, query history, usage and cost; Unity Catalog grantsSchema migrations generated with rollback SQL, for the owner to apply; Unity Catalog grants and revocations through its permissions API
Lake storageAzure Blob and ADLS Gen2Containers and dataset schemasRead: the fix lands in the pipeline that writes the files
StreamingApache Kafka, Kafka Connect, Schema RegistryTopics, consumer lag, connectors, schema versions and compatibilityRegister a compatible schema version (register_kafka_schema)
Transformationdbt (the manifest; dbt Cloud)Models, test definitions and column lineage from the manifestModel, test and docs changes as diffs for the owner to merge; dbt Cloud reruns queued by the owner's scheduler after approval
OrchestrationAirflow, Dagster, Prefect, Azure Data Factory, Managed Service for Apache Airflow (formerly Cloud Composer), AWS Step FunctionsDAG, job and flow runs, task states and schedulesReruns (trigger_airflow_dag on Airflow 2; on Airflow 3 the owner triggers the approved run; trigger_dagster_job, trigger_prefect_flow, trigger_adf_pipeline, trigger_step_function)
LineageOpenLineageYour emitters and backend stay as they areOpenLineage run events for Data Workers' own runs, sent to the endpoint you configure
Data qualityMonte Carlo, Great Expectations, SodaMonitors, suites, expectation results and check resultsRun a suite (run_monte_carlo_suite, run_gx_suite, run_soda_suite) to verify a fix
Observability and alertingDatadog, OpenTelemetry, New Relic, PagerDuty, Opsgenie, Slack and Microsoft Teams (alert channels)Metrics, traces, monitors and alertsSend and resolve alerts (send_pagerduty_alert, resolve_pagerduty_alert)
TicketingServiceNow, Jira Service ManagementJira Service Management tickets and their history; the ServiceNow incidents it opens or your flows hand itOpen tickets and keep their summary current (ServiceNow priority too), post diagnosis and receipt as Jira Service Management comments, link the receipt from Spellbook and the audit trail; your service agents resolve (create_servicenow_ticket, update_servicenow_ticket, create_jira_sm_ticket, update_jira_sm_ticket)
BITableau, LookerDashboards, datasources, Explores and the tables under themRead by design: the fix lands upstream, and the affected dashboards are named in the blast radius and the receipt
Code hostsGitHubPull requests and diffs that touch dataData review summaries on your pull requests; dbt docs changes as pull requests once your team turns on the GitHub target
Identity and accessOkta, Microsoft Entra IDUsers and groupsUsers and groups read as context for access requests; grants applied on Unity Catalog after approval and proposed for the owner elsewhere, each with its expiry recorded
CostAWS Cost Explorer, Snowflake usage, BigQuery job historyAWS spend by service and forecast; Snowflake credits by warehouse, query and dbt model; BigQuery account totalWarehouse settings drafted and cleanups proposed for the owner, checked against dependencies

Fivetran, Airbyte, Stitch, dlt, AWS DMS, Debezium, Redshift, Power BI, Microsoft Fabric, Sigma, Metabase, Superset, Elementary, Anomalo, Bigeye, MLflow, Weights & Biases, AWS Security Hub, Argo, Kestra, Mage, Temporal, ClickHouse, Oracle, SQL Server, GitLab, Jira, Linear, Hex, ThoughtSpot, the catalogs (DataHub, OpenMetadata, Microsoft Purview, Knowledge Catalog, Collibra, Alation and Atlan) and the lakehouse catalogs (Iceberg REST, Apache Polaris, Project Nessie, AWS Glue Data Catalog and Lake Formation, Hive Metastore) connect over their API or MCP server today, with approved facts recorded in the Context Wizard graph and catalog changes proposed for the owner, under the same approvals and receipts as everything in the table. The platform guides go deeper for each estate: Snowflake, Databricks, Google Cloud, AWS and Microsoft Fabric. The integration guides show the loop through one tool each: dbt, Airflow, Dagster and Prefect, Fivetran, OpenLineage, Tableau, Power BI and Looker and Monte Carlo. How the Connectors agent itself works is in Inside the Connectors Agent.

Reading, proposing and acting are three different permissions

The question a platform owner asks is what the agent may do once connected. Data Workers answers it per category and per domain.

Comparison matrix of Your estate and Data Workers on the outcomes a data leader buys

Reading is broad by design: the agents need the whole estate to find causes, so they read wherever you have issued a grant. Proposing is where Data Workers does most of its work: a migration with rollback SQL, a dbt diff, a catalog update, each with the blast radius from blast_radius_analysis attached. Acting is gated. Each domain sits on the autonomy ladder, and approvals go to a named person; an unanswered request expires and escalates, and never auto-grants.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous

At L0 manual and L1 observe agents read and report; at L2 propose every change waits for approval; L3 act reversibly applies changes that carry a rollback; L4 autonomous is for domains you trust. Most teams start every integration at L1 or L2. The detail is in How do approvals work for AI data agents? and Is it safe to let AI agents change production data?.

A worked example across six integrations

This is an illustration, not a customer event. The estate: Salesforce syncs through Fivetran into Snowflake, dbt Cloud builds the models from a GitHub repo, Airflow schedules the run, Tableau serves the revenue dashboard and ServiceNow holds data incidents. The finance domain runs at L2 propose.

TimeSystemWhat happens
02:10FivetranThe overnight sync lands opportunity.amount as a string after a Salesforce field change.
02:12Snowflake and dbtrun_quality_check fails on the landed table: amount is now a string; trace_cross_platform_lineage follows it through stg_opportunities to fct_revenue.
02:14TableauThe blast radius reaches the revenue workbook and its extract, refreshed at 06:00.
02:15ServiceNowcreate_servicenow_ticket opens a data incident with the cause and the blast radius.
02:18GitHubData Workers proposes a diff that casts amount in the staging model and tightens the source test, for the model owner to merge.
07:05SpellbookThe analytics engineer on call reviews the diff and the rollback path and approves.
07:12dbt Cloud and AirflowThe engineer merges the change, dbt CI passes, and trigger_airflow_dag reruns the DAG, which rebuilds the models in dbt Cloud.
07:20Tableau and ServiceNowThe extract's next refresh picks up the fixed model, revenue matches the source again, and update_servicenow_ticket updates the ticket summary with the receipt linked from Spellbook and the audit trail: diff, approver, checks and rollback path. The service agent resolves the ticket.

Six systems, one context, one approval and one record in the audit trail.

The alternatives buyers weigh

Every vendor now ships an MCP server or an agent, each well designed for its own system. Data Workers runs the loop across all of them.

Build it yourself. Wire vendor MCP servers into a coding agent. The pieces are good: the dbt MCP server reached v2.5.0 on Sep 28, 2026, and the Fivetran MCP server (v0.3.0, Aug 20, 2026) defaults to read scope, with writes opt-in. Its README is candid that write tools carry advisory instructions to confirm and "the server does not enforce confirmation". Enforcement, blast radius across systems, rollback and receipts become your team's work. Can't we just build this ourselves? maps that work.

Catalogs. The DataHub MCP server (v0.7.1, Sep 16, 2026) exposes search and lineage, with mutation tools behind an opt-in flag. A catalog holds the inventory; Data Workers runs the work and proposes approved facts back to its owner. See Data Workers vs a data catalog.

Observability. Monte Carlo's MCP server offers OAuth for agents and API keys for automations, and needs an Editor role or above (docs checked Oct 2, 2026). Its alert is the signal that starts the Data Workers loop. See Data Workers vs data observability.

Platform-native agents. Snowflake runs a managed MCP server, and Databricks managed MCP servers (Public Preview, docs updated Sep 21, 2026) reach Unity Catalog, AI Search and Genie Agents. Each governs its own platform well; Data Workers carries one context and one approval flow across them and everything around them.

Coding agents. Claude Code, Cursor and Codex are where engineers already work. Data Workers plugs in as MCP servers, so the coding agent writes code and Data Workers supplies the governed context and the change record.

The case for your CFO

The outcome: one platform that works with every tool the data team already pays for, with no migration and no new console per tool. The tools stay; the work between them gets done.

The risk story: agents read through grants you issue and can revoke, write only through approvals set per domain, and leave a receipt on every change with the diff, the approver and the rollback path. Your data stays in your systems; the hosted Conductor sees workflow metadata only. When Data Workers acts through a vendor's API or MCP server, that action goes through the same approval flow and lands in the same audit trail.

Why now: every tool is shipping its own agent. Without one layer across them, you get a dozen permission models and no single record of what changed.

The first win: connect the warehouse, dbt and the ticketing system at L1 observe, then let Data Workers propose fixes for the next schema-change incident. What stays the same: every tool, every contract, every permission system.

The path: start with a pilot at $7,500 one time, and the pilot is credited in full against the first year. Scale starts from $1,000 a month and Enterprise from $3,000 a month, billed annually, with unlimited seats, no usage meter, no markup on model spend, bring your own model and an Apache 2.0 core. See /pricing/, model the return in the ROI calculator, and read the ROI of agentic data operations.

The sentence to repeat upstairs: "Data Workers works with every tool we already pay for, and every change it makes has an owner and a receipt."

FAQ

How many connectors does Data Workers have? 50+ across the estate, named by category in the table above. Anything with an API or an MCP server connects over that today, so what matters is what Data Workers reads and changes in each system, which the table spells out.

We use a tool that isn't in the table. Does that block a pilot? No. Data Workers connects to it over its API or MCP server today, under the same approvals and receipts. Tell us the tool in the pilot scoping and we include it.

Does Data Workers replace the vendor's own MCP server? No. Keep it. Data Workers connects to it and adds what a single-vendor server leaves to you: context from every other system, enforced approvals, blast radius and a receipt.

Can ChatGPT or Claude call Data Workers directly? Yes. Every agent is an MCP server. Engineers add them to Claude Code, Cursor or Codex; hosted assistants reach the product's remote endpoint with OAuth through your identity provider. See Data Workers with your coding agents and AI assistants are rolled out. Now what?.

What credentials does it need, and where do they live? A scoped role per system, issued by you. The agents run in your infrastructure on every tier and hold the credentials there. The details are in Where does our data go? and Can Data Workers run in our VPC or air-gapped?.

Can it bring in our semantic layer, glossary or ontology? Yes. Context Wizard treats your existing context as a first-class source with provenance and joins it to lineage, quality and usage. See Bring your own context.

What does the first month of integrations look like? Warehouse, transformation and ticketing first, at L1 observe, then the rest of the estate. What a Data Workers pilot looks like and the first 90 days walk through it, and our guide to agentic data engineering with MCP covers the protocol side.

Sources

  • •Data Workers, Data Context Wizard ("50+ connectors"), checked Oct 2, 2026: https://dataworkers.io/product/data-context-wizard/
  • •Data Workers, Pricing (rate card; Spellbook Data Catalog in preview), checked Oct 2, 2026: https://dataworkers.io/pricing/
  • •Data Workers, Client setup (open-source core), checked Oct 2, 2026: https://dataworkers.io/opensource-docs/client-setup/
  • •Data Workers, Claude Code page (open-source core), checked Oct 2, 2026: https://dataworkers.io/on/claude-code/
  • •Data Workers open-source repo (connectors and tool registrations in dw-connectors, dw-context-catalog and dw-schema; every agent is a standard MCP stdio server), checked Oct 2, 2026: https://github.com/DataWorkersProject/dataworkers-claw-community
  • •dbt Labs, dbt MCP server releases (v2.5.0, Sep 28, 2026), checked Oct 2, 2026: https://github.com/dbt-labs/dbt-mcp/releases
  • •Fivetran, Fivetran MCP server README and releases (v0.3.0, Aug 20, 2026), checked Oct 2, 2026: https://github.com/fivetran/fivetran-mcp
  • •DataHub MCP server README and releases (v0.7.1, Sep 16, 2026), checked Oct 2, 2026: https://github.com/acryldata/mcp-server-datahub
  • •Monte Carlo, MCP Server documentation, checked Oct 2, 2026: https://docs.getmontecarlo.com/docs/mcp-server
  • •Snowflake, Snowflake-managed MCP server, checked Oct 2, 2026: https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-agents-mcp
  • •Databricks, Databricks managed MCP servers (Public Preview, last updated Sep 21, 2026), checked Oct 2, 2026: https://docs.databricks.com/aws/en/generative-ai/mcp/managed-mcp
  • •Tableau, Tableau MCP server releases (v4.13.3, Sep 23, 2026), checked Oct 2, 2026: https://github.com/tableau/tableau-mcp
  • •Microsoft, Power BI MCP servers overview, checked Oct 2, 2026: https://learn.microsoft.com/en-us/power-bi/developer/mcp/mcp-servers-overview