Product
Product12 min readBy The Data Workers Team

Data Workers + Claude Code, Cursor and Codex: The MCP Server for Data Engineering in Your Coding Agent

Wire Data Workers into Claude Code, Cursor and Codex over MCP: real config for .mcp.json, .cursor/mcp.json and config.toml, plus approvals and receipts.

Claude Code, Cursor and Codex are where your engineers work. Data Workers is the data team those coding agents call over MCP: the agent asks, Data Workers answers from one governed context graph, and every change to production data goes through an approval and leaves a receipt. Your engineers keep the tool they chose. The data work behind it becomes safe enough to hand to an agent.

Your team probably already looks like this. Analytics engineers write dbt models in Cursor. Platform engineers live in Claude Code in the terminal, editing Airflow DAGs. Someone runs Codex in the ChatGPT desktop app to refactor SQL. Each agent is excellent at the repository in front of it. The facts a data change depends on live outside that repository: which fct_ table finance treats as canonical, who owns stg_app_events, and which Looker dashboards break if a column changes type.

Data Workers is the agentic data platform that runs the whole data lifecycle, and it plugs into all three agents the same way, over MCP. Below: what it reads and sends back, a full incident run from Cursor, and the config for each tool.

Key takeaways

  • •Data Workers is the MCP layer for data engineering. Every agent in the Data-Agents Swarm is a standard stdio MCP server: catalog and lineage, incidents, quality, schema, pipelines, cost and governance.
  • •Answers come from governed context. Data Context Wizard holds lineage, owners, freshness and usage across Snowflake, dbt, Airflow, Kafka and BI, so the agent stops guessing from the repo.
  • •Every data change goes through approvals and receipts. Your coding agent's prompts and allowlists still gate each tool call; Data Workers adds approvals by domain and a receipt with the diff, checks and rollback path.
  • •Start read-only. Wire the server, let the agent answer questions for a week, then turn on proposals.

What Claude Code, Cursor and Codex do, and why engineers keep them

These three coding agents read the repository, run commands, edit files, open pull requests and hold a conversation about all of it. Engineers keep them because they're fast and meet people where they work: the terminal, the IDE and the desktop app. All three are MCP clients with real admin controls. We checked each vendor's MCP docs on October 2, 2026.

ToolWhere MCP servers are configuredApprovals and admin controls (October 2026)
Claude Codeclaude mcp add, with local, project (.mcp.json in the repo) and user scopesAsks before using project servers from .mcp.json; permission rules evaluated deny, then ask, then allow, written as mcp__<server>__<tool>; managed allowedMcpServers and deniedMcpServers
Cursor.cursor/mcp.json in the project or ~/.cursor/mcp.json globally; stdio, SSE and Streamable HTTPAsks for approval before using MCP tools by default; run modes and allowlisted tools; Enterprise MCP Allowlist; team MCP distribution to Cloud Agents
Codexcodex mcp add, or [mcp_servers.<name>] in ~/.codex/config.toml or a trusted project's .codex/config.tomlOne config shared by the ChatGPT desktop app, Codex CLI and IDE extension; enabled_tools and disabled_tools; default_tools_approval_mode per server and approval_mode per tool

The coding agent decides which servers load and asks a person before a tool runs. Data Workers builds on exactly that.

What Data Workers reads, and what it sends back

The coding agent sends Data Workers a question and the work in hand. Data Workers sends back context, a diagnosis or a proposal, and keeps the record.

From the coding agent into Data WorkersFrom Data Workers back into the coding agent
The question: "why did weekly active accounts drop?"Governed context: lineage, owners, freshness and usage for every asset involved
The file in hand: a dbt model, an Airflow DAG, a SQL queryA diagnosis with the root cause and the column-level blast radius
Who is asking, so autonomy and approvals apply per domainA proposed change: a diff for the owner to merge, or a reversible action at the level you set
The answer to the agent's own approval promptAn approval request routed to the named approver for that domain
The tools you enabled for that serverA receipt: the diff, the approver, the checks run, before and after values, the rollback path
What Data Workers reads from Claude Code, Cursor and Codex and what it writes back through Claude Code, Cursor and Codex

A handful of agents do most of the work behind these tools. The Data Context & Catalog agent answers lineage and ownership questions from Data Context Wizard, with tools such as trace_cross_platform_lineage, trace_column_lineage and blast_radius_analysis. The Incident Debugging agent runs diagnose_incident and get_root_cause across systems. The quality agent runs run_anomaly_sweep and run_quality_check. The Data Change Review agent posts assess_pr_impact on any pull request, whether a person or a coding agent wrote it. The Autonomous Data-Conductor sequences them and holds each change at the autonomy level you chose.

With only the repository in view, any agent has to infer lineage from code and query the warehouse to fill the gaps. With Data Workers attached, it asks the system that already knows the lineage, the owners and the last good value, and gets an answer with sources and timestamps.

One incident, from a question in Cursor

Here is a scenario many data teams will recognize. It's an illustration, not a customer case.

At 23:05 a mobile release starts sending account_id on the Kafka app_events topic as a UUID string instead of an integer. At 23:20 the events land in Snowflake raw.app_events. The dbt model stg_app_events uses try_cast on that column, so new IDs quietly become nulls. At 02:00 the Airflow DAG run builds dbt and every test passes, because nothing tests that column for nulls. Weekly active accounts in Looker drop by 19%.

StepWhere it runsWhat happensWho decides
1. DetectSnowflakeAt 02:12 the quality agent's anomaly sweep flags the drop in weekly active accounts and opens an incident.Data Workers, read-only
2. Diagnosedbt, Kafka, SnowflakeAt 02:20 Data Workers follows lineage from the Looker tile to fct_active_accounts, stg_app_events and the try_cast, matches the first null to the 4.12 release on app_events, and lists 2 models and 3 dashboards in the blast radius.Data Workers, read-only
3. AskCursorAt 08:45 an analytics engineer asks Cursor why the number dropped. Cursor calls Data Workers over MCP, and the diagnosis, lineage and owners appear in the chat with their sources.The engineer
4. ProposeGitHubThis team has turned on the GitHub pull-request target, so at 08:52 Data Workers opens the dbt fix as a pull request: map UUIDs through the account lookup table and add a not_null test. The PR carries the blast radius and the rollback path.Data Workers proposes
5. ReviewGitHub, dbt CIYour dbt CI runs on the PR exactly as it does for a person's change. The engineer reads the diff in Cursor.dbt CI
6. ApproveGitHub or SpellbookAt 09:05 the on-call analytics engineer approves and merges. Branch protection still applies.A named engineer
7. RerunAirflowAt 09:10 Data Workers triggers the night's build again through the orchestrator, within the approval from step 6.Approved in step 6
8. VerifySnowflake, LookerAt 09:40 Data Workers compares weekly active accounts with distinct accounts in raw.app_events, confirms the tile is right, and records the receipt.Data Workers, read-only
Incident timeline across the stack: what Claude Code, Cursor and Codex, your team and Data Workers each do, step by step

Cursor held the conversation, showed the diff and asked the engineer before each tool call. Data Workers supplied what sits outside any one repository: the overnight detection, the diagnosis across five systems, the pull request with its blast radius, the rerun and the check against the source. The same run works from Claude Code in a terminal or from Codex in the desktop app, because all three call the same Data Workers servers.

Why doesn't a coding agent just do this itself?

Claude Code, Cursor and Codex are built to be excellent general-purpose coding agents, and that focus is the right design. Their approval model belongs to the session: a person at the keyboard says yes to a tool call, and a policy file says which tools may run. That is exactly what you want from a coding agent.

Changing production data across systems is a different product category. It needs context about every system in the estate, including the ones that never appear in the repo. It needs blast-radius scoping before a change, approvals set by data domain and owner rather than by session, rollback for each step, receipts an auditor can read next quarter, and someone accountable for changes in Snowflake, Airflow and Kafka, which the coding agent vendor doesn't run. A coding agent that grew all of that would become a data platform with a code editor attached. Anthropic, Cursor and OpenAI made a sensible choice to stay general and let MCP servers bring the domain. Data Workers is that domain layer for data.

Why one data layer matters across three agents. Most data teams don't pick one coding agent. When each engineer wires their own Snowflake and dbt servers, you get several sets of credentials, several tool surfaces and no shared record of what changed. Data Workers gives every agent the same context, the same approvals and the same audit trail. If you're still choosing between the agents, our resource pages on Claude Code vs Cursor for data engineering and Claude Code vs OpenAI Codex cover that choice; this page is about the wiring, and it works the same whichever you pick.

Setup: wiring Data Workers into Claude Code, Cursor and Codex

Every Data Workers agent is a standard MCP server over stdio, so one install serves all three tools; only the config file differs. Start from the open-source repository. It runs every agent on a built-in sample estate, so engineers can drive the tools from their coding agent before anything touches production. In a Data Workers deployment, the same servers point at your Snowflake, dbt, Airflow and Looker and carry the autonomy levels, approvals and receipts described below.

Example: install the agents (Node.js 20+)

git clone https://github.com/DataWorkersProject/dataworkers-claw-community.git
cd dataworkers-claw-community
npm install --ignore-optional
export DW_HOME="$(pwd)"   # add this line to your shell profile

start-agent.sh <agent> in the repository root launches one agent as an MCP server. The examples wire three: dw-context-catalog (lineage, blast radius), dw-incidents (diagnosis, root cause) and dw-quality (checks, anomalies). Add dw-schema, dw-pipelines or dw-governance the same way.

Claude Code. Add the servers at user scope so they load in every project, or commit a project .mcp.json so the whole team gets them. Claude Code asks each person to approve project servers from .mcp.json the first time.

Example: Claude Code, user scope (run from the clone)

claude mcp add --scope user dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
claude mcp add --scope user dw-incidents -- "$(pwd)/start-agent.sh" dw-incidents
claude mcp add --scope user dw-quality -- "$(pwd)/start-agent.sh" dw-quality

Example: .mcp.json in your dbt repository

{
  "mcpServers": {
    "dw-catalog": {
      "command": "${DW_HOME}/start-agent.sh",
      "args": ["dw-context-catalog"]
    },
    "dw-incidents": {
      "command": "${DW_HOME}/start-agent.sh",
      "args": ["dw-incidents"]
    }
  }
}

Claude Code expands ${DW_HOME}, so each engineer's clone location stays out of the repository. Then scope what runs without a prompt. Read tools can be allowed; tools that change anything stay on ask, and anything you don't list prompts by default.

Example: .claude/settings.json

{
  "permissions": {
    "allow": [
      "mcp__dw-catalog__trace_cross_platform_lineage",
      "mcp__dw-catalog__blast_radius_analysis",
      "mcp__dw-incidents__diagnose_incident",
      "mcp__dw-incidents__get_root_cause"
    ],
    "ask": ["mcp__dw-incidents__remediate"]
  }
}

Run /mcp to confirm the servers are connected and see their tool counts.

Cursor. Put the servers in .cursor/mcp.json and commit it, or in ~/.cursor/mcp.json for every project. Cursor asks before using MCP tools by default, and you can expand each call to review its arguments.

Example: .cursor/mcp.json

{
  "mcpServers": {
    "dw-catalog": {
      "type": "stdio",
      "command": "${env:DW_HOME}/start-agent.sh",
      "args": ["dw-context-catalog"]
    },
    "dw-incidents": {
      "type": "stdio",
      "command": "${env:DW_HOME}/start-agent.sh",
      "args": ["dw-incidents"]
    }
  }
}

Restart Cursor after editing the file. On a Cursor Enterprise plan, admins add the start-agent.sh command pattern to the MCP Allowlist.

Codex. One command per server writes it to ~/.codex/config.toml, and that config is shared by the Codex CLI, the IDE extension and the ChatGPT desktop app.

Example: Codex (run from the clone)

codex mcp add dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
codex mcp add dw-incidents -- "$(pwd)/start-agent.sh" dw-incidents

Example: ~/.codex/config.toml

[mcp_servers.dw-catalog]
command = "/path/to/dataworkers-claw-community/start-agent.sh"
args = ["dw-context-catalog"]
startup_timeout_sec = 30
enabled_tools = ["search_across_platforms", "trace_cross_platform_lineage", "blast_radius_analysis", "assess_impact"]
default_tools_approval_mode = "prompt"

[mcp_servers.dw-catalog.tools.trace_cross_platform_lineage]
approval_mode = "auto"

enabled_tools keeps the surface small, "prompt" makes Codex ask before each tool runs, and the per-tool override lets a pure read like trace_cross_platform_lineage run without a prompt. Run /mcp in Codex to check the servers are listed.

One more line for all three. Add a short section to the repository's agent instructions (CLAUDE.md, Cursor rules or AGENTS.md) telling the agent to ask Data Workers before it queries the warehouse directly or edits a model it doesn't own. That one habit makes a coding agent careful with data.

Guardrails: approvals, and what each coding agent owns

What your coding agent stays responsible for. Which MCP servers load and who may add them; the prompt on every tool call, permission rules, approval modes and admin allowlists; the editor, the terminal and the engineer's code changes. Data Workers never routes around any of them.

What Data Workers enforces on top.

  • •Read-only start. New deployments observe and answer. You extend autonomy one domain at a time as the receipts earn trust.
  • •Autonomy per domain, L0 to L4. Lineage and documentation can move faster while revenue models stay at "propose".
  • •Approvals where they belong. Model and DAG changes are diffs your reviewers merge, so branch protection and your reviewers decide. Anything irreversible needs a named human.
  • •No self-approval. An agent can't approve its own change, whichever coding agent asked for it.
  • •Receipts and rollback. Every change records the diff, the approver, the blast radius, the checks run, the before and after values and the rollback path in a tamper-evident, hash-chained audit log. A merged fix reverts like any commit; a reversible action carries its undo.
  • •Least privilege. Data Workers acts with the grants you give it, through Snowflake's, dbt's, GitHub's and Airflow's own permission systems. The coding agent never needs warehouse write credentials of its own.
The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous

For the full safety argument, read is it safe to let AI agents change production data?. If your team is weighing wiring a warehouse MCP server into each agent by hand instead, build it ourselves with Claude Code and MCP servers? walks through what that takes.

How it fits together

How Data Workers fits with Claude Code, Cursor and Codex: your coding agent on top, Data Workers in the middle, your estate underneath

Your engineers work in Claude Code, Cursor or Codex, and review, roll back and audit in Spellbook Data Catalog, which is in preview. The Data-Agents Swarm does the work, Context Wizard keeps the one governed context graph, and the Conductor runs each fix end to end. Snowflake, dbt, Airflow, Kafka and Looker stay where they are. Nothing migrates, and Data Workers stores metadata and scrubbed facts, not copies of your tables.

What changes for your team

Analytics engineers stop spending the first hour of an incident rebuilding lineage by hand; the agent they already have open answers with owners, sources and a proposed fix. Platform engineers stop handing out warehouse credentials to every coding agent on every laptop; Data Workers carries the access. The on-call rotation moves from "find it" to "approve it". Data leaders get one record of every change, whichever agent asked for it, which makes audits a matter of reading receipts. For what this looks like at the leadership level, read the data leader's guide to data engineering agents; for the step-by-step climb, see the autonomous data platform playbook. Each tool also gets its own build-on page: you're on Claude, you're on Cursor and you're on Codex.

The case for your CFO

The outcome. The coding agent seats you already pay for start doing real data work. Engineers answer "why is this number wrong?" from governed context, and fixes reach production through one reviewed path, so finance gets the right number sooner and with a record of who changed what.

The risk story. At the observe level, the agents read and answer; they change nothing. At the propose level, they propose diffs and a named engineer merges. Acting levels are reversible and only for domains you choose. Two approval layers apply: the coding agent asks before each tool call, and Data Workers asks the domain's approver before any change. No agent approves its own work, and every change has a rollback path. Each receipt holds the diff, the approver, the blast radius, the checks run and the before and after values. There is no migration.

Why now. Your engineers already have coding agents, and those agents already reach for data through whatever MCP servers people install. One governed data layer now is cheaper than untangling a dozen private servers later.

The first win. Read-only answers in the coding agent for one domain, such as revenue models. Engineers stop guessing at lineage within the first week.

What stays the same. Claude Code, Cursor and Codex, your repositories, your CI, your reviewers and your warehouse permissions.

The pilot path. Start with a pilot on one domain: read-only first, then proposals as diffs for the owner to merge. The pilot is credited in full against the first year.

One sentence for upstairs: "Our engineers keep their coding agents; Data Workers gives those agents our governed data context and routes every data change through one approval flow with a receipt."

When the coding agent on its own is enough

If one engineer owns a small project, the data lives in a single database they understand end to end, and nothing they change feeds a number finance reports, a coding agent with a read-only database server can carry you. Once changes cross Kafka, the warehouse, dbt, the orchestrator and BI, or several engineers and agents touch the same models, one shared context and one approval flow start paying for themselves.

FAQ

Which MCP server should a data team add to Claude Code, Cursor or Codex? One that brings governed data context and an approval path for changes. Data Workers agents are standard stdio MCP servers that install the same way in all three, covering catalog, lineage, quality, incidents, schema, pipelines, cost and governance.

Does Data Workers replace our coding agent? No. Claude Code, Cursor and Codex stay the place engineers work. Data Workers is the data layer they call over MCP.

Will Data Workers bypass our approval prompts or allowlists? No. Every Data Workers tool call goes through the coding agent's own prompts, permission rules, approval modes and allowlists.

Can different engineers use different coding agents? Yes. Claude Code, Cursor and Codex all call the same Data Workers servers, so context, approvals and the audit trail are shared whichever agent asked.

How do we keep the tool list small? Add only the agents a team needs, then use Codex enabled_tools, Claude Code permission rules or Cursor's allowlisted tools, starting with read tools.

Does the coding agent need warehouse write credentials? No. Data Workers holds the grants you give it and acts through each system's own permissions after approval. Engineers' agents keep read context only.

Where do approvals and receipts show up? In the coding agent's chat as the run happens, on the pull request, and in Spellbook Data Catalog, where reviewers can approve, roll back and audit.

Sources

Sources for coding agent capabilities, checked October 2, 2026: Claude Code MCP docs (claude mcp add, scopes, .mcp.json, ${VAR} expansion, project server approval, /mcp), managed MCP (allowedMcpServers, deniedMcpServers) and permissions (deny, ask, allow order and mcp__ rules); Cursor MCP docs (config locations, type, ${env:NAME} interpolation, transports, approval by default, MCP Allowlist, team distribution); Codex MCP docs (codex mcp add, config.toml keys including startup_timeout_sec, enabled_tools, default_tools_approval_mode and per-tool approval_mode, trusted project config, shared config across the desktop app, CLI and IDE extension). Data Workers install steps and start-agent.sh come from the Data Workers open-source repository and the client setup docs, checked October 2, 2026.