Product
Product11 min readBy The Data Workers Team

You're on Coalesce: Your Team Designs the Pipeline There, Column by Column. Data Workers Runs Operations on It in Production

Coalesce is where your team builds Snowflake and Databricks transformations with column-level awareness. Data Workers catches the green Job that built wrong numbers and fixes it behind approvals.

Your data engineers build in Coalesce Transform. They pick a node type (Stage, Dimension, Fact, View), map columns in the grid, set joins in the Join tab and watch column lineage follow every field from source to mart. A new source column gets re-synced and propagated through every node that needs it. Marketplace packages bring node types for Snowflake, Databricks, BigQuery and Fabric. Where admins have turned it on, Copilot turns a plain-English request or a pasted query into nodes. With Coalesce 2.0, announced September 15, 2026, engineers also build locally with the coa CLI and their own coding agent, watch the result in Coalesce Desktop, and let the policy layer check every change before it ships. Jobs refresh deployed Environments on the Coalesce Scheduler or from Airflow. Data Workers does the operations work in production, behind approvals.

The costliest incidents on Coalesce are the ones where nothing failed: the Job is green, every test passed, and a number the business runs on is wrong because something changed outside the project. Data Workers watches what the Jobs build, traces a bad number to its cause, proposes the fix to each owner and, once they approve, verifies the rebuilt tables.

Key takeaways

  • •Coalesce keeps its job. Nodes, packages, Copilot, Jobs and the Scheduler stay as they are. Data Workers works on what the Jobs produce and the systems around them.
  • •Green Jobs get checked too. Volume and freshness checks on the tables your nodes build, plus baseline alerts on the business totals your team records, turn a passing Job with wrong numbers into a diagnosed incident the same night.
  • •Connected over Coalesce's API and MCP server today. Data Workers reads nodes, Jobs and run results; approved reruns are started by your Transform owner, under her own Coalesce permissions.
  • •Fixes land where each owner works. Node changes arrive as a diff to apply in Coalesce; every step leaves a receipt.
  • •Autonomy is set per domain. Start at L1 observe, move to L2 propose, and let a narrow class of work reach L3 act reversibly once the record supports it.

Coalesce is where your team designs the pipeline, column by column. Data Workers runs operations on it in production.

Coalesce builds the node, enforces the standard, tests the columns and refreshes the Environment on schedule. Operations means noticing a number is wrong while the Job is green, finding the change in another team's system that caused it, sizing the damage, agreeing a fix with each owner and proving the number. Here is one night at a distributor with both in place. This is an illustration, not a customer case.

TimeSystemWhat happens
Mon 14:00Oracle E-Business SuiteSupply chain opens a new inventory organization, DC7 (Dallas), and posts inter-org transfers that move 41,800 units across 1,260 SKUs out of DC2 (Memphis)
14:06Qlik ReplicateThe CDC task applies the changes to Snowflake. Rows for the new org arrive in RAW_EBS.MTL_ONHAND_QUANTITIES_DETAIL as ORGANIZATION_ID = 7120
Tue 02:00Coalesce TransformThe Scheduler runs the inventory_nightly Job on PROD. The Fact node FCT_INVENTORY_POSITION inner-joins on-hand rows to DIM_WAREHOUSE, which builds from a region map that planning ops maintains. Org 7120 is not in the map, so DC7's rows drop out. DC2's outbound transfers are in
02:19Coalesce TransformEvery test passes: INVENTORY_POSITION_KEY is unique and not null, and no column is empty. The Job succeeds and the Scheduler emails success
02:30Data WorkersNetwork on-hand units, a total the team records each night, fall 1.8% in a day against a baseline that rarely moves more than 0.3% a day. The replicated EBS on-hand total, recorded alongside it, rose 0.2%. The gap is 41,800 units, and the anomaly opens an incident
02:34Data Workers + SnowflakeFCT_INVENTORY_POSITION is fresh and its row count is in range, so the loss is in values, not a failed load. Lineage puts Monday's EBS changes and the Fact node's inner join to DIM_WAREHOUSE on the path. Blast radius: MART_DAYS_OF_SUPPLY, the ThoughtSpot "Network days of supply" Liveboard read at the 09:00 S&OP meeting, and the reorder candidates list, whose row count jumped from 31 to 195 overnight
02:38SlackData Workers sends the incident to the inventory analytics node owner and the planning ops owner of the region map, with a test query for on-hand units that match no warehouse. It tells the buyers' channel to hold expedite orders on the 164 newly flagged SKUs until the numbers are verified
07:40SpellbookThe node owner runs the test: all 41,800 missing units belong to org 7120, which has no row in the region map. The owners review the proposal: add 7120 → DC7, South to the map, and a node diff that switches the join to LEFT, buckets unmatched orgs as UNMAPPED and adds that test as a node test, so the next new org stops the Job instead of vanishing. Both approve
08:05Coalesce TransformPlanning ops adds the map row. The engineer applies the node change in the Coalesce App, commits and deploys to PROD. From her own Claude Code session she starts the rerun with Transform MCP's runs_start, scoped to FCT_INVENTORY_POSITION and the nodes downstream. The run completes at 08:21
08:24Data Workers + SnowflakeData Workers verifies: network on-hand is back in line with the replicated EBS total, the new UNMAPPED node test passed in the run, and the reorder list is back to 31 rows. It writes the receipt and posts the resolution in Slack
09:00ThoughtSpotS&OP reads network days of supply with DC7 in it. No expedite purchase orders go out on false shortages
Incident timeline across the stack: what Coalesce, your team and Data Workers each do, step by step

Every part of Coalesce did its job, and so did supply chain: opening a distribution center is the business working as planned. The break lived between a correct change in Oracle EBS and a reference table owned by a third team, and it showed up only as a total that moved when the source said it shouldn't. Data Workers watches that gap across systems and runs the incident until the number is right.

JobWhat Coalesce doesWhat Data Workers does
BuildingNode types, the mapping grid, packages, Copilot and local development with coa and Coalesce DesktopReads nodes and their columns into one context graph that also covers the ERP, replication and BI
StandardsPolicy checks in the build loop, column and node tests, built-in Quality tests with circuit breakingChecks volume and freshness on what each Job builds and watches the totals your team records, so a green Job with wrong numbers still raises an incident
RunningJobs on the Scheduler or an orchestrator, with retries and email notificationsReads run status and results, and warns downstream readers while an incident is open
DiagnosingRun results and column lineage inside the projectDiagnoses across systems: the Oracle EBS change, the Qlik Replicate load, the region map and the ThoughtSpot consumer
Rerunningruns_start and runs_retry on Transform MCP, under the caller's Coalesce permissionsProposes the scoped rerun with the fix; your Transform owner starts it, and Data Workers verifies the result
The recordCommits, deployments and run historyA receipt: cause, diff, approvers, run, verification and undo, in Spellbook and the audit trail

Why doesn't Coalesce just do this itself?

Because Coalesce is built to make transformation development fast, consistent and governed. Its units are the node, the Job and the Environment. Coalesce 2.0 puts governance into the build: testing and monitoring requirements, PII masking, ownership, documentation and modeling patterns are checked in the local CLI loop before code reaches production, whoever wrote it. That is the right place for standards.

Its AI follows the same scope on purpose. Copilot builds nodes in the workspace, for review before deploying. Transform MCP's 21 tools let an assistant read Environments, runs and nodes and start, retry or cancel runs as the signed-in user; deploying Workspaces and authoring node SQL stay in the CLI and the app. Coalesce's local development guide puts it plainly: use the CLI to build, use the MCP to observe.

The incident above sits outside that scope by design: the cause was in Oracle EBS, the data came through Qlik Replicate, the missing piece was planning ops' reference table, and the consumers sat in ThoughtSpot. Nothing in the Coalesce project was broken. Owning it means context about systems Coalesce does not own, routing each fix to its owner, scoping blast radius, keeping an undo and answering for changes across tools. That is a different product with a different liability: Data Workers. If you run Coalesce Quality, built on SYNQ, which Coalesce acquired in March 2026, its monitors and Scout cover incidents on the Coalesce side; you're on Coalesce Quality shows how Data Workers works alongside them. For how this holds up in production, read is it safe to let AI agents change production data.

Every tool owns a slice. Data Workers covers the whole lifecycle

Coalesce owns building transformations, and does it better than a general platform would. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the Coalesce project already there.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Coalesce goes deep on its own area
StageData WorkersCoalesceWhy we scored it this way
Catalog & Context96Column lineage across every node, AI-written descriptions, and Catalog on the same context layer since 2.0. Data Workers keeps one governed context graph that also covers the source systems, replication and BI.
Analytics & Insights83Not Transform's job: it builds the tables that BI reads. Data Workers answers data questions from governed definitions, with lineage behind every number.
Data Quality86Column and node tests, built-in Quality tests and circuit breaking stop a bad build. Data Workers adds volume checks on what each Job builds and baseline alerts on lateness and the totals your team records.
Observability & Incidents8.54Job notifications, run history and retries; incidents live in Coalesce Quality. Data Workers diagnoses the incident across systems, including Jobs that passed with wrong data.
Pipelines & Ingestion8.59Coalesce's home stage: node types, the mapping grid, column propagation, packages, Jobs and the Scheduler. Data Workers works through the runs your Transform owner approves.
Schema & Migration87Column propagation, Re-Sync Columns and Copilot-assisted SQL migration into nodes. Data Workers catches the change at the source first and plans migrations in approved waves.
Governance & Access8.55Workspace and Environment permissions, plus 2.0's policy checks in the build loop. Data Workers routes every production change to a named approver.
Security & Privacy83Coalesce protects your pipeline code; warehouse credentials stay in the Workspace connection. Data Workers' pull request review flags new columns whose names or annotations look sensitive, and leaves a receipt on every data change.
Cost / FinOps85Query tags on every Snowflake query name the Job, node and run; warehouse settings per node. Data Workers attributes Snowflake credits query by query and drafts the fix for the owner.
MLOps & Models7.52Cortex functions run inside pipelines; models are not watched. Data Workers keeps the data under your models fresh and correct.

How Coalesce and Data Workers work together

How Data Workers fits with Coalesce: your coding agent on top, Data Workers in the middle, your estate underneath

Data Workers connects to Coalesce over its API and MCP server today. It reads deployed nodes, Jobs, runs and run results into one context graph with the Snowflake tables the Jobs build, the sources underneath and the BI downstream. Snowflake is native, so null, uniqueness and row-count checks run straight against the warehouse, lateness is tracked with monitor_metrics baselines, and credits are attributed query by query; Oracle EBS, Qlik Replicate and ThoughtSpot connect over their APIs. Coalesce keeps the build: node edits, commits and deployments stay in the Coalesce App and coa, and Qlik Replicate tasks stay with their owner. An approved rerun is started by your Transform owner, so Coalesce's permissions decide who runs what. Data Workers then reads the run's status and verifies the result.

Engineers can run Transform MCP and the Data Workers agents side by side. Coalesce's server answers "what is deployed and what ran"; Data Workers answers "what did that run do to the data, and what's the fix". Transform MCP authenticates with your Coalesce refresh token. The Data Workers side follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry to your client.

# Example: Coalesce Transform MCP plus Data Workers agents in Claude Code
# Transform MCP (token from the Deploy tab or ~/.coa/config)
claude mcp add --transport http coalesce-transform "https://<your-app-host>/api/v1/mcp" --header "Authorization: Bearer <your-refresh-token>"

# Data Workers agents, from a clone of the open-source repo
claude mcp add --scope user dw-incidents -- "$(pwd)/start-agent.sh" dw-incidents
claude mcp add --scope user dw-quality -- "$(pwd)/start-agent.sh" dw-quality
claude mcp add --scope user dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
claude mcp add --scope user dw-connectors -- "$(pwd)/start-agent.sh" dw-connectors

List the tools with your client's own command (/mcp in Claude Code). In this incident: monitor_metrics (dw-incidents) flags the on-hand total; run_quality_check (dw-quality) reads row counts and nulls; diagnose_incident classifies it; trace_cross_platform_lineage and blast_radius_analysis (dw-context-catalog) map what it reached; send_slack_alert (dw-connectors) tells the buyers to hold; runs_start on Coalesce's server reruns the nodes under the engineer's permissions; remediate re-checks the quality assertions and escalates any failure to a person.

In production the agents run in your infrastructure and hold the warehouse credentials and model key. Your data stays in your systems; the hosted Conductor sees workflow metadata only. More in where does our data go.

One incident, L0 to L4, set per domain:

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. A planner notices at S&OP that network stock fell overnight, after buyers placed expedite orders.
  • •L1 observe. Data Workers flags the gap at 02:30 with the blast radius. Nothing changes; the morning starts from the answer.
  • •L2 propose. Data Workers proposes the map row, the node diff and the scoped rerun. Nothing reaches production until the named owners approve; an unanswered request expires and escalates, never auto-grants.
  • •L3 act reversibly. For a class with a clean record, such as telling the buyers' channel to hold orders on flagged SKUs while an incident is open, Data Workers acts under the standing approval your team set and sends anything outside it to a person.
  • •L4 autonomous. For a scoped domain, Data Workers flags every downstream reader on the first bad build, before the owners' fix lands.

What changes for your team

Coalesce gives your engineers a fast, governed way to build. Data Workers takes the operations work around it.

Six jobs that run on autopilot with Data Workers next to Coalesce, with a concrete example of each
  • •On-call starts from evidence. The blast radius, node diff and scoped rerun sit in one place.
  • •Other teams' changes stop being silent. A new ERP organization or an edited reference table shows up as a data effect, tied to the nodes it touched.
  • •Passing tests get a second look. Totals that move when the source says they shouldn't are watched every night.
  • •Engineering, planning and governance share one record. The node owner, the reference-data owner and the Liveboard's readers see the same incident and receipt in Spellbook Data Catalog (in preview).

Keep Coalesce, or consolidate?

Keep Coalesce if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

For most Coalesce teams the answer is keep it: nodes, packages and column propagation are where your engineers are fastest. What teams consolidate is the tooling around the Jobs: separate monitors watching the warehouse, alert scripts and runbooks that say "rerun the Job and check the dashboard". Data Workers covers those with one context and one approval flow. If your business users find data in Coalesce Catalog, see you're on Coalesce Catalog; if part of your estate builds in dbt, see you're on dbt. Native connectors are listed on the integrations page.

Weighing a build of your own on Transform MCP? Read build it ourselves with Claude Code and MCP servers: connecting servers is the easy part.

The case for your CFO

The outcome. The numbers your Coalesce pipelines build are right more often. When a change anywhere makes them wrong, it is caught the same night and corrected before the business acts, with a record of how it was checked. Above, that is the difference between a normal morning and a round of expedite orders for stock that already existed.

The risk story. At L0 and L1, agents only read. At L2 they propose and a named person approves; an unanswered request expires and escalates, never auto-grants. At L3 they act on reversible changes inside the domains you open; L4 is a later choice per domain. Reruns stay with your Transform owner. No agent can promote its own work, and an org-wide stop halts all autonomous dispatch. Every change carries a receipt: what changed, who approved it, the blast radius and how to undo it.

Why now. Coalesce 2.0 and coding agents mean more pipelines built faster, while other teams keep changing their systems. That means more green Jobs with wrong numbers. Someone has to own whether the output is right.

The first win. L1 on the Jobs behind planning and finance: a gap becomes a diagnosed incident by morning.

What stays the same. Coalesce, your nodes, packages and Jobs, Copilot and the policies you set, your warehouse and your on-call rota. For the numbers, see the ROI of agentic data operations.

The pilot path. Start with a pilot (pricing); the pilot is credited in full against the first year.

The sentence for upstairs: "Coalesce is where we build our pipelines to standard; Data Workers watches what they produce every night and fixes it with our approval when something upstream breaks it, so planning acts on the right number."

Getting started

Start with a pilot. Pick the Coalesce Jobs that feed the numbers leaders act on, give Data Workers read access to the Snowflake tables those Jobs build, record the totals that matter, and run at L1 for a few weeks. Then turn on L2 for one domain, and move a narrow class of actions to L3 once the receipts support it. The pilot path is on the pricing page, and the pilot is credited in full against the first year.

FAQ

How does Data Workers connect to Coalesce? Over Coalesce's API and MCP server today. It reads deployed nodes, Jobs, runs and run results, and checks the tables those Jobs build through its native Snowflake connection. Coalesce, Snowflake, the sources and the BI layer join one context graph.

Will Data Workers change our nodes, Jobs or deployments? No. Node changes arrive as a diff for the owner to apply in Coalesce, and commits and deployments stay with your team in the Coalesce App or coa.

Coalesce 2.0 already checks standards and runs tests. What does Data Workers add? 2.0's policies and tests make sure a pipeline is built right. Data Workers covers Jobs that pass every check and still produce wrong numbers because something changed outside the project, routes the fix to each owner and verifies the rebuilt tables.

Can Data Workers trigger Coalesce Jobs? Runs stay with the people Coalesce authorizes. Data Workers proposes the scoped rerun, your Transform owner starts it with Transform MCP or the app, and Data Workers verifies the result.

We also use Coalesce Catalog and Coalesce Quality. How does this fit? They keep their jobs. Data Workers works across all three and the systems outside Coalesce. See you're on Coalesce Catalog and you're on Coalesce Quality.

Where does our data go? The agents run in your infrastructure with your warehouse credentials and model key. Your data stays in your systems; the hosted Conductor sees workflow metadata only.

Sources

  • •Coalesce homepage (positioning; Transform, Catalog, Quality), https://coalesce.io/ (checked Oct 3, 2026)
  • •Coalesce, "Coalesce 2.0: The Agentic Data Engineering Platform" (Sep 15, 2026), https://coalesce.io/product-technology/coalesce-2-0-the-agentic-data-engineering-platform/ (checked Oct 3, 2026)
  • •Coalesce, "Coalesce Announces Acquisition of SYNQ and Launch of Coalesce Quality" (Mar 9, 2026), https://coalesce.io/company-news/coalesce-announces-acquisition-of-synq-and-launch-of-coalesce-quality/ (checked Oct 3, 2026)
  • •Coalesce Docs, About Transform MCP, https://docs.coalesce.io/docs/coalesce-ai/mcp/about-mcp.md (checked Oct 3, 2026)
  • •Coalesce Docs, Transform MCP Available Tools, https://docs.coalesce.io/docs/coalesce-ai/mcp/mcp-available-tools.md (checked Oct 3, 2026)
  • •Coalesce Docs, Integrate Claude with Transform MCP (token auth only), https://docs.coalesce.io/docs/coalesce-ai/mcp/integrate-mcp-claude.md (checked Oct 3, 2026)
  • •Coalesce Docs, Coalesce Copilot, https://docs.coalesce.io/docs/coalesce-ai/copilot.md, and Transform AI, https://docs.coalesce.io/docs/coalesce-ai.md (checked Oct 3, 2026)
  • •Coalesce Docs, Local AI Development, https://docs.coalesce.io/docs/coalesce-ai/local-development.md (checked Oct 3, 2026)
  • •Coalesce Docs, Column Lineage and Propagation, https://docs.coalesce.io/docs/build-your-pipeline/column-lineage.md, and How To Refresh Node Data, https://docs.coalesce.io/docs/build-your-pipeline/refresh-a-source-node.md (checked Oct 3, 2026)
  • •Coalesce Docs, Query Tags in Snowflake and Coalesce, https://docs.coalesce.io/docs/deploy-and-refresh/query-tag-snowflake.md (checked Oct 3, 2026)
  • •Coalesce Docs, Migrating SQL to Coalesce with Copilot, https://docs.coalesce.io/docs/guides/migrating-sql-to-coalesce-with-copilot.md (checked Oct 3, 2026)
  • •Coalesce Docs, Refreshing Your Pipeline Using the Coalesce Scheduler, https://docs.coalesce.io/docs/deploy-and-refresh/refresh/scheduling-jobs-in-coalesce.md, and Job Notifications and Monitoring, https://docs.coalesce.io/docs/deploy-and-refresh/monitoring-jobs.md (checked Oct 3, 2026)
  • •Coalesce Marketplace documentation index, https://docs.coalesce.io/llms-marketplace.txt (checked Oct 3, 2026)
  • •Data Workers open-source repository, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
  • •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)