Product
Product10 min readBy The Data Workers Team

You're on Dagster: Keep Every Asset Right, With Data Workers Doing the Repair Work Behind Approvals

Already on Dagster? It's where your team models data as assets and schedules them. Data Workers does the operations work when an asset goes wrong, behind approvals.

Your team thinks in assets. Every table, dbt model and feature set is a software-defined asset, with daily partitions for event data and Declarative Automation deciding what to materialize next. Blocking asset checks keep a bad partition from reaching the models downstream, and freshness policies say how current each asset must be. New projects start from Components and the dg CLI, Dagster+ Insights ties BigQuery and Snowflake usage to each asset, and the Dagster+ MCP server lets Claude Code or Cursor view runs, launch assets and read run logs.

Dagster is where your team models data as assets and schedules them. Data Workers does the operations work when an asset goes wrong, behind approvals. A failed blocking check is Dagster working as designed. What comes next is operations work: finding the cause outside the code location, planning the backfill, getting the owner's yes, rerunning what depends on it and proving the number before the standup.

There is also news in the room. On July 13, 2026, Prefect announced it is acquiring Dagster Labs. Dagster's site now carries the banner "Dagster is now part of Prefect" and says the combined company is expected to operate under the Prefect name from August 2026; Prefect's homepage FAQ says "Prefect acquired Dagster Labs in July 2026." Both say Dagster keeps its name and open source license, Dagster+ remains a supported commercial offering, and customers have no action to take. Nothing on this page depends on how that roadmap unfolds.

Key takeaways

  • •Dagster keeps its job. Assets, partitions, checks, automation and Dagster+ stay as they are.
  • •A failed check starts the work. Data Workers reads the failed run and its steps over Dagster's GraphQL API, through a native connector, and opens an incident with the cause traced past the code location.
  • •The repair lands where the cause lives. Data Workers writes the backfill plan and proposes the durable fix as a diff; a named owner approves before anything changes.
  • •Dagster runs every run. After approval, Data Workers queues runs of named jobs through Dagster's API, reads their status and verifies the numbers against the source.
  • •Neutral to the roadmap. Data Workers connects natively to Dagster and Prefect alike.
  • •Start with a pilot. One partitioned domain with blocking checks, read-only first, on the ladder from L0 manual to L4 autonomous.

Dagster is the asset graph. Data Workers is the repair loop behind every asset.

Dagster holds what should exist and when. Data Workers holds what happens when an asset goes wrong: one context across the source database, warehouse, dbt and the tools reading the asset, one approval flow and one audit trail, with Spellbook Data Catalog (in preview) as the place your team reviews it. The Dagster and Prefect integration guide covers the wiring; this page is about building on Dagster day to day.

A Monday night at a consumer subscription app on PostgreSQL, Dagster+ Hybrid, dbt and BigQuery, with Hex for the growth team's notebooks. Times are UTC. This is an illustration, not a customer case.

TimeSystemWhat happens
Mon 22:40PostgreSQLThe platform team starts a long index build on the app primary. The read replica that analytics jobs use falls behind
Tue 00:30Dagster+The schedule materializes partition 2026-09-28 of raw_sessions, which reads the day's sessions from the replica and loads them into BigQuery
00:41Dagster+The blocking asset check raw_sessions_complete fails: 2.31M rows against a trailing median of 2.84M, last session at 21:05. The run ends in failure, the dbt assets stg_sessions and fct_daily_active_users hold back for that partition, and an alert policy posts to #data-alerts in Slack
00:44Data WorkersIt reads the failed run and its step stats over Dagster's GraphQL API and opens an incident
00:52Data WorkersIt profiles app.sessions on the primary: 2.83M rows for the day, the last at 23:59. A read query on the replica shows it was replaying 21:05 when the asset ran, and has since caught up. The asset code and the check are fine; the source was behind
00:58Data WorkersIt lists the blast radius from lineage (stg_sessions, fct_daily_active_users, the agg_weekly_engagement rollup and the Hex "Growth standup" notebook the team recorded in the context graph) and proposes three things: a backfill of partition 2026-09-28 for raw_sessions and its two dbt assets, a run of growth_rollups_job after it, and a diff that makes raw_sessions check replica lag before it reads. The approval request reaches the asset's owner in Slack
06:50Dagster+fct_daily_active_users misses its freshness policy, as designed. The alert lands where the incident already sits, cause and plan attached
07:15SpellbookThe analytics engineer who owns the asset reviews the cause, the blast radius and the plan, approves, and launches the partition backfill in Dagster as the plan lays out
07:38Dagster+The backfill succeeds. raw_sessions_complete passes with 2.83M rows, the dbt partitions materialize and the freshness policy is met again
07:40Data WorkersIt queues growth_rollups_job with trigger_dagster_job through Dagster's API and reads the run status until it records success at 07:52
07:55BigQueryData Workers checks the partition's row count against the primary for the same day and the DAU total against the prior four Mondays, then writes the receipt
09:30HexThe growth standup opens on Monday's DAU: 412,300. The replica-lag diff is reviewed and merged by the owner later that morning
Incident timeline across the stack: what Dagster, your team and Data Workers each do, step by step

Dagster did its part exactly: the check refused a short partition, downstream assets held, and the freshness policy flagged the stale table on time. The cause sat in a PostgreSQL replica, a system Dagster reads from and doesn't run.

JobWhat Dagster doesWhat Data Workers does
The definitionDeclares raw_sessions, its daily partitions, its blocking check and the freshness policy downstreamReads the asset's lineage alongside the source table, dbt models and BI consumers
The signalFails the check, holds downstream assets, alerts Slack, flags freshnessPicks up the failed run and its steps and opens one incident for all of it
The causeShows which step failed and the run logsTraces it past the code location to the replica that was replaying 21:05
The planOffers backfills, reruns and automation for whoever decidesWrites the backfill plan, the follow-on run and the durable fix, with the blast radius, for a named owner
The runExecutes the backfill and the job, records every materializationQueues the approved run of a named job through Dagster's API and reads its status
The proofReruns the check during the backfillRe-checks counts against their baseline and keeps a receipt with cause, plan, approver and checks

Why doesn't Dagster just do this itself?

Because Dagster is built to run and record the graph, and its newest AI work follows that scope. Dagster+ AI, in early access preview, scans a deployment for failure patterns and creates Issues that summarize problems, identify root causes and "suggest solutions for you to review"; from an Issue, you can dispatch an agent to create a fix as a pull request. The Dagster+ MCP server, in beta, gives an agent the permissions of the person or service user behind it, including launching and deleting runs and editing alert policies. That is the right design for an orchestrator: fast for an engineer at the keyboard.

The replica in this incident isn't in the deployment. Neither is the BigQuery reconciliation, the Hex notebook or the decision about who approves a backfill on a revenue-adjacent table. Repairing it means reading systems Dagster doesn't own, scoping blast radius across them, getting a named person to approve, keeping an undo path and owning a receipt an auditor can read. Taking on liability for changes in your source database and warehouse is a different product. That product is Data Workers.

Every tool owns a slice. Data Workers covers the whole lifecycle

Dagster owns asset orchestration outright, with strong quality and cost views beside it. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the orchestrator you already run.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Dagster goes deep on its own area
StageData WorkersDagsterWhy we scored it this way
Catalog & Context95The Dagster+ catalog shows every asset with its owner, metadata and lineage across the graph. Data Workers joins that lineage to the warehouse, the source database and BI in one governed context.
Analytics & Insights84Dagster introduced Compass in September 2025 for asking questions about data in Slack. Data Workers answers from approved definitions across platforms and checks the number behind the answer.
Data Quality86Asset checks sit next to the data they guard, a blocking check stops bad data flowing downstream, and freshness policies state how current an asset must be. Data Workers also finds the cause of a failure and re-verifies after the fix.
Observability & Incidents8.56Alert policies route failures and freshness breaches, and Dagster+ AI (early access preview) opens Issues with likely root causes. Data Workers traces the cause outside the code location and runs the fix behind approval.
Pipelines & Ingestion8.59Dagster's home stage: software-defined assets, partitions and backfills, Declarative Automation, Components and the dg CLI. Data Workers queues approved runs through Dagster's API and leaves the scheduling to Dagster.
Schema & Migration83Assets carry table schema metadata and lineage, so a code change shows its reach. Data Workers detects schema changes in the source and generates migrations with rollback SQL for the owner.
Governance & Access8.53Dagster+ roles and the audit log govern who can change Dagster itself. Data Workers dry-runs a proposed grant on the data and routes it to a named approver.
Security & Privacy83SSO, SCIM and Hybrid deployment keep the data plane in your cloud. Data Workers flags new sensitive column names in pull request review and proposes masking for the owner to apply.
Cost / FinOps86Insights ties BigQuery and Snowflake usage to each asset next to its runs. Data Workers reads BigQuery spend from the Jobs API, attributes Snowflake credits to the dbt model behind them and drafts the fix for its owner.
MLOps & Models7.54Dagster runs training and feature pipelines as assets. Data Workers keeps the data under the models healthy.

How Dagster and Data Workers work together

People keep working where they work: Claude Code or Cursor to build assets, the Dagster UI to watch runs and launch backfills, Hex to read results. Spellbook is where the data team reviews each proposed repair with its cause, blast radius, approver and undo path; approval requests reach the approver in Slack or email. Underneath, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm does the work with 20+ specialist agents, and the Autonomous Data-Conductor runs each repair: detect, diagnose, fix, review, verify, remember.

How Data Workers fits with Dagster: your coding agent on top, Data Workers in the middle, your estate underneath

What comes in from Dagster. Data Workers connects to Dagster natively over its GraphQL API, against your OSS webserver or a Dagster+ deployment with an API token. It reads run status and the step stats of each run, so it knows which step failed and when. Context Wizard joins that to the PostgreSQL tables, BigQuery partitions and dbt manifest around the asset, through native connectors; Hex connects over its API today. Your agent can then call explain_table for definition, lineage and trust score, blast_radius_analysis for what a change touches, get_quality_score for the data itself and diagnose_incident for the incident.

What goes back to Dagster. One thing, deliberately: after a named person approves, Data Workers queues a run of a named job with trigger_dagster_job and reads its status until it finishes. Dagster schedules and executes it like any other run. Partition backfills stay in your team's hands: Data Workers writes the plan (which partitions, which assets, in what order) and the owner launches it. Asset and check code changes are proposed as a diff for the owner to merge, or opened as a pull request when your team turns on the GitHub pull-request target. Alert policies, automation conditions and RBAC stay with your team.

Setup over MCP today. Keep the Dagster+ MCP server for your engineers and add Data Workers beside it in the same client. For Data Workers, clone the open-source repository and add start-agent.sh entries, as the client setup docs show.

// Example: .mcp.json for Claude Code
{
  "mcpServers": {
    "dagster-plus": {
      "type": "http",
      "url": "https://mcp.agent.dagster.cloud/mcp"
    },
    "dw-context-catalog": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-context-catalog"]
    },
    "dw-incidents": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-incidents"]
    },
    "dw-quality": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-quality"]
    },
    "dw-connectors": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-connectors"]
    }
  }
}

The Dagster+ server authenticates with OAuth (type /mcp, select Authenticate) and runs with that engineer's Dagster+ permissions; /mcp also lists every tool. Then "why did raw_sessions fail its check last night?" gets the run logs from Dagster+ and the cause, blast radius and source comparison from Data Workers, in one answer.

One request, L0 to L4. The autonomy ladder is set per domain.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. Connected, not acting. On-call works the alert by hand.
  • •L1 observe. Every failed check arrives with cause, blast radius and plan; nothing changes, and the log records what each action would have needed.
  • •L2 propose. Data Workers drafts the backfill plan, the follow-on run and the diff; a named owner approves in Spellbook first.
  • •L3 act reversibly. For proven classes, such as a downstream rollup rerun after an approved repair, Data Workers queues the run, verifies and keeps the receipt.
  • •L4 autonomous. For a scoped, trusted class in one domain, it repairs and verifies on its own and posts the receipt.

An unanswered approval request expires and escalates; it never auto-grants. No agent can promote its own work, and an org-wide stop halts all autonomous dispatch. More in autonomy levels L0 to L4, is it safe to let AI agents change production data, how approvals work and who owns the agents. The agents run in your infrastructure and hold your warehouse credentials and model key; your data stays in your systems, and the hosted Conductor sees workflow metadata only (where does our data go).

What changes for your team

Six jobs that run on autopilot with Data Workers next to Dagster, with a concrete example of each

Much of a Dagster on-call week comes after the check fails: run logs at 1 a.m., asking the platform team what changed, choosing partitions to backfill, convincing the growth lead the number is right. With Data Workers on top, those jobs run on autopilot at the level you set.

  • •Incidents. A failed asset check arrives with its cause traced past the code location and a repair plan attached.
  • •Data quality. Check failures turn into fixes and approved runs, confirmed when the check passes again.
  • •Cloud spend. BigQuery spend is read from the Jobs API and Snowflake credits are tied to the dbt model behind each asset, with the fix drafted for its owner.
  • •Access. Grant requests get a dry run of effective privileges and sensitive columns before the owner approves.
  • •Audits. Every repair carries the run, the partition, the approver, the checks and the undo path.
  • •Migrations. Source schema changes come with migrations and rollback SQL for the owner to apply.

The 1 a.m. archaeology becomes a review of a plan that is already written.

Keep Dagster, or consolidate?

Keep Dagster if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

Most teams keep Dagster: their assets, checks and automation are written in it. What teams consolidate is the tooling around it: the backfill runbook for source outages, the script that compares warehouse counts to the app database, the second alerting console, the spreadsheet of who owns which asset. If the acquisition has you weighing your orchestrator, Data Workers connects natively to Dagster, Prefect, Airflow and dbt Cloud alike, so your repair loop, receipts and approvals stay put whichever way you go. Building this yourself on the Dagster+ MCP server? Read build it ourselves with Claude Code and MCP servers: the connection is easy; cross-system context, approvals and rollback are the work.

The case for your CFO

The outcome: breaks Dagster's checks catch get repaired at the cause the same night, so the numbers the business reads each morning are complete and right.

The risk story is plain. Data Workers reads Dagster with an API token whose role your team picks. Every repair shows its blast radius, goes to a named owner, runs through Dagster like any other run, is verified against the source and leaves a receipt: trigger, plan, before and after checks, approver and undo path. Autonomy is set per domain from L0 manual to L4 autonomous. Zero migration: Dagster, dbt, BigQuery and PostgreSQL stay where they are.

Why now: orchestrators are adding AI that diagnoses and launches runs, and the market is consolidating. A record of who approved what, held outside any one vendor, is cheaper to set up before agents change production on their own. The first win is one domain where every failed check arrives explained and planned. What stays the same: your assets, your checks, your Dagster+ contract and your review process. For the numbers, see the ROI of agentic data operations. Start with a pilot (pricing); the pilot is credited in full against the first year.

The sentence to repeat upstairs: "Dagster runs our data on schedule and stops bad data at the door; Data Workers repairs what got stopped, with an owner approving every fix and a receipt for each one."

Getting started

Start with a pilot. Pick one partitioned domain where blocking asset checks already run, such as product events or billing, connect Data Workers to Dagster with a Viewer-role token beside your Dagster+ MCP server, and let Data Workers explain every failed check before you turn on the first approved run. Plans are on the pricing page, and the pilot is credited in full against the first year.

FAQ

Is the Data Workers connection to Dagster native, or does it go through MCP? Native. The connector talks to Dagster's GraphQL API, on Dagster OSS or Dagster+, with an API token: it reads run status and step stats and, after approval, queues runs of named jobs. Keep the Dagster+ MCP server for your engineers, under their own permissions, beside it.

Does Data Workers launch backfills? Data Workers writes the backfill plan: which partitions, which assets and in what order, with the blast radius. Your owner approves and launches it in Dagster. Data Workers then queues the follow-on runs of named jobs and verifies the results against the source.

What does the Prefect acquisition mean for Data Workers on Dagster? Nothing in how they work together. Both companies say Dagster keeps its name and license, Dagster+ remains supported and customers have no action to take, and Data Workers connects natively to both.

Dagster+ AI already finds root causes. Why add Data Workers? Use it: Dagster+ AI (early access preview) looks across your deployment and can dispatch an agent to fix code as a pull request. Data Workers covers causes outside the deployment, such as a source database, a warehouse table or an upstream sync, and carries the repair through approval, rerun and verification.

Can Data Workers change our asset checks or alert policies? It proposes check and asset code changes as a diff for the owner to merge. It doesn't edit alert policies, automation conditions or Dagster+ roles; those stay with your team.

Sources

  • •Dagster, Dagster and Prefect ("Dagster is now part of Prefect"; Prefect name expected from August 2026; no action required; July 13, 2026), https://dagster.io/prefect (checked Oct 3, 2026)
  • •Dagster blog, Prefect is Acquiring Dagster, Nick Schrock (July 13, 2026), https://dagster.io/blog/prefect-is-acquiring-dagster (checked Oct 3, 2026)
  • •Prefect, Prefect acquires Dagster Labs, letter from Jeremiah Lowin (July 13, 2026), https://www.prefect.io/prefect-acquires-dagster (checked Oct 3, 2026)
  • •Prefect, homepage FAQ ("Prefect acquired Dagster Labs in July 2026"), https://www.prefect.io/ (checked Oct 3, 2026)
  • •Dagster docs, Dagster+ MCP server (Dagster+ only, beta; permissions table), https://dagster.io/docs/getting-started/ai-tools/dagster-mcp (checked Oct 3, 2026)
  • •Dagster docs, Proactive monitoring (Dagster+ AI, early access preview), https://dagster.io/docs/guides/labs/dagster-ai/proactive-monitoring (checked Oct 3, 2026)
  • •Dagster docs, Testing assets with asset checks (blocking checks), https://dagster.io/docs/guides/test/asset-checks (checked Oct 3, 2026)
  • •Dagster docs, Asset freshness policies (docs say "under active development"; pricing lists the new API as GA; no status label used here), https://dagster.io/docs/guides/observe/asset-freshness-policies (checked Oct 3, 2026)
  • •Dagster docs, Components and the dg CLI, https://dagster.io/docs/guides/build/components (checked Oct 3, 2026)
  • •Dagster docs, Insights, https://dagster.io/docs/guides/observe/insights (checked Oct 3, 2026)
  • •Dagster docs, User roles and permissions, https://dagster.io/docs/deployment/dagster-plus/authentication-and-access-control/rbac/user-roles-permissions (checked Oct 3, 2026)
  • •Dagster blog, Introducing Compass (September 10, 2025), https://dagster.io/blog/introducing-compass (checked Oct 3, 2026)
  • •Dagster, Pricing (feature list), https://dagster.io/pricing (checked Oct 3, 2026)
  • •Dagster releases (1.13.25, October 1, 2026), https://github.com/dagster-io/dagster/releases (checked Oct 3, 2026)
  • •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)
  • •Data Workers open-source repository, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)