Product
Product12 min readBy The Data Workers Team

You're on Prefect: Your Team Writes and Runs the Flows. Data Workers Does the Operations Work When One Goes Wrong

Prefect runs your Python flows, retries and work pools. Data Workers catches the flow run that completed with bad data, fixes the cause and queues the rerun after approval.

Your team writes pipelines as ordinary Python. A function becomes a flow with one decorator, each step a task with another, and a deployment puts the flow on a schedule. Work pools decide where runs execute: a Kubernetes cluster with a worker polling it, an ECS push pool, or Prefect's managed infrastructure. Retries and timeouts sit on the task decorator. Automations watch for a flow run entering Crashed or a deployment going quiet and send a notification, pause a work pool or run another deployment. Prefect Cloud adds AI log summaries that turn a failed run into one line. Prefect 3.8.7 shipped on September 26, 2026, and Prefect describes itself as "durable orchestration for data, ML, and agents". Prefect is where your team writes and runs Python workflows. Data Workers does the operations work when a flow goes wrong, behind approvals.

The hardest Prefect incidents end in Completed. A retry that succeeded on the fourth attempt looks healthy; the question is what the first three attempts left in the warehouse. Data Workers answers it: it watches the tables your flows write, traces a bad number back to the flow run and task attempts behind it, proposes the cleanup and code fix to the owner, and after approval queues the rebuild through Prefect.

Key takeaways

  • •Prefect keeps its job. Flows, deployments, work pools, retries and automations stay where your team built them. Data Workers works on what the runs produce.
  • •Green runs get checked too. Uniqueness, volume and null checks on the tables your flows write turn a completed run with duplicate or missing rows into an incident the same night.
  • •Connected natively over the Prefect REST API. Data Workers reads flow runs and task runs, including attempts per task, on Prefect Cloud or self-hosted, and queues a flow run from a named deployment after approval.
  • •Every fix goes through a named person. Cleanups and code changes go to the owner to approve and apply; reruns are queued only after approval, each with a receipt.
  • •Autonomy is set per domain. Start at L1 observe, move to L2 propose, and let a narrow class of reruns reach L3 act reversibly once the record supports it.

Prefect is where your flows live. Data Workers is the operations engineer when one goes wrong.

Prefect does what an orchestrator should: it schedules the run, provisions infrastructure through the work pool, retries what you told it to retry and records every state change. The operations work around a bad run is different. Someone has to notice the data is wrong while the run is green, connect it to yesterday's infrastructure change, size the damage, clean up the warehouse safely, fix the code, rebuild in the right order and prove the numbers. Here is one night with both in place. This is an illustration, not a customer case.

TimeSystemWhat happens
Thu 16:30Prefect CloudThe platform team moves the billing_events_load/prod deployment to a new Kubernetes work pool, k8s-prod-v2, with a lower CPU limit per pod
Fri 01:00Prefect Cloud + PostgreSQLThe scheduled flow run starts and extracts the day's events from the billing Postgres database
01:05Prefect Cloud + BigQueryThe load_to_bigquery task submits a BigQuery load job, then hits its 300-second timeout on the slower pod. The task fails and Prefect retries it (retries=3). The load job keeps running in BigQuery and commits
01:19Prefect Cloud + BigQueryThe fourth attempt finishes inside the timeout. The flow run records Completed. raw.billing_events now holds four copies of the same batch
01:42Prefect Cloud + dbtThe dbt_finance_build/prod deployment runs dbt against BigQuery and rebuilds fct_invoices on the duplicated events. Completed
02:10Data WorkersThe nightly volume check on raw.billing_events comes in at 3.4 times the 30-day norm, and the uniqueness check on event_id fails
02:16Data Workers + Prefect REST APIData Workers reads the flow run and its task runs: the flow completed, and load_to_bigquery ran four times. BigQuery job history shows four committed load jobs for one batch. Blast radius: fct_invoices, two finance models and the Sigma "Daily revenue" workbook, read over Sigma's API
02:18PagerDuty + SlackData Workers raises a PagerDuty incident on the data service with the diagnosis, and sends the approval request to the billing pipeline's owner in Slack
07:40Spellbook + BigQuery + GitHubThe owner reviews three proposals in Spellbook: a cleanup that keeps one row per event_id, with row counts before and after; a diff that makes the load idempotent (load to a staging table, then MERGE on event_id); and a queued rebuild of dbt_finance_build/prod. She approves, applies the cleanup in BigQuery, merges the diff and raises the task timeout on the deployment
07:52Prefect CloudData Workers queues a flow run from dbt_finance_build/prod through the Prefect API. A worker on the pool picks it up; Prefect records Completed at 08:09
08:14BigQuery + PostgreSQL + PagerDutyData Workers verifies: event_id is unique, daily event counts and invoice totals in BigQuery match the Postgres source. It resolves the PagerDuty incident with the receipt
09:00SigmaThe finance stand-up opens the revenue workbook on verified numbers
Incident timeline across the stack: what Prefect, your team and Data Workers each do, step by step

Every part of Prefect did its job. A task that exceeds its timeout is marked failed, the retries ran, and the fourth attempt succeeded. The duplicates came from the gap between a client-side timeout and a BigQuery load job that kept running server-side, and they surfaced only in the warehouse. Prefect's transactions docs say it plainly: rollback hooks must clean up external side effects, because "Prefect does not automatically undo those operations." Data Workers is the layer that looks at those side effects across systems.

JobWhat Prefect doesWhat Data Workers does
The runSchedules the deployment, provisions the pod through the work pool, enforces the timeout, retries the taskReads the flow run and task runs over the REST API, including attempts per task
The signalRecords states (Completed, Failed, Crashed, Late) and fires automations on themChecks the data itself: uniqueness, volume and nulls on the tables the flow writes, so a green run with bad rows still raises an incident
The diagnosisAI log summaries in Prefect Cloud explain a failed run in one lineJoins the task attempts, the warehouse job history and lineage into one cause, with the blast radius across models and workbooks
The fixRuns whatever code and configuration the team deploysProposes the cleanup and the idempotent-load diff to the owner, who approves and applies them
The rerunRuns the flow run it is given on the work poolQueues the rebuild from the named deployment after approval, then reads its status
The proofKeeps the run history and logsVerifies against the source and writes a receipt: what changed, who approved it, how it was checked, how to undo it

Why doesn't Prefect just do this itself?

Because Prefect is built to make your code run reliably, and its design follows from that job. Its unit is the flow run. It knows whether a task raised, how many times it ran and on what infrastructure. It does not know that raw.billing_events feeds fct_invoices, that a finance workbook reads the result, or what a correct day of billing events looks like. That boundary is sensible: Prefect runs Python that can touch anything, so it leaves the meaning of those side effects to the people who wrote the code.

Prefect's AI features follow the same scope. AI log summaries explain one run's logs. The Prefect MCP server, in beta, is "primarily designed for reading data and monitoring", with changes made through the prefect CLI or SDK, and its docs warn that an assistant with terminal access could run destructive CLI commands outside the server. Prefect Horizon governs which agents reach which MCP tools, with OAuth 2.1, RBAC and audit logging. That is careful, well-scoped work.

Cleaning up a warehouse table, making a load idempotent and deciding which downstream deployments to rebuild in what order is a different product. It needs a context graph of every table, model and workbook, blast-radius scoping, approvals that name a person, a recorded undo and receipts an auditor can read. That is the product Data Workers is. For how that holds up in production, read is it safe to let AI agents change production data.

Every tool owns a slice. Data Workers covers the whole lifecycle

Prefect owns writing and running Python workflows, and does it better than a general platform would. Around it sit a warehouse, dbt, a catalog, a quality tool, BI and a pager. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the Prefect deployments already there.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Prefect goes deep on its own area
StageData WorkersPrefectWhy we scored it this way
Catalog & Context92Not Prefect's job: it knows flows, deployments and runs, not tables, metrics or owners. Data Workers keeps one governed context graph of tables, models, lineage and owners.
Analytics & Insights81Not Prefect's job: dashboards show run health, not business numbers. Data Workers answers data questions from governed definitions with lineage behind every number.
Data Quality83Flows can call any test library, and transactions with keys help authors make writes idempotent. Data Workers runs uniqueness, volume and null checks on the tables flows write.
Observability & Incidents8.55States such as Crashed and Late, automations and Cloud AI log summaries explain a run. Data Workers diagnoses the data incident across systems, fixes the cause and verifies it.
Pipelines & Ingestion8.59Prefect's home stage: Python flows and tasks, retries, deployments, work pools and workers on Kubernetes, ECS or managed infrastructure. Data Workers queues approved reruns through its API.
Schema & Migration82Not Prefect's job: a flow writes whatever schema its code emits. Data Workers catches schema changes and assesses blast radius before downstream models break.
Governance & Access8.54Workspace roles, service accounts and Horizon's tool-level RBAC govern who can do what in Prefect and over MCP. Data Workers routes every data change to a named approver.
Security & Privacy83Prefect secures its own workspace and the MCP gateway; the data a flow writes is classified elsewhere. Data Workers leaves a receipt on every data change.
Cost / FinOps82Not Prefect's job: concurrency limits shape compute, while warehouse spend sits in the warehouse. Data Workers sums BigQuery spend from the Jobs API and traces Snowflake credits to the dbt model behind them.
MLOps & Models7.55Strong fit for ML pipelines and agent workflows written in Python. Data Workers keeps the data under those models healthy.

How Prefect and Data Workers work together

How Data Workers fits with Prefect: your coding agent on top, Data Workers in the middle, your estate underneath

Data Workers connects to Prefect natively over the Prefect REST API, with an API key and your workspace URL on Prefect Cloud, or your server URL on self-hosted Prefect. It reads flow runs (state, start and end, run time) and task runs (state and how many times each ran), and after approval it creates a flow run from a named deployment, which Prefect queues for a worker like any other run. That is the whole write surface in Prefect, by design: deployments, schedules, work pools, automations and roles stay with your team. The step-by-step wiring, and how it works next to Dagster, is in Data Workers + Dagster and Prefect.

Engineers in a coding agent can run Prefect's MCP server and the Data Workers agents side by side: Prefect's answers "what ran and what failed", Data Workers' answers "what did that run do to the data, and what's the fix". This follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry to your client.

# Example: Prefect's MCP server plus Data Workers agents in Claude Code
# Prefect MCP server (beta, read-only by design), using your active Prefect profile
claude mcp add prefect -- uvx --from prefect-mcp prefect-mcp-server

# Data Workers agents, from a clone of the open-source repo
claude mcp add --scope user dw-incidents -- "$(pwd)/start-agent.sh" dw-incidents
claude mcp add --scope user dw-quality -- "$(pwd)/start-agent.sh" dw-quality
claude mcp add --scope user dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
claude mcp add --scope user dw-connectors -- "$(pwd)/start-agent.sh" dw-connectors

List the tools with your client's own command (/mcp in Claude Code). In this incident: monitor_metrics and run_quality_check catch the volume jump and the uniqueness failure; diagnose_incident names the cause; blast_radius_analysis and trace_cross_platform_lineage (dw-context-catalog) map what the duplicates reached; trigger_prefect_flow (dw-connectors) queues the approved rebuild; remediate (dw-incidents) re-checks the quality assertions and escalates any failure to a person.

In production the agents run in your infrastructure and hold the warehouse credentials and model key. Your data stays in your systems; the hosted Conductor sees workflow metadata only. More in where does our data go.

One incident, L0 to L4, set per domain:

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. Finance flags the revenue jump at 10:00. An engineer sees Completed in Prefect and finds the duplicates in BigQuery an hour later.
  • •L1 observe. Data Workers flags the duplicates at 02:10 with the diagnosis and blast radius. Nothing changes; the morning starts from the answer.
  • •L2 propose. Data Workers proposes the cleanup, the diff and the rebuild. Nothing reaches production until the named owner approves; an unanswered request expires and escalates, never auto-grants.
  • •L3 act reversibly. For a class with a clean record, such as rebuilding a dbt deployment after an approved fix, Data Workers queues the run itself, verifies, and sends any failed check to a person. Warehouse cleanups stay with the owner.
  • •L4 autonomous. For a scoped domain, Data Workers catches the volume jump right after the load and holds the downstream rebuild until the owner's fix lands, so bad numbers never reach the workbook.

What changes for your team

Prefect gives your engineers a fast way to write and run pipelines. Data Workers takes the operations work after: watching what runs produce, and running the incident when one produces the wrong thing.

Six jobs that run on autopilot with Data Workers next to Prefect, with a concrete example of each
  • •On-call starts from a cause. The paged engineer finds the flow run, the task attempts, the warehouse damage, the proposed fix and the rebuild in one place.
  • •Infrastructure changes stop being silent risks. A work pool move or a new CPU limit shows up as a data effect, tied to the runs it touched.
  • •Retries get a second look. A task that needed four attempts is a signal, and Data Workers checks what the attempts wrote.
  • •Platform and data teams share one record. The team that owns the work pools and the team that owns the tables see the same incident and receipt in Spellbook Data Catalog (in preview).

Keep Prefect, or consolidate?

Keep Prefect if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

For most Prefect teams the answer is keep it. Your flows are Python your engineers wrote, and Prefect runs them well. What teams consolidate is the tooling around the runs: a separate observability tool, scripts that diff row counts after loads, and runbooks that say "rerun and check the dashboard". On the vendor question, Prefect's own words are clear: its July 13, 2026 announcement that it was acquiring Dagster Labs says "Dagster and Dagster+ are here to stay", and its customer FAQ says the Prefect product "continues to work the same way it does today." Data Workers works through each product's public API either way. If you run Dagster too, see you're on Dagster; for Airflow, you're on Airflow; for the dbt side of every fix here, you're on dbt.

If you are weighing building this yourself on Prefect's MCP server and a coding agent, read build it ourselves with Claude Code and MCP servers. Connecting servers is the easy part; the context graph, approvals, undo and receipts are the work.

The case for your CFO

The outcome. Revenue, billing and model numbers built by Prefect flows are right more often. When a flow writes bad data, it is caught the same night and corrected before the business opens, with a record of what went wrong and how it was checked.

The risk story. At L0 and L1, agents only read. At L2 they propose and a named person approves; an unanswered request expires and escalates, never auto-grants. At L3 they act on reversible changes inside the domains you open; L4 is a later choice per domain. Warehouse cleanups always go to the owner to apply. No agent can promote its own work, and an org-wide stop halts all autonomous dispatch. Every change carries a receipt: what changed, who approved it, the blast radius and how to undo it.

Why now. Coding agents now write flows for your engineers, and Prefect runs data, ML and agent workflows on one platform. More flows means more runs that complete with the wrong data. The orchestrator answers whether a run finished; someone has to own whether the output is right.

The first win. L1 on the flows that feed finance: every completed run gets its output checked, and a duplicate or short load becomes a diagnosed incident by morning.

What stays the same. Prefect, your flows, deployments and work pools, your dbt project, your warehouse and your on-call rota. For the numbers, see the ROI of agentic data operations.

The pilot path. Start with a pilot (pricing); the pilot is credited in full against the first year.

The sentence for upstairs: "Prefect runs our pipelines; Data Workers checks what they produce and fixes it with our approval when a run gets it wrong, so finance sees the right number."

Getting started

Start with a pilot. Pick the Prefect deployments that feed the numbers leaders read, give Data Workers a Prefect API key and read access to the tables those flows write, and run at L1 for a few weeks: every completed run gets its output checked, and every incident arrives with a cause and a blast radius. Then turn on L2 for one domain, and let reruns of a downstream deployment move to L3 once the receipts show the agents were right. The pilot path is on the pricing page, and the pilot is credited in full against the first year.

FAQ

Does Data Workers have a native Prefect connector? Yes. It connects over the Prefect REST API, to Prefect Cloud with an API key or to a self-hosted Prefect server. It reads flow runs and task runs, including how many times each task ran, and after approval creates a flow run from a named deployment, which Prefect queues for a worker.

Prefect already has retries, automations and AI log summaries. What does Data Workers add? Those cover the run: whether it finished, why it failed, who to notify. Data Workers covers the data the run produced: whether it is unique, complete and on time against its baseline, what it reached downstream, how to fix the cause, and the proof afterwards.

Will Data Workers change our deployments, work pools or automations? No. In Prefect it only reads runs and queues flow runs from deployments you name, after approval. Work pools, timeouts and automations stay with your team; Data Workers points to the setting involved, and a person changes it.

Does Data Workers delete the duplicate rows itself? No. It proposes the cleanup, with row counts before and after and the query to run, and the owner approves and applies it. Code fixes arrive as a diff for the owner to merge.

Does Prefect's acquisition of Dagster change anything? Not for this setup. Prefect's customer FAQ says the Prefect product continues to work the same way, and Data Workers connects natively to both Prefect and Dagster through their public APIs.

Where does our data go? The agents run in your infrastructure with your warehouse credentials and model key. Your data stays in your systems; the hosted Conductor sees workflow metadata only.

Sources

  • •Prefect, homepage: "Durable orchestration for data, ML, and agents"; products including Prefect Cloud, open source, FastMCP and Horizon, https://www.prefect.io/ (checked Oct 3, 2026)
  • •Prefect blog, the Dagster acquisition announcement by Jeremiah Lowin (July 13, 2026), https://www.prefect.io/blog/prefect-is-acquiring-dagster-labs (checked Oct 3, 2026)
  • •Prefect, customer FAQ on the Dagster acquisition, https://www.prefect.io/prefect-dagster-faq (checked Oct 3, 2026)
  • •Prefect releases: 3.8.7 (Sept 26, 2026), https://github.com/PrefectHQ/prefect/releases (checked Oct 3, 2026)
  • •Prefect docs, How to write and run a workflow (task timeout behavior) and Python SDK prefect.tasks ("If the task exceeds this runtime, it will be marked as failed"), https://docs.prefect.io/v3/how-to-guides/workflows/write-and-run, https://docs.prefect.io/v3/api-ref/python/prefect-tasks, and Retries, https://docs.prefect.io/v3/how-to-guides/workflows/retries (checked Oct 3, 2026)
  • •Prefect docs, Work pools, https://docs.prefect.io/v3/concepts/work-pools (checked Oct 3, 2026)
  • •Prefect docs, States, https://docs.prefect.io/v3/concepts/states (checked Oct 3, 2026)
  • •Prefect docs, Automations, https://docs.prefect.io/v3/concepts/automations (checked Oct 3, 2026)
  • •Prefect docs, Transactions, https://docs.prefect.io/v3/advanced/transactions (checked Oct 3, 2026)
  • •Prefect docs, AI log summaries (Prefect Cloud only), https://docs.prefect.io/v3/how-to-guides/ai/flow-run-log-summaries (checked Oct 3, 2026)
  • •Prefect docs, Use the Prefect MCP server (beta, read-focused), https://docs.prefect.io/v3/how-to-guides/ai/use-prefect-mcp-server, and releases (0.0.1-beta.13, Sept 1, 2026), https://github.com/PrefectHQ/prefect-mcp-server/releases (checked Oct 3, 2026)
  • •Prefect Horizon: Deploy, Gateway, Registry, Remix, https://www.prefect.io/horizon, and "AI agent representation comes to Horizon" (June 24, 2026), https://www.prefect.io/blog/ai-agent-representation-comes-to-horizon (checked Oct 3, 2026)
  • •Data Workers open-source repository (trigger_prefect_flow in dw-connectors; tool registrations in dw-incidents, dw-quality, dw-context-catalog), https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
  • •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)