Data Workers for ML engineers
What Data Workers changes in an ML engineer's week: upstream data incidents traced to the features and models they reach, the cause behind a drifting feature, a fix proposed before the retrain reads bad data, while you keep the models and every approval.
Data Workers takes the upstream data work out of an ML engineer's week: when a source table breaks, the incident arrives traced to the features and models it reaches, with the training window it touched and a fix proposed before the next retrain reads the bad rows. You keep owning the models, the training runs and every approval; each agent action runs at the autonomy level your team sets, waits for a named approver where that level asks for one, and leaves a receipt.
Data Workers is the agentic data platform. It works from your coding agent, your warehouse and Spellbook, its catalog and approvals app.
Key takeaways
- •Upstream breaks, traced to your models.
get_root_causeandblast_radius_analysisfollow a broken table through your feature tables to the models registered on them. - •Drift with a cause attached. When your monitoring tool flags a drifting feature, Data Workers traces the feature tables upstream to the data change that explains it; your monitoring tool keeps measuring the drift.
- •A fix before the retrain. The fix arrives as a diff and a backfill plan for your approval, so you decide whether the retrain waits.
- •Schema changes flagged before they spread. The dbt manifest diff, change review on pull requests,
run_quality_checkandassess_impactcatch a renamed or retyped column before it becomes a quietly null feature. - •Data access through approvals.
provision_accessturns "I need this table for an experiment" into a scoped, policy-checked request for the data owner. - •You keep the job. Models, features, experiments and autonomy levels stay yours, set per domain from L0 manual to L4 autonomous.
Your week today
The planned work is features, evaluation and deployment. The unplanned work is about data you do not own:
- •Training data breaks. A source team changes an event payload; the feature table builds green with a column mostly null, and the model learns from it.
- •Freshness, skew and schema changes. The online feature lags the offline one, or an upstream rename turns into silent nulls three hops away.
- •Drift investigations. A metric sags and the first day goes to asking whether the world changed or the pipeline did.
- •Retraining pipelines. Scheduled retrains in Vertex AI, SageMaker or Databricks read whatever is in the table at 10:00.
- •Labels, access and RAG data. Labels arrive skewed, experiment data waits on a grant, and documents get embedded with duplicates and personal data in them.
The trust gap is in the data. In Stack Overflow's 2025 Developer Survey (49,009 developers overall, July 2025), the top frustration with AI tools, at 66%, is "AI solutions that are almost right, but not quite." In ML, almost-right training data produces almost-right models, and no one sees it until production.
The same week with Data Workers
Data Workers runs the data operations around your models. It reads your warehouse, orchestrator, dbt project and streams natively, keeps lineage, quality and ownership in one governed context graph, and hands you proposals with evidence.

What it takes off your week:
- •Upstream incidents, traced to features and models.
run_quality_checkand the lateness baselines your team records withmonitor_metricswatch the tables your features read against SLAs you set withset_sla;get_root_causetraces a break to the change behind it. With your models registered in the context graph,blast_radius_analysislists the feature tables, models, retrains and scoring outputs downstream, with owners. - •Drift investigations with lineage. When Arize, SageMaker Model Monitor, W&B or your own monitor flags a feature, Data Workers answers whether the pipeline moved: the dbt manifest diff, orchestrator runs,
run_quality_checkresults andmonitor_metricsbaselines on the feature tables, traced through the context graph to the change behind them. Your monitoring tool measures the drift; Data Workers finds the upstream cause and fixes the data, so "why did this model move?" has its cause next to the alert. - •A fix before retraining reads it. The fix comes back as a diff for you to merge, plus a backfill plan for the affected partitions. After you approve, Data Workers queues the rebuild through Airflow, Dagster or Prefect (dbt Cloud jobs rerun on your scheduler) and re-checks the table against its quality assertions and the feature's
monitor_metricsbaseline. Whether the retrain waits is your call. - •Schema changes. A rename is caught in the dbt manifest diff, in change review or by a failing
run_quality_check;assess_impactsays what it breaks;generate_migrationwrites the change with rollback SQL for the owner to apply. - •Feature and training data checks.
run_quality_checkchecks null rates, key uniqueness and minimum row counts on the feature and training tables (Snowflake, Postgres, BigQuery with a service-account key), andmonitor_metricsbaselines watch the values that matter, such as a feature's null share or mean. The ML agent writes the point-in-time join SQL for a leakage-safe training set and can emit a Feast-compatible feature view definition, both for the owner to run. - •Label and RAG data quality. On the scores and vectors you pass in, the ML agent ranks unlabelled rows by model uncertainty and flags duplicate and outlier embeddings before an index is built. Pull request review flags new columns whose names look personal before they reach a training set or a retrieval corpus.
- •Model lineage and access. Register a model with its training tables and features in the context graph, and lineage and blast radius reach it; MLflow, Weights & Biases or SageMaker stay your system of record for runs and versions.
provision_accesschecks policy and records a column-level grant with its expiry (nothing is granted when the verdict is review); on Unity Catalog the approved grant is applied, on Snowflake and BigQuery it is proposed for the owner.
What you still own and decide:
- •The models. Design, features, training, evaluation and promotion. Retraining, serving and rollback stay in your ML platform, run by you.
- •The approvals. They go to a named person; an unanswered request expires and escalates, never auto-grants, and no agent can promote its own work (how approvals work).
- •The autonomy settings. Each domain sits on L0 manual, L1 observe, L2 propose, L3 act reversibly or L4 autonomous; Data Workers ships observe-only (autonomy levels, explained). The org-wide stop halts all autonomous dispatch.
- •The receipt. Every change records the diff, approver, time, blast radius and rollback path in a hash-chained audit log (
verify_global_hash_chainchecks it). It doubles as reproducibility evidence.
How Data Workers fits the tools you already use
Your coding agent. Data Workers agents are MCP servers, so they sit inside Claude Code, Cursor, Codex or Copilot. Ask "what feeds plan_tier and did anything upstream change this week?" and the answer comes from the context graph, with sources. Setup follows the client setup docs.
# Example: register the ML, quality and lineage agents in Claude Code
git clone https://github.com/DataWorkersProject/dataworkers-claw-community.git
cd dataworkers-claw-community && npm install --ignore-optional
claude mcp add dw-ml -- "$(pwd)/start-agent.sh" dw-ml
claude mcp add dw-quality -- "$(pwd)/start-agent.sh" dw-quality
claude mcp add dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalogThe core runs on a built-in sample estate; in a deployment the same agents point at your estate and carry the approvals and receipts. List the tools with /mcp; Data Workers with your coding agents covers each client.
Your ML platform. MLflow, Weights & Biases, SageMaker and Vertex AI keep training, tracking and serving, and your team's assistant can use their APIs or MCP servers side by side with Data Workers today (you're on MLflow). Data Workers' work rests on the tables and pipelines its own connectors read, plus the model lineage you register. Feast, Tecton and the Databricks feature store work the same way: their offline tables in Snowflake or BigQuery get Data Workers' checks natively, and in Databricks the owner runs the check while Data Workers follows the dbt manifest and orchestrator runs, among 50+ connectors (on Databricks, on Snowflake, integrations).
Pipelines, streams and context. Airflow, Dagster, Prefect, dbt, Kafka and Schema Registry connect natively, so the jobs that build your features and the streams that feed them are in the same graph. Feature definitions, label rules and "which table is authoritative" live in Data Context Wizard with named owners (bring your own context). Spellbook Data Catalog (in preview) holds the proposals, approvals and receipts.
The metrics you are judged on
| Metric | How Data Workers moves it | Where you see it |
|---|---|---|
| Model quality in production | Upstream breaks caught at the table, before a retrain or a scoring run reads them | Incidents that reached a model vs caught upstream |
| Drift incidents | The upstream change behind a flagged feature, with lineage attached | Drift alerts from your monitoring tool, each with a cause in the receipt |
| Time to deploy | Fewer days lost to data chases; access requests through approvals | Time from data request to approved grant, from the receipts |
| Reproducibility | Every data fix recorded with its diff, approver and time | The audit trail for the tables a training run read |
| Warehouse cost of features | Snowflake credits attributed to each feature model through dbt query tags; fixes drafted for the owner | Spend per feature model; the design target is 25 to 40% lower warehouse spend |
Take a baseline first: how to measure AI data agents and the ROI calculator.
A worked example: a renamed event field and a 10:00 churn retrain
This is an illustration, not a customer case. A mobile app publishes app_events to Kafka with Schema Registry; the events land in BigQuery; Airflow runs a dbt build of features.churn_user_daily at 06:00; a Vertex AI pipeline retrains the churn model at 10:00 and logs runs to MLflow; the CRM scores accounts from the model. The churn model is registered in the context graph with its training tables and features. The ML data domain runs at L2 propose.
| Time | System | What happened |
|---|---|---|
| 18:05 | Schema Registry | The mobile team registers a new app_events version: plan_tier becomes subscription_tier |
| 06:00 | Airflow, BigQuery | The nightly feature build completes green |
| 06:10 | Data Workers | run_quality_check: the plan_tier null rate breaches its threshold; get_root_cause points upstream to app_events, and the stream's owner confirms the 18:05 schema version |
| 06:12 | Data Workers | Blast radius: two features, the churn model, the 10:00 retrain and CRM scoring |
| 06:14 | Data Workers | From the registered lineage and the Airflow runs, names the training window the bad partitions fall in |
| 06:16 | Data Workers | Proposes a staging diff mapping subscription_tier to plan_tier and a backfill plan for the partitions since 18:20, and recommends holding the retrain |
| 06:17 | Slack | The approval request reaches the model owner with the diagnosis and the receipt |
| 08:40 | ML engineer | Reviews in Cursor, merges the diff, pauses the Vertex AI schedule, approves the backfill in Spellbook |
| 08:45 | Data Workers, Airflow | Queues the feature rebuild and backfill through Airflow |
| 09:30 | Data Workers | Re-checks the feature table: quality passes; the plan_tier null share is back inside its baseline |
| 09:35 | Jira SM | Opens a ticket for the mobile team, linking the receipt |
| 10:00 | Vertex AI, MLflow | The ML engineer resumes the schedule; the retrain runs on corrected features and MLflow logs the run |

Without Data Workers, the retrain reads a feature that is 60% null and the first clue is a retention campaign aimed at the wrong accounts a week later. With it, the ML engineer starts the day with the cause, the affected models and a fix to approve, and the retrain runs on clean data on schedule.
How to bring it to your team
- •Try it on the sample estate. Ask the questions you answer when a feature looks wrong.
- •Pick one model that has been bitten by upstream data.
- •Register its lineage. Connect the warehouse read-only and record the model's training tables and features in the graph.
- •Start at L1 observe. For two weeks the agents read, trace and flag, and the permission ladder logs what each action would have needed; you review that record and its receipts against what happened to the model.
- •Move the domain to L2 propose and share the record with your data engineering partners (how to get your team to trust AI agents).
For risk questions, send is it safe to let AI agents change production data and where your data goes: the agents run in your infrastructure, your data stays in your systems, and the hosted Conductor sees workflow metadata only. Build it ourselves and the ROI of agentic data operations cover the rest. Then start a pilot (what a pilot looks like; $7,500 one-time, credited in full against the first year, on /pricing/). Your partners have their own pages: data engineers and data platform engineers. For the wider picture: MLOps tools moving to AI agents and inside the MLOps and models agent.
FAQ
Will Data Workers retrain or deploy my models? No. Training, promotion and deployment stay in your ML platform and with you. Data Workers works on the data around the models: it traces upstream breaks, finds the data change behind a drifting feature, proposes fixes and records what changed.
We already have MLflow and Weights & Biases. Why add this? Keep them; they track experiments and models. Data Workers covers what sits upstream (tables, schemas, pipelines, owners), with approvals and receipts, and your team's assistant can use their APIs or MCP servers side by side with Data Workers today.
Can it stop a retrain from reading bad data? It gets the diagnosis, the affected models and a proposed fix to you before the scheduled retrain, and recommends a hold. Pausing or running the retrain stays your decision, in your ML platform.
Is it safe to let agents touch training data? At L2 propose every change reaches you as a diff first, backfills run through your orchestrator after approval, and every step lands in a hash-chained audit log. How Data Workers handles PII goes deeper.
Does it help with LLM and RAG pipelines? Yes, on the data side: privacy flags in pull request review before data reaches a corpus, duplicate and outlier detection across embeddings, and lineage and baselines on the tables an index is built from.
Sources
- •Stack Overflow, 2025 Developer Survey, AI section (third-party; 49,009 responses, July 2025; 66% "almost right, but not quite"): https://survey.stackoverflow.co/2025/ai (checked Oct 2, 2026)
- •Google Cloud, Vertex AI Pipelines, schedule a pipeline run (schedules can be paused and resumed; page last updated 2026-10-01): https://cloud.google.com/vertex-ai/docs/pipelines/schedule-pipeline-run (checked Oct 2, 2026)
- •Data Workers client setup documentation: https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)
- •Data Workers public repository (every tool named on this page is registered there): https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)
- •Data Workers product repository,
data-workers-agent-swarmmain @ 0c2491e3: agents/dw-ml, agents/dw-context-catalog, agents/dw-governance, core/enterprise (checked Oct 2, 2026) - •Data Workers product pages: https://dataworkers.io/product/autonomous-data-conductor/ , https://dataworkers.io/product/data-agents-swarm/ , https://dataworkers.io/product/spellbook-data-catalog/ (checked Oct 2, 2026)
- •Data Workers security and pricing: https://dataworkers.io/security/ , https://dataworkers.io/pricing/ (checked Oct 2, 2026)