Industry
Industry7 min readBy The Data Workers Team

Data Workers for ML engineers

What Data Workers changes in an ML engineer's week: upstream data incidents traced to the features and models they reach, the cause behind a drifting feature, a fix proposed before the retrain reads bad data, while you keep the models and every approval.

Data Workers takes the upstream data work out of an ML engineer's week: when a source table breaks, the incident arrives traced to the features and models it reaches, with the training window it touched and a fix proposed before the next retrain reads the bad rows. You keep owning the models, the training runs and every approval; each agent action runs at the autonomy level your team sets, waits for a named approver where that level asks for one, and leaves a receipt.

Data Workers is the agentic data platform. It works from your coding agent, your warehouse and Spellbook, its catalog and approvals app.

Key takeaways

  • •Upstream breaks, traced to your models. get_root_cause and blast_radius_analysis follow a broken table through your feature tables to the models registered on them.
  • •Drift with a cause attached. When your monitoring tool flags a drifting feature, Data Workers traces the feature tables upstream to the data change that explains it; your monitoring tool keeps measuring the drift.
  • •A fix before the retrain. The fix arrives as a diff and a backfill plan for your approval, so you decide whether the retrain waits.
  • •Schema changes flagged before they spread. The dbt manifest diff, change review on pull requests, run_quality_check and assess_impact catch a renamed or retyped column before it becomes a quietly null feature.
  • •Data access through approvals. provision_access turns "I need this table for an experiment" into a scoped, policy-checked request for the data owner.
  • •You keep the job. Models, features, experiments and autonomy levels stay yours, set per domain from L0 manual to L4 autonomous.

Your week today

The planned work is features, evaluation and deployment. The unplanned work is about data you do not own:

  • •Training data breaks. A source team changes an event payload; the feature table builds green with a column mostly null, and the model learns from it.
  • •Freshness, skew and schema changes. The online feature lags the offline one, or an upstream rename turns into silent nulls three hops away.
  • •Drift investigations. A metric sags and the first day goes to asking whether the world changed or the pipeline did.
  • •Retraining pipelines. Scheduled retrains in Vertex AI, SageMaker or Databricks read whatever is in the table at 10:00.
  • •Labels, access and RAG data. Labels arrive skewed, experiment data waits on a grant, and documents get embedded with duplicates and personal data in them.

The trust gap is in the data. In Stack Overflow's 2025 Developer Survey (49,009 developers overall, July 2025), the top frustration with AI tools, at 66%, is "AI solutions that are almost right, but not quite." In ML, almost-right training data produces almost-right models, and no one sees it until production.

The same week with Data Workers

Data Workers runs the data operations around your models. It reads your warehouse, orchestrator, dbt project and streams natively, keeps lineage, quality and ownership in one governed context graph, and hands you proposals with evidence.

Comparison matrix of Your ML stack and Data Workers on the outcomes a data leader buys

What it takes off your week:

  • •Upstream incidents, traced to features and models. run_quality_check and the lateness baselines your team records with monitor_metrics watch the tables your features read against SLAs you set with set_sla; get_root_cause traces a break to the change behind it. With your models registered in the context graph, blast_radius_analysis lists the feature tables, models, retrains and scoring outputs downstream, with owners.
  • •Drift investigations with lineage. When Arize, SageMaker Model Monitor, W&B or your own monitor flags a feature, Data Workers answers whether the pipeline moved: the dbt manifest diff, orchestrator runs, run_quality_check results and monitor_metrics baselines on the feature tables, traced through the context graph to the change behind them. Your monitoring tool measures the drift; Data Workers finds the upstream cause and fixes the data, so "why did this model move?" has its cause next to the alert.
  • •A fix before retraining reads it. The fix comes back as a diff for you to merge, plus a backfill plan for the affected partitions. After you approve, Data Workers queues the rebuild through Airflow, Dagster or Prefect (dbt Cloud jobs rerun on your scheduler) and re-checks the table against its quality assertions and the feature's monitor_metrics baseline. Whether the retrain waits is your call.
  • •Schema changes. A rename is caught in the dbt manifest diff, in change review or by a failing run_quality_check; assess_impact says what it breaks; generate_migration writes the change with rollback SQL for the owner to apply.
  • •Feature and training data checks. run_quality_check checks null rates, key uniqueness and minimum row counts on the feature and training tables (Snowflake, Postgres, BigQuery with a service-account key), and monitor_metrics baselines watch the values that matter, such as a feature's null share or mean. The ML agent writes the point-in-time join SQL for a leakage-safe training set and can emit a Feast-compatible feature view definition, both for the owner to run.
  • •Label and RAG data quality. On the scores and vectors you pass in, the ML agent ranks unlabelled rows by model uncertainty and flags duplicate and outlier embeddings before an index is built. Pull request review flags new columns whose names look personal before they reach a training set or a retrieval corpus.
  • •Model lineage and access. Register a model with its training tables and features in the context graph, and lineage and blast radius reach it; MLflow, Weights & Biases or SageMaker stay your system of record for runs and versions. provision_access checks policy and records a column-level grant with its expiry (nothing is granted when the verdict is review); on Unity Catalog the approved grant is applied, on Snowflake and BigQuery it is proposed for the owner.

What you still own and decide:

  • •The models. Design, features, training, evaluation and promotion. Retraining, serving and rollback stay in your ML platform, run by you.
  • •The approvals. They go to a named person; an unanswered request expires and escalates, never auto-grants, and no agent can promote its own work (how approvals work).
  • •The autonomy settings. Each domain sits on L0 manual, L1 observe, L2 propose, L3 act reversibly or L4 autonomous; Data Workers ships observe-only (autonomy levels, explained). The org-wide stop halts all autonomous dispatch.
  • •The receipt. Every change records the diff, approver, time, blast radius and rollback path in a hash-chained audit log (verify_global_hash_chain checks it). It doubles as reproducibility evidence.

How Data Workers fits the tools you already use

Your coding agent. Data Workers agents are MCP servers, so they sit inside Claude Code, Cursor, Codex or Copilot. Ask "what feeds plan_tier and did anything upstream change this week?" and the answer comes from the context graph, with sources. Setup follows the client setup docs.

# Example: register the ML, quality and lineage agents in Claude Code
git clone https://github.com/DataWorkersProject/dataworkers-claw-community.git
cd dataworkers-claw-community && npm install --ignore-optional
claude mcp add dw-ml -- "$(pwd)/start-agent.sh" dw-ml
claude mcp add dw-quality -- "$(pwd)/start-agent.sh" dw-quality
claude mcp add dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog

The core runs on a built-in sample estate; in a deployment the same agents point at your estate and carry the approvals and receipts. List the tools with /mcp; Data Workers with your coding agents covers each client.

Your ML platform. MLflow, Weights & Biases, SageMaker and Vertex AI keep training, tracking and serving, and your team's assistant can use their APIs or MCP servers side by side with Data Workers today (you're on MLflow). Data Workers' work rests on the tables and pipelines its own connectors read, plus the model lineage you register. Feast, Tecton and the Databricks feature store work the same way: their offline tables in Snowflake or BigQuery get Data Workers' checks natively, and in Databricks the owner runs the check while Data Workers follows the dbt manifest and orchestrator runs, among 50+ connectors (on Databricks, on Snowflake, integrations).

Pipelines, streams and context. Airflow, Dagster, Prefect, dbt, Kafka and Schema Registry connect natively, so the jobs that build your features and the streams that feed them are in the same graph. Feature definitions, label rules and "which table is authoritative" live in Data Context Wizard with named owners (bring your own context). Spellbook Data Catalog (in preview) holds the proposals, approvals and receipts.

The metrics you are judged on

MetricHow Data Workers moves itWhere you see it
Model quality in productionUpstream breaks caught at the table, before a retrain or a scoring run reads themIncidents that reached a model vs caught upstream
Drift incidentsThe upstream change behind a flagged feature, with lineage attachedDrift alerts from your monitoring tool, each with a cause in the receipt
Time to deployFewer days lost to data chases; access requests through approvalsTime from data request to approved grant, from the receipts
ReproducibilityEvery data fix recorded with its diff, approver and timeThe audit trail for the tables a training run read
Warehouse cost of featuresSnowflake credits attributed to each feature model through dbt query tags; fixes drafted for the ownerSpend per feature model; the design target is 25 to 40% lower warehouse spend

Take a baseline first: how to measure AI data agents and the ROI calculator.

A worked example: a renamed event field and a 10:00 churn retrain

This is an illustration, not a customer case. A mobile app publishes app_events to Kafka with Schema Registry; the events land in BigQuery; Airflow runs a dbt build of features.churn_user_daily at 06:00; a Vertex AI pipeline retrains the churn model at 10:00 and logs runs to MLflow; the CRM scores accounts from the model. The churn model is registered in the context graph with its training tables and features. The ML data domain runs at L2 propose.

TimeSystemWhat happened
18:05Schema RegistryThe mobile team registers a new app_events version: plan_tier becomes subscription_tier
06:00Airflow, BigQueryThe nightly feature build completes green
06:10Data Workersrun_quality_check: the plan_tier null rate breaches its threshold; get_root_cause points upstream to app_events, and the stream's owner confirms the 18:05 schema version
06:12Data WorkersBlast radius: two features, the churn model, the 10:00 retrain and CRM scoring
06:14Data WorkersFrom the registered lineage and the Airflow runs, names the training window the bad partitions fall in
06:16Data WorkersProposes a staging diff mapping subscription_tier to plan_tier and a backfill plan for the partitions since 18:20, and recommends holding the retrain
06:17SlackThe approval request reaches the model owner with the diagnosis and the receipt
08:40ML engineerReviews in Cursor, merges the diff, pauses the Vertex AI schedule, approves the backfill in Spellbook
08:45Data Workers, AirflowQueues the feature rebuild and backfill through Airflow
09:30Data WorkersRe-checks the feature table: quality passes; the plan_tier null share is back inside its baseline
09:35Jira SMOpens a ticket for the mobile team, linking the receipt
10:00Vertex AI, MLflowThe ML engineer resumes the schedule; the retrain runs on corrected features and MLflow logs the run
Incident timeline across the stack: what Your ML stack, your team and Data Workers each do, step by step

Without Data Workers, the retrain reads a feature that is 60% null and the first clue is a retention campaign aimed at the wrong accounts a week later. With it, the ML engineer starts the day with the cause, the affected models and a fix to approve, and the retrain runs on clean data on schedule.

How to bring it to your team

  • •Try it on the sample estate. Ask the questions you answer when a feature looks wrong.
  • •Pick one model that has been bitten by upstream data.
  • •Register its lineage. Connect the warehouse read-only and record the model's training tables and features in the graph.
  • •Start at L1 observe. For two weeks the agents read, trace and flag, and the permission ladder logs what each action would have needed; you review that record and its receipts against what happened to the model.
  • •Move the domain to L2 propose and share the record with your data engineering partners (how to get your team to trust AI agents).

For risk questions, send is it safe to let AI agents change production data and where your data goes: the agents run in your infrastructure, your data stays in your systems, and the hosted Conductor sees workflow metadata only. Build it ourselves and the ROI of agentic data operations cover the rest. Then start a pilot (what a pilot looks like; $7,500 one-time, credited in full against the first year, on /pricing/). Your partners have their own pages: data engineers and data platform engineers. For the wider picture: MLOps tools moving to AI agents and inside the MLOps and models agent.

FAQ

Will Data Workers retrain or deploy my models? No. Training, promotion and deployment stay in your ML platform and with you. Data Workers works on the data around the models: it traces upstream breaks, finds the data change behind a drifting feature, proposes fixes and records what changed.

We already have MLflow and Weights & Biases. Why add this? Keep them; they track experiments and models. Data Workers covers what sits upstream (tables, schemas, pipelines, owners), with approvals and receipts, and your team's assistant can use their APIs or MCP servers side by side with Data Workers today.

Can it stop a retrain from reading bad data? It gets the diagnosis, the affected models and a proposed fix to you before the scheduled retrain, and recommends a hold. Pausing or running the retrain stays your decision, in your ML platform.

Is it safe to let agents touch training data? At L2 propose every change reaches you as a diff first, backfills run through your orchestrator after approval, and every step lands in a hash-chained audit log. How Data Workers handles PII goes deeper.

Does it help with LLM and RAG pipelines? Yes, on the data side: privacy flags in pull request review before data reaches a corpus, duplicate and outlier detection across embeddings, and lineage and baselines on the tables an index is built from.

Sources

  • •Stack Overflow, 2025 Developer Survey, AI section (third-party; 49,009 responses, July 2025; 66% "almost right, but not quite"): https://survey.stackoverflow.co/2025/ai (checked Oct 2, 2026)
  • •Google Cloud, Vertex AI Pipelines, schedule a pipeline run (schedules can be paused and resumed; page last updated 2026-10-01): https://cloud.google.com/vertex-ai/docs/pipelines/schedule-pipeline-run (checked Oct 2, 2026)
  • •Data Workers client setup documentation: https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)
  • •Data Workers public repository (every tool named on this page is registered there): https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)
  • •Data Workers product repository, data-workers-agent-swarm main @ 0c2491e3: agents/dw-ml, agents/dw-context-catalog, agents/dw-governance, core/enterprise (checked Oct 2, 2026)
  • •Data Workers product pages: https://dataworkers.io/product/autonomous-data-conductor/ , https://dataworkers.io/product/data-agents-swarm/ , https://dataworkers.io/product/spellbook-data-catalog/ (checked Oct 2, 2026)
  • •Data Workers security and pricing: https://dataworkers.io/security/ , https://dataworkers.io/pricing/ (checked Oct 2, 2026)