You're on DataRobot: DataRobot Builds, Deploys and Governs the Models. Data Workers Owns Whether the Data They Depend On Is Right
On DataRobot? Keep it for AutoML, deployments, retraining policies and agent governance. Data Workers traces data drift to the upstream change behind it and fixes the tables your models score and retrain on, behind approvals.
Your models live in DataRobot, which now calls itself the Unified Agent Workforce Platform for Enterprise. Data scientists build in AutoML, compare the ranked candidates and register the winner. Deployments score on a schedule through batch prediction jobs that read from BigQuery or Snowflake and write predictions back. The Data Drift tab plots each feature's PSI against its importance, Accuracy tracks how predictions held up, and up to five retraining policies per deployment rebuild a model on a schedule or when drift or accuracy status changes, then add it as a challenger, save it or request a replacement through your approval policy. Model risk reads the compliance documentation DataRobot generates, and the same platform builds, moderates and governs agents, on-premise or air-gapped. The tables your models score and retrain on are built upstream, in dbt models, BigQuery tables and nightly DAGs another team owns. Data Workers owns whether that data is right: it catches the upstream change, traces it to every deployment it reaches, fixes it through the data owner and leaves a receipt.
Key takeaways
- •DataRobot keeps its job: AutoML, the registry, deployments, drift and accuracy monitoring, retraining policies, challengers, governance and agent moderation stay where they are.
- •Data Workers owns the tables your deployments read: scoring and training tables get checks and the metric baselines your team records, and an upstream change is traced through dbt lineage to every model registered in the context graph.
- •A drift alert gets a cause and a fix: Data Workers finds the upstream change behind it, proposes the fix as a diff for the named data owner, queues the rebuild through your orchestrator after approval and re-checks the table.
- •The ML owner keeps the model decisions: rerunning a batch job, promoting or deleting a challenger and replacing a champion stay in DataRobot.
- •DataRobot connects over its API or its Global MCP server today, in your team's client next to Data Workers' MCP servers.
DataRobot builds, deploys and governs the models. Data Workers owns whether the data they depend on is right.
Here is a Friday at a consumer lender. The early payment default model, epd-risk, is a DataRobot deployment. A batch prediction job runs at 03:00, reads ml_scoring.epd_features_daily from BigQuery and writes scores to ml_scoring.epd_predictions, which builds the collections team's morning queue. The deployment has a retraining policy with a Drift status trigger: when drift turns red, DataRobot takes a new snapshot of the AI Catalog dataset that queries ml_training.epd_examples, retrains on it and adds the new model as a challenger. Both tables come from dbt models in the lending-dbt repo, built on BigQuery by the Airflow 2 DAG lending_dbt_nightly from loan servicing data that Airbyte copies out of Postgres. One feature matters more than most: days_since_last_payment. For a loan with no payment yet it is missing, and the model learned that missing means "new and unproven". Months ago the ML team registered the model in Data Workers' context graph with its scoring, training and prediction tables. The DAG's last task computes two numbers in BigQuery after each build, the share of missing days_since_last_payment and the share of zeros, and records them with monitor_metrics, which compares each new value with the history recorded so far. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| Thu 18:10 | GitHub + dbt | The collections dashboard shows blanks for new loans, and a reviewer complains before Monday's ops review. An analyst pushes a hotfix straight to main in lending-dbt, so no pull request review runs: int_loan_payments now wraps days_since_last_payment in coalesce(..., 0). A loan with no payment yet now reads the same as one paid today |
| Fri 01:00 | Airflow + dbt | lending_dbt_nightly runs dbt build on BigQuery. int_loan_payments, ml_scoring.epd_features_daily and ml_training.epd_examples are table models, so two years of training history are rewritten with zeros too |
| 01:25 | Data Workers + BigQuery | The DAG's last task records a missing share of 0.0% against a typical 6.8%, and a zero share of 9.1% against 2.3%. monitor_metrics flags both against their recorded history. run_quality_check passes: loan_id has no nulls and stays distinct, and the table clears the row-count floor. The load is complete; the meaning of a column changed. Data Workers posts the finding to the data on-call channel in Teams, naming the registered epd-risk model downstream and recommending a hold on its drift-triggered retrain. Nobody reads it overnight |
| 03:00 | DataRobot | The batch prediction job scores the zero-filled table. About 3,100 loans with no payment yet now look current and fall out of the high-risk tier |
| 03:40 | DataRobot | The Data Drift tab turns red: days_since_last_payment, the second most important feature, shows its missing-value bin empty and a PSI far over threshold. The deployment owner gets the drift notification the team configured. The Drift status policy fires, snapshots the rewritten training table, retrains and adds the result as a challenger |
| 07:30 | Collections | The collections lead sees the morning queue short by about 3,000 new loans and asks the ML team why |
| 07:55 | Data Workers | The ML engineer reads the Data Workers card, then asks from Cursor what changed upstream. The dbt manifest from the 01:00 build shows int_loan_payments changed, and blast_radius_analysis from it returns the scoring table, the training table, the registered epd-risk model and the mart_collections dbt model. The engineer finds the coalesce in the analyst's commit on main |
| 08:10 | Spellbook + Teams | Data Workers proposes a diff: int_loan_payments keeps the missing value, and the coalesce moves into mart_collections, where the dashboard reads it. The plan rebuilds the four models; the undo, a BigQuery time travel restore of the rebuilt tables, is written into it for the owner to run if needed. The approval request reaches the lending data owner by email |
| 08:38 | Spellbook | The data owner reviews the diff, the blast radius and the undo, merges the diff and approves the rebuild |
| 08:41 | Airflow | Data Workers queues a run of lending_dbt_nightly on Airflow 2, with the four models in the run's conf, and reads its status; it succeeds at 09:06 |
| 09:15 | Data Workers + BigQuery | The run's last task records the missing share back at 6.8% and the zero share at 2.3%, and monitor_metrics raises no flag; run_quality_check passes. Data Workers writes the receipt and posts it in Spellbook and Teams |
| 09:30 | DataRobot | In DataRobot, the ML owner reruns the batch prediction job, deletes the challenger trained on the zero-filled snapshot and attaches the receipt to the deployment's record for model risk. The collections queue rebuilds by 10:00 |

DataRobot did its job well. It saw the feature move within an hour of scoring and turned the deployment red. Then its retraining policy did its job too, which was the risk: it rebuilt the model on the same rewritten history, and a challenger trained on bad data sat waiting for promotion. The cause was a one-line change in another team's repo, made for a dashboard. One diff fixed the dashboard's need and the model's input at once; the data owner approved the data change, and the ML owner kept every model decision.
| Job | What DataRobot does | What Data Workers does |
|---|---|---|
| Building | AutoML builds and ranks candidate models | Keeps the training tables those candidates learn from right |
| Scoring | Batch prediction jobs read the warehouse and write predictions back | Checks the scoring table before the job reads it, with the checks and baselines your team sets |
| Drift | The Data Drift tab measures PSI against the training baseline and sets a status | Finds the upstream change behind the drift and fixes it at the source |
| Retraining | Policies retrain on a schedule or on drift or accuracy status and add a challenger | Tells the ML owner when the retraining data is bad and recommends the hold; the hold is the owner's call |
| Governance | Replacement approvals, a central registry and automatic compliance documentation | Records the data change behind a model event: diff, approver, undo, re-check |
| The fix | Retrains or replaces when the ML owner decides | Proposes the data fix as a diff for its owner and queues the rebuild through the orchestrator after approval |
| The proof | The next drift window and accuracy readout | Re-checks the table, confirms the baseline and writes a receipt |
Why doesn't DataRobot just do this itself?
Because DataRobot is built for the people who own the models and agents, and its view starts where data arrives. That is the right focus for an enterprise AI platform. The Data Drift tab compares what a deployment scores against what it trained on, and it can tell you that days_since_last_payment lost its missing values. Retraining policies, challengers and replacement approvals act inside the model lifecycle, and the Global MCP server gives an assistant DataRobot's predictive AI tools. DataRobot measures the data as it lands, and does it well.
Why the data changed lives elsewhere: a dbt model in another team's repo, a DAG, a warehouse table that also feeds a dashboard. Fixing it means changing that team's code, rebuilding tables DataRobot does not own and getting a data owner's approval for a change that moves a business report too: a different owner, different credentials and a different liability. An AI platform that rewrote its upstream tables would be the consumer editing its supplier. Data Workers is the product built for that side of the line, and it builds on the DataRobot estate you already run.
Every tool owns a slice. Data Workers covers the whole lifecycle
DataRobot owns building, deploying and governing models and agents. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail.

| Stage | Data Workers | DataRobot | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 5 | DataRobot's registry is a central hub for models, LLMs, agents, tools and vector databases. Data Workers joins the tables under each model to dbt, orchestrator runs and quality results in one governed graph. |
| Analytics & Insights | 8 | 6 | AutoML, insights and shareable apps help analysts and data scientists explore and build. Data Workers answers questions from governed context and checks the numbers behind them. |
| Data Quality | 8 | 4 | Exploratory data analysis and data quality checks run on the data a project ingests. Data Workers checks the source tables after every build and proposes the fix where the failure starts. |
| Observability & Incidents | 8.5 | 5 | AI Observability tracks drift, accuracy and service health per deployment, and agent moderation logs guard failures. Data Workers detects, diagnoses, fixes and verifies data incidents across systems, with a receipt for each. |
| Pipelines & Ingestion | 8.5 | 3 | Data connections and batch prediction jobs read from and write back to the warehouse. Data Workers keeps the upstream jobs right and queues reruns through Airflow, Dagster, Prefect or ADF after approval. |
| Schema & Migration | 8 | 2 | Upstream table and dbt changes sit outside DataRobot's view until the data arrives. Data Workers scores a change against everything downstream and proposes the fix with its undo. |
| Governance & Access | 8.5 | 7 | Policies defined once, approval workflows for model replacement and automatic compliance documentation govern the model lifecycle. Data Workers runs the data access queue and proposes each grant for its owner. |
| Security & Privacy | 8 | 6 | Guards cover PII leakage, prompt injection and hallucinations, and the platform runs on-premise or air-gapped. Data Workers flags sensitive column names in pull request review and proposes masking for the owner. |
| Cost / FinOps | 8 | 4 | Compute orchestration and resource monitoring track what deployments use. Data Workers attributes Snowflake spend to the dbt model and sums BigQuery spend from the Jobs API. |
| MLOps & Models | 7.5 | 9 | The home stage: AutoML, the registry, deployments, drift and accuracy monitoring, retraining policies with challengers, and agent build, evaluation and moderation. Data Workers keeps the data under those models right. |
How DataRobot and Data Workers work together
Data scientists stay in DataRobot, ML engineers in Cursor, Claude Code or VS Code, and data engineers in dbt and Airflow. Spellbook Data Catalog (in preview) is where everyone looks: each finding, proposed change, blast radius, approver and rollback, with approval requests in Slack or email and alert cards in Teams. Data Context Wizard keeps one governed context graph, the Data-Agents Swarm's specialist agents do the work, and the Autonomous Data-Conductor runs each fix end to end.

What Data Workers reads. BigQuery natively with a service-account key, where run_quality_check covers nulls, distinct IDs and a minimum row count (on Snowflake and Postgres too). The default null check passes up to 10% nulls, so a shift in missing values belongs in a baseline. The dbt manifest, which is how an upstream change is traced to the tables a deployment reads. Airflow, Dagster, Prefect and ADF runs. And the metrics your team records with monitor_metrics: a DAG task computes a value, such as a feature's missing share or a daily row count, sends it after each build, and Data Workers flags it against the recorded history (after ten points). Your team registers each model once with its tables, so lineage and blast radius reach it. DataRobot connects over its REST API or Global MCP server today, in your team's client; Data Workers' own agents work from what its 50+ connectors read.
What Data Workers writes, and where. To DataRobot, nothing. Data fixes go to the data owner as diffs to merge, and Data Workers opens the pull request itself when your team turns on the GitHub pull-request target. After approval, rebuilds are queued through your orchestrator (on Airflow, version 2), with the undo written into the plan for the owner to run if needed.
Setup today. Give Data Workers read access to the BigQuery datasets and dbt project behind your deployments, and Airflow credentials scoped to the DAGs that build them. For your assistant, clone the open-source repository and add start-agent.sh entries to your client config, as the client setup docs show. DataRobot's Global MCP server, provisioned on your DataRobot instance, sits alongside with a DataRobot API key as the bearer token.
// Example: .cursor/mcp.json
{
"mcpServers": {
"dw-context-catalog": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-context-catalog"]
},
"dw-incidents": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-incidents"]
},
"datarobot-mcp": {
"url": "https://{DATAROBOT_URL}/api/v2/genai/globalmcp/mcp",
"headers": { "Authorization": "Bearer <DATAROBOT_API_TOKEN>" }
}
}
}Restart Cursor and check the servers in its MCP settings. An ML engineer can then ask "why did epd-risk drift overnight, and did anything change in the tables it reads?" and get the deployment's status from DataRobot's server and the lineage and findings from Data Workers in one answer.
One request, L0 to L4. The ladder is set per domain, so credit risk can sit at a different level from marketing.

- •L0 manual. Connected, not acting. The team chases a red drift status by hand.
- •L1 observe. Data Workers reports each upstream change with the deployments it reaches, and logs what each action would have needed.
- •L2 propose. It drafts the dbt diff and rebuild plan with blast radius and undo; the data owner approves first.
- •L3 act reversibly. For proven classes, such as queuing a scoring table's rebuild after its fix merges, it acts, re-checks and records.
- •L4 autonomous. For a scoped, trusted class in one domain, it runs the loop end to end. Retraining, challengers and replacements stay with the ML owner at every level.
The same pattern holds on Amazon SageMaker AI, Domino Data Lab and MLflow. The ML engineer's view is in Data Workers for ML engineers. For background, see data quality for ML and AI model governance. On trust: safety, approvals, the autonomy levels and where your data goes.
What changes for your team

ML teams on DataRobot lose days to data they don't own: a red drift status with no visible cause, a challenger nobody trusts, a batch job rerun twice, a data team that answers tomorrow. With Data Workers, the six jobs above run on autopilot at the level you set, and every data fix reaches model risk with its approver, diff and undo.
After this incident the team made two changes. lending-dbt now requires pull requests on main, so Data Workers reviews each change and shows its blast radius before it merges. And the ML owner set the drift-triggered policy to save the model instead of adding a challenger, so a person reads the Data Workers finding first.
Data scientists keep their week for what only people decide: which model to build and which challenger earns promotion.
Keep DataRobot, or consolidate?
Keep DataRobot if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
DataRobot is your model and agent platform, and it stays. What teams consolidate is the tooling around it: notebooks that compare this week's feature table with last week's, scripts that query the warehouse when drift turns red, spreadsheets of which dbt model feeds which deployment, and email threads with model risk about why a challenger appeared. If you are weighing building this layer on DataRobot's API and a coding agent, read build it ourselves with Claude Code and MCP servers: reading a deployment's drift is the easy part; cross-system context, approvals and rollback are the work.
The case for your CFO
The outcome: models score and retrain on data that is right. A broken upstream table becomes an approved fix within the morning, before a collections queue, a credit decision or a challenger is built on it.
The risk story is plain. Data Workers changes nothing in DataRobot; retraining, challengers and replacements stay with the ML owner. Every data change shows its blast radius, goes to a named approver, is re-checked and leaves a receipt with the undo. Unanswered requests expire and escalate; they never auto-grant. No agent can promote its own work, and an org-wide stop halts all autonomous dispatch. The agents run in your infrastructure, next to DataRobot even on-premise; your data stays in your systems, and the hosted Conductor sees workflow metadata only. Zero migration.
Why now: retraining policies make model refresh automatic, so a bad upstream change reaches a new model faster than it used to, and model risk teams want the data side of every model event documented too. The first win is one deployment's scoring and training tables, read-only, with a report of every upstream change that reaches them. What stays the same: DataRobot, your deployments, your approval policies and your compliance documentation. Start with a pilot; the path is on the pricing page, and the pilot is credited in full against the first year. See the ROI of agentic data operations.
The sentence to repeat upstairs: "DataRobot builds and governs our models; Data Workers makes sure the data they score and retrain on is right, and every fix comes with an approval and a receipt."
Getting started
Start with a pilot. Pick the deployment whose predictions drive the most decisions, register it with its tables, connect Data Workers read-only to the warehouse and dbt project behind it, add a DAG task that records missing shares for its top features with monitor_metrics, and let it report upstream changes for a few weeks before turning on the first fix class. Plans are on the pricing page, and the pilot is credited in full against the first year.
FAQ
Does Data Workers read DataRobot deployments, drift or the registry? They connect over DataRobot's REST API or its Global MCP server today, in your team's client next to Data Workers. Data Workers' own work rests on what it reads natively: the warehouse tables, the dbt manifest, orchestrator runs and the metrics your team records. Your team registers each model with its tables once, so lineage reaches it.
Does Data Workers measure data drift instead of DataRobot? No. DataRobot measures drift against the training baseline, and it should keep doing so. Data Workers watches the tables before the deployment reads them, with its checks and the metrics your team records, and when DataRobot's drift turns red it finds the upstream cause and fixes it through the data owner.
Does Data Workers retrain, promote or replace models? No. Retraining policies, challengers, replacement approvals and batch jobs stay in DataRobot with the ML owner. Data Workers fixes the data, queues the table rebuild after the data owner approves and tells the ML owner which deployments and retraining data the change reached.
We run DataRobot on-premise. Does that change anything? No. Data Workers' agents run in your infrastructure on every tier and hold the warehouse credentials and model key. Your data stays in your systems; the hosted Conductor sees workflow metadata only, never rows, credentials or model keys.
What if our scoring data lives in Databricks? Data Workers runs its table checks on BigQuery, Snowflake and Postgres. For tables built in Databricks, it works from the dbt manifest, orchestrator runs and the metrics your team records with monitor_metrics, and the owner runs the check in Databricks.
Sources
- •DataRobot, homepage (Unified Agent Workforce Platform for Enterprise), https://www.datarobot.com/ (checked Oct 3, 2026)
- •DataRobot, AI Governance product page, https://www.datarobot.com/product/ai-governance/ (checked Oct 3, 2026)
- •DataRobot, "DataRobot Gives Enterprises Full Control Over Where and How Their AI Runs" (Jul 22, 2026), https://www.datarobot.com/newsroom/press/datarobot-gives-enterprises-full-control-over-where-and-how-their-ai-runs/ (checked Oct 3, 2026)
- •DataRobot docs, Data Drift tab (updated Sep 9, 2026), https://docs.datarobot.com/en/docs/classic-ui/mlops/monitor/data-drift.html (checked Oct 3, 2026)
- •DataRobot docs, Retraining (triggers and model actions, updated Sep 9, 2026), https://docs.datarobot.com/en/docs/workbench/nxt-console/nxt-mitigation/nxt-retraining.html (checked Oct 3, 2026)
- •DataRobot docs, Retraining settings (new snapshot at trigger, updated Sep 9, 2026), https://docs.datarobot.com/en/docs/classic-ui/mlops/deployment-settings/retraining-settings.html (checked Oct 3, 2026)
- •DataRobot docs, Batch prediction jobs (sources and destinations, updated Sep 9, 2026), https://docs.datarobot.com/en/docs/classic-ui/predictions/batch/batch-dep/batch-pred-jobs.html (checked Oct 3, 2026)
- •DataRobot docs, MCP overview and MCP clients (Global MCP, updated Sep 9, 2026), https://docs.datarobot.com/en/docs/agentic-ai/agentic-mcp/agentic-mcp-overview.html and https://docs.datarobot.com/en/docs/agentic-ai/agentic-mcp/agentic-mcp-clients.html (checked Oct 3, 2026)
- •DataRobot docs, August 2026 release announcements (updated Sep 30, 2026), https://docs.datarobot.com/en/docs/release/cloud-history/2026-announce/august2026-announce.html (checked Oct 3, 2026)
- •DataRobot docs, Moderation events, https://docs.datarobot.com/en/docs/agentic-ai/agentic-monitor/agent-moderation.html (checked Oct 3, 2026)
- •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)
- •Data Workers open-source repository, agent tools and start-agent.sh, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)