You're on MLflow: MLflow Keeps the Record of Every Run and Model Version. Data Workers Keeps the Data Under Them Right
On MLflow for experiment tracking, the model registry and GenAI tracing? Keep it. Data Workers makes sure the tables your models train and score on are right, behind approvals.
Your models live in MLflow. Every training job logs a run with its parameters, metrics and artifacts, mlflow.log_input() records the dataset with a digest and a row-count profile, and the Model Registry links each version to the run that produced it. Aliases tell everyone which version is @champion and which is @challenger; on Databricks, MLflow 3 registers to Unity Catalog by default. On the GenAI side, MLflow 3 traces your agents on OpenTelemetry, scores them with LLM judges, and since April it identifies issues in traces automatically. MLflow tracks experiments, models and traces very well. Data Workers owns the other question: whether the data those models train and score on is right, and when it isn't, getting it fixed through its owner, behind approvals, with a receipt.
Dataset tracking captures the data when a run logs it; watching the source table between runs is code you write. When a dbt change rewrites the training table overnight, the next run logs a new digest and a better metric, and nothing says why.
Key takeaways
- •MLflow keeps its job. Tracking, the registry and aliases, evaluation, tracing, the AI Gateway and the MCP server stay where they are.
- •Data Workers watches the tables, not the runs. Baselines on your training and feature tables, quality checks on Snowflake, and lineage from the dbt manifest catch the change before a retrain reads it.
- •The model is a node in the graph. Register each production model once with its training tables and features in the Context Wizard graph, and blast radius reaches it from any upstream dbt model.
- •Fixes go to data owners. Data Workers proposes the dbt change as a diff, queues the rebuild through your orchestrator after a named person approves, and re-checks. The ML owner retrains and moves the alias in MLflow.
- •Every fix leaves a receipt: trigger, evidence, diff, approver, checks and undo path.
- •Start with a pilot on two production models, read-only first, from L0 manual toward L4 autonomous.
MLflow keeps the record of every run and model version. Data Workers keeps the data under them right.
MLflow answers what was trained, on what, with which result, and which version serves. Data Workers answers what changed in the data, what it reaches, who owns the fix and whether it held. Here is how they meet when a training table changes silently under a registered model.
A payments company scores card transactions with a fraud model, fraud_scorer, registered in a self-hosted MLflow 3.16 tracking server; version 30 holds the @champion alias. A Dagster job builds dbt models in Snowflake at 02:00, including ml.fraud_training_examples, a 180-day window of transactions joined to each merchant's risk tier. At 05:00 another Dagster job retrains the model, logs the run and the dataset to MLflow, and registers a new version as @challenger. The team registered fraud_scorer in Data Workers' context graph with its training table and features, and the nightly job posts two numbers to Data Workers' metric monitor (monitor_metrics): rows in the training window and the share of rows from high-risk-tier merchants. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| Mon 16:20 | GitHub and dbt | An analytics engineer merges a pull request that simplifies dim_merchants for a dashboard: it now builds from stg_merchants, the current-state feed, instead of the snp_merchants snapshot. Data Workers' pull request review lists ml.fraud_training_examples among the models dim_merchants feeds; every dbt test passes, and the reviewer merges a dashboard cleanup |
| Tue 02:00 | Dagster and Snowflake | The nightly job runs dbt build. ml.fraud_training_examples now joins each historical transaction to the merchant's tier as of today, including tiers raised after chargebacks. The label leaks into a feature |
| 02:40 | Data Workers | The job posts its numbers. Rows sit inside their baseline; the high-tier share jumps from 3.8% to 11.6%. monitor_metrics flags the anomaly |
| 02:43 | Snowflake | Data Workers runs run_quality_check on the training table: the null check, the txn_id distinct ratio and the row floor all pass. The load is complete; the change is in the values |
| 02:47 | Slack and Spellbook | Data Workers posts the finding to data on-call and the fraud model owner. trace_cross_platform_lineage from the dbt manifest shows dim_merchants now built from stg_merchants, changed in Monday's merge. blast_radius_analysis reaches the training table and the registered fraud_scorer, and the scoring view, which needs today's tier. It recommends holding the 05:00 retrain |
| 05:00 | Dagster and MLflow | Nobody has read the card yet. The retrain runs, logs a new dataset digest and an AUC of 0.981 against 0.912 for version 30, and registers version 31 as @challenger |
| 08:15 | MLflow | The fraud model owner sees a 7-point jump in the run comparison, reads the Slack card, keeps @champion on version 30 and tags version 31 validation_status: rejected |
| 08:40 | GitHub and dbt | Data Workers proposes the fix as a diff for the analytics engineer: the training model joins snp_merchants on the transaction time between dbt_valid_from and dbt_valid_to, and dim_merchants stays current for the dashboard. Its pull request review shows the change reaches the training table only |
| 09:30 | Spellbook and Dagster | The engineer merges. The data owner approves the rebuild in Spellbook; Data Workers queues the Dagster job and reads the run status |
| 10:20 | Snowflake | The rebuilt job posts its numbers: the high-tier share is 3.9%, inside its baseline; run_quality_check passes |
| 10:35 | Spellbook | Data Workers writes the receipt: alert, checks, lineage, diff, approver, rebuild, before and after, undo path. The approved fact "fraud training reads merchant tier as of transaction time" lands in the graph |
| 11:30 | MLflow | The owner retrains. Version 32 logs an AUC of 0.914, goes through the team's evaluation and takes @champion when the owner decides |

Every tool did its job: dbt built what the model said, Dagster ran on time, and MLflow recorded the run, the digest and the suspicious metric exactly. The problem lived between them: a join that was point-in-time correct on Monday and leaked the label on Tuesday, with no test failing and no column changing. A point-in-time join error is not something run_quality_check detects: nulls, uniqueness and minimum rows all pass on a complete table. It comes from Data Workers' review of the dbt diff on the pull request, a monitor_metrics baseline on a feature's distribution, or the ML owner. Here review showed the reach on Monday and the baseline caught the values on Tuesday. Data Workers traced the cause, routed the fix to its owner and left the model decisions to the person who owns the model.
| Job | What MLflow does | What Data Workers does |
|---|---|---|
| The record | Logs every run: parameters, metrics, artifacts, the dataset digest and profile | Keeps the baselines your jobs post on the training and feature tables after every build |
| The model | Holds each version with aliases, tags and a link to its run | Holds the model as a graph node joined to its tables, features and their owners |
| The signal | Shows the metric jump when someone compares runs | Flags the table change before the retrain reads it |
| The cause | Shows which dataset a run used | Traces the change through the dbt manifest to the model that changed |
| The fix | The owner retrains, evaluates and moves the alias | Proposes the dbt diff; queues the Dagster rebuild after approval; re-checks |
| The proof | Run history and the registry's version lineage | A receipt per fix: evidence, approver, checks, before and after, undo path |
Why doesn't MLflow just do this itself?
Because MLflow is the system of record for ML work, and it is built to be open, light and everywhere. It runs on a laptop, a Kubernetes cluster or inside Databricks, and asks nothing of your data stack: a run logs what the training code hands it. That design made it the default tracker across clouds and frameworks. Its newer intelligence, the MLflow Assistant and automatic issue identification, works where MLflow's data lives: in runs, traces and evaluations, and so does its MCP server.
Fixing a training table is a different product with a different liability: warehouse and orchestrator credentials, lineage across another team's dbt models, a named data owner's approval, a rollback path and proof the fix held. An experiment tracker that started rewriting dbt models in someone else's repo would stop being the neutral record every ML team trusts. Data Workers is the product on the other side of that line.
Every tool owns a slice. Data Workers covers the whole lifecycle
Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the tools already there.

MLflow's home stage is MLOps & Models, where it leads by design; Data Workers keeps the tables under it right.
| Stage | Data Workers | MLflow | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 4 | MLflow catalogs runs, models, prompts and datasets as logged. Data Workers joins tables, lineage, quality, owners and registered models in one governed graph. |
| Analytics & Insights | 8 | 3 | MLflow compares runs and charts metrics for ML teams. Data Workers answers questions about the data from governed context. |
| Data Quality | 8 | 3 | MLflow records a dataset's digest and profile at log time and does not watch the source. Data Workers checks nulls, keys and row counts and keeps baselines on the tables. |
| Observability & Incidents | 8.5 | 5 | MLflow traces agents and flags trace issues automatically. Data Workers detects, diagnoses, fixes and verifies data incidents, with a receipt for each. |
| Pipelines & Ingestion | 8.5 | 2 | MLflow logs what pipelines produce; it does not run them. Data Workers queues reruns through your orchestrator after approval and reviews pipeline changes. |
| Schema & Migration | 8 | 2 | MLflow stores model signatures. Data Workers shows a dbt change's reach and writes migrations with rollback for the owner. |
| Governance & Access | 8.5 | 5 | MLflow 3.13 added role-based access control; Unity Catalog governs models on Databricks. Data Workers routes each data change to a named approver and keeps the record. |
| Security & Privacy | 8 | 3 | MLflow secures its own server and Gateway. Data Workers proposes classifications and masking for data owners to apply. |
| Cost / FinOps | 8 | 3 | The AI Gateway tracks LLM spend and budgets. Data Workers attributes Snowflake spend to dbt models through query tags. |
| MLOps & Models | 7.5 | 9.5 | MLflow's home stage: tracking, evaluation, the registry with aliases, tracing. Data Workers keeps the training and feature tables under each model right. |
Scores are directional measures of scope, not benchmarks.
How MLflow and Data Workers work together
Your people stay where they are: ML engineers in MLflow and notebooks, analytics engineers in dbt, platform engineers in Dagster, and Claude, Cursor or GitHub Copilot for questions and code. Spellbook Data Catalog (in preview) is where the data team looks: each finding, proposed change, blast radius, approver and rollback, with approval requests reaching people in Slack or email. Underneath, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm does the work with 20+ specialist agents, and the Autonomous Data-Conductor runs each fix (detect, diagnose, fix, review, verify, remember) behind per-domain guardrails.

Where MLflow fits. MLflow stays the system of record for runs, models, aliases and traces; it connects over its REST API or MCP server today. Data Workers reads natively what sits under the model: Snowflake checks, the dbt manifest, GitHub pull requests and Dagster runs, among 50+ connectors. Your team, or its coding agent, registers each production model once with its training tables, features and prediction table, and from then on lineage and blast radius reach it. See Inside the MLOps & Models Agent for the agent itself.
What Data Workers writes, and where. To MLflow, nothing: runs, versions, aliases and tags belong to the model owner. dbt fixes go to their owner as a diff to merge; Data Workers opens the pull request when your team turns on the GitHub pull-request target. Rebuilds are queued through your orchestrator after approval; the owner runs any recorded undo. Approved facts land in the Context Wizard graph.
Setup today. Clone the open-source repository and add start-agent.sh entries to your client config, as the client setup docs show. Your client can run the MLflow MCP server next to them, so one assistant sees your traces and your data operations; Data Workers' agents never call it.
// Example: .mcp.json for Claude Code
{
"mcpServers": {
"dw-context-catalog": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-context-catalog"]
},
"dw-quality": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-quality"]
},
"dw-incidents": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-incidents"]
},
"mlflow": {
"command": "uv",
"args": ["run", "--with", "mlflow[mcp]>=3.5.1", "mlflow", "mcp", "run"],
"env": { "MLFLOW_TRACKING_URI": "http://mlflow.internal:5000" }
}
}
}List the tools with your client's own command (/mcp in Claude Code), then ask "which dbt models feed ml.fraud_training_examples, and what else reads dim_merchants?"
One request, L0 to L4. The autonomy ladder is set per domain, so the fraud domain can sit at a different level from marketing analytics.

- •L0 manual. Connected, not acting. The team finds the leak when production precision drops.
- •L1 observe. Data Workers flags each baseline break on a training table with its blast radius and owners, and logs what each action would have needed.
- •L2 propose. Data Workers drafts the dbt diff and the rebuild plan; owners approve before anything runs.
- •L3 act reversibly. For proven classes, such as queuing the rebuild after a fix merges, Data Workers acts, verifies and records it.
- •L4 autonomous. For a scoped, trusted class, Data Workers runs the loop end to end and posts the receipt. Retraining and aliases stay with the ML owner at every level.
Read more on safety, how approvals work, the autonomy levels and where your data goes.
What changes for your team

Teams on MLflow already made ML work reproducible. What still costs them time is the data under the record: the metric jump nobody can explain, the day in another team's dbt repo, the ticket bouncing between the ML platform and analytics. With Data Workers, those jobs run on autopilot at the level you set.
- •Incidents. A change to a training or feature table arrives traced to its cause, fixed through its owner and recorded.
- •Data quality. Nulls, keys, row floors and baselines are checked on the tables your models read.
- •Cloud spend. Snowflake spend is tied to each dbt model through query tags, including the nightly training-table rebuilds.
- •Access. Grant changes a fix needs arrive as proposals; Unity Catalog grants apply after approval.
- •Audits. Each fix links trigger, approver, checks and undo path, ready to attach to the model version.
- •Migrations. Each wave is planned with its parity checks, and the owner signs off before it closes.
The ML engineers guide covers the role in depth. Running other ML tools too? See You're on Weights & Biases, You're on Amazon SageMaker and You're on Tecton.
Keep MLflow, or consolidate?
Keep MLflow if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
Nearly every team keeps MLflow: it is open source, built into Databricks, and the registry is where model decisions live. What they consolidate is the tooling around it: row-count asserts in training scripts, notebooks that diff dataset versions, and spreadsheets mapping features to dbt models. Background: MLflow alternatives for model management, data lineage for ML features, data quality for ML, automating model retraining pipelines and, on Databricks, the Data Workers on Databricks guide. Building this yourself? Read build it ourselves with Claude Code and MCP servers: reading a baseline is easy; approvals, rollback and receipts are the work.
The case for your CFO
The outcome: models train and score on data that is right, and an upstream break is caught before a retrain ships it. MLflow keeps the record of what was trained; Data Workers makes sure what it was trained on deserves the record.
The risk story is plain. Data Workers changes nothing in MLflow: it never logs runs, registers versions or moves aliases. dbt fixes arrive as diffs your engineers merge; rebuilds are queued only after a named approver says yes; retraining stays with the model owner. Each action leaves a receipt. Unanswered requests expire and escalate; they never auto-grant. No agent can promote its own work. Autonomy is set per domain from L0 manual to L4 autonomous, and an org-wide stop halts all autonomous dispatch. The agents run in your infrastructure; your data stays in your systems, and the hosted Conductor sees workflow metadata only. Zero migration.
Why now: retraining is scheduled, coding agents ship dbt changes faster, and a leaked label looks like progress until production numbers fall. One bad version in front of customers, or model risk, costs far more than catching it in the table. The first win is two production models registered with their training tables, baselines on those tables and review on the dbt repo, read-only, with a report of every change that would have reached a retrain. What stays the same: MLflow, your registry workflow, your training code and your owners. Start with a pilot (pricing); the pilot is credited in full against the first year. For the numbers, see the ROI of agentic data operations.
The sentence to repeat upstairs: "MLflow records every model we ship; Data Workers makes sure the data under each one is right, with an approval and a receipt for every fix."
Getting started
Start with a pilot. Pick the two models your business feels first, register each with its training tables and features in the context graph, connect Data Workers read-only to Snowflake, the dbt project, pull requests and Dagster runs, post baselines after every build, and let it report what it would have caught before you turn on the first fix class. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.
FAQ
Does Data Workers read our MLflow runs and registry? Data Workers works on the tables under them, not on the runs. MLflow stays the system of record for runs, versions and aliases, and connects over its REST API or MCP server today. Your team registers each production model in the context graph with its training tables and features once.
Does Data Workers detect model drift? That job stays with your model monitoring and MLflow evaluations. Data Workers watches the tables under the model, with quality checks and the baselines your team posts, and traces a change to its cause.
Can Data Workers retrain the model or move the champion alias? Retraining, evaluation and aliases stay with the model owner in MLflow. Data Workers can recommend holding a retrain while a training table is wrong; the hold is the owner's call.
We run MLflow on Databricks with Unity Catalog models. Does this still work? Yes. On Databricks, Data Workers reads the dbt manifest and your orchestrator's runs and applies Unity Catalog grants after approval; table checks there stay with the owner. Snowflake and Postgres tables get its quality checks directly.
How is this different from the MLflow MCP server? The MLflow MCP server, experimental by its own label, works on traces and assessments. Data Workers' MCP servers work on the data estate: lineage, checks, baselines, approvals and receipts. Run both in one client.
Does this help our GenAI apps traced in MLflow? Yes, on the data side. When an agent's traces show failures because the documents table behind its retrieval stopped refreshing or doubled, a baseline your team posts flags it, and Data Workers traces it and queues the rebuild after approval. The prompt fix stays in MLflow.
Where does our data go? The agents run in your infrastructure and hold the credentials. Your data stays in your systems; the hosted Conductor sees workflow metadata only (table names, proposals, approvals), never rows, credentials or model keys.
Sources
- •MLflow, home page: open source AI platform for agents, LLMs and models; Apache 2.0, Linux Foundation, https://mlflow.org/ (checked Oct 3, 2026)
- •MLflow releases: 3.9.0 (Jan 30, 2026) to 3.16.0 (Sep 3, 2026), including automatic issue identification (3.11.1), role-based access control (3.13.0) and the MCP Registry (3.15.0), https://mlflow.org/releases (checked Oct 3, 2026)
- •MLflow GitHub releases: v3.16.1 (Sep 17, 2026) and v2.11.5 (Sep 23, 2026), https://github.com/mlflow/mlflow/releases (checked Oct 3, 2026)
- •MLflow docs, MLflow MCP Server (experimental; MLflow 3.5.1+; trace and assessment tools), https://mlflow.org/docs/latest/genai/mcp/ (checked Oct 3, 2026)
- •MLflow docs, Model Registry (aliases, tags, version lineage to runs), https://mlflow.org/docs/latest/ml/model-registry/ (checked Oct 3, 2026)
- •MLflow docs, dataset tracking (
mlflow.data,log_input, digest, profile), https://mlflow.org/docs/latest/ml/dataset/ (checked Oct 3, 2026) - •Databricks docs, Manage model lifecycle in Unity Catalog (updated Sep 11, 2026), https://docs.databricks.com/aws/en/machine-learning/manage-model-lifecycle/ (checked Oct 3, 2026)
- •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)
- •Data Workers open-source repository, agent tools and start-agent.sh, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
- •Data Workers pricing, pilot and hosted Conductor, https://dataworkers.io/pricing/ (checked Oct 3, 2026)