Product
Product12 min readBy The Data Workers Team

You're on Weights & Biases: W&B Tracks the Experiments, Artifacts and Evaluations. Data Workers Owns Whether the Data Behind Them Is Right

Running W&B, now part of CoreWeave Forge? Keep it. Data Workers checks the training tables behind every dataset version, traces upstream changes to your models and gets the data fix approved before the next sweep.

Your research team lives in Weights & Biases. Every training run logs its config, metrics and system stats to W&B Models; sweeps fan out hyperparameter searches overnight; datasets and checkpoints are logged as versioned Artifacts, and the Registry holds model and dataset versions with lineage between the runs that made and used them. W&B Weave covers LLM and agent traces, evaluations and monitors. CoreWeave completed its acquisition of Weights & Biases on May 5, 2025, and since the CoreWeave Forge launch on Sep 30, 2026, W&B Models sits inside Forge next to ARIA, the research agent that analyzes experiment history and proposes what to try next, and Agent Lens for agent traces. W&B tracks the experiments, artifacts and evaluations. Data Workers owns whether the data behind them is right: it checks the warehouse tables a dataset version is built from, traces an upstream change to the models that read it, and gets the fix approved and rebuilt before the next sweep trains on it.

Key takeaways

  • •W&B keeps its job. Runs, sweeps, Artifacts, the Registry, Weave, ARIA and Agent Lens stay with your research and ML platform teams.
  • •Data Workers owns the data under the dataset version. It watches the baselines your team records on training tables, checks keys and nulls on Snowflake, Postgres and BigQuery (with a service-account key), and reads the dbt manifest and orchestrator runs.
  • •It traces the change to the model. Register each model once with its training tables in the context graph, and dbt lineage plus that registration show which models an upstream change reaches. Data Workers does not read the W&B Registry.
  • •Fixes go to owners, behind approvals. dbt changes are drafted as diffs for the model owner; rebuilds are queued through your orchestrator after a named person approves. Retraining and promotion stay in W&B.
  • •Every fix leaves a receipt with the baselines before and after, the lineage, the diff and the approver.
  • •Start with a pilot on the training tables one team's sweeps read, read-only first, and climb from L0 manual toward L4 autonomous.

W&B tracks the experiments, artifacts and evaluations. Data Workers owns whether the data behind them is right.

W&B answers what was trained, on which dataset version, with which config, and how it scored. Data Workers answers whether the data in that version was right, who reads it and who signs off on the repair. When an upstream table changes quietly, W&B records the new version faithfully and the sweep learns from it.

Here is a Thursday at a podcast app with a personalized episode ranker. The iOS and Android apps send events through RudderStack, which writes each event type to its own table in BigQuery. A Prefect flow, ranking_nightly, runs the dbt Core project at 02:00: stg_listens reads the completion events, int_listen_labels marks each impression as completed or not, and ml.ranking_examples is the training table. The flow's last task logs the new partition as the W&B Artifact ranking-examples, and a scheduled sweep trains candidate rankers on the latest version at 03:00. During the pilot the team had the flow record two metrics with monitor_metrics after each build (completion events per platform per day, and the positive-label share), and registered the ranker with its training table in the Data Workers context graph. This is an illustration, not a customer case.

TimeSystemWhat happens
Wed 16:30iOS app and RudderStackRelease 9.2 ships with a tracking cleanup: Episode Completed is now sent as Playback Completed. RudderStack writes the new event to a new table, ios_app.playback_completed, and ios_app.episode_completed stops growing. Android is unchanged
Thu 02:00Prefect and dbt Coreranking_nightly runs dbt build and succeeds. Every dbt test passes: keys are unique, nothing is null. stg_listens simply finds half as many iOS completions for Wednesday, and ml.ranking_examples labels the rest of those listens as skips
02:15W&BThe flow logs ranking-examples:v14, with the positive-label share in the Artifact's metadata, as it does every night
02:16Data Workers and BigQueryThe flow's post-run step records the two metrics. monitor_metrics flags iOS completions for Wednesday at 512,000 against a weekday baseline near 1.2M, and the positive-label share at 27% against a baseline near 44%. Android completions sit on their baseline. run_quality_check, running on BigQuery with a service-account key, passes nulls on user_id and episode_id and uniqueness on example_id
02:18Data Workers, dbt, SlackThe dbt manifest shows no model change since Monday and the Prefect run succeeded, so the gap sits upstream of dbt, in the iOS source. blast_radius_analysis on ios_app.episode_completed returns stg_listens, int_listen_labels and ml.ranking_examples from the manifest, and, through the context graph, the ranker the team registered, with its note that the 03:00 sweep reads the latest version. A Slack card goes to the analytics engineering owner and the ML on-call recommending a hold on the sweep; stopping it is the ML owner's call
03:00W&BNobody is awake. The sweep agents on the team's GPU nodes train 24 candidate rankers on v14, as scheduled
08:40Claude Code, W&B MCP server, Data WorkersThe ML engineer reads the Slack card before linking the best run to the Registry. In Claude Code she asks the W&B MCP server to diff ranking-examples v13 and v14: row count nearly flat, positive-label share down from 44% to 27%. In the same session Data Workers returns the open incident with its baselines, lineage and the hold recommendation. She tags the 24 runs and links none of them
09:05RudderStackThe analytics engineer checks the iOS source's events in RudderStack and finds Playback Completed arriving since release 9.2. The mobile lead confirms the rename and that Android follows in 9.3
09:12SpellbookData Workers proposes one change set with the readers attached: a dbt diff that has stg_listens read both iOS tables as one completion event, a rebuild of Wednesday's partition of the three models, and a note for the mobile team to add the new name to the tracking plan before Android switches
09:20Spellbook and GitHubThe analytics engineering owner reviews the diff and the readers in Spellbook, approves, and merges it in the dbt repository
09:21PrefectAfter that approval, Data Workers queues a ranking_nightly run for Wednesday's partition through Prefect
10:25BigQuery, W&BThe run completes. iOS completions read 1.19M and the positive-label share 43.8%, both back on baseline; run_quality_check passes again, and the distinct ratio on example_id holds, so no listen counts twice. The team's flow, not Data Workers, logs ranking-examples:v15 to W&B
10:40W&B and SpellbookThe ML engineer reruns the sweep on v15 and notes on v14 where the receipt is: baselines before and after, lineage, diff, approver, Prefect run
Incident timeline across the stack: what W&B, your team and Data Workers each do, step by step

Nothing in W&B was wrong: the Artifact recorded exactly what the pipeline produced, the sweep ran on schedule, and the dataset diff showed the change the moment someone asked. Every dbt test passed too; the rows were unique and complete, just mislabeled. What was missing was a check on the data behind the version and someone to follow the change to the owner who could fix it. Data Workers caught it from the baselines the team records before the sweep started, traced it to the registered ranker, put the fix to the right owner and closed the loop with a receipt. Stopping the sweep, the Registry and the next experiment stayed with the ML engineer.

JobWhat W&B doesWhat Data Workers does
RecordingLogs runs, configs, metrics and system stats; versions datasets and checkpoints as ArtifactsKeeps a receipt for every data fix, with checks, lineage and approver
LineageRegistry lineage between the runs that made and used each Artifactdbt manifest lineage from the source table to the training table, plus the models your team registers in the context graph
DetectionShows what changed between two dataset versions when someone compares themBaselines and key, null and row-count checks on the training tables after every build
DiagnosisARIA analyzes experiment data and explains what drove a change in resultsFollows the change upstream through the dbt manifest and orchestrator runs to the system that caused it
The fixNew experiments, sweeps and retrainingA dbt diff for the data owner, then a rebuild queued through your orchestrator after approval
PromotionRegistry versions and the handoff to deploymentRecommends holding the promotion; the call stays with the ML owner
Agents and LLM appsWeave traces, evaluations and monitors; Agent Lens on production tracesChecks the tables those agents read and retrieve from

Why doesn't W&B just do this itself?

Because W&B is built around one unit of work, the experiment, and it is very good at it. It records what your training code logs (run, config, metrics, the dataset version by reference or checksum, lineage between runs and Artifacts) and sits downstream of the warehouse by design. ARIA works over experiment history, Agent Lens over agent traces, and the W&B MCP server reads runs, Artifacts, the Registry and Weave traces, with a read-only mode that drops its write tools (creating reports, logging analyses and, in an opt-in profile, sending work to ARIA). That focus is right for an experiment platform whose buyer is the head of research or the ML platform and whose job is to make every result reproducible and comparable.

Knowing that a mobile release renamed an event, that RudderStack wrote it to a new table, that stg_listens reads only the old one, and that a two-line dbt change plus a one-partition rebuild fixes three models, is data-engineering knowledge. Acting on it crosses the event pipeline, the warehouse, a dbt repository and an orchestrator, each owned by a different team, and an experiment platform that rewrote the analytics team's models would carry the liability for every dashboard that reads them. Data Workers is built for that job, with approvals, rollback and receipts.

Every tool owns a slice. Data Workers covers the whole lifecycle

W&B owns its slice: experiment tracking, artifacts, the registry and evaluations. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, W&B goes deep on its own area
StageData WorkersW&BWhy we scored it this way
Catalog & Context94Artifacts and the Registry keep versioned datasets and models with lineage between the runs that made and used them. Data Workers joins tables, dbt lineage, runs, owners, quality results and the models your team registers in one graph.
Analytics & Insights84Reports and workspace panels compare runs and explain results to the research team. Data Workers answers questions about the business data from governed context and checks the numbers behind them.
Data Quality83W&B shows the dataset version a run used and whatever the training code logs about it; it does not check the warehouse table that produced it. Data Workers checks keys, nulls, row counts and the baselines your team records.
Observability & Incidents8.56Weave monitors score live agent behavior and Agent Lens clusters trace failures. Data Workers detects, diagnoses and fixes data incidents across systems, with a receipt for each.
Pipelines & Ingestion8.52W&B receives what the pipeline logs at the end. Data Workers reads orchestrator runs, queues reruns after approval and reviews pipeline changes before they merge.
Schema & Migration82W&B versions the files and tables you log, not warehouse schemas or dbt models. Data Workers reviews model changes from the dbt manifest, assesses impact and plans migrations in approved waves.
Governance & Access8.55Teams, projects and Registry roles govern who sees and promotes what. Data Workers dry-runs grants, applies approved Unity Catalog changes and drafts the rest for their owners.
Security & Privacy84Dedicated and self-managed deployments keep experiment data in your perimeter. Data Workers keeps a receipt for every data change and proposes masking for its owner.
Cost / FinOps83W&B records system metrics such as GPU use per run; spend lives with the cloud. Data Workers ties Snowflake spend to dbt models and reads AWS Cost Explorer by service.
MLOps & Models7.59.5W&B's home stage: experiment tracking, sweeps, Artifacts, the Registry, Weave evaluations and ARIA, now inside CoreWeave Forge. Data Workers keeps the training and feature data under those models healthy.

How W&B and Data Workers work together

Researchers stay in W&B, notebooks and Claude Code, Cursor or Claude. Spellbook Data Catalog (in preview) is where the data team sees each finding, diff, approver and rollback; approval requests arrive in Slack or email. Underneath, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm (20+ specialist agents) does the work, and the Autonomous Data-Conductor runs each fix behind per-domain guardrails.

How Data Workers fits with W&B: your coding agent on top, Data Workers in the middle, your estate underneath

What W&B keeps. Runs, sweeps, Artifacts, Reports, Automations, the Registry, Weave, ARIA and Agent Lens, with training, evaluation, promotion and the deployment handoff. Data Workers connects to W&B over its API or MCP server today. The W&B MCP server runs hosted by W&B, on your dedicated or self-managed instance, or locally, and your assistant uses it side by side with Data Workers' MCP servers: one question to W&B about what changed between two dataset versions, one to Data Workers about why. Data Workers' agents don't call the W&B MCP server; their checks rest on what Data Workers reads itself.

What Data Workers reads. It reads the dbt manifest for lineage, checks training tables with run_quality_check (nulls, ID uniqueness and a row floor on Snowflake and Postgres, and BigQuery with a service-account key), records the baselines your team chooses with monitor_metrics, and reads Airflow, Dagster, Prefect and ADF runs, among 50+ connectors. You register each model once with its training tables and features in the context graph, and from then on the model is a node that trace_cross_platform_lineage and blast_radius_analysis reach; nothing is read from the W&B Registry. RudderStack connects over its API or MCP server today. The MLOps & Models agent keeps its own versioned record of models; W&B stays the system of record for experiments.

What Data Workers changes. dbt changes are drafted as diffs for their owners to merge. Rebuilds and backfills are queued through your orchestrator after a named person approves, with the undo written down for the owner before the run. Event pipelines, ingestion and deletes stay with their owners. Nothing in W&B changes: retraining, sweeps, logging new Artifact versions and Registry promotion are your ML team's calls.

Setup. Clone the open-source repository and add start-agent.sh entries to your client config, as the client setup docs show. Add the W&B MCP server to the same client, in read-only mode if you prefer.

// Example: .mcp.json for Claude Code
{
  "mcpServers": {
    "dw-incidents": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-incidents"]
    },
    "dw-context-catalog": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-context-catalog"]
    },
    "dw-quality": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-quality"]
    },
    "dw-connectors": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-connectors"]
    }
  }
}

List the tools with /mcp in Claude Code, then ask "ranking-examples v14 has fewer positive labels than v13. What changed upstream of ml.ranking_examples, and what else reads it?"

One dataset version, L0 to L4. The ladder is set per domain.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. A researcher notices odd sweep results days later and asks the data team.
  • •L1 observe. Data Workers watches the baselines, reports gaps with the models they reach, and logs what each action would have needed.
  • •L2 propose. It drafts the dbt fix and the rebuild; the owner approves before anything runs.
  • •L3 act reversibly. For proven classes, such as rebuilding a partition once its source is back on baseline, it queues the step, re-checks and records.
  • •L4 autonomous. For a scoped class in one domain, it runs the loop end to end and posts the receipt. Sweeps, retraining and promotion stay with your ML team.

More on safety, approvals, who owns the agents, autonomy levels, where your data goes and integrations, plus the ML engineer guide. Siblings: MLflow, Comet, Arize, and for this incident's data side RudderStack, dbt and Prefect.

What changes for your team

Six jobs that run on autopilot with Data Workers next to W&B, with a concrete example of each

ML teams lose time at the seam between the experiment and the table: the sweep that looked better for the wrong reason, the ticket bouncing between research and data engineering. With Data Workers, those jobs run on autopilot at the level you set.

  • •Incidents. An upstream change to a training table arrives traced to the model, owner named, fix drafted.
  • •Data quality. Training and feature tables are checked against baselines, keys and row counts after every build.
  • •Audits. Each fix links the approver, the diff and the dataset versions it touched.
  • •Cloud spend. Snowflake credits per dbt model and the BigQuery account total sit on the receipt.
  • •Access. A grant a pipeline needs arrives dry-run, with the sensitive columns it reaches.
  • •Migrations. Each wave is planned with its checks; the owner signs off before it closes.

Researchers keep researching, and the analytics engineers get a precise, approved change instead of "the labels look off".

Keep W&B, or consolidate?

Keep W&B if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

Nothing here replaces W&B's tracking, sweeps, Registry or Weave, and the move into CoreWeave Forge does not change that. What teams consolidate is the glue around it: notebooks that sanity-check a table before a sweep, Slack threads reconstructing which upstream change moved a dataset version, spreadsheets of who approved which fix. Building it yourself? Calling an MCP server is the easy part; the checks, lineage, approvals and receipts are the work. See also data quality for ML and MLflow alternatives.

The case for your CFO

The outcome: every model your team trains in W&B learns from checked data, so GPU hours go to real improvements and no promotion rests on a quietly broken table.

The risk story is plain. Data Workers changes nothing in W&B and never trains, stops a sweep, promotes or deploys a model. It checks the training tables, traces the readers and drafts each data fix for its owner; rebuilds are queued through your orchestrator after a named person approves. Unanswered requests expire and escalate, never auto-grant. No agent can promote its own work. An org-wide stop halts all autonomous dispatch. The agents run in your infrastructure with your credentials; your data stays in your systems, and the hosted Conductor sees workflow metadata only. Zero migration.

Why now: experiment platforms have agents of their own. ARIA proposes the next experiments and sweeps run unattended, so more training starts without a person looking at the table. The faster that loop turns, the more a bad partition costs in GPU time and in models that look better than they are. The first win is one team's training tables, read-only, with baselines on the columns that drive labels. W&B, dbt, your orchestrator and your research workflow stay as they are. The ROI guide shows the math and the flat-fee explainer covers pricing without a usage meter.

The sentence to repeat upstairs: "W&B tells us which dataset version a model learned from; Data Workers makes sure that version was right, with an approver and a receipt for every fix."

Getting started

Start with a pilot. Pick the training tables one team's sweeps read most, connect Data Workers read-only to the warehouse, the dbt project and your orchestrator, register the models with their training tables, record baselines on row counts and label shares, and watch a few weeks of builds before turning on the first fix class. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.

FAQ

Does anything change now that W&B is part of CoreWeave Forge? Not for Data Workers. W&B Models, the Registry and Weave carry on inside Forge, and Data Workers works on the warehouse side, whichever cloud trains your models.

Does Data Workers connect to Weights & Biases, and does it read our runs or Artifacts? It connects over W&B's API or MCP server today; your assistant uses the W&B MCP server side by side with Data Workers'. W&B stays the system of record for experiments: Data Workers reads the warehouse, the dbt manifest and orchestrator runs with your credentials, and reaches the models you register in its context graph.

Does Data Workers measure model drift or run evaluations? That is W&B's job, with Weave evaluations and monitors, and your training code's. Data Workers finds and fixes the upstream data change behind a shift, so the next run trains on the right data.

Will Data Workers retrain, stop a sweep or promote a model? No. It recommends holding a sweep or a promotion when the data behind it is wrong; the ML owner decides and acts in W&B.

How does Data Workers know a training table is wrong if every dbt test passes? From the baselines your team records with monitor_metrics, such as completions per platform or the positive-label share, plus key, null and row-count checks with run_quality_check. Your dbt tests keep running.

Does this work for Weave and RAG applications? Yes, for the data side. When a Weave evaluation drops because the documents table behind a retrieval index stopped refreshing, a baseline your team records on its rows per day flags it; Data Workers traces the readers and queues the refresh through your orchestrator after approval. Re-indexing and evaluation stay with your AI engineers.

Sources

  • •CoreWeave, completed acquisition of Weights & Biases (May 5, 2025), https://www.coreweave.com/news/coreweave-completes-acquisition-of-weights-biases-2 (checked Oct 3, 2026)
  • •CoreWeave, CoreWeave Forge launch with ARIA, Agent Lens and Registry (Sep 30, 2026), https://www.coreweave.com/news/coreweave-forge-launches-turning-the-ai-loop-production-run-into-a-better-model-and-agent (checked Oct 3, 2026)
  • •CoreWeave Forge product page, https://www.coreweave.com/forge (checked Oct 3, 2026)
  • •CoreWeave Forge docs home (docs.wandb.ai redirects here), https://docs.coreweave.com/forge-home (checked Oct 3, 2026)
  • •Weights & Biases site (Forge banner), https://wandb.ai/site (checked Oct 3, 2026)
  • •W&B MCP server README (tools, read-only mode, hosted and local deployment), https://github.com/wandb/wandb-mcp-server (checked Oct 3, 2026)
  • •Data Workers product repository (origin/main): baselines, quality checks, lineage, model registration, orchestrator triggers, approvals (checked Oct 3, 2026)
  • •Data Workers open-source repository (dataworkers-claw-community main): tool registrations (checked Oct 3, 2026)
  • •Data Workers client setup docs, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)