You're on Comet: Opik Scores the Agent and Comet Tracks the Runs. Data Workers Answers for the Data They Read
On Comet for experiments and Opik for agent traces and evals? Keep both. Data Workers finds the upstream data change behind a failing eval and fixes it, behind approvals.
Your team works in two Comet products. ML engineers log training runs to Comet Experiment Management, with versioned datasets and a registry of model stages. AI engineers trace every agent call into Opik, the open-source observability and evaluation platform: online evaluation rules score production traces with LLM-as-a-judge metrics such as Answer Relevance, Test Suites hold regression cases as plain-language assertions, and Diagnostics groups repeated trace failures into one issue with a root cause. Ollie, the assistant built into Opik, suggests the fix and changes nothing without your click. Agent Optimizer tunes prompts and tool descriptions; Cost Intelligence shows where your coding agents spend tokens.
Comet tracks the experiments and Opik scores the agent, and both do it well. But when a Test Suite drops overnight and the agent's code did not change, the answer usually sits in a table another team owns: the corpus the agent retrieves from, the features a run trained on. Data Workers owns that part of the job. It catches the upstream change, traces it to the agent or model it reaches, proposes the fix to the data owner, follows it through approval and leaves a receipt.
Key takeaways
- •Comet keeps its job. Experiment Management, the registry, Opik tracing, Test Suites, Diagnostics, Ollie and the Opik MCP server stay where they are.
- •Data Workers works on the tables, not the traces. Checks on corpus and training tables, baselines your team records and dbt manifest lineage catch the upstream change, often before the next Test Suite run.
- •The agent is a node in the graph. Register each production agent or model once with its tables in the Context Wizard graph, and blast radius reaches it from any upstream dbt model.
- •Fixes go to data owners. Data Workers proposes the dbt change as a diff, routes the rebuild to a named approver, follows the run and re-checks. Suites, re-indexing and retraining stay with the AI or ML owner.
- •Every fix leaves a receipt with the trigger, evidence, diff, approver, checks and undo path, ready to link from the Opik issue.
Comet tracks the experiments and Opik scores the agent. Data Workers answers for the data both of them read.
Opik answers what the agent did and how it scored. Data Workers answers why the data changed, what else it reaches, who owns the fix and whether the fix held. Here is how they meet when a failing eval is really a dedupe that ignores language.
A software company runs a support agent that answers from its Zendesk help center. A Dagster job, kb_nightly, starts at 01:00: an asset pulls articles for every enabled locale from the Zendesk API into Snowflake, dbt build keeps one row per article in int_help_articles and splits passages into kb.article_chunks, and a final asset re-embeds changed chunks into pgvector in Postgres. Opik traces the agent, with an Answer Relevance rule and a 180-item policy Test Suite that CI runs each morning. The team registered the agent in Data Workers' context graph with its chunk and embedding tables, set a 1% null threshold on the chunk_text check, and has the job send one metric after each build: the share of chunks in English, steady at 98% to 100% for months. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| Mon 15:40 | Zendesk | Support operations enables French and German and publishes translations of 212 of the 640 articles for the EU launch. Each translation keeps its article's ID and carries a later edit time |
| Tue 01:00 | Dagster, Snowflake, dbt | kb_nightly lands 852 rows in raw_zendesk.help_articles, one per article and locale. int_help_articles keeps the most recently edited row per article_id, so 212 English bodies are replaced by translations. The chunk count barely moves. Every dbt test passes |
| 01:14 | Data Workers, Snowflake | monitor_metrics flags the English share of kb.article_chunks at 67%. run_quality_check passes: nulls on chunk_text under 1%, the chunk_id distinct ratio above 99%. An incident opens |
| 01:16 | Slack, Spellbook | The finding reaches the knowledge-base data owner and the AI on-call. trace_cross_platform_lineage over the dbt manifest shows the chunks built from int_help_articles and the raw table, with no model change since last week, so the change sits in the source data. blast_radius_analysis reaches the support agent the team registered with its embedding table. Data Workers recommends holding the re-embed |
| 01:20 | Dagster, Postgres | Nobody is awake. The embedding asset re-embeds about 1,900 changed chunks. For 212 articles the agent now retrieves French or German passages |
| 08:10 | Opik | The morning CI run of the policy Test Suite passes 71% of items against 94% on Monday. Answer Relevance on production traces has slid since 01:30 |
| 08:45 | Opik | The AI engineer clicks Run diagnostic. Diagnostics groups 140 traces into one high-severity issue: French or German context for English questions. The Ollie fix suggests a prompt line telling the agent to answer only from English context |
| 08:55 | Claude Code | Before approving it she asks in one session. The Opik MCP server returns the issue and sample traces; Data Workers returns the open incident with its baseline, lineage and the held re-embed. She leaves the prompt change unapproved |
| 09:25 | Spellbook, GitHub | Data Workers proposes one change set: a dbt diff that keys int_help_articles on article_id and locale and has kb.article_chunks take en-us rows for the English agent, then a kb_nightly run to rebuild and re-embed. The undo, a Snowflake Time Travel restore of both tables, is written into the plan for the owner. The approval goes to the data owner, with the AI engineer copied |
| 09:40 | Spellbook, Dagster | The owner reviews the diff, the blast radius and the undo, merges in the dbt repository, approves and launches the kb_nightly run. Data Workers follows the run; it succeeds at 10:21 |
| 10:25 | Data Workers, Snowflake | The English share is back at 99% and run_quality_check passes. Data Workers writes the receipt: baseline, checks, lineage, diff, approver, run, before and after, undo path. The approved fact "the article key is article ID plus locale; the English agent reads en-us only" lands in the graph |
| 10:40 | Opik | The AI engineer reruns the Test Suite in Opik: 94%. She marks the Diagnostics issue Resolved and drops the prompt line |

Every tool did its job, and Opik caught the regression precisely. Ollie's suggestion was reasonable for what the traces showed; approved, it would have hidden the symptom and left 212 articles missing from the English corpus. The problem lived between the tools: a dedupe that was right on Monday and wrong on Tuesday, with no check failing. Data Workers caught the table, routed the fix to its owner and left the agent decisions with the agent's owner.
| Job | What Comet does | What Data Workers does |
|---|---|---|
| The signal | Evaluation rules, Test Suites and Diagnostics on what the agent logs | Checks and baselines on the tables the agent and runs read, after every build |
| The run | Tracks experiments, dataset versions and registry stages | Holds the agent or model as a graph node joined to its tables and owners |
| The cause | Ollie reads the traces to find where the answer went wrong | Traces the change through the dbt manifest to the upstream model and its owner |
| The fix | Ollie suggests code or prompt edits; the AI owner approves them | Proposes the dbt diff, routes the rebuild for approval, follows the run, re-checks |
| The proof | Test Suite pass rates, scores and resolved issues | A receipt per fix: evidence, approver, checks, before and after, undo path |
Why doesn't Comet just do this itself?
Because Comet is built for the people who own the model and the agent, and its design follows from that. Its agents work where Comet's data lives: Diagnostics reads traces, Ollie reads traces and, with opik connect, your code, and proposes edits that change nothing until you click. The Opik MCP server writes traces, scores, prompt versions and Test Suites, and by Comet's own description nothing can be deleted through it. That focus is why AI teams trust Opik's scores: it judges the agent without becoming part of the pipeline that feeds it.
Fixing the corpus table is a different product with a different liability: warehouse and orchestrator credentials, lineage across another team's dbt models, a named owner's approval, a rollback path and proof the fix held. An evaluation tool that edited staging models would stop being the neutral judge. Data Workers is the product on the other side of that line: the one that changes data, carefully.
Every tool owns a slice. Data Workers covers the whole lifecycle
Each point tool adds another console, another contract and another handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and it builds on the tools already there.

Comet's home stage is MLOps & Models, where it leads by design.
| Stage | Data Workers | Comet | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 4 | Projects, workspaces, Artifacts and the registry organize runs, datasets and model versions. Data Workers joins tables, lineage, quality, owners and registered models and agents in one governed graph. |
| Analytics & Insights | 8 | 4 | Experiment views and Opik dashboards compare runs, prompts and scores. Data Workers answers questions about the data itself from governed context. |
| Data Quality | 8 | 3 | Opik judges the agent's answers, including hallucination and answer relevance. Data Workers checks the tables the agent and training runs read, and fixes what fails where it starts. |
| Observability & Incidents | 8.5 | 6 | Diagnostics groups recurring trace failures into issues with a root cause and a suggested fix. Data Workers detects, diagnoses, fixes and verifies data incidents, with a receipt for each. |
| Pipelines & Ingestion | 8.5 | 2 | Comet logs what runs produce; it does not run the pipelines that build training data or the corpus. Data Workers queues reruns through your orchestrator after approval. |
| Schema & Migration | 8 | 2 | Comet versions datasets and models. Data Workers shows a dbt change's reach and writes migrations with rollback for the owner. |
| Governance & Access | 8.5 | 4 | Workspaces and roles govern who sees which project; the hosted Opik MCP server signs people in through the browser. Data Workers routes each data change to a named approver and keeps the record. |
| Security & Privacy | 8 | 4 | Comet runs as SaaS, in your VPC or on-premises, and Opik is open source. Data Workers proposes masking for data owners to apply, behind approvals. |
| Cost / FinOps | 8 | 5 | Cost Intelligence attributes coding-agent token spend and sets policies on it. Data Workers attributes Snowflake spend to dbt models and reads AWS Cost Explorer. |
| MLOps & Models | 7.5 | 9 | The home stage: experiment tracking, the registry, Opik tracing, Test Suites, LLM-as-a-judge metrics, Ollie, Diagnostics and Agent Optimizer. Data Workers keeps the data under those runs and agents right. |
Scores are directional measures of scope, not benchmarks.
How Comet and Data Workers work together
Your people stay where they are: ML engineers in Comet, AI engineers in Opik, analytics engineers in dbt, and Claude Code, Cursor or VS Code for questions. Spellbook Data Catalog (in preview) is where the data team looks: each finding, proposed change, blast radius, approver and rollback, with approval requests reaching people in Slack or email. Underneath, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm does the work with 20+ specialist agents, and the Autonomous Data-Conductor runs each fix (detect, diagnose, fix, review, verify, remember) behind per-domain guardrails.

Where Comet fits. Comet stays the system of record for runs, model versions, traces, scores and issues. Data Workers connects to Comet over its API or MCP server today: your client runs the Opik MCP server side by side with Data Workers' MCP servers. Data Workers reads natively what sits under the agent: Snowflake, BigQuery and Postgres checks, the dbt manifest, GitHub pull requests, and Dagster, Airflow, Prefect or Azure Data Factory runs, among 50+ connectors. Your team registers each agent or model once with its tables, and from then on lineage and blast radius reach it. See Inside the MLOps & Models Agent and the ML engineers guide.
What Data Workers writes, and where. To Comet, nothing. dbt fixes go to the model's owner as a diff to merge, and Data Workers opens the pull request when your team turns on the GitHub pull-request target. Reruns go through your orchestrator after approval; the owner runs any recorded undo. Approved facts land in the Context Wizard graph.
Setup today. Clone the open-source repository and add start-agent.sh entries to your client config, as the client setup docs show. Add the Opik MCP server next to them so one assistant sees traces and data operations; Data Workers' agents never call it. Opik Cloud users can use Comet's hosted server with browser sign-in instead of the local entry.
// Example: .mcp.json for Claude Code
{
"mcpServers": {
"dw-context-catalog": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-context-catalog"]
},
"dw-quality": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-quality"]
},
"dw-incidents": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-incidents"]
},
"opik-mcp": {
"type": "stdio",
"command": "uvx",
"args": ["opik-mcp"],
"env": { "OPIK_API_KEY": "<your-key>", "OPIK_WORKSPACE": "<your-workspace>" }
}
}
}List the tools with your client's own command (/mcp in Claude Code), then ask "Opik says the policy Test Suite dropped overnight. Which dbt models build kb.article_chunks, did any baseline on them break, and what else reads them?"
One request, L0 to L4. The ladder is set per domain.

- •L0 manual. Connected, not acting. The team finds the dedupe after the suite fails.
- •L1 observe. Data Workers flags each failed check or baseline break with its blast radius and owners, and logs what each action would have needed.
- •L2 propose. Data Workers drafts the dbt diff and the rebuild plan; owners approve before anything runs.
- •L3 act reversibly. For proven classes, such as queuing a rerun after a fix merges, Data Workers acts, verifies and records it.
- •L4 autonomous. For a scoped, trusted class in one domain, Data Workers runs the loop end to end and posts the receipt. Prompts, suites, re-indexing and retraining stay with the AI or ML owner at every level.
Read more on safety, how approvals work, the autonomy levels, who owns the agents and where your data goes.
What changes for your team

Teams on Comet know the moment a score drops. What costs them is the hour after: reading another team's dbt repo, the prompt patch that hides a data problem. With Data Workers, these jobs run on autopilot at the level you set.
- •Incidents. A failing eval arrives with its upstream cause, the owner and a proposed fix.
- •Data quality. Corpus and training tables get checked after every build, plus your baselines.
- •Cloud spend. Snowflake spend is tied to the dbt models that rebuild your corpus, through query tags.
- •Access. A request for training data arrives as a dry run, with the sensitive columns it would reach.
- •Audits. Each data fix links the alert, the approver, the diff and the undo path.
- •Migrations. Each wave is planned with its parity checks, and the owner signs off before it closes.
Running other ML tools too? See You're on MLflow, You're on Weights & Biases and You're on LangSmith. For the vector store side, see You're on Pinecone; for the orchestrator, You're on Dagster.
Keep Comet, or consolidate?
Keep Comet if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
Nearly every team keeps Comet: its experiments, registry, traces and Test Suites are where model and agent decisions get made. What teams consolidate is the tooling around it: hand-written asserts in embedding jobs, notebooks that join Opik exports to warehouse tables, spreadsheets mapping agents to dbt models, and the data observability seat the AI team never logs into. For background, see data pipelines for LLMs, agentic RAG for data engineering, MLflow alternatives, Data Workers vs data observability and what an agentic data platform is. Building this layer yourself? Read build it ourselves with Claude Code and MCP servers: reading a check is the easy part; approvals, rollback and receipts are the work.
The case for your CFO
The outcome: when an agent or model goes wrong because of its data, the cause is found and fixed through its owner the same morning, and nobody ships a prompt patch over a broken table. Opik tells you the score dropped; Data Workers makes sure the data under it is right again, with proof.
The risk story is plain. Data Workers changes nothing in Comet. dbt fixes arrive as diffs your engineers merge; reruns happen only after a named approver says yes; re-indexing and retraining stay with the owner. Each action leaves a receipt with the trigger, evidence, diff, approver, checks and undo path. Unanswered requests expire and escalate; they never auto-grant. No agent can promote its own work. Autonomy is set per domain from L0 manual to L4 autonomous, and an org-wide stop halts all autonomous dispatch. The agents run in your infrastructure; your data stays in your systems, and the hosted Conductor sees workflow metadata only. Zero migration.
Why now: every LLM-ops tool now ships an agent that proposes prompt and code fixes, and none of them owns the table underneath, so data problems get patched in the prompt. A support agent answering from the wrong corpus costs real money in escalations. The first win is one production agent registered with its tables, checked and baselined read-only, with a report of every upstream change that would have reached it. What stays the same: Comet, Opik, your rules, suites and owners. Start with a pilot; the path is on the pricing page, and the pilot is credited in full against the first year. For the numbers, see the ROI of agentic data operations.
The sentence to repeat upstairs: "Opik tells us when our agents get worse; Data Workers fixes the data behind it, with an approval and a receipt for every change."
Getting started
Start with a pilot. Pick the agent your customers feel first, register it with its tables in the context graph, connect Data Workers read-only to your warehouse, dbt project and orchestrator, record a baseline for the property the agent depends on (language mix, documents per category), and let it report what it would have caught before you turn on the first fix class. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.
FAQ
Does Data Workers read our Opik traces or Comet experiments? No. Comet stays the system of record, and evaluation stays with Opik; your team reads them there. Data Workers connects over Comet's API or MCP server today: your client runs the Opik MCP server next to Data Workers' MCP servers, so one assistant sees both. Data Workers works on the tables under the agent and the model.
Why didn't the quality check catch the translations? Because the table was complete. The null check passed (by default it allows up to 10% nulls; teams set a stricter threshold per check, here 1% on chunk_text), the chunk_id distinct ratio stayed above 99%, and the row count barely moved. The change was in the values, which is what a baseline on a metric your team cares about, here the English share, is for.
Should we approve Ollie's fix or Data Workers' fix? They answer different questions, and you keep both decisions. Ollie proposes changes to the agent's code and prompts; Data Workers proposes changes to the data it reads. An open incident on the agent's tables tells the AI owner to check the data before patching the prompt.
Can Data Workers re-index embeddings or retrain models? Those stay with the owner. In the example, the re-embed is a step in the team's own Dagster job: the owner approved and launched the run, and Data Workers followed it and re-checked the tables. Retraining stays with the ML owner.
Does this help classic ML runs in Comet Experiment Management? Yes. Register the model with its training tables; when a run regresses after a refresh, Data Workers traces the upstream dbt change and proposes the fix, and the ML engineer reruns.
Where does our data go? The agents run in your infrastructure and hold the credentials. Your data stays in your systems; the hosted Conductor sees workflow metadata only, such as table names, proposals and approvals, never rows, credentials or model keys.
Sources
- •Comet, home page, https://www.comet.com/site/ (checked Oct 3, 2026)
- •Comet, ML Experiment Tracking, https://www.comet.com/site/products/ml-experiment-tracking/ (checked Oct 3, 2026)
- •Opik docs, Online Evaluation rules, https://www.comet.com/docs/opik/production/online-evaluation/rules (checked Oct 3, 2026)
- •Opik docs, Evaluation Concepts (Test Suites), https://www.comet.com/docs/opik/evaluation/concepts (checked Oct 3, 2026)
- •Opik docs, Ollie, https://www.comet.com/docs/opik/ollie (checked Oct 3, 2026)
- •Opik docs, Diagnostics, https://www.comet.com/docs/opik/tracing/diagnostics (checked Oct 3, 2026)
- •Opik docs, MCP tools reference, https://www.comet.com/docs/opik/mcp-server/tools (checked Oct 3, 2026)
- •Opik docs, MCP advanced setup, https://www.comet.com/docs/opik/mcp-server/advanced-setup (checked Oct 3, 2026)
- •Opik docs index (Agent Optimizer, Cost Intelligence), https://www.comet.com/docs/opik/llms.txt (checked Oct 3, 2026)
- •Opik MCP server repository, release 0.2.38 (Sep 30, 2026), https://github.com/comet-ml/opik-mcp (checked Oct 3, 2026)
- •Opik repository, Apache-2.0, release 2.2.88 (Oct 2, 2026), https://github.com/comet-ml/opik (checked Oct 3, 2026)
- •Zendesk Help Center API, Articles, https://developer.zendesk.com/api-reference/help_center/help-center-api/articles/ (checked Oct 3, 2026)
- •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)
- •Data Workers open-source repository, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
- •Data Workers pricing, https://dataworkers.io/pricing/ (checked Oct 3, 2026)