Product
Product11 min readBy The Data Workers Team

You're on Anomalo: It Finds What No Rule Caught. Data Workers Fixes It and Proves It

Anomalo detects data quality issues with unsupervised ML. Data Workers diagnoses across systems, fixes behind approvals, verifies downstream and leaves a receipt.

Your data team runs on Anomalo, "the autonomous data system for the agentic enterprise." Unsupervised checks learn the shape of every table and flag what changed without anyone writing a rule first. The Table Observability agent watches availability, freshness and schema. When a check fails, root-cause analysis points to the segment where the change sits, with sample rows. Since the April 2026 "Self-Driving Data" relaunch, Table Observability, Data Quality and AIDA, the conversational analyst, are generally available, and the Insights and Documentation agents, in preview at launch, are now marked live on Anomalo's homepage. Unstructured monitoring scores documents for PII and erasure requests. On Databricks, findings surface in Unity Catalog and customers can buy Anomalo with their existing Databricks commitments (June 16, 2026); on Snowflake, Anomalo runs as a Native App or Connected App, payable from committed Snowflake capacity (June 9, 2026). Anomalo is the smoke alarm that learns the room. Data Workers is the crew.

Anomalo detects quality issues with ML. Data Workers diagnoses, fixes and verifies behind approvals. Anomalo's next agent, the Data Issue First Responder, is listed as coming soon; its documented job is to follow your runbooks and start workflows in ServiceNow and Jira. Somebody still has to find the cause upstream, change the model, rebuild, confirm the numbers and write it down. That is Data Workers' job.

Key takeaways

  • •Anomalo keeps its job. Unsupervised checks, Table Observability, root-cause segments, AIDA and unstructured monitoring stay exactly where they are.
  • •A finding becomes a governed fix. Anomalo's finding reaches Data Workers through the on-call, whose assistant can use Anomalo's MCP server side by side with Data Workers; Data Workers traces the failing segment to its cause across Segment, dbt, Airflow and Snowflake, and proposes the full repair with its blast radius.
  • •A named owner approves; Data Workers verifies. The domain owner approves in Spellbook, Data Workers queues the rebuild through your orchestrator, checks quality, totals and load lag against its baseline downstream, and records a receipt.
  • •The ticket carries the evidence. Data Workers opens the Jira Service Management or ServiceNow ticket with the receipt link, or comments it on the Jira Service Management ticket your team or First Responder opened; the service agent closes it.
  • •Autonomy per domain. Each data domain climbs from L0 manual to L4 autonomous at the pace your team sets.

Anomalo is the smoke alarm. Data Workers is the crew.

Here is a Tuesday with both in place, across Segment, Snowflake, dbt run by Airflow, Census and Braze. Snowflake, dbt and Airflow are native connections for Data Workers; Segment, Census and Braze connect over their APIs or MCP servers today. This is an illustration, not a customer case.

TimeSystemWhat happens
Mon 17:40SegmentA web release changes the consent banner; campaign fields stop reaching track events from iOS Safari, though the raw page URL still carries the UTM parameters
Tue 02:00SnowflakeThe Segment warehouse load lands Monday's events with utm_source null for that segment
04:00Airflow + dbtThe nightly DAG run rebuilds stg_sessions and the incremental fct_attribution; paid social sessions on iOS Safari fall into "direct"
06:00CensusThe scheduled sync sends a paid-social retargeting audience to Braze, short by the sessions attribution lost
07:30AnomaloThe unsupervised check on events.sessions flags 14% of rows with a null utm_source, concentrated in device = iOS Safari, with sample rows
07:34Data WorkersTakes the finding the on-call's assistant reads from Anomalo, traces lineage from events.sessions through fct_attribution to the Census sync the team recorded in the context graph, and sees the Braze campaign that uses the audience sends at 11:00
07:50Data WorkersProposes the repair: a diff to stg_sessions that falls back to the UTM parameters in the page URL, a test on the utm_source null rate per device, a rebuild of two days of fct_attribution partitions, and a note for the web team on the release
08:20SpellbookThe marketing data owner reviews the diff, the blast radius and the rebuild plan, approves and merges the diff
08:35Airflow + dbtData Workers queues the dbt rebuild of the affected partitions through Airflow; the new test passes
09:05SnowflakeData Workers verifies the null rate and load lag against their baselines; the owner's query matches paid-session totals to the raw page URLs
10:00Census + BrazeThe scheduled sync sends the full audience before the 11:00 send
10:30AnomaloThe next check run passes; the receipt is linked on the Jira Service Management ticket
Incident timeline across the stack: what Anomalo, your team and Data Workers each do, step by step

Anomalo did what it is built to do: it found a change nobody had written a rule for, on one device segment of one column. What changed is everything after the finding. The cause sat in a staging model two hops upstream. The fact that mattered most, a Braze campaign about to send to a short audience, was in the proposal. A named owner approved one change set, and the campaign went out on the right audience.

JobWhat Anomalo doesWhat Data Workers does
DetectionUnsupervised checks learn each table and flag availability, freshness, schema and content changes without rulesWatches freshness, volume and schema itself, so many breaks are caught before a check runs
Root causePoints to the segment where the change sits, with sample rowsJoins that segment to dbt code, DAG runs, recent releases and incident history to find what changed and what else it touched
The fixAIDA explains; First Responder (coming soon) will follow your runbooks and start ServiceNow or Jira workflowsProposes the complete repair (dbt diff, test, partition rebuild, upstream note) with blast radius, routed to a named owner
Running itYour team changes the model and reruns the jobsThe owner approves and merges in Spellbook; Data Workers queues the rebuild through your orchestrator
VerificationThe next check run passesChecks freshness, quality, volume and totals on every downstream model the fix touched
The recordThe failed check, its history and the segment viewA receipt: the cause, the diff, who approved it, what it touched, how it was verified and how to undo it

Why doesn't Anomalo just do this itself?

Because Anomalo made a sensible choice about where its product ends. Its strength is breadth: it watches hundreds of tables with no rules because it only reads them. Its MCP server, shipped as a Gemini CLI extension in December 2025, reads failing checks and table status, and "Gemini can only access metadata and data quality summaries for which the authenticated user already has permissions." Even the agent Anomalo describes for the next step keeps to routing: First Responder will "assess impact and criticality and then follows any established runbooks or policies", "initiating workflows in tools like ServiceNow and JIRA and appropriately escalating to human team members." Its unstructured workflows that redact PII or drop conflicting documents change a dataset Anomalo prepares, not your production models.

That is the right line for a monitoring product. Changing a dbt model that feeds attribution and rebuilding incremental partitions before a campaign sends is a different product with a different liability. The change has to be checked against every model, exposure and sync downstream, approved by the domain owner, given an undo before it runs, rebuilt through your orchestrator and proven afterwards. Anomalo would have to own changes in dbt, Airflow and Snowflake, systems it watches but does not run. Data Workers is built for exactly that job.

Every tool owns a slice. Data Workers covers the whole lifecycle

Anomalo owns detection and goes deep on rule-free checks. Each point tool adds another console, another contract and another handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Anomalo goes deep on its own area
StageData WorkersAnomaloWhy we scored it this way
Catalog & Context94The Documentation agent drafts dataset descriptions and Unity Catalog shows Anomalo's findings. Data Workers keeps one governed context graph of definitions, owners, lineage, quality and usage across every platform.
Analytics & Insights86AIDA answers data questions in plain language. Data Workers answers business questions from governed definitions with lineage behind every number.
Data Quality89Anomalo's home stage: unsupervised checks learn each table's shape and flag what nobody wrote a rule for, down to the segment. Data Workers also writes, runs and repairs the checks and dbt tests behind them.
Observability & Incidents8.58.5Table Observability covers availability, freshness and schema, and the root-cause view points to the segment that moved. Data Workers diagnoses across systems, fixes with approval and verifies the fix.
Pipelines & Ingestion8.53Anomalo watches the tables pipelines write. Data Workers queues reruns and backfills through your orchestrator, with approvals.
Schema & Migration83Table Observability flags schema changes after they land. Data Workers catches schema changes, scores their blast radius and plans migrations in waves.
Governance & Access8.53Findings respect each user's permissions. Data Workers runs the access request queue: it dry-runs each grant, routes it to a named approver and records its expiry.
Security & Privacy85Unstructured checks flag PII and erasure requests in documents. Data Workers' pull request review flags new columns whose names or annotations look sensitive, and leaves a receipt on every data change.
Cost / FinOps82Anomalo does not manage warehouse spend. Data Workers traces Snowflake credits to the dbt model behind them and drafts the fix for its owner.
MLOps & Models7.54Unstructured monitoring readies documents for AI, and Experiment Evaluation is listed as coming soon. Data Workers keeps the data under your models fresh and correct.

Anomalo leads on Data Quality and ties on observability, as it should. For the side-by-side view across the category, read Bigeye, Anomalo, Sifflet and Soda vs Data Workers and Data Workers vs data observability; for the head-to-head questions buyers ask, see Data Workers vs Anomalo and Anomalo alternatives for data observability.

How Anomalo and Data Workers work together

Anomalo stays on top, where checks fail and AIDA explains. Spellbook Data Catalog (in preview) is where the data team looks: each proposed change, who approved it and how to roll it back. Between them, Data Context Wizard keeps one governed context graph across the estate, the Data-Agents Swarm does the work with 20+ specialist agents, and the Autonomous Data-Conductor runs each fix end to end: detect, diagnose, fix, review, verify, remember.

How Data Workers fits with Anomalo: your coding agent on top, Data Workers in the middle, your estate underneath

Anomalo to Data Workers. Anomalo connects over its API or MCP server today: your team's assistant reads failed checks, the segments behind them and table status side by side with Data Workers, whose own checks rest on Snowflake, dbt and Airflow. Snowflake, dbt and Airflow connect natively; Segment, Census and Braze over their APIs. A new finding starts a diagnosis: diagnose_incident and get_root_cause work through the evidence, trace_cross_platform_lineage and blast_radius_analysis map every model downstream and every sync the team has recorded, and get_incident_history checks whether the same break has happened before. The repair runs through remediate, with code changes proposed as a diff for the owner to merge and reruns queued through your orchestrator. Verification uses run_quality_check, monitor_metrics and get_quality_score, and every step lands in get_audit_trail.

Data Workers back to your team. Anomalo's MCP server is built for reading, so the receipt lives in Spellbook and the audit trail, linked from the ticket: Data Workers opens a Jira Service Management or ServiceNow ticket for the domain with the link, or comments it on the Jira Service Management ticket your team or First Responder opened. The service agent closes the ticket.

Side by side in one client. Anomalo ships its MCP server as a Gemini CLI extension; add Data Workers beside it. The client setup guide documents the path: clone the open-source repo and add one start-agent.sh entry per agent. The Gemini CLI guide covers the policy file that makes remediate always ask.

# Example: Anomalo's Gemini CLI extension, with your instance and API key in ~/.gemini/.env
echo 'ANOMALO_INSTANCE_HOST=https://<your-instance>.anomalo.com/' >> ~/.gemini/.env
echo 'ANOMALO_API_SECRET_TOKEN=<your-anomalo-api-key>' >> ~/.gemini/.env
gemini extensions install https://github.com/datagravity-ai/anomalo-gemini-extension --auto-update
{
  "mcpServers": {
    "dw-context-catalog": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-context-catalog"]
    },
    "dw-incidents": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-incidents"]
    },
    "dw-quality": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-quality"]
    }
  }
}

The second block is an example ~/.gemini/settings.json. Ask "why is utm_source null on iOS this morning, and what does the fix touch?" and the client calls both: Anomalo returns the failing check and the table status; Data Workers returns the cause in the staging model, the downstream attribution partitions and the Census sync to Braze, and a proposed repair with its blast radius. Anomalo's key scopes what the client can read in Anomalo; Data Workers' guardrail decides whether a change may run in that domain, with a named approver. For a shared endpoint, the Data Workers remote server serves /mcp with an API key (bearer) or OAuth tokens from your identity provider, verified through JWKS.

One Anomalo finding, L0 to L4, set per domain.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. An engineer reads the segment view, finds the staging model, edits it and reruns the DAG by hand.
  • •L1 observe. Data Workers posts the diagnosis and the Braze audience at risk. Nothing changes.
  • •L2 propose. Data Workers proposes the diff, the test and the rebuild. Nothing runs until the marketing data owner approves in Spellbook.
  • •L3 act reversibly. For change classes with a proven record, such as rerunning a failed load or rebuilding a partition, Data Workers queues the step through the orchestrator, verifies it and records the receipt.
  • •L4 autonomous. For a scoped domain with a long clean record, Data Workers handles the repeatable steps end to end and posts the receipt; code changes still go to the owner as a diff.

For the safety model, read is it safe to let AI agents change production data, how approvals work and how to roll back an AI agent change. On where data lives: the agents run in your infrastructure, your data stays in your systems, and the hosted Conductor sees workflow metadata only.

What changes for your team

Anomalo made every table watched. Data Workers gives that team a crew.

Six jobs that run on autopilot with Data Workers next to Anomalo, with a concrete example of each
  • •Incidents. An Anomalo finding arrives with a diagnosis, a blast radius and a proposed fix, and closes with a receipt.
  • •Data quality. Each fix leaves a dbt test on the model that broke, so the repeat is caught before the next check run.
  • •Cloud spend. Snowflake credits are traced to the dbt model behind them, and each fix goes to its owner drafted.
  • •Access. A request for warehouse access becomes a dry-run, time-boxed grant proposal the data owner approves.
  • •Audits. Anomalo records the failed check; Data Workers records what changed, who approved it and how to undo it.
  • •Migrations. A move to a new warehouse or lakehouse runs in approved waves, with parity checks planned and tracked for each wave and Anomalo watching both sides.

The analytics engineer on call changes most: they review proposals instead of hunting for the staging model. See the data incident response playbook, beyond data observability: autonomous resolution and who owns the agents.

Keep Anomalo, or consolidate?

Keep Anomalo if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

Most teams keep Anomalo where rule-free coverage on hundreds of tables and unstructured document checks matter. What they consolidate is the stack around it: a second quality tool alerting on the same break, hand-written backfill scripts and a runbook wiki. If you are weighing building the fix layer yourself on Anomalo's MCP server and a coding agent, read build it ourselves with Claude Code and MCP servers: the MCP calls are the easy part; the context graph, approvals and rollback are the work. The same pattern holds across the category: see you're on Monte Carlo, you're on Bigeye, you're on Soda, Data Workers integrations and what is an agentic data platform.

The case for your CFO

The outcome: you already pay Anomalo to find bad data before the business does. Data Workers turns each finding from a morning of tracing and hand-run backfills into a diagnosis on arrival and a complete fix with one approval, so attribution and campaign audiences are right before anyone acts on them.

The risk story is plain. Autonomy is set per domain: at L1 Data Workers only reads; at L2 it changes nothing until a named person approves; at L3 it applies changes it can undo. Code changes go to the model's owner as a diff, every time. An unanswered approval request expires and escalates, never auto-grants. No agent can promote its own work. Every change carries a receipt: the cause, the diff, who approved it, what it touched downstream, how it was verified and how to undo it. An org-wide stop halts all autonomous dispatch. Zero migration: Anomalo, Segment, Snowflake, dbt, Airflow, Census and Braze stay where they are.

Why now: Anomalo already finds the change, and its own next agent routes the work to a ticket. The slow part is now fixing, approving, rebuilding and proving the fix across systems. The first win is read-only: every Anomalo finding in one domain, such as marketing attribution, gets a diagnosis and a blast radius. What stays the same: your checks, notification routing, warehouse permissions and dbt review. For the numbers, see the ROI of agentic data operations.

The sentence to repeat upstairs: "Anomalo tells us the data changed and where; Data Workers fixes the cause with an approval and proves it held."

Getting started

Start with a pilot. Pick one domain whose findings already live in Anomalo, such as marketing attribution, connect Data Workers to Anomalo, Snowflake, dbt and Airflow, and run at L1 so every finding gets a diagnosis and a blast radius. Then turn on partition rebuilds and reruns at L2 with owner approval. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.

FAQ

How does Data Workers connect to Anomalo? Over its API or MCP server today, reading failed checks, the segments behind them and table status. You can also run Anomalo's Gemini CLI extension beside Data Workers and ask both in one conversation.

Anomalo's Data Issue First Responder is coming. Do we still need Data Workers? Anomalo describes First Responder as following your runbooks and starting workflows in ServiceNow and Jira, with escalation to people. That routes the work. Data Workers does the work the ticket describes: the cause, the diff, the rebuild, the verification and the receipt.

Anomalo shows the root-cause segment. What does Data Workers add? The segment says where the change sits. Data Workers follows it upstream into dbt code, DAG runs and releases to find why, and downstream to find what else it touched.

Does Data Workers change anything in Anomalo? No. Checks, thresholds, notification routing and unstructured workflows stay under your Anomalo admins. Data Workers reads Anomalo and keeps its own receipt in Spellbook, linked from the ticket.

We run Anomalo inside Snowflake or with Unity Catalog on Databricks. Does that change the setup? No. Data Workers connects natively to Snowflake and to Databricks with Unity Catalog, and its agents run in your infrastructure. Anomalo keeps surfacing findings where it does today, whether you pay for it from Snowflake capacity or Databricks commitments, and the on-call brings each finding to Data Workers, with Anomalo's MCP server in the same client.

What about unstructured data? Document quality is Anomalo's ground, and its unstructured checks stay as they are. Data Workers keeps the structured tables and pipelines around those documents healthy.

Sources

  • •Anomalo, homepage ("The autonomous data system for the agentic enterprise"; Live and Coming Soon agent labels), https://www.anomalo.com/ (checked Oct 3, 2026)
  • •Anomalo, "Introducing Self-Driving Data and the new Anomalo agentic platform" (Apr 2, 2026; agent statuses, First Responder), https://www.anomalo.com/blog/introducing-self-driving-data-and-the-new-anomalo-agentic-platform/ (checked Oct 2, 2026)
  • •Anomalo, "Self-Driving Data for the Databricks Data Intelligence Platform" (Jun 16, 2026; Unity Catalog integration), https://www.anomalo.com/blog/self-driving-data-for-the-databricks-data-intelligence-platform/ (checked Oct 2, 2026)
  • •Anomalo, "Databricks Customers Can Now Purchase Anomalo Using Their Existing Databricks Commitments" (Jun 16, 2026), https://www.anomalo.com/blog/databricks-customers-can-now-purchase-anomalo-using-their-existing-databricks-commitments/ (checked Oct 3, 2026)
  • •Anomalo, "Anomalo is Now MCD-Eligible on Snowflake Marketplace" (Jun 9, 2026; Native App and Connected App, committed capacity), https://www.anomalo.com/blog/anomalo-is-now-mcd-eligible-on-snowflake-marketplace/ (checked Oct 3, 2026)
  • •Anomalo, Unstructured data monitoring, https://www.anomalo.com/unstructured/ (checked Oct 2, 2026)
  • •Anomalo, "Anomalo announces Gemini CLI extension" (Dec 9, 2025; MCP server, permissions), https://www.anomalo.com/blog/anomalo-announces-gemini-cli-extension-data-quality-intelligence-meets-ai-enabled-development/ (checked Oct 2, 2026)
  • •Anomalo Gemini CLI extension repository, https://github.com/datagravity-ai/anomalo-gemini-extension (checked Oct 2, 2026)
  • •Data Workers open-source repository, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)
  • •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)