Product
Product9 min readBy The Data Workers Team

How is Data Workers different from data observability?

Data observability detects and alerts. Data Workers resolves: it detects, diagnoses, fixes through your approvals, verifies the fix held and leaves a receipt, across pipelines, quality, schema, cost and access.

Data observability detects and alerts: monitors, anomaly detection, lineage-based impact and, increasingly, AI agents that triage alerts and suggest or draft fixes. Data Workers resolves: it detects too, then diagnoses the cause, fixes it at the source through your approvals, verifies the fix held and leaves a receipt, across pipelines, quality, schema, cost and access.

Observability is the smoke alarm. Data Workers is the crew, and it works with the observability tool you already run from day one.

Key takeaways

  • •Detection and resolution are different jobs. Observability owns detection and triage, and the best tools do it very well. The work after the alert is where Data Workers starts.
  • •Observability vendors now ship agents, and several draft or execute fixes. Monte Carlo's Remediation skill executes fixes in an engineer's session, Elementary opens pull requests, Bigeye drafts remediation SQL and Acceldata applies fixes after human approval. Each keeps a person confirming every step, which is the right design for a monitoring product.
  • •Data Workers runs the whole loop across every tool. One context, one approval flow and one audit trail, with autonomy set per domain from L0 to L4.
  • •Every fix is verified and leaves a receipt. The trigger, the cause, the diff, the blast radius, the approver, the checks with before and after values, and the undo path.
  • •Zero migration. Your monitors, warehouse, dbt project and orchestrator stay where they are.

What observability does, and does well

Observability tools watch the estate. They learn what normal looks like for freshness, volume, schema and field health, alert when something drifts, and use lineage to show what sits downstream. This year every major vendor added AI on top.

Monte Carlo's Agentic Operations section is the most complete. The Operations agent is the chat front door and runs with the user's own permissions. It routes to the Troubleshooting agent, which works through hundreds of hypotheses across lineage, dbt and Airflow failures and code changes (its page lists preview caveats), the Triage agent, which scores every new alert HIGH, MEDIUM or LOW by default, and the Cost and PII agents (PII in preview). Its MCP server lets an assistant investigate alerts, explore lineage and create and manage monitors. Its open-source Agent Toolkit adds a Remediation skill that "Proposes and executes fixes for data-quality alerts after assessing blast radius, or escalates with full context when uncertain," with sensible rules: present the plan first, never modify data without explicit confirmation and a rollback plan, never mark an alert fixed before verifying.

The rest of the field has moved the same way. Bigeye's bigAI (public preview) writes Suggested Resolutions, and its AI Remediation SQL (beta) drafts an UPDATE or DELETE for the affected rows; Bigeye "generates the SQL but does not execute it on your behalf." Sifflet's Sage finds the root cause and Forge recommends next steps and routes the incident. Soda's root cause analysis agent (private preview) "only proposes fixes." Anomalo lists its Data Issue First Responder as coming soon. Elementary's Triage & Resolution agent opens pull requests with proposed fixes, and "Every action the agent takes goes through a pull request or requires explicit confirmation." Acceldata's Incident Management agent proposes a remediation that is applied with "human approval before any fix is applied" and a full audit trail. Datadog, which acquired Metaplane in April 2025, pairs quality monitoring with jobs monitoring for Airflow, dbt, Spark and Databricks and routes each issue to the right owner.

That is a strong detection category, and each design choice above is right for a monitoring product. Our roundup of Bigeye, Anomalo, Sifflet and Soda and the Monte Carlo comparison go tool by tool. Our guide to what data observability is covers the basics.

What Data Workers does after the alert

Data Workers is the agentic data platform. It reads your observability alerts as signals and runs the loop from there.

Detect. Data Workers reads alerts from Monte Carlo and Soda through its connectors; Bigeye, Anomalo, Elementary and other tools connect over their API or MCP server today, and your team's assistant can use those servers side by side with Data Workers. It also runs its own checks: run_quality_check, get_quality_score and monitor_metrics.

Diagnose. The incident agent runs diagnose_incident and get_root_cause, the context graph traces the cause across ingestion, dbt, the orchestrator and the warehouse with trace_cross_platform_lineage, and get_incident_history shows whether it has happened before.

Scope. Before any change, blast_radius_analysis lists every model, DAG, dashboard and export the fix touches, at column level.

Fix through approvals. The fix lands at the source: a dbt pull request, a DAG rerun, a backfill, a grant, a cleanup. remediate runs a known playbook only above a high confidence bar and routes anything new to a person. Your team sets autonomy per domain on the ladder L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous. At L2, a named owner approves each change in Spellbook Data Catalog (in preview) or on the pull request. Unanswered requests expire and escalate; they never auto-grant. How approvals work has the detail.

Verify. After the change, Data Workers re-runs the assertion set for the affected tables (quality, contract, volume and schema, plus the lateness baseline) and clears the incident only when the checks pass.

Leave a receipt. Each action is written to a tamper-evident audit trail (a SHA-256 hash chain, readable with get_audit_trail) with its trigger, diff, blast radius, approver, checks and undo path. How to roll back an AI agent change shows the undo side.

The same loop covers more than incidents. The cost agent traces Snowflake credits to the query and dbt model behind them and drafts the fix for the owner, and the governance agent runs provision_access for least-privilege, column-level grants with a 90-day expiry, and check_policy.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous

One alert, taken to a verified fix

An illustration, not a customer case. The estate runs Postgres, Fivetran, BigQuery, dbt on Airflow, Datadog Data Observability and Looker. The logistics domain is set to L2 propose.

TimeSystemWhat happens
Tue 18:20PostgresAn app release starts storing weight_kg as text, such as "12.5 kg"
18:45FivetranThe sync lands weight_kg as STRING in raw.shipments
Wed 02:00Airflow + dbtThe shipments_daily DAG run fails casting "12.5 kg"; fct_shipments is not refreshed
05:00DatadogA freshness alert fires on fct_shipments; lineage points to the failed job
05:01Data WorkersReads the alert and opens an incident
05:04Data WorkersDiagnosis: a column type change in Postgres, not an outage
05:05Data WorkersBlast radius: fct_shipments, two Looker Explores and the carrier cost export
05:08Data WorkersDrafts a dbt pull request that parses the unit and adds a test; records the undo path (revert the PR, rebuild)
05:09SpellbookRoutes the approval to the logistics owner
07:40Analytics engineerReviews the diff, blast radius and test run, and approves
07:44Airflow + dbtThe PR merges and a DAG run rebuilds fct_shipments
08:02Data WorkersVerifies: the table is fresh, and row counts and weights match the source; the engineer confirms dbt tests pass
08:03DatadogThe freshness monitor recovers on its next check
08:04SpellbookReceipt: alert, cause, diff, approver, checks, undo path, and the app team that owns the Postgres table
08:30LookerOperations opens correct shipping costs
Incident timeline across the stack: what Observability tools, your team and Data Workers each do, step by step

Datadog did its job at 05:00, and did it well: it caught the stale table and showed the failed job. The next three hours used to belong to whoever was on call: the DAG log, the cast error, the trace to Fivetran and Postgres, the fix, the review, the rerun. With Data Workers, the on-call engineer spends that time on one approval.

Observability vs Data Workers, step by step

Comparison matrix of Observability tools and Data Workers on the outcomes a data leader buys
StepObservability toolsData Workers
DetectTheir home job: monitors, ML thresholds, checks and contractsReads their alerts and runs its own quality, schema and metric checks
TriageMonte Carlo's Triage agent scores every alert; others group and scoreScores by blast radius across every downstream system
Root causeTroubleshooting and RCA agents over lineage and logsTraced across ingestion, dbt, orchestration and the warehouse
Fix at the sourceProposals, per-session skills, pull requests, drafted SQLdbt PR, rerun, backfill, grant or cleanup, behind approval
ApproveA person confirms each fix in a session, a PR or the vendor's appA named owner, with rules set per domain, L0 to L4
VerifyThe monitor or check reruns; Monte Carlo's skill re-checks before marking fixedBefore and after checks on every affected table
ReceiptsAlert comments and incident historyHash-chained receipt with diff, approver, checks and undo
CostRecommendations (Monte Carlo: "recommends; it does not act")Traces Snowflake credits to the dbt model behind them and drafts the fix for the owner
AccessAccess to the tool; Monte Carlo's PII agent in previewLeast-privilege grants with expiry, and current classification

Why doesn't the observability tool just do this itself?

Focus and risk. An observability tool reads almost everything you run, from metadata, metrics and query logs, which makes it easy to approve: its blast radius is close to zero. Its own writes stay mostly inside its own objects, such as alerts, comments and monitors.

Writing to production data across systems is a different product category. It needs blast-radius scoping before every change, approvals that vary by domain, rollback, a receipt an auditor can read, context about every other system, and liability for changes inside tools the vendor doesn't own: your dbt repo, your Airflow deployment, your warehouse grants, your Fivetran connectors. The vendors that go furthest keep it close: Monte Carlo's Remediation skill runs in one engineer's coding-agent session with that engineer's confirmation, Elementary's fixes arrive as pull requests, Bigeye's SQL is copied and run by a person, and Acceldata applies a fix after a person approves it. Those are sensible choices for a detection product. The cross-system write path, per-domain autonomy and verified receipts are the product Data Workers is built around.

Keep your observability tool, or consolidate?

Keep your observability tool if you love it; Data Workers works with it from day one. Its monitors stay, its alerts route where they route today, and Data Workers writes its diagnosis and receipt back where the tool allows it, as the Monte Carlo integration shows. The Monte Carlo playbook walks the climb stage by stage.

Many teams consolidate once Data Workers runs that slice too. Its own checks cover freshness, quality, schema and metrics, and you can make that choice domain by domain. Each point tool adds another console, contract and handoff; Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail. The same logic applies to catalogs (Data Workers vs a data catalog); the full definition is in what is an agentic data platform.

The case for your CFO

The outcome: alerts end in verified fixes instead of morning tickets. Stale tables are rebuilt before the business opens its dashboards, and the hours engineers spent tracing alerts through four tools go back to the data products the business asked for.

The risk story: agents act within limits your team sets per domain. At L1 they read; at L2 a named owner approves; at L3 they take only reversible actions. Every change carries its blast radius, checks and undo path, no agent can promote its own work, and an org-wide stop halts all autonomous dispatch. Agents run in your infrastructure; your data stays in your systems, and the hosted Conductor sees workflow metadata only. The safety page and the security and deployment page cover the detail.

Why now: observability vendors and coding agents are already moving toward remediation, one session at a time. The choice is between per-engineer fixes and one governed loop.

The first win: take the top recurring alert class from your observability tool to verified fixes at L2 during the pilot. What stays the same: your monitors, warehouse, dbt project, orchestrator and dashboards.

Start with a pilot: $7,500 one-time, and the pilot is credited in full against the first year. After it, Scale starts from $1,000/month and Enterprise from $3,000/month, billed annually, with unlimited seats, no usage meter, no markup on model spend, bring your own model, and an Apache 2.0 core. See pricing, the ROI calculator and the ROI page. If you're weighing a build, read build it ourselves with Claude Code.

The sentence to repeat upstairs: "Our observability tool tells us what broke; Data Workers fixes it through our approvals, proves the fix held and leaves a receipt."

FAQ

Is Data Workers a data observability tool? No. Data Workers detects issues, but its job is resolution: diagnosing, fixing through approvals, verifying and recording. It reads your observability tool's alerts as signals.

Monte Carlo now has a Remediation skill. Isn't that the same thing? It is a real step: it executes fixes inside one engineer's coding agent, with that engineer's confirmation for each action. Data Workers runs one loop for the whole team across every system, with approvals set per domain, verification and a receipt on every change. See Monte Carlo vs Data Workers.

Will Data Workers create more alerts? It reduces the work behind them. Data Workers scores each alert by blast radius and fixes the cause at the source, so the same class of alert has a reason to stop recurring.

Does Data Workers need write access to our observability tool? No. It can start read-only. Where the tool allows it, Data Workers can also post its diagnosis and receipt on the alert, where people already look.

Which observability tools does it work with? Data Workers reads Monte Carlo, Soda and Datadog through native connectors; Bigeye, Anomalo, Elementary, Sifflet, Acceldata and others connect over their API or MCP server today, used by your team's assistant side by side with Data Workers.

Can we try it before buying? Yes. The Apache 2.0 core runs in your coding agent on your own model key; see the open-source docs and our guide to moving beyond observability to autonomous resolution.

Sources