Product
Product11 min readBy The Data Workers Team

You're on New Relic: It Watches the Stream. Data Workers Fixes the Data Behind the Alert

New Relic and Autopilot catch Kafka lag and the deploy behind it. Data Workers finds the data the lag left wrong, fixes it with approval and proves the fix.

Your engineers live in New Relic. Services, hosts and Kafka clusters are entities. Alert conditions written in NRQL open issues, and since February 16, 2026 the notifications people act on are called alert events (the old "Incidents"). Queues and Streams shows consumer lag per topic, Change Tracking pins each deploy to the chart, and New Relic Autopilot (formerly SRE Agent, announced July 21, 2026) triages the alert, finds the likely cause and recommends the next step, right in a Slack thread. New Relic is where application and infrastructure telemetry is watched and alerts fire. Data Workers is where data incidents get diagnosed, fixed and verified.

New Relic will tell you that consumer lag on orders-cdc crossed 50,000 and that the 14:02 deploy caused it. The warehouse table behind the stream is a different system, and an hour of orders can go missing there after the lag clears and every service is green. Data Workers finds that gap, repairs it with a named approver, checks the numbers and leaves a receipt on the alert event.

Key takeaways

  • •New Relic keeps its job. Alert conditions, alert events, Queues and Streams, Change Tracking and Autopilot stay where they are.
  • •A green service can hide a wrong table. When lag drains, a streaming job can drop late events. An hourly row-count baseline in Data Workers catches the gap in the table itself.
  • •Native for alert events. Over NerdGraph, Data Workers raises an alert event when it finds data impact and closes it once the fix is verified.
  • •Two MCP servers, one client. Run Data Workers beside the New Relic MCP server (part of New Relic Ground Truth, launched September 15, 2026) in Claude Code or Codex: New Relic answers what fired, Data Workers fixes what broke in the data.
  • •Autonomy per domain. Each domain climbs from L0 manual to L4 autonomous at your pace, with approvals, rollback and receipts.

New Relic is the smoke alarm. Data Workers is the crew.

New Relic sees the symptom in your infrastructure. Data Workers fixes the data behind it and proves the fix. Here is a Wednesday across an order service, Kafka, Flink, Snowflake, New Relic and ThoughtSpot. This is an illustration, not a customer case.

TimeSystemWhat happens
Wed 14:02Order serviceDeploy v4.18 adds a larger line-items payload to every order event; New Relic records the change event
14:19New RelicAn NRQL alert condition on orders-cdc consumer lag (over 50,000) opens an issue; the alert event pages the on-call SRE in Slack
14:24New Relic AutopilotTies the lag to deploy v4.18 using change events and the Flink job's metrics, and recommends a rollback for review
14:31Order serviceThe SRE approves; the rollback runs
15:12New RelicLag drains to zero; the issue closes. Every service is green
15:40Data WorkersThe hourly row-count baseline the team records with monitor_metrics flags analytics.orders_hourly far below its expected volume for 14:00 to 15:00, and the team's count of raw.orders_cdc puts the gap at 23,418 events: Flink's event-time windows dropped events that arrived after allowed lateness while the backlog drained
15:46Data WorkersMaps the blast radius (two Snowflake tables, the ThoughtSpot "Orders today" Liveboard and the hourly finance export) and raises an alert event in New Relic that links the data impact to the closed lag issue
15:52Data WorkersProposes the fix: a backfill of the missing hour from raw.orders_cdc with an undo step recorded, plus a diff to the orders-enrich Flink job that routes late events to a side output, for its owner to merge
16:05SpellbookThe commerce data owner reviews the row counts and approves the backfill
16:09SnowflakeData Workers starts the backfill through the orchestrator; the undo step, recorded before the run, is saved with the receipt
16:15ThoughtSpotData Workers re-checks the Liveboard's tables: the hourly row count is back on its baseline and the quality checks pass
16:18New RelicData Workers closes the alert event it raised, with the receipt linked
Incident timeline across the stack: what New Relic, your team and Data Workers each do, step by step

New Relic did everything it was built to do: it alerted on the lag at 14:19, Autopilot named the cause, and the rollback ended the infrastructure incident. The data incident started where that one ended. Without Data Workers, the first sign would have been a finance analyst asking why Wednesday's 14:00 revenue dipped. With it, the gap was fixed and verified before anyone read the hour.

JobWhat New Relic doesWhat Data Workers does
DetectionAlert conditions in NRQL watch lag, throughput, errors and latency across entitiesChecks the landed tables and tracks rows per load and load lag against baselines, so a gap is caught when the service looks healthy
CoordinationIssues group alert events; workflows notify Slack and paging toolsRaises its own alert event with the data impact, so the on-call sees it where they already look
InvestigationAutopilot correlates traces, logs, metrics and change events into a likely causeTraces the break into the data: which rows, which tables, which Liveboards, which owners
The fixAutopilot recommends remediation, such as a rollback, for an engineer to approveProposes the data change (backfill with undo, rerun, schema change), routes it to a named owner and runs it after approval
VerificationThe alert condition recovers and the issue closesRe-checks the numbers: the row-count baseline holds again and quality checks pass
The recordIssue history and the alert event timelineA receipt for the data change: cause, diff, approver, what it touched, how it was verified, how to undo it

Why doesn't New Relic just do this itself?

Because New Relic built its AI around the telemetry it owns, which is the right focus for a platform that watches every service you run. Its own docs are clear about where Autopilot stops: it can "recommend remediation, such as a deployment rollback, for you to review, evaluate, and approve", and "Autopilot recommends actions but does not take them." In Slack it works "using your New Relic permissions." The New Relic MCP server follows the same design: its tool reference lists query and analysis tools such as execute_nrql_query, list_recent_issues, analyze_deployment_impact, list_change_events and analyze_kafka_metrics, every call is governed by the permissions of the user's OAuth profile, and administrators decide which MCP capabilities are available through Feature Control Manager. Every tool in the reference reads or analyzes; none writes to your systems.

That is a sensible line for an observability vendor. A data fix is a different job with a different liability. The missing orders go back into a production Snowflake table with an undo step saved, after the table's owner approves. Lineage has to show every Liveboard and export that reads it. The Flink job owner merges the lateness change. Someone checks the totals and writes a receipt an auditor can read. None of those systems belong to New Relic, and none of that work is telemetry. Data Workers is built for that job.

Every tool owns a slice. Data Workers covers the whole lifecycle

New Relic owns one slice and owns it well: telemetry is watched there and alerts fire there. Each point tool adds another console, another contract and another handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on New Relic where your engineers already look.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, New Relic goes deep on its own area
StageData WorkersNew RelicWhy we scored it this way
Catalog & Context92Entities and service maps describe services, hosts and queues, not tables. Data Workers keeps one governed context graph of definitions, owners, lineage, quality and usage across every platform.
Analytics & Insights84NRQL and Dashboards answer questions about telemetry, and New Relic AI turns questions into NRQL. Data Workers answers business data questions from governed definitions with lineage behind every number.
Data Quality82Custom events can carry a data check if someone sends them. Data Workers writes, runs and repairs the checks on the tables themselves.
Observability & Incidents8.59.5New Relic's home stage: alert conditions, issues and alert events, Change Tracking, and Autopilot investigating the likely cause. Data Workers diagnoses the data incident across lineage, fixes it with approval and verifies it.
Pipelines & Ingestion8.53Queues and Streams shows Kafka lag and throughput, and Pipeline Control governs telemetry ingestion. Data Workers repairs the data pipeline itself: backfills and reruns, with approvals.
Schema & Migration81Change Tracking records deployments, not table schemas. Data Workers catches schema changes, scores their blast radius and plans migrations in waves.
Governance & Access8.52Roles and Feature Control Manager govern who sees and uses New Relic. Data Workers runs the warehouse access request queue with time-bound, approved grants.
Security & Privacy83Security Rx flags runtime risk in instrumented services. Data Workers' pull request review flags new columns whose names or annotations look sensitive, and leaves a receipt on every data change.
Cost / FinOps83Cloud Cost Intelligence shows cloud spend in real time. Data Workers traces Snowflake credits to the dbt model behind them and drafts the fix for its owner.
MLOps & Models7.53AI Observability watches AI applications, latency and token usage. Data Workers keeps the data under your models fresh and correct.

New Relic leads where it should. For the broader split between watching systems and fixing data, read Data Workers vs data observability and what is an agentic data platform.

How New Relic and Data Workers work together

New Relic stays on top, where alert events fire and Autopilot investigates. Spellbook Data Catalog (in preview) is where the data team reviews, approves and rolls back changes. Between them, Data Context Wizard keeps one governed context graph across Kafka, Snowflake and the jobs between them, the Data-Agents Swarm does the work with 20+ specialist agents, the Autonomous Data-Conductor runs each fix end to end (detect, diagnose, fix, review, verify, remember), and per-domain guardrails hold approvals, receipts and rollback.

How Data Workers fits with New Relic: your coding agent on top, Data Workers in the middle, your estate underneath

Data Workers to New Relic, natively. Data Workers connects to New Relic over NerdGraph, New Relic's GraphQL API, with two actions: send_newrelic_alert raises an alert event when it finds data impact, and resolve_newrelic_alert closes it once the fix is verified. Engineers never leave New Relic to learn that a green stream left a wrong table.

The data work. diagnose_incident and get_root_cause work through the evidence, trace_cross_platform_lineage and blast_radius_analysis map everything downstream of the topic, and get_incident_history checks whether the same lag has dropped events before. Fixes run through remediate (playbooks such as backfill_data), verification uses run_quality_check and get_quality_score, and every step lands in get_audit_trail. Kafka and Snowflake are native. Flink and ThoughtSpot connect over their APIs or MCP servers today. Kafka consumer lag stays New Relic's job: Queues and Streams measures it, and Data Workers picks up the data impact.

Data Workers in New Relic's own view. export_otel_spans emits Data Workers' agent activity as OpenTelemetry spans, which New Relic ingests like any other trace.

Side by side in one client. Add Data Workers beside the New Relic MCP server. Every Data Workers agent is an MCP server; the client setup guide documents the path: clone the open-source repo and add one start-agent.sh entry per agent.

# Example: New Relic MCP server plus Data Workers in Claude Code
# New Relic's documented endpoint (EU: mcp.eu.newrelic.com); admins govern MCP access in Feature Control Manager
claude mcp add --transport http --scope user newrelic https://mcp.newrelic.com/mcp/
claude mcp login newrelic

# Data Workers agents over stdio, from a clone of the open-source repo
claude mcp add --scope user dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
claude mcp add --scope user dw-incidents -- "$(pwd)/start-agent.sh" dw-incidents
claude mcp add --scope user dw-quality -- "$(pwd)/start-agent.sh" dw-quality

Ask "why is the 14:00 orders number low?" and the client calls both: New Relic returns the lag issue, the deploy and the Kafka metrics; Data Workers returns the missing events, the blast radius and a proposed backfill. New Relic enforces its own permissions on its tools; Data Workers' guardrail decides whether a data change may run in that domain, with a named approver. To trim New Relic's toolsets, send its include-tags header (for example alerting). For a shared endpoint, the Data Workers remote server serves /mcp with an API key (bearer) or OAuth tokens from your identity provider, verified through JWKS.

One alert event, L0 to L4. The same lag incident on orders-cdc, at each level of the autonomy ladder, set per domain.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. After the rollback, an engineer counts rows and writes the backfill SQL by hand.
  • •L1 observe. Data Workers raises an alert event with the baseline result and blast radius. No data changes.
  • •L2 propose. Data Workers proposes the backfill and the Flink diff. Nothing runs until the commerce data owner approves in Spellbook.
  • •L3 act reversibly. For change classes with a proven record, such as backfilling a missed window from the raw table, Data Workers runs the fix with its undo step saved, verifies it and records the receipt.
  • •L4 autonomous. In a scoped domain such as the orders stream, every closed lag issue triggers a baseline re-check, and any gap is backfilled through the orchestrator, with the undo recorded first, and re-checked before the hour is reported.

For the safety model, read is it safe to let AI agents change production data, how approvals work and autonomy levels L0 to L4. On where data lives: the agents run in your infrastructure, your data stays in your systems, and the hosted Conductor sees workflow metadata only.

What changes for your team

New Relic made every service visible. Data Workers gives the data team a crew for what those services carry.

Six jobs that run on autopilot with Data Workers next to New Relic, with a concrete example of each
  • •Incidents. A pipeline alert event gets a data-impact check, a fix and a receipt.
  • •Data quality. Each incident becomes a recorded baseline on the table that broke.
  • •Cloud spend. Snowflake credits are traced to the query and dbt model behind them, and each fix goes to its owner drafted.
  • •Access. Warehouse access becomes a scoped, time-boxed grant the data owner approves.
  • •Audits. New Relic records the alert event; Data Workers records what changed in the data and who approved it.
  • •Migrations. A warehouse or streaming move runs in approved, parity-checked waves.

Keep New Relic, or consolidate?

Keep New Relic if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

For almost every New Relic customer the answer is to keep it: services, Kafka and the on-call workflow run there. What teams consolidate is the stack around the warehouse: a separate data observability contract, a data-quality tool, a stale catalog and per-team reconciliation scripts. Weighing a build on the New Relic MCP server and a coding agent? Read build it ourselves with Claude Code and MCP servers: the MCP calls are the easy part; the context graph, approvals and rollback are the work. The same pattern holds across the incident stack: see you're on Datadog, you're on Dynatrace and you're on incident.io, and the full list in Data Workers integrations.

The case for your CFO

The outcome: you already pay New Relic to see infrastructure problems early, and it does. Data Workers makes sure the business numbers that infrastructure carries are right after every incident, so a 14:31 rollback does not become a wrong revenue hour in tomorrow's report.

The risk story is plain. Autonomy is set per domain: at L1 Data Workers changes no data; at L2 it changes nothing until a named person approves; at L3 it applies changes it can undo. An unanswered approval request expires and escalates, never auto-grants. No agent can promote its own work. Every change carries a receipt: the cause, the diff, who approved it, what it touched downstream, how it was verified and how to undo it. An org-wide stop halts all autonomous dispatch. Zero migration: New Relic, Kafka, Flink, Snowflake and ThoughtSpot stay where they are.

Why now: Autopilot finds the likely cause of an infrastructure alert quickly, so the slow part has moved to the data the incident left behind, several systems away from the telemetry. The first win is read-only: every closed lag or error issue on a pipeline entity gets a baseline check and a blast radius posted as an alert event. What stays the same: your alert conditions, workflows, on-call rotations, warehouse permissions and code review. For the numbers, see the ROI of agentic data operations.

The sentence to repeat upstairs: "New Relic tells us the system recovered; Data Workers proves the data did, with an approval and a receipt for every change."

Getting started

Start with a pilot. Pick one streaming domain whose alert conditions already live in New Relic, such as the orders topic behind the revenue Liveboard, connect Data Workers to New Relic, Kafka and Snowflake, and run at L1 so every closed lag issue gets a baseline check and a blast radius. Then turn on the first write class at L2, such as backfills of missed windows with owner approval. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.

FAQ

Does Data Workers have a New Relic connector? Yes. Data Workers connects natively to New Relic over NerdGraph: send_newrelic_alert raises an alert event when it finds data impact and resolve_newrelic_alert closes it once a fix is verified. For telemetry such as NRQL results, change events and Kafka metrics, it runs beside the New Relic MCP server in the same client.

New Relic renamed incidents to alert events. Does that change anything? No. The rename took effect across New Relic's UI on February 16, 2026, and New Relic states that existing alert conditions, workflows and permissions keep working. Data Workers raises and closes alert events through the same API.

How is this different from New Relic Autopilot? Autopilot investigates infrastructure and application issues and recommends remediation, such as a rollback, for a person to review and act on. Data Workers works on the data: re-checking a table against its baseline, backfilling a missed window, proposing a streaming job change, each with approval, rollback and a receipt. In the incident above, Autopilot pointed to the service fix and Data Workers fixed the orders.

Can New Relic alert on stale tables directly? It alerts on any NRQL query over data you send it, so teams that push custom freshness events get freshness alerts. Data Workers tracks load lag and row counts against baselines in the warehouse itself, then fixes and verifies.

Does Data Workers change anything in our New Relic account? Only the alert events it raises: it opens one when it finds data impact and closes it when the fix is verified. Your alert conditions, policies, workflows and dashboards stay as they are.

Do we need New Relic's Kafka integration for this to work? No. Data Workers finds the data gap from the warehouse side, with row-count baselines on the tables the stream fills. With Queues and Streams in place, New Relic measures the lag and the data impact lands beside it as an alert event.

Sources

  • •New Relic, What's new: Accelerate incident investigation with New Relic Autopilot (Jul 21, 2026), https://docs.newrelic.com/whats-new/2026/07/whats-new-07-21-autopilot/ (checked Oct 2, 2026)
  • •New Relic Docs, Autopilot overview (formerly SRE Agent; "recommends actions but does not take them"), https://docs.newrelic.com/docs/agentic-ai/autopilot/overview/ (checked Oct 2, 2026)
  • •New Relic, What's new: Connect AI to trusted operational context with New Relic Ground Truth (Sep 15, 2026; launch note, no GA or preview label), https://docs.newrelic.com/whats-new/2026/09/whats-new-09-15-ground-truth/ (checked Oct 2, 2026)
  • •New Relic Docs, MCP server overview (Feature Control Manager toggles, MCP server on by default, observability:read scope), https://docs.newrelic.com/docs/agentic-ai/mcp/overview/ (checked Oct 2, 2026)
  • •New Relic Docs, MCP server setup (endpoints, Claude Code commands, include-tags header), https://docs.newrelic.com/docs/agentic-ai/mcp/setup/ (checked Oct 2, 2026)
  • •New Relic Docs, MCP tool reference (35 GA and 3 preview read and analysis tools, permissions), https://docs.newrelic.com/docs/agentic-ai/mcp/tool-reference/ (checked Oct 2, 2026)
  • •New Relic, What's new: "Incidents" renamed to "Alert events" (effective Feb 16, 2026), https://docs.newrelic.com/whats-new/2026/01/whats-new-01-23-incidents-to-alert-events/ (checked Oct 2, 2026)
  • •New Relic, What's new (MCP server public preview Nov 5, 2025; Compound Alerts Sep 22, 2026; Pipeline Control processors Sep 30, 2026), https://docs.newrelic.com/whats-new/ (checked Oct 2, 2026)
  • •New Relic, Platform (Queues and Streams, Change Tracking, Security Rx, Cloud Cost Intelligence, AI Observability), https://newrelic.com/platform (checked Oct 2, 2026)
  • •Data Workers open-source repository, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)
  • •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)
  • •Data Workers, Data-Agents Swarm, https://dataworkers.io/product/data-agents-swarm/ (checked Oct 2, 2026)