Product
Product12 min readBy The Data Workers Team

You're on Segment: It Collects and Routes Every Customer Event. Data Workers Owns Whether What They Become Is Right

Twilio Segment collects and routes your customer events. Data Workers checks the warehouse tables, dbt models and audiences built from them, and gets fixes approved.

Your team runs Twilio Segment. Your web and mobile sources send Track, Identify and Page calls, and Connections routes them to your destinations and to warehouse syncs in Snowflake or BigQuery, where each Track event gets its own table next to tracks, identifies and users. Protocols holds the tracking plan, with a violation when a live event doesn't match. Source Schema controls block what doesn't belong. Unify resolves identities into profiles, Profiles Sync writes them back to the warehouse, the Data Graph relates them to your entity tables, and Engage turns it all into audiences and journeys, including Linked Audiences that marketers build on warehouse data without waiting for the data team.

That last part is where the risk moves. A Linked Audience reads the tables your team builds in dbt. Segment delivers events faithfully and checks their shape; the models in between belong to you. When an app release changes what an event carries, Protocols can stay quiet, the profile can stay right, and a dbt model can still be wrong by the next audience run. Segment collects and routes your customer events. Data Workers owns whether the tables, models and audiences built from them are right, and gets the fix approved before a campaign sends.

Key takeaways

  • •Segment keeps its job. Sources, destinations, the tracking plan, profiles and audiences stay with your team. Data Workers works on what lands in the warehouse and what reads it.
  • •A clean tracking plan isn't a clean model. Protocols checks event shape at collection. Data Workers checks the event tables and dbt models built from them, so a change Protocols has no rule for still becomes a diagnosed incident.
  • •Audiences get a check before they run. When an entity table behind a Linked Audience goes wrong, Data Workers maps the audience in the blast radius and proposes a hold to its owner before the next scheduled run.
  • •Connected over Segment's Public API today. Data Workers reads Segment's warehouse tables natively; your team reads its configuration over the Public API. It never edits a tracking plan, a block or an audience.
  • •Autonomy is set per domain. Start at L1 observe on the tables behind audiences and revenue, move to L2 propose, and open L3 act reversibly for narrow classes once the record earns it.

Segment collects and routes your customer events. Data Workers owns whether what they become is right.

Segment's job is to collect every event, resolve who sent it and deliver it everywhere your teams need it. The job after delivery is different: notice that a model built on those events has gone wrong, tie it to a release, size the damage, hold what would send the mistake to customers, get the cause fixed, rebuild and prove the result. Here is one Tuesday at a meal-kit subscription company. This is an illustration, not a customer case.

TimeSystemWhat happens
Tue 14:00iOS appRelease 7.4 rolls out with guest checkout first: shoppers sign in on the confirmation screen, so identify() now fires a few seconds after Order Completed. Segment's SDK adds the userId only to calls made after identify(), so those order events carry an anonymous_id and no user_id
14:02SegmentEvery Order Completed property matches the tracking plan, so Protocols records no violation. Identity Resolution merges each anonymous visit into the right profile once the Identify call arrives
15:00SnowflakeThe scheduled warehouse sync loads ios.order_completed: 1,380 of 4,212 new rows have a null user_id
15:30dbt CloudThe hourly job builds fct_orders, which joins orders to customers on user_id, and drops those rows. dim_customer_orders, the entity table the Data Graph relates to each profile, now shows hundreds of customers who ordered an hour ago as lapsed. Every dbt test passes
15:41Data WorkersThe null share on ios.order_completed.user_id, a metric the team records with monitor_metrics, reads 32.8% against a 1.5% baseline; run_quality_check confirms the nulls; the dbt manifest shows no model change. Data Workers opens an incident
15:48Data WorkersDiagnosis: every null row comes from app version 7.4, and in the Profiles Sync user_identifiers view each anonymous_id sits on a profile that gained a user_id seconds later. Blast radius: fct_orders, dim_customer_orders, the Linked Audience "Lapsed subscribers" (daily run at 18:00, activation to a Customer.io win-back journey with a discount code), the Tableau "Mobile orders" workbook and finance's daily orders extract
15:52SlackData Workers posts the incident to #growth-data and sends the approval request to the analytics engineering owner, with the lifecycle marketing manager and the iOS lead tagged
16:20SpellbookThe owner reviews four proposals: a hold on the "Lapsed subscribers" audience until the rebuild is verified; a diff to stg_ios__order_completed that fills a missing user_id from Profiles Sync identity data on anonymous_id; a dbt test on null user_id share by app version; and the rebuild order. She approves all four
16:25SegmentThe lifecycle marketing manager disables the audience in Segment
16:35GitHub + dbt CloudThe owner merges the diff and queues the approved dbt Cloud job, recorded in Spellbook, which rebuilds the staging model, fct_orders and dim_customer_orders
16:58Data WorkersData Workers verifies: iOS order volume in fct_orders since 14:00 is back in its normal range, null user_id share is back to baseline, the recorded count of lapsed customers in dim_customer_orders is back in yesterday's range, and the owner reports the new test passing. The receipt records the cause, approvals, run, checks and the undo
17:20SegmentThe marketing manager re-enables the audience, which starts a run right away; the win-back offer reaches only customers who really lapsed, and the 18:00 schedule carries on
Wed 11:00iOS appRelease 7.4.1 moves identify() back before Order Completed; the dbt test stays as a guard
Incident timeline across the stack: what Segment, your team and Data Workers each do, step by step

Every part of Segment did its job: the event matched the plan, Identity Resolution put every order on the right profile, and the sync delivered every row. What broke was a team-owned model that assumed every order has a user_id, and an audience that trusted it. Catching it takes context Segment was never meant to hold: which dbt model reads the event table, which entity the Data Graph maps, and which audience runs at 18:00.

JobWhat Segment doesWhat Data Workers does
The collectionCollects Track, Identify, Page and Screen calls from every source and routes them to 550+ destinations and the warehouseReads the event tables Segment writes, natively in Snowflake or BigQuery; connects to Segment over its Public API
The contractValidates events against the tracking plan and raises Protocols violations; Source Schema blocks events, properties and traitsChecks the data the plan can't describe: null keys, volume, freshness and recorded metrics on the landed tables, so a change with no violation still raises an incident
The identityResolves anonymous and known activity into one profile and syncs it to the warehouseUses Profiles Sync tables as evidence in the diagnosis and as the source of the proposed fix
The audienceBuilds Linked Audiences on warehouse entities and activates them on a run scheduleMaps audiences that read a damaged entity table into the blast radius and proposes a hold to the owner before the next run
The fixApplies whatever tracking plan, block or audience setting its owners chooseProposes the dbt diff, the test and the rebuild order to a named owner; records the approval for the dbt Cloud run the owner queues
The proofKeeps violation history and Delivery Overview for every event and destinationRe-checks the tables and writes a receipt: what changed, who approved it, how it was checked, how to undo it

Why doesn't Segment just do this itself?

Because Segment is built to collect and deliver customer data faithfully, and its quality tools are scoped to the event. Protocols validates an event's shape against the plan. Segment's own docs say you "can't require a Track call because Segment is unable to verify when a Track call should be fired", and the order in which a release sends Identify and Track calls is outside what a tracking plan can describe. Source Schema controls are, in Segment's words, a first-line defense; blocked data isn't recoverable by default, so blocking is a careful tool, not an incident response. Identity Resolution is right to keep the profile correct and stop there.

Segment's AI points the same way. Predictions builds propensity models that predict whether a customer will purchase, churn or perform another conversion event, and Generative Audiences drafts audience conditions from a plain-language prompt for a marketer to review; each carries an AI Nutrition Facts label. The Twilio MCP server (public beta) gives coding agents search over Twilio's API specs and docs, Segment docs included, and is read-only by design: it "does not execute API calls on your behalf."

Owning whether a dbt model, a Tableau workbook and a Linked Audience are right after a release is a different product: a context graph of every table and consumer, blast radius across vendors, named approvers, a recorded undo and receipts an auditor can read. That is Data Workers. More in is it safe to let AI agents change production data.

Every tool owns a slice. Data Workers covers the whole lifecycle

Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the Segment workspace already there.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Segment goes deep on its own area
StageData WorkersSegmentWhy we scored it this way
Catalog & Context94The Data Graph relates Segment profiles to entity tables in the warehouse. Data Workers keeps one governed context graph of every table, model, owner and consumer across the estate.
Analytics & Insights86Engage audiences, journeys and Predictions turn events into targeting. Data Workers answers data questions from governed definitions with lineage behind every number.
Data Quality85Protocols validates each event against the tracking plan and Source Schema blocks what doesn't belong. Data Workers checks the tables built from those events: nulls and volume, plus lateness and the other metrics your team records.
Observability & Incidents8.54Delivery Overview and violation alerts show what reached each destination. Data Workers diagnoses the data incident behind a clean delivery, proposes the fix and verifies it.
Pipelines & Ingestion8.59Segment's home stage: SDKs, the Tracking API, 550+ destinations, warehouse syncs, Profiles Sync and Reverse ETL. Data Workers plans the reruns downstream for the owner.
Schema & Migration83Typewriter generates type-safe tracking libraries from the tracking plan, keeping event shapes in step with code. Data Workers traces a changed event through dbt, BI and the audiences that read it.
Governance & Access8.56Workspace roles and Source Schema controls govern who changes what in Segment. Data Workers routes every data change to a named approver.
Security & Privacy86Blocking events, properties and traits keeps unwanted data out at the source. Data Workers leaves a receipt on every data change.
Cost / FinOps83MTU accounting prices collection; blocked events drop out of it. Data Workers traces Snowflake credits to the dbt model behind them through query tags.
MLOps & Models7.54Predictions builds purchase, churn and conversion propensity models. Data Workers keeps the data under models and agents healthy.

How Segment and Data Workers work together

How Data Workers fits with Segment: your coding agent on top, Data Workers in the middle, your estate underneath

Data Workers connects to Segment over its Public API today. What it needs from each warehouse sync sits in the warehouse: the per-event tables, identifies, users and the Profiles Sync tables and views are ordinary tables it reads natively. Through the Public API it reads sources and tracking plans with a token your Workspace Owner issues. Tracking plans, Source Schema blocks, profiles, audiences and activations stay with your team: Data Workers never changes them. When an audience reads a damaged table, it proposes the hold, and the audience owner disables and re-enables it in Segment. The rest of this incident runs on native connections: Snowflake, dbt Cloud, Tableau (read), GitHub and Slack, among 50+ connectors that also cover BigQuery, Databricks, Airflow and Dagster. Customer.io connects over its API today; in this incident Data Workers leaves it alone, and the audience owner decides what reaches it.

The Twilio MCP server can sit next to the Data Workers agents in one client: an engineer looks up how Profiles Sync works while Data Workers explains what this afternoon's events did to the models. The Data Workers side follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry.

# Example: the Twilio MCP server (docs search, public beta) plus Data Workers agents in Claude Code
claude mcp add --transport http twilio-docs https://mcp.twilio.com/docs

# Data Workers agents, from a clone of the open-source repo
claude mcp add --scope user dw-incidents -- "$(pwd)/start-agent.sh" dw-incidents
claude mcp add --scope user dw-quality -- "$(pwd)/start-agent.sh" dw-quality
claude mcp add --scope user dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
claude mcp add --scope user dw-schema -- "$(pwd)/start-agent.sh" dw-schema
claude mcp add --scope user dw-connectors -- "$(pwd)/start-agent.sh" dw-connectors

List the tools with your client's own command (/mcp in Claude Code). In this incident: monitor_metrics (dw-incidents) tracks the null share the team records against its baseline; run_quality_check (dw-quality) runs the null and volume checks; diagnose_incident names the cause; blast_radius_analysis and trace_cross_platform_lineage (dw-context-catalog) map what the dropped orders reached; remediate re-checks the quality assertions afterwards and escalates any failure to a person.

In production the agents run in your infrastructure and hold the warehouse credentials and model key.

One incident, L0 to L4, set per domain:

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. Tuesday's iOS buyers get a win-back discount at 18:00; an analyst finds the null user_id rows on Wednesday.
  • •L1 observe. Data Workers flags the null spike at 15:41 with the cause and the audience in the blast radius. Nothing changes.
  • •L2 propose. Data Workers proposes the hold, the diff, the test and the rebuild; nothing moves until the named owner approves.
  • •L3 act reversibly. For a class with a clean record, Data Workers queues the rebuilds itself and sends any failed check to a person.
  • •L4 autonomous. For a scoped domain, Data Workers checks event tables after every sync, so the owner's fix is waiting before the next audience run.

What changes for your team

Six jobs that run on autopilot with Data Workers next to Segment, with a concrete example of each
  • •App releases get a downstream check. When a release changes what an event carries, the data team hears within the hour, with the models and audiences it reaches.
  • •Marketing gets a hold, not a surprise. A Linked Audience that reads a damaged table gets a proposed hold before its run, sent to the person who owns it.
  • •The tracking plan and the models stay in step. Protocols owns event shape; Data Workers covers the tables built on top.
  • •One record for data, product and marketing. Everyone sees the same incident and receipt in Spellbook Data Catalog (in preview), linked from the Slack thread.

For the warehouse side of customer data, see how Salesforce Data 360 teams keep segments true and how retail and e-commerce teams keep opt-outs in every audience.

Keep Segment, or consolidate?

Keep Segment if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

For most Segment teams the answer is keep it: SDKs on every platform, an enforced tracking plan, identity resolution and self-serve audiences are hard to replace. What teams consolidate is the tooling around the warehouse: a separate tool watching event tables, hand-written null checks in dbt, and "check the dashboard before the campaign" runbooks. Many estates run more than one customer-data tool; see you're on RudderStack, you're on Hightouch and you're on Census. One incident record spans them, and the integrations page lists what Data Workers reads and changes in each category.

Weighing a build on the Public API and a coding agent? Read build it ourselves with Claude Code and MCP servers: reading a tracking plan from a chat is easy; the context graph, approvals, undo and receipts are the work.

The case for your CFO

The outcome. When an app release quietly breaks the tables behind revenue reporting and audiences, it is caught within the hour and corrected before a campaign sends or a leader reads the number, with a record of how it was checked.

The risk story. At L0 and L1, agents only read. At L2 they propose and a named person approves; an unanswered request expires and escalates, never auto-grants. At L3 they act on reversible changes inside the domains you open; L4 is a later choice per domain. Segment settings stay with their owners. No agent can promote its own work, and an org-wide stop halts all autonomous dispatch. Every change carries a receipt: what changed, who approved it, the blast radius and the undo.

Why now. Linked Audiences put warehouse tables in front of marketers by design, so a wrong model reaches customers on the next scheduled run.

The first win. L1 on the event tables behind your audiences and revenue models: a release that changes what an event carries becomes a diagnosed incident before the next audience run.

What stays the same. Nothing migrates: Segment, your tracking plan and audiences, your dbt project, warehouse and destinations. For the numbers, see the ROI of agentic data operations.

The sentence for upstairs: "Segment gets our customer events everywhere they need to go; Data Workers makes sure what we build from them is right, and gets it fixed with our approval before a campaign sends."

Getting started

Start with a pilot. Pick the event tables that feed revenue models and your key audiences, give Data Workers read access to those schemas, and run at L1 for a few weeks. Then turn on L2 for one domain, and open L3 for a narrow class once the receipts show the agents were right. The pilot path is on the pricing page, and the pilot is credited in full against the first year.

FAQ

How does Data Workers connect to Segment? Over Segment's Public API today, with a token your Workspace Owner creates. It reads the tables Segment writes, including per-event and Profiles Sync tables, natively in Snowflake or BigQuery.

Protocols shows no violations. How can the data be wrong? Protocols validates each event against the tracking plan: names, property types, required properties and permitted values. It can't describe when a call fires or what a team's model assumes about it. In the example above every event matched the plan; the dbt model's join on user_id broke. Data Workers checks the landed tables and the models built on them.

Will Data Workers change our tracking plan, block events or edit audiences? No. Those stay with their owners in Segment. Data Workers proposes a hold or a plan change with the reason, and the owner decides. Its own changes land in your dbt project as diffs the owner merges.

Can't we just block bad events in Source Schema? Blocking is a good first line for events that shouldn't exist. By default blocked data isn't recoverable, and in rare cases a block takes up to six hours to reach every destination, so it isn't a fit for a valid event that a model reads wrong. The fix there belongs in the model.

Does the Twilio MCP server let agents run our Segment workspace? No. The Twilio MCP server is in public beta, searches Twilio's API specs and docs (Segment docs included) and is read-only by design. Data Workers runs next to it in the same client and works from the event tables Segment lands in your warehouse.

Where does our customer data go? The Data Workers agents run in your infrastructure and read your warehouse there. Your data stays in your systems; the hosted Conductor sees workflow metadata only. See where does our data go.

Sources

  • •Twilio Segment product page (Connections, Unify, Engage, Protocols; "550+ destinations"), https://www.twilio.com/en-us/segment (checked Oct 3, 2026)
  • •Twilio docs, Protocols Tracking Plan (violations, data types, "can't require a Track call"; updated Mar 25, 2026), https://www.twilio.com/docs/segment/protocols/tracking-plan/create (checked Oct 3, 2026)
  • •Twilio docs, Source Schema controls (blocking, recovery, six-hour note; updated Jun 25, 2026), https://www.twilio.com/docs/segment/connections/sources/schema (checked Oct 3, 2026)
  • •Twilio docs, Warehouse Schemas (tracks, identifies, users, per-event tables; updated Jul 15, 2026), https://www.twilio.com/docs/segment/connections/storage/warehouses/schema (checked Oct 3, 2026)
  • •Twilio docs, Profiles Sync tables and materialized views (user_identifiers, id_graph; updated Jun 12, 2026), https://www.twilio.com/docs/segment/unify/profiles-sync/tables (checked Oct 3, 2026)
  • •Twilio docs, Data Graph (updated Jun 26, 2026), https://www.twilio.com/docs/segment/unify/data-graph (checked Oct 3, 2026)
  • •Twilio docs, Linked Audiences (run schedules, enabling starts a run, activations; updated Oct 1, 2026), https://www.twilio.com/docs/segment/engage/audiences/linked-audiences (checked Oct 3, 2026)
  • •Twilio docs, Predictions Nutrition Facts label (updated Jul 2, 2026), https://www.twilio.com/docs/segment/unify/traits/predictions/predictions-nutrition-facts (checked Oct 3, 2026)
  • •Twilio docs, Segment Public API (updated Feb 24, 2026), https://www.twilio.com/docs/segment/api/public-api (checked Oct 3, 2026)
  • •Twilio docs, Generative Audiences (updated Jul 15, 2026), https://www.twilio.com/docs/segment/engage/audiences/generative-audiences (checked Oct 3, 2026)
  • •Twilio docs, Spec: Identify and Best Practices for Identifying Users (anonymousId, userId added to subsequent calls; updated Feb 27 and Mar 25, 2026), https://www.twilio.com/docs/segment/connections/spec/best-practices-identify (checked Oct 3, 2026)
  • •Twilio docs, Twilio MCP server (public beta, read-only; updated Aug 31, 2026), https://www.twilio.com/docs/ai/mcp (checked Oct 3, 2026)
  • •Data Workers open-source repository (tool registrations in dw-incidents, dw-quality, dw-context-catalog, dw-schema, dw-connectors), https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
  • •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)