Industry
Industry8 min readBy The Data Workers Team

The agentic data platform for SaaS and tech companies

How Data Workers runs data operations for SaaS teams: product schema changes caught before dbt breaks, one approved ARR and NRR definition, usage data checked before the billing export, and warehouse cost managed as COGS, with approvals and receipts.

For a SaaS data team, Data Workers is the agentic data platform that keeps product, revenue and usage data right while the product ships every day: its agents catch the schema change before dbt breaks, hold one approved definition of ARR and NRR with a named owner, and check the warehouse usage tables before your billing export reads them. Every change goes through the approval level you set per domain and leaves a receipt, and the agents run in your own infrastructure next to the warehouse you already have.

In SaaS, the app database changes with every release, and in a usage-priced business the warehouse is part of the billing system.

Key takeaways

  • •Product schema changes stop being incidents. Data Workers catches a breaking change in the CDC stream's schema versions or the pull request, finds the dbt models and dashboards it reaches, and reviews the pull request before it merges.
  • •One ARR, one NRR. The Data Context Wizard holds each revenue metric's approved definition and its canonical table, and routes a conflicting definition to a person for review.
  • •Usage data is checked before it bills. When price depends on metered usage, Data Workers checks the usage tables against their baseline and traces what reads them, including the billing export.
  • •Warehouse cost is managed like COGS. Snowflake credits are attributed to the query and the dbt model behind them, and the fix arrives drafted, with its dependencies checked, for the owner to apply.
  • •The controls your customers ask about. Named approvals, expiring access, privacy checks on pull requests and a hash-chained audit trail, with the agents in your environment.

The SaaS data reality

The typical stack: Postgres or MySQL app databases feed the warehouse through change data capture (Debezium on Kafka, Fivetran or Airbyte), product events arrive through Segment or RudderStack, billing lives in Stripe or Zuora and bookings in Salesforce or HubSpot. It all lands in Snowflake, BigQuery or Databricks, where dbt models it on an Airflow, Dagster or Prefect schedule and Hex, Looker or Mode presents it.

The operations that hurt:

  • •Fast-moving product schemas. A column rename rides the CDC stream, and a dbt model fails at 02:00 or, worse, succeeds with nulls.
  • •Metric drift. Finance, sales and product each have an ARR and an NRR, and they disagree at the board meeting.
  • •Usage-based billing. A duplicate event is a wrong invoice, and a wrong invoice is a customer escalation.
  • •Warehouse cost as COGS. Compute sits in gross margin. The dbt Labs 2026 State of Analytics Engineering report found 57% of respondents report increased warehouse and compute spend, against 36% reporting increased team budgets.
  • •AI speeds up the writing, not the running. The same report found 72% prioritise AI-assisted coding but only 24% prioritise AI-assisted pipeline management, and 71% are concerned about hallucinated or incorrect data reaching stakeholders.

Your coding agents already write the dbt models; Data Workers runs the operations around them. The Data Engineering for SaaS playbook and AI for data infrastructure in SaaS go deeper on the stack.

Three use cases

1. Product schema changes from the app database. A release renames plan_tier to plan_code in Postgres, and Debezium carries it onto a Kafka topic. The streaming agent sees the new schema version on the topic, assess_impact finds the dbt models and Looker Explores downstream, and check_compatibility tests the new shape against its consumers and flags it as breaking. When the change arrives as a pull request, Data Workers reviews it with a blast radius and an impact diff, and blocks a risky change from merging. generate_migration writes the migration with its rollback SQL for the owner to apply. Start this domain at L2 propose. Engineers in Cursor or Claude Code get the same checks over MCP (Data Workers with your coding agents), and Data Workers + dbt covers the dbt side.

2. Revenue metrics: ARR, NRR and usage-based billing. Stripe's Data Pipeline syncs Stripe data to Snowflake, Redshift, Databricks or BigQuery, so subscriptions, bookings and usage meet in one place; the hard part is the definition. define_business_rule records the calculation with its author (ARR from active subscriptions at month end, excluding one-time services, for example), and mark_authoritative marks fct_arr_monthly as the canonical table for revenue, with who designated it and why. When a new model or dashboard defines ARR differently, the Context Wizard routes the contradiction to a person for review instead of picking a winner, and no agent can promote its own definition to authoritative. run_quality_check, get_anomalies and set_sla watch the usage and revenue tables, as the worked example shows.

3. Warehouse cost as a line in gross margin. Data Workers' cost agent attributes Snowflake credits to the query and the dbt model, project and run behind them through query tags, reads BigQuery spend from the Jobs API, and drafts cost-safe warehouse defaults (x-small, 60-second auto-suspend, a resource monitor) for the platform owner. Every change it drafts comes with downstream dependencies checked, and the owner applies it. The design target is 25 to 40% lower warehouse spend; the ROI of agentic data operations shows how to model it for your estate.

A worked example: double-counted usage before the month-end invoice

This is an illustration, not a customer case. A company prices its API by call volume. RudderStack sends product events to Databricks, dbt models them on a Dagster schedule, and a Dagster job the billing team owns exports monthly usage to Stripe at 06:00 on the first of the month. The billing-data domain runs at L2 propose.

TimeSystemWhat happened
23:10RudderStackA mobile SDK release resends a day of api_call events with new message IDs
23:40DatabricksThe duplicates land in the raw tracks table
00:20dbt, DagsterThe nightly run passes its tests; fct_usage_daily doubles for 40 accounts
00:25Data Workersmonitor_metrics flags the usage volume against its baseline
00:31Data Workersblast_radius_analysis finds the 06:00 Dagster export job to Stripe and fct_nrr_monthly
00:38Data WorkersProposes a dedupe step in the staging model as a diff, with dry-run row counts and a rollback path
00:40PagerDutyData Workers pages the billing on-call engineer with the proposal link
01:05SpellbookThe billing data owner reviews the diff, approves and merges it
01:10Data Workers, DagsterQueues the dbt rerun through Dagster; usage is back within its baseline and the checks pass
06:00StripeThe usage export runs on corrected numbers and invoices finalize
Incident timeline across the stack: what Your SaaS stack, your team and Data Workers each do, step by step

The receipt records the anomaly, blast radius, diff, approver and time in the hash-chained audit log. Nobody issued 40 credit notes.

Governance: what your customers and regulators ask

Data Workers gives your team the controls and the evidence; your auditors and counsel decide how they fit your program. This is not legal advice.

SOC 2 is what your customers ask for. The AICPA Trust Services Criteria (2017, points of focus revised 2022) are an attestation framework, not a law. For agent changes, an auditor samples what they sample for an engineer: who approved, what the actor could reach, and whether the log can be trusted. Data Workers routes each change to a named approver per domain; an unanswered request expires and escalates, never auto-grants. provision_access grants least-privilege, column-level access with a 90-day expiry by default, and grants nothing when the verdict is review. Every tool call lands in a SHA-256 hash-chained log that verify_global_hash_chain checks end to end and generate_audit_report exports for the period. See the SOC 2 data governance checklist and how approvals work.

GDPR and CCPA for user data. GDPR has applied since 25 May 2018, and its Article 5(1)(c) requires personal data to be "adequate, relevant and limited to what is necessary". The California Privacy Protection Agency's regulations on automated decisionmaking technology, risk assessments and cybersecurity audits have an effective date of January 1, 2026. Your team's classification marks the columns that hold user identifiers, and lineage shows the tables built from them, so a deletion or access request starts from a current map. What agents write to the context graph is PII-scrubbed first, and an erasure request purges a person from Data Workers' own graph with a tombstone in the audit chain. How Data Workers handles PII and SOX, HIPAA, GDPR and the EU AI Act go deeper.

Where it runs. The agents run in your infrastructure on every tier, with your warehouse credentials and model key; Enterprise runs in your VPC, your cloud or on-premise. Your data stays in your systems; the hosted Autonomous Data-Conductor sees workflow metadata only, never rows, credentials or model keys. See Data Workers in your VPC or air-gapped and where does our data go?.

What changes for your team

Six jobs that run on autopilot with Data Workers next to Your SaaS stack, with a concrete example of each

Each job runs at the level you set on one ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous, set per domain. Data Workers ships observe-only. A sensible first quarter moves schema detection and freshness to L3 and keeps anything that feeds billing or board metrics at L2 propose. Autonomy levels L0 to L4 and who owns the agents explain the choices, and will Data Workers replace my data team? answers the question your analytics engineers will ask. See also Data Workers for analytics engineers, for FinOps leads, and the sibling pages for financial services and retail and e-commerce.

The alternatives SaaS teams weigh

  • •Build it with your coding agents. Claude Code, Cursor and Codex write dbt models well; the operating layer (approvals, blast radius, rollback, audit) remains. Build it ourselves prices that path.
  • •Platform-native agents. Snowflake, Databricks and Google Cloud each ship capable agents for their own platform. Keep them; Data Workers works across all of them (Snowflake, Databricks, Google Cloud).
  • •Observability tools. A monitor says the usage table looks wrong; Data Workers finds the cause, proposes the fix and verifies it (the split).
  • •Billing tools. Stripe and Zuora are the system of record for invoices; Data Workers makes sure the warehouse side of the billing loop is right.

Why don't these tools do this themselves? Focus. Each is right to do one job well, and writing to production data across systems, with approvals, rollback and receipts, is a different product. Is it safe to let AI agents change production data? walks through the risk. The integrations page lists the native connectors, 50+ in all, including Postgres, Snowflake, BigQuery, Databricks, dbt, Kafka, Dagster, Airflow, Prefect, Looker and PagerDuty. Data Workers connects to Segment, RudderStack, Stripe and Salesforce over their APIs or MCP servers today.

The case for your CFO

The outcome is revenue numbers you can defend and a slower-growing warehouse bill. In SaaS, data errors show up as a wrong invoice, an ARR figure that moves between the board deck and the finance model, or compute that eats gross margin; Data Workers catches them before they reach a customer or the board.

The risk story: billing and revenue can stay at L2 propose, where a named owner approves every change and no agent can promote its own work. Every change has a receipt with the diff, approver, blast radius and rollback path. Nothing migrates: your warehouse, dbt, orchestrator, BI and billing tools stay.

Why now: your engineers already write data code with AI, and the operating work around it grows. The first win is the usage and revenue domain at L2 for one month-end close, with schema detection on the app database alongside it.

Start with a pilot ($7,500 one-time, credited in full against the first year). Scale is from $1,000 a month and Enterprise from $3,000 a month, billed annually, with unlimited seats, no usage meter, no markup on model spend and your own model. See pricing and the ROI calculator. The sentence for upstairs: "Agents check our revenue and usage data before it reaches a customer's invoice, and every change they make has an approver and a receipt."

FAQ

We are a 6-person data team. Is this for us? Yes. Small teams carry the most on-call load per person, and unlimited seats let finance and product partners use Spellbook (in preview) too.

Does Data Workers send anything to Stripe or Salesforce? No. Stripe and Salesforce stay the systems of record, and your billing export stays your job. Data Workers checks the warehouse side, including the Stripe data your Data Pipeline syncs there, changes it only through approved diffs, and connects to Stripe and Salesforce over their APIs or MCP servers today.

Can it handle CDC from MySQL as well as Postgres? Postgres connects natively. MySQL changes reach Data Workers through the CDC path you already run (Debezium on Kafka, Fivetran or Airbyte); Kafka, its Schema Registry and the warehouse tables are where the schema agent watches them.

Will our SOC 2 auditor accept agent changes? Your auditor decides. What Data Workers gives them is the evidence they sample for any change: a named approver, scoped access that expires, and a tamper-evident log they can export.

Does it help with GDPR and CCPA deletion requests? Lineage shows the tables built from the columns your team has classified as personal, so your team scopes the request, and it purges the person from Data Workers' own graph. Warehouse deletion stays with your process.

Sources

  • •dbt Labs, 2026 State of Analytics Engineering Report: https://www.getdbt.com/resources/state-of-analytics-engineering-2026 (checked Oct 2, 2026)
  • •Stripe, Data Pipeline ("Sync Stripe data to Snowflake, Redshift, Databricks, BigQuery..."): https://docs.stripe.com/stripe-data/data-pipeline (checked Oct 2, 2026)
  • •RudderStack, Databricks Delta Lake destination: https://www.rudderstack.com/docs/destinations/warehouse-destinations/delta-lake/ (checked Oct 2, 2026)
  • •AICPA, 2017 Trust Services Criteria for Security, Availability, Processing Integrity, Confidentiality, and Privacy (with revised points of focus, 2022): https://www.aicpa-cima.com/resources/download/2017-trust-services-criteria-with-revised-points-of-focus-2022 (checked Oct 2, 2026)
  • •GDPR, Regulation (EU) 2016/679, Article 5, EUR-Lex: https://eur-lex.europa.eu/eli/reg/2016/679/oj (checked Oct 2, 2026)
  • •California Privacy Protection Agency, CCPA updates, cybersecurity audits, risk assessments and ADMT regulations (Effective Date: January 1, 2026): https://cppa.ca.gov/regulations/ccpa_updates.html (checked Oct 2, 2026)
  • •Data Workers public repository tools (detect_schema_change, assess_impact, check_compatibility, generate_migration, blast_radius_analysis, define_business_rule, mark_authoritative, run_quality_check, get_anomalies, set_sla, scan_pii, provision_access, verify_global_hash_chain, generate_audit_report): https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)
  • •Data Workers product repository, data-workers-agent-swarm main @ 0c2491e3: agents/dw-review (pull request review), agents/dw-cost (Snowflake metering and per-query attribution with dbt query tags, provision-time guardrails), connectors BigQuery job history, core/contradiction-sentinel (contradictions to human review), connectors/enterprise PagerDuty, core/enterprise (approvals, promotion guard, audit log, rtbf) (checked Oct 2, 2026)
  • •Data Workers product pages: https://dataworkers.io/product/data-context-wizard/ , https://dataworkers.io/product/spellbook-data-catalog/ , https://dataworkers.io/product/autonomous-data-conductor/ (checked Oct 2, 2026)
  • •Data Workers security page: https://dataworkers.io/security/ (checked Oct 2, 2026)
  • •Data Workers pricing: https://dataworkers.io/pricing/ (checked Oct 2, 2026)