Product
Product9 min readBy The Data Workers Team

What happens in the first 90 days?

A 90-day path from pilot to production with Data Workers: weeks 1-2 observe, weeks 3-6 propose with approvals, weeks 7-12 act reversibly in the domains your team picks, with measures reviewed at each step.

The first 90 days with Data Workers follow a staged plan from pilot to production in three steps: weeks 1 and 2 connect and read your estate at L1 observe, weeks 3 to 6 put agents to work proposing fixes for real incidents and for the dbt models behind your Snowflake credits that your people approve at L2 propose, and weeks 7 to 12 let agents act reversibly at L3 in the one or two domains your team picks. Every step ends with a review of the same measures, and L4 autonomous is earned later, one domain at a time, on the record the first quarter builds.

The plan starts with the pilot, with a forward-deployed engineer alongside your team, and the week counts are a design target set with you, not a fixed schedule. Here is each phase: what your team does, what the agents may do, and how you decide to move up.

Key takeaways

  • •Weeks 1 and 2 are read-only. Data Workers connects to your warehouses, dbt, orchestrator and catalog with least-privilege read grants, builds the context graph, and sets quality and cost baselines across every engine.
  • •Weeks 3 to 6 are propose and approve. Agents diagnose real incidents, trace Snowflake credits to the dbt models behind them and propose fixes. A named owner approves each one, and every approved change lands with a receipt.
  • •Weeks 7 to 12 are reversible action in one or two domains. Your team picks the domains, usually late loads or access requests. Everything else stays at propose.
  • •Every phase ends with the same measures. Time from alert to verified fix, proposals approved unchanged, how often the agents' suggested fix matched your team's, overrides and rollbacks, approval wait and coverage.
  • •L4 is earned per domain, later. Autonomy rises only when the receipts support it, and your team can lower any domain at any time.

The plan on one page

First-quarter autonomy roadmap starting from Your data stack

Autonomy in Data Workers is set per domain on one ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous. The first quarter climbs it deliberately. The timing above is a design target; the pace in your estate follows the receipts.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous

The autonomy levels explainer covers each rung, and the pilot page covers the commercial shape.

Weeks 1 and 2: connect and read (L1 observe)

What your team does. Name a pilot lead and one owner per domain you care about. Grant read access with least privilege: on Snowflake, a role with USAGE on a warehouse, SELECT on the schemas in scope and read on SNOWFLAKE.ACCOUNT_USAGE for cost; on Databricks, SELECT on the catalogs in scope; on BigQuery, metadata viewer and data viewer on the datasets in scope; a read token for your catalog. Credentials stay in your environment, and you bring your own model key. Every system is confirmed with a live connection test before anyone calls it connected.

What Data Workers does. Data Context Wizard builds one governed context graph from your warehouses, dbt project, orchestrator, BI and the catalog you already run. The quality agent sets baselines for volume, freshness, uniqueness and nulls on the tables your owners name, and the cost agent reads Snowflake metering and attributes credits to the query and dbt model behind them. Your engineers query all of it from the coding agent they already use: trace_cross_platform_lineage follows a column from source to dashboard, explain_table returns definition, lineage and trust score, and get_quality_score reports quality against the new baselines.

What agents may do. Read and report. Nothing writes. Agents diagnose and suggest fixes, and the permission ladder logs what approval each action would have needed, so the record starts building before anyone approves anything. Your team reviews it against what it actually did.

Measures at the end of week 2. Coverage: how many of the tables and pipelines in scope have lineage, an owner and a quality baseline. Your team's first review of suggested fixes against what it did. A list of open incidents and cost findings for weeks 3 to 6 to work on.

Weeks 3 to 6: propose (L2)

What Data Workers does. The Autonomous Data-Conductor runs its loop on live signals: detect, diagnose, fix, review, verify and remember. When a quality check fails or a pipeline breaks, the incident agent runs diagnose_incident and get_root_cause, blast_radius_analysis lists every downstream model, dashboard and export the fix would touch, and the agent writes a proposal: the change, dry-run counts, the blast radius and the rollback plan. Cost findings arrive the same way, with dependencies checked first.

What your team does. Owners approve, edit or reject in the coding agent they already use, or in Spellbook Data Catalog on Enterprise (in preview). Approval requests carry timeouts and escalation, so nothing waits forever and nothing proceeds without a named person. An agent can't approve its own proposal.

Measures at the end of week 6. Proposals approved unchanged, approval wait, time from alert to verified fix, how often suggested fixes matched your team's, and every override with the reason. These numbers decide which domains go to L3.

Weeks 7 to 12: act reversibly in the domains you pick (L3)

Choosing the domains. Pick one or two where the week 3 to 6 record is strongest and the changes are easy to reverse. Late and failed loads are a common first choice: rerunning a failed Dagster or Airflow task and rebuilding the dbt models behind it is routine and reversible. Access requests that fit an existing policy are another. Finance models, PII tags and masking, and migrations stay at propose, by design, until their own records earn more.

What changes. In the L3 domains, the agent acts, verifies and writes the receipt, and the owner reviews after the fact. Credentials are scoped per agent. A rollback path is recorded before every change. A global kill switch halts every agent at once, and your team can lower any domain back to L2 at any time.

Measures at the end of week 12. The same six, now with a quarter of history. Your team reviews the agents' records and receipts against the bar it set. A reasonable starting bar to leave observe-only is at least 100 runs, 95% agreement with your team, under 1% errors and under 5% human overrides; your team owns the decision for each domain.

Six numbers to report monthly as a Your data stack estate climbs the autonomy levels, and which way each should move

The measurement page goes deeper on each measure and how to set a baseline both sides agree on.

An illustration: a team's first incident, end to end

This is an illustration, not a customer case. It's week 4 of a pilot. The team runs Salesforce into BigQuery through Airbyte, dbt models orchestrated by Dagster, and Looker for sales reporting. The revenue domain is at L2 propose.

Incident timeline across the stack: what Your data stack, your team and Data Workers each do, step by step

At 23:40 on Tuesday an admin runs a refresh on the Airbyte Salesforce connection and keeps existing records. The Opportunity stream syncs in Incremental | Append mode, so the refresh writes every opportunity again on top of the rows already there. By 23:55 every opportunity row sits in raw.sf_opportunity twice. At 01:15 the nightly Dagster run materializes fct_bookings, and the run is green, because nothing failed.

At 01:20 the uniqueness check on opportunity Id fails and bookings land at twice the baseline set in week one. That baseline is why the problem surfaces at 01:20 instead of at the 09:00 sales standup. By 01:22 the context graph has traced the duplicates to the refresh and listed the blast radius: four dbt models, two Looker Explores and a board export. At 01:25 the agent proposes a dbt diff that deduplicates the staging model on Id, with dry-run row counts and a revert as the rollback, and routes it to the revenue owner.

At 08:05 the analytics engineer reads the diff and the blast radius and approves. The change merges at 08:10 and the job rebuilds the four affected models. At 08:40 the verify step confirms one row per opportunity Id and bookings totals match the Salesforce report, and the receipt records the cause, the diff, the approver, the checks and the rollback path. The sales standup opens on correct numbers.

How other paths handle the first 90 days

Buyers weigh several paths, and each is good at its job.

  • •Build it yourself. A coding agent plus vendor MCP servers gives you a working demo in the first week. The quarter then goes into lineage across engines, approvals, rollback and receipts. The build-vs-buy page lays out that estimate.
  • •Observability. Monte Carlo's agentic onboarding has agents propose a monitoring plan, and "nothing deploys until you approve it." Monte Carlo puts setup and coverage on a first domain at 30 days end to end, with monitors live by week 3. Its alerts feed the Data Workers loop, which carries the incident through fix and verification.
  • •Catalogs. Atlan's remote MCP server offers a read-only mode, and its SQL tools pass a validator that blocks DDL and DML. Good governance for metadata; changes to the data itself are a different product.
  • •Platform-native agents. Snowflake's managed MCP server and Databricks managed MCP servers (Public Preview) are governed for their own platform. Data Workers works across them, with one approval flow.
  • •Coding agents. Claude Code, Cursor and Codex stay where your engineers work. Data Workers is what they call over MCP when the work touches production data.

The case for your CFO

The goal for the first quarter is a data team that fixes production data problems with an audit trail, spending less engineering time on them, without migrating anything. The plan is staged so the risk is small at every step: two weeks read-only, then a month where every change needs a named person's approval, then reversible action only in the one or two domains with the strongest record.

The risk story is concrete. At L1 agents read. At L2 they propose, and an owner approves anything that changes data. At L3 they act inside a scope your team set, every change carries a rollback path and a receipt with its diff, approver and checks, and a kill switch stops every agent at once. Data Workers runs in your infrastructure, and your warehouse, dbt project, orchestrator, BI and catalog stay exactly as they are.

Why now: your engineers already run coding agents against production, and a staged quarter puts approvals and receipts around that work.

The first win the plan aims for is incident triage at L2 inside the first month: a real incident, diagnosed across systems, fixed after one approval and verified.

The path is a pilot: $7,500 one-time with a forward-deployed engineer, run over the pilot window on /pricing/, and the pilot is credited in full against the first year. After it, Scale starts from $1,000/month and Enterprise from $3,000/month, billed annually, with unlimited seats, no usage meter and no markup on model spend. See pricing and the ROI calculator, and the ROI page for the model behind it.

The sentence to repeat upstairs: "Our plan takes us from read-only to reversible fixes in one or two domains in a quarter, every change approved or receipted, and we only move further when the record says so."

FAQ

How much of my team's time does the first quarter take? A pilot lead, plus one owner per domain who reviews proposals. The forward-deployed engineer does the connection and configuration work alongside your team.

What access do you need on day one? Read-only, least-privilege grants on the systems in scope. Write-scoped credentials come later, only for the domains you move to L3. The security and deployment page covers where data and credentials live.

When does an agent first change production data? In weeks 3 to 6, and only after a named owner approves the specific change. Agents act without a prior approval only in the L3 domains you choose in weeks 7 to 12, and only reversibly.

What happens if an agent is wrong? At L2, the owner rejects the proposal and the override is recorded. At L3, the recorded rollback path undoes the change and the domain can drop back to L2. The safety page walks through each gate.

Can we stay at L2 forever? Yes. Autonomy is your team's decision per domain. Many domains, such as finance models, are well served at propose and approve.

What if our catalog or observability tool stays? It stays. Data Workers reads the catalog you already run and takes alerts from your observability tool into its loop. If you're starting from AI assistants your company already rolled out, the assistants hub and bring your own context show how that context plugs in. For the guardrail patterns behind approvals, see approval gates for data agents and audit trail requirements.

What happens after day 90? Teams that continue move to Scale or Enterprise, with the $7,500 pilot fee credited in full against the first year, and the climb continues one domain at a time, including L4 where a domain's record earns it.

Sources

  • •Data Workers, Pricing: Pilot Program $7,500 one-time with a forward-deployed engineer, credited in full against the first year; Scale and Enterprise tiers. Rate card as published September 10, 2026; checked October 2, 2026.
  • •Data Workers, client setup and what the platform adds. Checked October 2, 2026.
  • •Data Workers, open-source core on GitHub: public tool registrations for trace_cross_platform_lineage, explain_table, blast_radius_analysis, diagnose_incident, get_root_cause and get_quality_score. Checked October 2, 2026.
  • •Airbyte, Refreshes: "Refresh stream and retain records" is available on Incremental | Append streams. Checked October 2, 2026.
  • •Monte Carlo, Agentic Onboarding: approve-before-deploy and 30 days for a first domain. Checked October 2, 2026.
  • •Monte Carlo, Cost agent docs: recommends and does not act on your behalf. Checked October 2, 2026.
  • •Atlan, Remote MCP security: read-only mode and the SQL validator. Checked October 2, 2026.
  • •Snowflake, Snowflake-managed MCP server. Checked October 2, 2026.
  • •Databricks, Databricks managed MCP servers: Public Preview, Unity Catalog permissions. Checked October 2, 2026.