Product
Product8 min readBy The Data Workers Team

What does a pilot look like?

A Data Workers pilot is $7,500 one-time, credited in full against the first year. It runs with a forward-deployed engineer and every agent live on your estate, and ends with a readout against measures you agree at the start.

A Data Workers pilot costs $7,500 one-time, and the fee is credited in full against your first year. It runs over the pilot window published on /pricing/ with a forward-deployed engineer, puts every agent live on your own estate, and ends with a readout against measures your team and ours agree before anything connects.

The pilot is the Start plan on the rate card, built for one decision: whether agents should run your data operations, judged on your incidents, your spend and your team's hours.

Key takeaways

  • •Fixed price, fully credited. $7,500 one-time, credited in full against the first year on Scale or Enterprise. Unlimited seats, no usage meter, no markup on model spend.
  • •Your estate, every agent. /pricing/ lists "Every agent live against your data, inside your coding agent" and "The Autonomous Data-Conductor, running", with a forward-deployed engineer.
  • •Five steps. Scoping call, connect and observe, propose on real incidents and cost, act reversibly in one or two domains, then a readout and a decision.
  • •Measures first. Estimated hours returned, incidents resolved and verified, warehouse spend and approval rate are agreed on the scoping call and baselined before agents propose anything.
  • •Risk stays small. Read-only grants first, a named owner approves every change at L2, and only the domains you pick act reversibly at L3.

What the pilot includes

The Start plan card on /pricing/ reads "$7,500 one time", "six to twelve week pilot", "For data teams putting agents into a production stack for the first time." The page's structured data calls it "a fixed-price pilot with a forward-deployed engineer" and says it "runs six to twelve weeks with a forward-deployed engineer, every agent live on your data, and governed writes with a receipt on every change." The Start card lists, word for word:

  • •"A pilot program with deployment support"
  • •"With a forward deployed engineer dedicated to build your custom needs"
  • •"Every agent live against your data, inside your coding agent"
  • •"The Autonomous Data-Conductor, running"
  • •"Data-Layer Context Graph and Data Context Wizard"
  • •"Connected to your existing data stack on day one, even custom tools"
  • •"Approval workflows and audit trails, configured with your team"

The plan comparison on the same page lists, for Start, the living context graph as "Built for you", connectors as "Wired up for you", "Governed writes, with a receipt on every change" as included, seats as "Unlimited" and the place it runs as "Your infrastructure". The agents run in your environment and hold your warehouse credentials and model key. Your data stays in your systems; the Conductor, hosted by us, sees workflow metadata only.

So the forward-deployed engineer is dedicated to building your custom needs, and the agents do their work on your stack.

The shape of a pilot, step by step

First-quarter autonomy roadmap starting from Your data stack

The chart shows a common scope. Your team picks the areas and the one or two domains that act; the rest stay at propose by design. Autonomy is set per domain on one ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous

1. Scoping call. Your pilot lead, the domain owners and Data Workers agree on three things: the systems in scope, the domains that matter most, and the measures the readout will use, with the bar each one has to clear. Nothing connects until those are written down.

2. Connect and observe (L1). Your team grants least-privilege read access: on Snowflake a role with USAGE on a warehouse, SELECT on the schemas in scope and read on SNOWFLAKE.ACCOUNT_USAGE for cost; on Databricks SELECT on the catalogs in scope; on BigQuery metadata viewer and data viewer. Each system passes a live connection test before it counts as connected. Data Context Wizard builds the context graph across warehouses, dbt, the orchestrator, BI and the catalog you already run. Quality baselines (volume, freshness, uniqueness, nulls) and cost baselines are set here, and they become the "before" numbers for the readout. Agents diagnose and suggest fixes, and the permission ladder logs what approval each action would have needed; your team reviews that record against what it did.

3. Propose on real work (L2). The Autonomous Data-Conductor runs its loop on live signals: detect, diagnose, fix, review, verify and remember. Incidents get diagnose_incident, get_root_cause and blast_radius_analysis before any fix is written, and the proposal carries the change, dry-run counts, the blast radius and a rollback plan. Cost findings arrive the same way, dependencies checked first. A named owner approves, edits or rejects each one inside the coding agent your team already uses. Requests time out and escalate; nothing auto-grants.

4. Act reversibly in one or two domains (L3). Where the propose record is strongest and changes are easy to undo, the agents act, verify and write the receipt, and the owner reviews after the fact. Late and failed loads are a common first choice. Credentials are scoped per agent, a rollback path is recorded before every change, and the org-wide stop halts all autonomous dispatch at once.

5. Readout and decision. The readout compares the agreed measures against the baselines, using the receipts, approval records and audit log the pilot produced. Then your team decides: continue on Scale or Enterprise with the fee credited, extend a domain's autonomy, or stop.

The first 90 days page carries the same plan past the pilot, phase by phase.

What the readout measures

Six numbers to report monthly as a Your data stack estate climbs the autonomy levels, and which way each should move

Set the success criteria before anything connects, and write a target for each. These are the measures we recommend; the bars are yours.

  • •Hours returned. Hand-work hours on incidents, tickets and cleanup the agents now carry, estimated from the incidents and requests they closed against the time your team logged for the same classes in the baseline.
  • •Incidents resolved and verified. Incidents closed with a fix whose post-change checks passed, each with a receipt. Pair it with time from alert to verified fix.
  • •Warehouse spend. Snowflake credits on the models whose drafted fixes owners approved, before and after. Data Workers' design target for cost savings is 25 to 40%; it is a target, and your bar can be set on your own baseline.
  • •Approval rate. The share of proposals a named owner approved without edits. Track overrides and rollbacks next to it, each with its reason.

Your team reviews the agents' records and receipts against the bar it sets. For moving a domain beyond observe-only, a reasonable starting bar is at least 100 runs, 95% agreement with your team, under 1% errors and under 5% human overrides. The measurement page covers each measure and how to set a baseline both sides accept.

An illustration: one pilot, start to finish

This is an illustration, not a customer case. A data team runs Fivetran into Snowflake, dbt models orchestrated by Airflow, and Looker for reporting. On the scoping call they pick two domains to act reversibly, late loads and access requests, and keep finance models at propose.

StepWhat happensLevel
Scoping callOwners named for loads, access and finance. Bars written for hours returned, incidents verified, spend and approval rate.L0 manual
Connect and observeRead grants on Snowflake, dbt and Airflow pass live tests. Context graph built; freshness and cost baselines set. Ladder logging what each load fix would have needed.L1 observe
ProposeA Fivetran connector adds a renamed column; nulls spike in a dbt staging model. Data Workers traces it, lists two Looker Explores in the blast radius, and proposes a diff the owner merges. A Snowflake warehouse's credits are traced to the dbt models behind them, and a drafted auto-suspend setting goes to its owner, dependencies checked.L2 propose
Act reversiblyAn Airflow DAG run fails on a warehouse timeout. The agent reruns the failed task, rebuilds the dbt models behind it, verifies freshness and writes the receipt. The owner reviews in the morning.L3 act reversibly
ReadoutThe four measures against their bars, read from receipts and approval records. The team chooses Scale and keeps finance at propose.Decision

Every change in that run carries a receipt: the cause, the diff, the approver, the checks and the rollback path.

How other ways of evaluating agents compare

Buyers weigh several paths, and each is good at its job.

  • •Build it yourself. A coding agent plus vendor MCP servers gives you a working demo in a week. The pilot question then becomes lineage across engines, approvals, rollback and receipts, which is what the build-vs-buy page prices out. The Data Workers open-source core is free to clone and read before any money moves.
  • •Observability onboarding. Monte Carlo's Agentic Onboarding has agents propose a monitoring plan, and "Nothing deploys until you approve it"; its FAQ says "Setup and coverage on your first domain runs 30 days end to end, with monitors live and detecting by week 3." Its alerts feed the Data Workers loop, which carries an incident through fix and verification.
  • •Catalogs. Atlan's remote MCP server offers a read-only mode, and its SQL tools pass a validator that blocks DDL and DML. Good governance for metadata; changes to the data itself are a different product.
  • •Platform-native agents. Databricks managed MCP servers (Public Preview) and the Snowflake-managed MCP server are governed for their own platform. A Data Workers pilot runs across them with one approval flow.
  • •Coding agents. Claude Code, Cursor and Codex stay where your engineers work, and the pilot runs inside them. The coding agents page shows the setup.

The case for your CFO

The pilot buys a decision on evidence. For $7,500, credited in full against the first year, your data team finds out on its own estate whether agents can carry incidents, cleanup and requests that engineers handle by hand today, with estimated hours returned, spend and approval rate read against bars set before anything connects.

The risk story is concrete. Agents start read-only. At L2 they propose, and a named owner approves anything that changes data. At L3, only in the one or two domains your team picks, they act reversibly, with a rollback path recorded first and a receipt on every change. No agent can promote its own work, and an org-wide stop halts all autonomous dispatch. Your warehouse, dbt project, orchestrator, BI and catalog stay exactly as they are; nothing migrates.

Why now: your engineers already run coding agents against production. A pilot puts approvals, receipts and a measured readout around that work.

The first win to aim for is one real incident diagnosed across systems, fixed after one approval and verified, early in the propose step.

After the readout, Scale starts from $1,000/month and Enterprise from $3,000/month, billed annually, with unlimited seats, no usage meter, no markup on model spend and your own model. On Scale at its published floor the first year lists at $12,000, and the pilot fee counts toward it. See pricing, the ROI calculator and the ROI page.

The sentence to repeat upstairs: "We pay a fixed $7,500 that counts toward year one, and we only continue if the readout clears the bars we set before it started."

FAQ

How long does the pilot run? /pricing/ describes it as a "six to twelve week pilot". The scoping call sets the plan for your estate, and the readout lands at the end of the window.

What does the forward-deployed engineer do? The Start card on /pricing/ lists "With a forward deployed engineer dedicated to build your custom needs", and the page's structured data describes "a fixed-price pilot with a forward-deployed engineer". The plan comparison lists the context graph as "Built for you", connectors as "Wired up for you", and approval workflows and audit trails "configured with your team".

How much of my team's time does it take? A pilot lead, plus one owner per domain in scope who reviews proposals and joins the readout.

What access do you need, and where does our data go? Read-only, least-privilege grants first; write-scoped credentials only for the domains you move to L3. The agents run in your infrastructure. The security and deployment page covers what the hosted Conductor sees.

What if an agent gets something wrong during the pilot? At L2 the owner rejects the proposal and the override is recorded. At L3 the recorded rollback path undoes the change and the domain can drop back to propose. The safety page walks through each gate, and approval gates for data agents covers the patterns.

Do we need Spellbook Data Catalog for the pilot? No. Approvals happen in the coding agent your team already uses. Spellbook, in preview, comes with Enterprise. If your company already rolled out AI assistants, the assistants hub and bring your own context show how they connect.

What happens after the readout? Teams that continue move to Scale or Enterprise with the $7,500 credited in full against the first year, and the climb continues one domain at a time.

Sources

  • •Data Workers, Pricing: Start plan card and plan comparison, visible copy ("six to twelve week pilot", "Term: Six to twelve weeks", "Runs in: Your infrastructure", "Built for you", "Wired up for you"), and FAQPage structured data ("a fixed-price pilot with a forward-deployed engineer"; "credited in full against your first year"). Rate card as published September 10, 2026; fetched October 2, 2026.
  • •Data Workers, Should you use Data Workers? An evaluation guide: Pilot Program $7,500 one-time with a forward-deployed engineer, credited in full against year one. Fetched October 2, 2026.
  • •Data Workers, Pricing and licensing docs: $7,500 credited in full against the first year. Fetched October 2, 2026.
  • •Data Workers, open-source core on GitHub: registrations for diagnose_incident, get_root_cause and blast_radius_analysis. Checked October 2, 2026.
  • •Monte Carlo, Agentic Onboarding. Fetched October 2, 2026.
  • •Atlan, Remote MCP security. Fetched October 2, 2026.
  • •Databricks, Databricks managed MCP servers: Public Preview, page updated September 21, 2026. Fetched October 2, 2026.
  • •Snowflake, Snowflake-managed MCP server. Fetched October 2, 2026.