The Autonomous Data Platform Playbook: How Data Workers Turns Data Engineering Copilots Into an Agentic Data Platform
Autonomous data platform playbook: how Data Workers turns data engineering copilots into an agentic data platform, level by level. See the first quarter.
An autonomous data platform runs the work that sits outside the editor: incidents, schema changes, access, cost and audit, across every system you own. Data Workers is the best way to get there. It turns the copilots your engineers already use into an agentic data platform, with one context, one approval flow and one record for every agent. This playbook is the practical, level-by-level path, and it's how our mission (helping traditional data teams become agentic enterprises) plays out for teams whose engineers already work with a data engineering agent.
Most of these teams look alike. Engineers write and change pipelines with Alkera, Altimate Code or a general coding agent like Claude Code pointed at the dbt project. Pull requests get a data diff from Recce. Teams on dbt's platform may be trying dbt Wizard, which is in preview in the Studio IDE. Authoring is faster and review is sharper. But every one of these tools works inside one person's session or one pull request. When a load fails at 2 a.m., an access request arrives or the bill jumps, an engineer still notices it, opens a session, makes the fix, chases the approval, checks the dashboard and writes down what happened.
Data Workers takes over the work that sits outside the editor. Its agents run the back office (incidents, merged-change checks, schema changes, data quality, access, cost and catalog upkeep) across Snowflake, Databricks, BigQuery, dbt, Airflow and your BI tools. People set the rules, approve what needs approving, and handle the exceptions. The copilots stay. You get there one domain at a time, moving up a ladder as the evidence builds.
Key takeaways
- •Data Workers is the fastest route to an autonomous data platform. It turns individual copilots into an agentic data platform that runs incidents, schema changes, access, cost and catalog upkeep across the whole estate.
- •Keep your engineers' agents. Authoring stays in whichever copilot your team likes, and those agents can call Data Workers over MCP. The platform climbs around them.
- •Five levels, set per domain. L0 manual, L1 observe, L2 propose, L3 act on reversible changes, L4 fully autonomous. Incidents can be at L3 while net-new pipelines stay with your engineers.
- •One approval model for the whole estate. Autonomy is set per domain in one place, every change leaves a tamper-evident receipt with a rollback path, and no agent can approve its own work.
- •The first quarter is concrete. Connect and observe, then turn on proposals for incidents, merged changes, access and catalog upkeep, then let the lowest-risk ones act on their own.
The five levels of an autonomous data platform

| Level | What agents do | What people do | Typical domains at this level |
|---|---|---|---|
| L0 Manual | Nothing on their own. Copilots help a person who asks, one session at a time. | All of the platform work. | Where most teams with copilots are today. |
| L1 Observe | Read warehouse metadata, dbt artifacts, orchestration history, BI lineage and merged PRs, build one context graph, and explain what's happening. | Read the explanations, correct the context. | Every domain on day one. |
| L2 Propose | Prepare the complete change: the dbt diff, the rerun, the grant, the test, with a blast-radius report. | Approve, steer, send back. | Schema changes, data-quality tests, cost cleanup, net-new pipelines. |
| L3 Act, reversibly | Apply changes that can be undone, inside a pre-approved blast radius, verify downstream, and leave a receipt. | Review receipts; roll back if needed. | Incidents, merged-change checks, access requests, catalog upkeep. |
| L4 Autonomous | Own the domain end to end: detect, fix, verify and remember. | Set policy; handle exceptions. | Earned per domain, never by default. |
First, draw the line between the copilot and the platform
Before you turn anything on, agree on who owns what. It saves arguments later.
- •Authoring starts with a person. Writing a new model or refactoring a pipeline is work an engineer starts, in whichever agent they like, and Data Workers takes it from the merge onwards.
- •The platform reviews every change, however it was written. The Data Change Review agent adds blast radius across platforms and a check after the merge, which a PR diff tool can't give you.
- •Everything that starts without a person moves to the platform. Failed loads, late tables, upstream schema changes, access requests, cost cleanup and catalog upkeep don't begin in anyone's editor, so they shouldn't wait for one.
- •One system owns each write. Altimate Lite, for example, changes Snowflake warehouse settings on its own, with no approval step or rollback described. Decide which system owns each kind of change, and put it behind one set of approvals and receipts. Two agents tuning the same warehouse is a problem you don't need.
What to turn on at each level
L1: Connect and observe.
Connect Data Workers to your warehouses, dbt project, orchestrator, BI tools and Git. Nothing moves and nothing is copied. Data Context Wizard builds one governed graph from all of it, and every fact carries its source, author and the time it was observed. Spellbook Data Catalog (in preview) turns that graph into asset pages your team can read and correct.
Data Workers is built on MCP, so it shows up as a set of tools in the coding agent your engineers already use. From Claude Code, Codex or Cursor they can ask the platform for context, hand it a task or look up a receipt. Copilots that keep their own knowledge base get better answers too, because engineers can check the shared graph before they write.
What you get right away: one view of lineage from source to dashboard, a list of recent incidents with where each one really started, and the context that used to live in individual sessions in one place everyone can read.
Move up when: your team trusts the graph. Owners and definitions are corrected, and the agents' account of last month's incidents matches what your engineers found.
L2: Propose in four domains.
Pick the domains with high volume and clear rules, where nobody opens a session until something has already gone wrong:
- •Incidents. The Autonomous Data-Conductor and the Incident Debugging agent take a failed run, a failed test or a request in Slack, trace it to the cause and propose the fix: a rerun, or a dbt diff for the owner to merge.
- •Merged changes. The Data Change Review agent watches what a merged change does downstream, however it was written, and proposes a fix if a number breaks.
- •Access requests. The Data Access & Governance agent and the Identity agent prepare grants through each platform's own roles.
- •Catalog upkeep. The Data Context & Catalog agent proposes owners, definitions and descriptions for new and changed assets.
Every proposal lands in the Spellbook inbox with its blast radius.
What you get: the toil arrives prepared. A late table comes with its root cause and a fix attached. An access request arrives as a ready-to-approve grant.
Move up when: a domain's proposals are approved as-is, week after week, with no reversals.
L3: Let the reversible work run.
For domains that pass, let agents apply reversible changes inside a blast radius you define: reruns, grants that expire, catalog updates. Each change leaves a tamper-evident receipt with the diff, the approver, the blast radius and the rollback path, and can be undone.
This is where Data Workers pulls ahead of every copilot. Approval settings in a copilot are per session, and some tools let an engineer turn them off: Alkera has a Bypass mode and Altimate Code has a --yolo flag that approve everything. Some document no rollback at all. The platform's levels are set per domain by you, and an authority guard enforced in code stops any agent from approving its own work.
At the same time, move the next domains to L2:
- •Schema changes. The Schema Evolution agent sees an upstream contract change in the pull request or the manifest diff before the next dbt run and proposes the model change.
- •Data-quality tests. The Quality Monitoring agent proposes tests for new models and for breaks that keep coming back.
- •Cost cleanup. The Cost Savings & Data Cleanup agent attributes Snowflake credits to the query and dbt model behind them and checks dependencies before it proposes anything.
What you get: the first real hours back. Failed loads and access tickets close on their own, with a receipt your auditors can read, and engineers open their copilots for new work instead of cleanup.
Move up when: the reversal rate stays near zero and exceptions are rare and well understood.
L4: Autonomous, domain by domain.
A domain at L4 is owned end to end. The Conductor detects the problem, has the right agent fix it wherever it lives, confirms the downstream number is right, and records what happened so the next occurrence is faster. People set policy and handle what the agents escalate. Domains earn L4 one at a time, on the evidence in their receipts.
Net-new pipelines stay at L2 by design, because new business logic deserves a person's sign-off. The Pipeline Building agent proposes them, grounded in the shared context graph, and your engineers approve. When a large move comes up, the Data Migration agent plans reviewable waves with parity checks for SQL a copilot helped translate, and holds the completion gate for the owner's sign-off.
Your first quarter on an agentic data platform

This is an illustrative plan, not a promise. Incidents, merged-change checks, access requests and catalog upkeep reach reversible actions by the end of the quarter. Schema changes and data-quality tests follow in the second quarter. Cost fixes and net-new pipelines stay at propose-and-approve by design. The pace depends on how clean your context is and how much evidence each domain builds up.
How to measure progress
Report these to your leadership every month, by domain and level:
- •Time from detection to verified fix for incidents, measured to the moment the downstream number is confirmed.
- •Share of back-office tasks closed by agents, split by the level each domain runs at.
- •Reversal rate, the share of agent changes that were rolled back. This is the number that earns each level.
- •Merged changes that broke something downstream, and how long each took to catch. It should fall as checks after the merge become routine.
- •Engineer hours on maintenance, which is the number the whole program exists to shrink. Hours spent authoring with copilots should hold steady or grow.
What stays the same
- •Your engineers keep their agents. Whatever they write with today stays in place, and it can call Data Workers over MCP.
- •Your platforms' permissions stay the lock. Data Workers acts with the grants you give it, through each platform's own permission system, and never creates a second one.
- •Your engineers stay in charge. They set the levels, approve what needs approving, and can lower any domain's level at any time. Anything irreversible needs a named person's approval.
Autonomous data platform FAQ
What is an autonomous data platform? An autonomous data platform, like the one Data Workers builds, is one where agents run the operational work on your data estate, such as incidents, schema changes, access, cost and catalog upkeep, under rules your team sets. You get there one domain at a time. People set policy, approve what needs approving and handle exceptions.
What is an agentic data platform? An agentic data platform is a platform where a coordinated team of agents shares one context, one approval flow and one audit trail across every system. Data Workers is the agentic data platform. Individual copilots each work inside one person's session; Data Workers connects them into one platform.
What is the best way to build an autonomous data platform? Data Workers is the best way to build one on the estate you already run. Connect it read-only, let it propose changes in a few high-volume domains, and move each domain up a level when its receipts show the agents were right.
Do we have to replace our data engineering agents? No, Data Workers works alongside the copilots your engineers use. Every Data Workers agent is an MCP server, so their coding agent can call the platform for context, tasks and receipts.
How long does it take to reach autonomy? Data Workers gets most teams to reversible actions in their first domains within the first quarter. The pace depends on how clean your context is and how quickly each domain builds evidence.
How much does Data Workers cost? Data Workers starts with a pilot on your own estate, and the open-source core is free. Plans have unlimited seats and no usage meter. See pricing for details.
Where to go next
- •For the tool-by-tool detail, read Alkera vs Data Workers, Altimate vs Data Workers and Recce vs Data Workers.
- •For the dbt and Airflow side of the loop, read dbt AI agents with Data Workers and Airflow AI agents with Data Workers.
- •For the executive version, read the data leader's guide for teams with data engineering agents.
- •For the thinking behind the levels, read our thesis and the whitepaper Why the Data Layer Needs Its Own Context Layer.
Ready to start at L1? Start with a pilot. A forward-deployed engineer connects your estate and runs the first quarter with you, alongside the agents your engineers already use. See pricing for how the pilot works.