What do the autonomy levels mean, from L0 to L4?
L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous, set per domain. What the agent may do at each level, what a person does, what gets recorded, and the evidence that moves a domain up or down.
The five autonomy levels in Data Workers are L0 manual, L1 observe, L2 propose, L3 act reversibly and L4 autonomous, and you set one per domain, so finance can sit at L2 while freshness fixes in marketing run at L3. Each level is defined by four things: what the agent may do, what a person still does, what gets recorded, and what evidence moves the domain up a level or back down.
That last part is what makes the ladder useful. Autonomy in Data Workers, the agentic data platform, is a setting your team holds and moves on the record, one domain at a time, with receipts behind every step up and a clear trigger for every step down.
Key takeaways
- •Five levels, one per domain. L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous. A domain is a slice of work such as freshness, schema changes, access requests or cost, owned by a named person.
- •Each level answers four questions. What the agent may do, what a person does, what is recorded, and what evidence moves the domain.
- •Every deployment starts observe-only. Agents read, trace and explain, and the permission ladder logs what each action would have needed, so your team can review the record against what it did.
- •Some actions always need a person. With ladder enforcement turned on, deletes, elevated access, spend and other irreversible actions go to a named approver at every level, L4 included.
- •Levels move down as easily as up. A failed check, a rollback, a budget cap or a runaway run rate pulls a domain back, and your team can lower or pause any domain at any time.
The ladder, level by level
The Autonomous Data-Conductor ships observe-only, and you open the rope domain by domain: which agents can write, how far a change may spread and what still comes to a person. The dial goes all the way to fully autonomous.

L0 manual. People do the work and agents stay out of the domain. Teams use it for freezes, pending audits or hand-run migrations, and your own change process does the recording. The evidence to move up is a decision: the owner agrees the agent may watch.
L1 observe. Agents read, trace lineage, diagnose and explain. They answer from the governed context, run run_quality_check, diagnose_incident and get_root_cause, and diagnose_incident suggests the fix for each incident. The permission ladder logs what approval each action would have needed; nothing writes. What is recorded: findings, suggested fixes and the ladder's verdicts. The evidence to move up is agreement, judged by your team: it reviews the agents' records and receipts against what on call actually did, against a bar it sets. A reasonable starting bar: at least 100 incidents reviewed, the suggested fix right at least 95% of the time, errors under 1% and human overrides under 5%.
L2 propose. Agents prepare the change and a named person approves it. A proposal carries the blast radius from blast_radius_analysis, a dry run where the playbook supports one, and the rollback plan. It lands in the approvals inbox in Spellbook (in preview), where the owner approves, edits, sends it back or rejects it. No agent can approve its own work, and an approval nobody answers expires or escalates; it never grants itself. What is recorded: the proposal, the approver, the outcome and the post-checks, in the tamper-evident audit chain. The evidence to move up: proposals approved unchanged and every applied fix passing its checks. How approvals work for AI data agents goes deeper.
L3 act reversibly. Agents apply changes that can be undone, inside the scope your team set, then verify and report. The rollback path is recorded before the change runs, and if a multi-step fix fails partway, Data Workers rolls back the completed steps and escalates. The owner reviews receipts after the fact and can undo any change from Spellbook. What is recorded: a receipt with the diff, the rollback path, the checks and the time. The evidence to move up: rare rollbacks, passing checks, and owner reviews that find nothing to change. How to roll back an AI agent change walks the undo path.
L4 autonomous. Agents close the loop end to end in the domain: detect, diagnose, fix, verify and remember, across a wider blast radius. People set goals and budgets, read the receipts and take the escalations. What is recorded: every receipt, plus the guardrail events around them. L4 is earned per domain, and the most severe action class still goes to a person.

How the ladder works under the hood
Every tool in Data Workers carries a consequence class. Reversible enrichment such as descriptions and tags is the mildest, then semantic assertions such as marking a table authoritative, then compensable changes to the live stack such as fixing a pipeline or altering a schema. The most severe class covers deletes, elevated access, money and anything irreversible. A tool not in the registry is treated as a live-stack change, so an unknown tool never slips into the mildest class.
The level and the class together decide whether a call acts or asks. Teams turn on ladder enforcement and set each domain's level; from then on every tool call resolves to act or to route to approval, and routed calls land in the same queue people already review, tagged with their class, policy and blast radius. The most severe class resolves to a person for every input: no level, mode or setting lifts it. Before enforcement is on, the ladder runs in shadow, logging the verdict it would have given for each call, so you have a record to review before you rely on it.
Levels change at runtime without a restart. A new tenant routes every governed write to approval until onboarding grants more, and every change of mode is an audited event naming who made it.
What moves a domain down is just as concrete. A global kill switch halts every agent, and a per-tenant pause stops one estate. A write-side anomaly breaker pauses a routine whose run rate jumps past twice its trailing average. A budget governor drops work to propose-only when a cost cap is reached, and spend always needs a person. A stale heartbeat pauses dispatch. These guardrails only ever tighten. Every change goes to a SHA-256 hash-chained audit log that auditors read through get_audit_trail or in Spellbook.
One worked example: a domain climbs from L1 to L3 in a quarter
This is an illustration, not a customer case or a measured result. The domain is freshness and failed-run recovery for revenue reporting. Fivetran loads Salesforce and Stripe into Snowflake, Airflow schedules dbt, and the sales team reads Looker's Revenue Explore every morning. The head of data owns the domain; an analytics engineer is on call.

| Weeks | Level | What the agent did | What people did | Evidence reviewed |
|---|---|---|---|---|
| 1 to 4 | L1 observe | Diagnosed every late sync and failed DAG run and suggested the rerun or backfill | Fixed each one by hand, as before | Owner reviewed 118 suggested fixes: 114 matched the on-call fix, 2 overrides with reasons |
| 5 to 9 | L2 propose | Sent each fix to Spellbook with blast radius, dry run and rollback plan | Approved 60 of 63 unchanged, edited 3 | Every applied fix passed its row-count checks and lateness baselines; the owner confirmed its dbt tests |
| 10 | L3 for reruns and single-partition backfills | Applied those fixes, verified them, wrote receipts | Reviewed receipts the next morning; schema fixes stayed at L2 | Receipts, rollback paths, check results |
| 11 | Backfills back to L2 | A backfill failed its quality re-check; Data Workers escalated it with the recorded undo, and the owner ran the undo | Owner moved backfills to L2 and fixed the check; reruns stayed at L3 | The failed check, the escalation and the undo receipt |
| 13 | Quarter review | Reruns at L3, backfills back at L3 after two clean weeks at L2 | Read the quarter's receipts with the owner | Revenue Explores fresh by 07:00 on review days |
The owner moved the level every time, on evidence in the receipts. The domain moved by action type, so reruns climbed faster than backfills. And the step down in week 11 cost an escalation and one recorded undo, not a bad number in Monday's revenue review. Who owns the agents covers the owner's job, and the first 90 days with Data Workers lays out the same climb as a pilot plan.
How this compares with published autonomy frameworks
SAE J3016, the driving levels, as an analogy. SAE's driving automation taxonomy (J3016_202104, revised April 30, 2021) defines six levels, from Level 0 (no driving automation) to Level 5 (full driving automation), and the operational design domain: the set of conditions in which a system is designed to do the driving. A Data Workers domain plays that role, so an agent can be at L3 for freshness reruns and L2 for schema changes at once. The difference: a driving feature's level is fixed by its design, while a data domain's level is set by its owner, moves on the record, and comes back down when a check fails.
Coding agents, such as Claude Code permission modes. Claude Code sets a permission mode per session, from Manual (config value default, only reads run without asking) through acceptEdits, plan, auto (a classifier reviews actions instead of the user) and dontAsk to bypassPermissions, for isolated containers and VMs only. From v2.1.283, auto mode is the starting mode for interactive terminal and VS Code sessions, and administrators can turn off auto and bypass modes in managed settings (docs checked Oct 2, 2026). That is a well-designed control for one person's session on a codebase. Production data changes need the level held per domain, a named approver who isn't the session's user, and a receipt across every system touched. Run Data Workers as an MCP server inside the coding agent and both apply; build it ourselves with Claude Code and MCP servers prices the alternative.
Platform-native agents, such as Genie Code approvals. Databricks Genie Code offers "Ask first" and "Auto-approve" modes plus per-tool choices, and Auto-approve, where an AI classifier checks each action against the user's request, is the default for every chat on first use. Its docs say auto-approve "is a productivity feature, not a security boundary" and advise keeping it off "when working with production data" (agent mode docs, updated Sep 25, 2026). A candid, sensible design for an assistant in one workspace. Data Workers sets the level per domain across Databricks and everything around it, with the approver and the receipt outside the session.
Observability tools. Monte Carlo's cost agent "recommends; it does not act on your behalf", and its docs say "Monte Carlo is read-only by design" (checked Oct 2, 2026). That is a permanent L1, the right posture for a smoke alarm. Data Workers takes the alert as input and runs the fix at whatever level the domain has earned.
For the controls themselves, see is it safe to let AI agents change production data, where your data goes, and our resources on approval gates for data agents and human-in-the-loop approvals.
The case for your CFO
The outcome is data work done without waiting for a person, in exactly the domains where the record says it is safe, so engineers spend their time on what needs judgment.
The risk story is the ladder itself. At L1 agents read. At L2 a named person approves every change. At L3 agents act only on changes that can be undone, with the rollback path recorded first. With ladder enforcement on, deletes, elevated access and spend go to a person at every level, and the kill switch, tenant pause, anomaly breaker and budget governor only ever tighten.
Why now: agents already reach production through coding assistants and platform agents. A per-domain level gives you one policy and one audit trail instead of a mode chosen in each chat session.
The first win is one domain, usually freshness or failed-run recovery, at L2 inside the first month. Your warehouse, dbt, orchestrator, BI and coding agents stay as they are; nothing migrates.
Start with a pilot: $7,500 one-time, and the pilot is credited in full against the first year. Scale is from $1,000 a month and Enterprise from $3,000 a month, billed annually, with unlimited seats, no usage meter and no markup on model spend. See pricing, the ROI calculator and our ROI of agentic data operations. The sentence for upstairs: "Agents earn autonomy one domain at a time, on receipts, and the riskiest changes always come to a person."
FAQ
Can an agent raise its own autonomy level? No. A person sets the level, per domain and down to a single agent or operation type, and every change of mode is an audited event naming who made it. The guardrails that act on their own, such as the anomaly breaker and budget governor, only lower autonomy.
What evidence should we require before moving to L3? At L1, suggested fixes that your team's review finds match what it did; then L2 proposals approved unchanged with every fix passing its checks. The starting bar above is yours to tighten.
What makes a domain drop a level? A failed post-check, a rollback, an owner's decision, or a guardrail: the kill switch, a tenant pause, the anomaly breaker or a budget cap. The step down is recorded like any other change.
Does L4 mean nobody reviews anything? No. People set goals and budgets, read receipts and take escalations, and the most severe action class still goes to a named approver.
Where should we start? L1 observe everywhere, then one domain at L2 propose. The first 90 days shows the plan week by week.
How is this different from Claude Code's auto mode or Genie Code's auto-approve? Those modes govern one person's session in one tool. The Data Workers level is held per domain by its owner, applies to every agent and system in that domain, and leaves a hash-chained receipt for every change.
Sources
- •Data Workers product repository,
data-workers-agent-swarmmain @ 0c2491e3 (Oct 2, 2026): core/enterprise/src (classes, permission-ladder, agent-middleware, autonomy, auto-mode-store, autonomy-controller, kill-switch, approval-workflow, audit), agents/dw-conductor/src/conductor-governor.ts, agents/dw-orchestration/src/remediation-saga.ts. Checked Oct 2, 2026. - •Data Workers public repository tools (
check_freshness,diagnose_incident,get_root_cause,blast_radius_analysis,get_audit_trail): https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026) - •Autonomous Data-Conductor: https://dataworkers.io/product/autonomous-data-conductor/ (checked Oct 2, 2026)
- •Spellbook Data Catalog: https://dataworkers.io/product/spellbook-data-catalog/ (checked Oct 2, 2026)
- •SAE International, J3016_202104, Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles: https://www.sae.org/standards/content/j3016_202104/ (revised Apr 30, 2021; checked Oct 2, 2026)
- •U.S. House Subcommittee on Highways and Transit, hearing memo listing the six SAE levels and citing J3016_202104, revised April 30, 2021: https://docs.house.gov/meetings/PW/PW12/20220202/114362/HHRG-117-PW12-20220202-SD001.pdf (Feb 2, 2022; checked Oct 2, 2026)
- •CEDR, report on distributed ODD awareness, defining the operational design domain per SAE J3016 (April 2021 edition): https://www.cedr.eu/docs/view/645cd354de69a-en (March 2023; checked Oct 2, 2026)
- •Anthropic, Claude Code permission modes: https://code.claude.com/docs/en/permission-modes (checked Oct 2, 2026)
- •Databricks, Genie Code agent mode: https://docs.databricks.com/aws/en/genie-code/agent-mode (updated Sep 25, 2026; checked Oct 2, 2026)
- •Monte Carlo, Cost agent: https://docs.getmontecarlo.com/docs/cost-agent (checked Oct 2, 2026)