How do we get our team to trust agents?
Trust in AI data agents is earned per domain, with evidence: start at L1 observe, review every change at L2 propose, let approval records and verified outcomes decide L3, and keep your team in charge of the ladder.
Your team comes to trust agents the way it trusts a new colleague: one domain at a time, on evidence it can check. With Data Workers you start at L1 observe so the team sees what the agents would have done, move to L2 propose so every change is reviewed with its evidence, let the approval record and verified outcomes decide when a domain earns L3 act reversibly, and keep rollback, receipts and the ladder itself in your team's hands.
The playbook is short: pick the first domain, write down the bar for moving up, review receipts every week, publish wins and misses, and make your most skeptical engineers the reviewers. Data Workers, the agentic data platform, is built to support every one of those steps.
Key takeaways
- •Trust is per domain. Freshness fixes can earn L3 while finance models stay at observe, on one ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous.
- •Evidence comes first, permission second. At L1 agents suggest fixes and the ladder logs what each action would have needed, and your team reviews that record against what it actually did; every L2 proposal carries its blast radius, diff, checks and rollback path.
- •The bar is written before the work starts. Your team sets it and reviews the agents' records and receipts against it. A reasonable starting bar to leave observe-only is at least 100 runs, 95% agreement, under 1% errors and under 5% human overrides.
- •Publish the misses, and let skeptics review. Rejections and their reasons sit in the same record as approvals, and the engineers who doubt the agents most give the sign-off that convinces everyone else.
- •The team keeps the dial. Any owner can lower a domain, pause the agents or roll back a change, and every change to a level is recorded.
Why trust is the real adoption question
These are third-party findings, not Data Workers figures. The 2025 Stack Overflow Developer Survey (49,009 responses, fielded May 29 to June 23, 2025; the latest edition) found that "more developers actively distrust the accuracy of AI tools (46%) than trust it (33%)," with only 3% highly trusting the output, and experienced developers the most cautious (2.6% highly trust, 20% highly distrust). The top frustration, cited by 66%, was "AI solutions that are almost right, but not quite." On AI agents specifically, 87% agreed they are concerned about the accuracy of the information agents provide, and 81% had concerns about the security and privacy of data.
That skepticism is healthy. Lee and See's 2004 review in Human Factors argues that the goal is calibrated trust, where reliance matches what the automation can actually do. Parasuraman and Riley (1997) named both failure modes: misuse, when people lean on automation past its abilities, and disuse, when they ignore automation that works. Dietvorst, Simmons and Massey found in 2015 that people lose confidence in an algorithm faster than in a person after seeing both make the same mistake, and in 2018 that people will use an imperfect algorithm when they can modify its forecasts, even slightly.
The lesson for a data team: trust grows when people can see the agent's record, errors included, and keep real control over what it does. Data Workers is built around exactly those two things.
How Data Workers earns trust, rung by rung

L1 observe: let the team see what the agents would have done. The Autonomous Data-Conductor "ships observe-only," and teams "open the rope domain by domain." Agents diagnose each incident and suggest the fix, and the permission ladder logs what approval each action would have needed; nothing executes. After a few weeks the domain owner reviews the agents' suggested fixes against the engineers' fixes and counts how often they matched.
L2 propose: every change reviewed with its evidence. At L2 the agent sends the change to the domain's named owner, carrying the blast radius from blast_radius_analysis across engines and dashboards, the diff, dry-run counts, the checks that ran and the rollback path recorded before anything runs. The owner can approve it, steer it, send it back or reject it: the "even slightly modify" finding, built into the workflow. Unanswered requests expire and escalate; nothing is auto-granted. How approvals work walks through each step.
L3 act reversibly: the record decides. Moving a domain up is your team's decision, made on the record. Every approval ticket records its approver, status, timestamps and the reason for any denial, and every applied change leaves a receipt with its diff, checks and rollback path in a SHA-256 hash-chained audit trail that reviewers read through get_audit_trail or in Spellbook Data Catalog (in preview). The team sets the counts from those records next to the bar it wrote, so the review is about numbers both sides can see.
The team stays in charge of the ladder. In Spellbook, owners "point the agents at a domain, scope how far they range, pause anything mid-flight," and can "roll it back after the fact, with the full trace of what the agent saw and why it acted." No agent can promote its own work to authoritative or canonical, and a fix whose checks fail goes to a person instead of being marked resolved. Teams turn on ladder enforcement, and from then on irreversible actions such as deletes, elevated access and spend wait for a named person at every level, L4 included. The org-wide stop halts all autonomous dispatch, and the two-person admin kill switch revokes the tenant's keys, tokens and sessions. The autonomy levels explainer covers what moves a domain down as well as up.
The playbook: five habits that build trust
1. Pick the first domain. Choose work that is frequent, well understood and easy to reverse, such as late loads and failed runs. Leave close-critical finance models at L0 or L1 for now. One domain with a clean quarter convinces more people than five with a muddy one.
2. Write the bar for moving up. Before the agents touch anything, the domain owner writes down what earns L2 and what earns L3: a minimum number of runs, an agreement rate, an override ceiling, a rollback ceiling. Start from Data Workers' defaults and tighten them as your team wants. A bar written in week one can't be moved to fit the result.
3. Review receipts every week. Half an hour, the owner and reviewers, the week's receipts on screen: what was proposed, approved unchanged, edited, or rejected and why. The measurement page lists the numbers to read and where each comes from.
4. Publish wins and misses. Post a short weekly note with links to receipts. A proposal the reviewer caught is as useful as one that landed cleanly, because it shows the review works. Hiding misses is how teams slide into misuse and disuse.
5. Make skeptics the reviewers. Ask the engineer who trusts the agents least to own the first domain's review and co-write the bar. Their objections become checks, and when they vote to move a domain up, the rest of the team believes it.

The plan above is an illustration: some domains climb fast, most stay at propose, and finance stays at observe until the next quarter. That spread is what calibrated trust looks like. The first 90 days page gives the week-by-week schedule.
An illustration: one skeptic, one domain, one quarter
This is an illustration, not a customer case or a measured result. A team runs Fivetran into Snowflake, Airflow and dbt, and Looker for revenue reporting. Their staff engineer has watched coding agents produce answers that were almost right, and says so.

In week 1 the head of data picks one domain, freshness and failed-run recovery, and sets it to L1 observe with read-only grants. She asks the staff engineer to write the bars. For leaving observe-only he takes the starting bar from the measurement guide: 100 runs, 95% match with what the on-call engineer did, under 5% overrides. For L3 he writes his own: 40 approved proposals, nine in ten approved unchanged, no rollbacks.
For four weeks the agents suggest a fix for each late Fivetran sync and failed Airflow DAG run. At the week 4 review the record clears the bar, and the staff engineer still finds two plans that would have rerun a dbt model before the upstream sync landed. Both are logged as misses and posted in the team channel next to the matches. In week 5 the domain moves to L2 propose with a write role scoped to reruns and one partition.
From then on each fix arrives in the Spellbook inbox with its blast radius, dry-run counts and rollback path. In week 7 the staff engineer rejects a backfill that would have run ahead of a late Fivetran sync and records why; that goes into the weekly note too. By week 11 the reruns clear his L3 bar, so he proposes moving reruns to L3 while backfills stay at L2, and the head of data agrees. In week 12 a rerun lands, its row-count check and lateness baseline pass, and the receipt is written before anyone opens Looker. The week 13 quarter review is a walk through receipts, and the Revenue Explore is fresh by 07:00.
The skeptic didn't change his standards. He wrote them down, the agents were measured against them, and the evidence did the rest.
How other paths approach trust
Each option below is a sensible design for its own job. The difference is how much evidence the team sees before an agent acts, and who holds the dial.
- •Build it yourself. A coding agent plus vendor MCP servers makes a quick demo; observe-only logging, approval records, receipts and the review against your bar then become your own project. The build-vs-buy page lays out that work.
- •Coding agents. Cursor's run modes include Auto-review, and its docs say plainly that "Auto-review is not a security boundary" (checked Oct 2, 2026), the right scope for a developer's session in a repository. Coding agents connect to Data Workers over MCP when the work touches production data.
- •Platform-native agents. Databricks Genie Code's docs say "the default mode for every chat is Auto-approve" on first use, call it "a productivity feature, not a security boundary," and advise keeping it off "when working with production data" (checked Oct 2, 2026). The approver is the person in the chat, inside one platform.
- •Observability. Monte Carlo's agentic onboarding has agents propose a monitoring plan, and "nothing deploys until you approve it" (checked Oct 2, 2026). Its alerts feed the Data Workers loop, which carries an incident through fix and verification.
- •Catalogs. Atlan's MCP guidance says "you can start read-only," removing every write, admin and lifecycle tool while keeping search and lineage (checked Oct 2, 2026). That is good governance for metadata; changes to the data itself are a different product.
Data Workers works alongside all of them and adds what trust needs for data changes: a record per domain, built from the agents' suggested fixes, the ladder's log, approvals and verified receipts, with your team deciding each step.
The case for your CFO
The outcome is a data team that hands repetitive operational work to agents at the pace it trusts, with every step backed by a record. Domains that earn L3 get faster fixes; the rest get complete, ready-to-approve proposals.
The risk story is the ladder. At L1 agents read and suggest fixes, and the ladder logs what each action would have needed. At L2 a named owner approves every change with the evidence attached. At L3 agents act only on changes that can be undone, rollback path recorded first, receipt after. With ladder enforcement on, deletes, elevated access and spend wait for a person at every level. Your data stays in your systems and the hosted Conductor sees workflow metadata only (where our data goes). Nothing is migrated: your warehouse, dbt project, orchestrator and BI stay as they are.
Why now: your engineers already use AI tools they don't fully trust. A visible, per-domain record turns that doubt into a decision process instead of a standoff.
The first win is one domain moving from observe to propose with a written bar and a weekly review inside the first month. Start with a pilot: $7,500 one-time, and the pilot is credited in full against the first year. Scale is from $1,000 a month and Enterprise from $3,000 a month, billed annually, with unlimited seats, no usage meter and no markup on model spend. See pricing, the ROI calculator and the ROI of agentic data operations.
The sentence for upstairs: "Our agents earn each permission on a record our own engineers set the bar for, and any of them can take it back."
FAQ
Our senior engineers think AI agents are unreliable. Where do we start? At L1 observe, with the most skeptical engineer writing the bar and running the weekly review. Nothing changes in production while the record builds, and their concerns become the checks.
What evidence does the team see before an agent changes anything? At L1, your team's review of the agents' suggested fixes against what it did, then, at L2, proposals with the blast radius, diff, dry-run counts, checks and rollback path. The safety page covers every gate.
Who decides when a domain moves up? The domain owner, at a review, against the bar written in advance, reading the agents' records and receipts.
What happens when an agent gets something wrong? At L2 the owner rejects or edits the proposal and the reason is recorded. At L3, if a step of the fix fails, Data Workers runs the recorded rollback and hands the incident to a person; if the fix completes but its checks still fail, it escalates to a person, who can roll it back from Spellbook. The owner can drop the domain back to L2. See how to roll back an AI agent change.
Can some domains stay at L2 forever? Yes. Finance models, access requests and PII tags and masking are often best served at propose and approve. The ladder is your team's decision per domain.
Will this make our team less necessary? No. Your team sets the bar, approves the changes and owns the outcomes; the agents take the repetitive operational load. The team page covers how each role shifts, and our human-in-the-loop resource and audit trail requirements go deeper.
Sources
- •Stack Overflow, 2025 Developer Survey, AI section (latest edition; fielded May 29 to June 23, 2025; 49,009 responses): https://survey.stackoverflow.co/2025/ai (checked Oct 2, 2026)
- •Lee, J. D. and See, K. A., "Trust in Automation: Designing for Appropriate Reliance," Human Factors 46(1), 50-80, 2004: https://doi.org/10.1518/hfes.46.1.50_30392 (checked Oct 2, 2026)
- •Parasuraman, R. and Riley, V., "Humans and Automation: Use, Misuse, Disuse, Abuse," Human Factors 39(2), 230-253, June 1997: https://doi.org/10.1518/001872097778543886 (checked Oct 2, 2026)
- •Dietvorst, B. J., Simmons, J. P. and Massey, C., "Algorithm aversion: People erroneously avoid algorithms after seeing them err," Journal of Experimental Psychology: General 144(1), 114-126, 2015: https://doi.org/10.1037/xge0000033 (checked Oct 2, 2026)
- •Dietvorst, B. J., Simmons, J. P. and Massey, C., "Overcoming Algorithm Aversion: People Will Use Imperfect Algorithms If They Can (Even Slightly) Modify Them," Management Science 64(3), 1155-1170, March 2018: https://doi.org/10.1287/mnsc.2016.2643 (checked Oct 2, 2026)
- •Databricks, Genie Code agent mode (last updated Sep 25, 2026): https://docs.databricks.com/aws/en/genie-code/agent-mode (checked Oct 2, 2026)
- •Cursor, Run modes: https://cursor.com/docs/agent/security/run-modes (checked Oct 2, 2026)
- •Monte Carlo, Agentic Onboarding: https://montecarlo.ai/platform/agentic-onboarding (checked Oct 2, 2026)
- •Atlan, Remote MCP security: https://docs.atlan.com/product/capabilities/atlan-ai/references/mcp-security (checked Oct 2, 2026)
- •Data Workers, Autonomous Data-Conductor: https://dataworkers.io/product/autonomous-data-conductor/ (checked Oct 2, 2026)
- •Data Workers, Spellbook Data Catalog: https://dataworkers.io/product/spellbook-data-catalog/ (checked Oct 2, 2026)
- •Data Workers, open-source core (
blast_radius_analysis,get_audit_trail): https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026) - •Data Workers, pricing: https://dataworkers.io/pricing/ (checked Oct 2, 2026)