You Built on AWS. Now Make Your Whole Estate Agentic & Autonomous with Data Workers
A guide for heads of data on AWS: why your platform team is still the human control plane between Redshift, Glue, S3 Tables and Lake Formation, and how to put the data back office on autopilot.
AWS is where your data lives and your analytics run. Data Workers is the agentic data platform that runs the whole estate, AWS included. This guide is for the leader who owns that estate and wants the back office to run itself.
You built on AWS for good reasons. S3 holds everything, and S3 Tables now keeps your Iceberg tables compacted without a team babysitting them. Glue catalogs and transforms. Lake Formation decides who sees which column. Redshift serves the warehouse, and SageMaker Unified Studio is pulling analytics, notebooks and the catalog into one place.
But look at how the work actually gets done. A Glue job fails at 3 a.m. because an Aurora column was renamed upstream. A request for a finance table waits in a queue until a data owner gets to it. The Redshift bill jumps and nobody can say which workgroup did it. Some transformations live in dbt, some schedules in Airflow, and the dashboards may be in Tableau as well as Amazon Quick (formerly QuickSight). Each fix means a platform engineer opening four consoles, pinging three people and doing it by hand.
Your data platform has a control-plane problem, and the control plane is your team. People are the connective tissue between excellent services that don't talk to each other, carrying context from one to the next by hand. The expensive part isn't knowing what needs doing. It's the doing: the building, the changing, the checking and the fixing.
What AWS solved, and what it didn't
AWS solved the infrastructure. Storage, compute, the catalog and the permission model are scalable, secure and yours to configure. AWS has also put an assistant into almost every service. The SageMaker Data Agent drafts a query that an analyst accepts and runs. The Spark troubleshooting agent explains a failed EMR job and recommends a fix without changing anything. SageMaker Catalog suggests descriptions that a person accepts one by one. Bedrock AgentCore gives your developers the parts to build agents of their own.
What AWS doesn't do is operate your estate as a whole. Each assistant works inside one service and one person's session, and none of them owns a business outcome from the source change to the right number on the dashboard. That isn't a flaw. Each AWS service team builds the best assistant for its own service, and taking responsibility for changes across Redshift, Glue, Lake Formation and the tools AWS doesn't run is a different product with a different risk.
Why now: AWS is pulling analytics into SageMaker Unified Studio, standardising tables on Iceberg in S3 Tables, and putting an API and an MCP server behind most of its services. Everything an agent platform needs to work across your estate is already in place, so the first results come from your own data within a pilot.
AWS is where your data lives and your analytics run. Data Workers is the agentic data platform that runs your whole data estate, AWS included.
Or, in one sentence for your team: AWS gave every service its own assistant; Data Workers gives the platform one crew that fixes, approves and records the work across all of them.
What goes on autopilot
This is the part that changes your week. Each of these is a queue your platform team runs by hand today.

None of this requires replacing anything. Data Workers works inside the services and tools you already run, reasons across all of them, and leaves an auditable record of every change it makes. Your data stays in your accounts.
What changes for your organization
Here's what that looks like in a normal week.

- •Your platform team stops being the human control plane. The queues that used to wait for a platform engineer (access, schema changes, failed loads, cleanup and catalog upkeep) move without them. Engineers spend their time on the work only they can do.
- •Incidents stop becoming meetings. A broken number is traced to its cause in whatever service or tool it started in, fixed there and checked downstream, and the record is there when someone asks what happened.
- •Spend is managed continuously. Idle compute, unused tables and orphaned data are found and cleaned up as they appear, with a dependency check before anything is archived.
- •Audits get easier. Every change has a receipt: who or what made it, why, what it touched and how to undo it. Evidence builds up as a side effect of the work.
- •The migrations AWS puts on your roadmap get smaller. Moving tables into S3 Tables, retiring old functions or bringing a legacy warehouse into Redshift becomes a series of approvals instead of a dedicated project.
- •Everyone works from one shared understanding of your data. Definitions, owners and lineage across every service and tool live in one governed place, and nothing becomes official until a person approves it.
One incident, start to finish
This is an illustration, not a customer case.
- •A product team renames a column in an Aurora database overnight, and DMS replicates the change into S3.
- •The Glue job that reads it keeps running, but writes blanks into the orders table in S3 Tables.
- •Redshift loads the table on schedule, and nothing alerts.
- •By 9 a.m. the revenue dashboard in Amazon Quick has dropped, and each AWS assistant can explain its own piece but none of them owns the fix.
Today that's a morning of Slack threads. With Data Workers, the rename is traced to its source, the fix to the Glue job arrives ready to approve, the table is backfilled, and the dashboard number is confirmed before finance logs in. A receipt records every step. AWS's tools explain the symptom. Data Workers closes the ticket.
You choose how far and how fast
Our thesis is that data teams will climb from people working alongside a coding agent, to people governing a team of agents, to a largely self-running agentic enterprise. You don't have to jump. You choose the altitude, one area at a time.

Every area starts with agents watching and explaining. Then they propose changes for your team to approve. Then they make reversible changes on their own and leave a receipt. Only areas that have earned it run fully on their own, and any area can be dialled back at any time. No agent can approve its own work.
The detailed, stage-by-stage version, with a first-quarter plan, is in our AWS autonomy playbook.
Why doesn't AWS do this itself?
Focus and risk. AWS builds the best primitive for each job and an assistant that helps a person use it, and when you want agents that act, it hands you the parts to build and govern them yourself. Taking responsibility for changes across Glue, Redshift, S3, your dbt project and your dashboards is a different product with a different risk: blast-radius checks, approvals, rollback and a receipt for every change. That's the product we built.
Why now: AWS is putting an assistant in every service and opening its APIs to agents, so everything an agent platform needs to work across your estate is in place, and the first results come from your own data within a pilot.
How it fits with AWS
- •IAM and Lake Formation stay the lock. Every change goes through the roles and grants you give Data Workers, inside your accounts. There's never a second permission system, and its calls show up in CloudTrail like anyone else's.
- •Your AWS services keep doing their jobs. Redshift, Glue, S3 Tables, EMR and Athena run exactly as they do today. Nothing is migrated, and no tables are copied out. Data Workers stores metadata and scrubbed facts about your data.
- •Your people keep their tools. Analysts keep SageMaker Unified Studio. Business users keep Amazon Quick. Engineers ask and approve from the coding agent they already use, including Kiro, and teams building their own agents on Bedrock AgentCore can call Data Workers' agents as tools.
- •Your models stay your choice. Data Workers can run on the models you already use in Amazon Bedrock.
When you don't need Data Workers
Be honest with yourself about these:
- •Your estate is a handful of AWS services, with no dbt, Airflow, second warehouse or outside BI tool in the critical path.
- •Your only need is help writing queries and notebooks, which the SageMaker Data Agent handles well.
- •Your platform team isn't the bottleneck.
If all three are true, keep your budget. If any of them isn't, the rest of this guide is about you.
The case for your CFO
The outcome: your data team's hours go to new work while the back office runs itself. Access is granted the day it's asked for, broken numbers are fixed and checked before the business sees them, and data spend falls because the waste behind it is removed instead of paid for.
The risk story: agents start by watching. Each area earns the right to propose, then to make reversible changes, on its own record. Every change goes through the permissions you grant and leaves a receipt saying who or what made it, why, what it touched and how to undo it. Anything irreversible needs a named person's approval, and nothing is migrated to start.
What stays the same: your AWS services, your dashboards and the people who use them. The path is a pilot, and the pilot is credited in full against the first year.
The sentence to repeat upstairs: AWS gave every service its own assistant; we're adding one crew that fixes, approves and records the work across all of them.
Where to start
Pick one area where your team is the bottleneck. Most AWS teams start with one of three: access requests waiting on Lake Formation and SageMaker Catalog approvals, the failed Glue and Step Functions loads that end the same way every time, or Redshift and S3 spend. Each is high-volume, rule-bound and easy to measure.
Start with a pilot. A forward-deployed engineer connects your estate and runs the first areas alongside your team, so you see results on your own data before you commit. See pricing for how the pilot works.
If your architects want the detail, send them Data Workers on AWS. It covers what each AWS service does, what Data Workers adds, and the four products that do the work: Data Context Wizard keeps one governed graph across Glue, SageMaker Catalog, Redshift and everything outside AWS; the Data-Agents Swarm makes the changes; the Autonomous Data-Conductor runs each fix from detection to a verified result; and Spellbook Data Catalog (in preview) is where your team approves, audits and rolls back. The thinking behind them is in our thesis.
Sources
- •AWS, Generate SQL with the Data Agent (SageMaker Unified Studio), checked Oct 2, 2026: https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/sql-query-data-agent.html
- •AWS, Using machine learning and generative AI in SageMaker Unified Studio (AI metadata recommendations), checked Oct 2, 2026: https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/autodoc.html
- •AWS, Approve or reject a subscription request in SageMaker Unified Studio, checked Oct 2, 2026: https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/approve-reject-subscription-request.html
- •AWS What's New, Spark troubleshooting agent for Amazon EMR on EKS (July 10, 2026), checked Oct 2, 2026: https://aws.amazon.com/about-aws/whats-new/2026/07/amazon-emr-eks-spark-troubleshooting/
- •AWS, Amazon S3 Tables features, checked Oct 2, 2026: https://aws.amazon.com/s3/features/tables/
- •AWS, What is Amazon Bedrock AgentCore, checked Oct 2, 2026: https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/what-is-bedrock-agentcore.html
- •AWS, Amazon Quick, checked Oct 2, 2026: https://aws.amazon.com/quick/quicksight/