Data Workers on AWS: The Agentic Data Platform on Top of Redshift, Glue, S3 Tables and SageMaker
AWS is where your data lives and runs. Data Workers runs the operations across Redshift, Glue, S3 Tables, Lake Formation and SageMaker, with approvals and receipts.
AWS is where your data lives and runs. Data Workers is the agentic data platform that runs the operations on top of it: it finds the cause of a broken number, fixes it where it started, checks the result downstream and leaves a receipt, across every AWS service you use and the tools around them.
Here is the estate most AWS data teams run. Aurora and other operational databases replicate through AWS DMS or zero-ETL into an S3 landing zone. Glue jobs on Spark turn that into Apache Iceberg tables, more and more of them in S3 Tables, registered in the Glue Data Catalog and locked down with Lake Formation. Redshift serves the warehouse, with dbt models built on it and Airflow DAGs on MWAA or Step Functions running the schedule. Athena and EMR handle the ad hoc and heavy work. Amazon Quick (formerly QuickSight) dashboards sit on top, and the newer teams work in SageMaker Unified Studio. Often there's a Snowflake account or a Databricks workspace next to all of it.
AWS gives that estate excellent services and a growing set of AI features. The SageMaker Data Agent writes code and SQL in notebooks. Glue Data Quality now recommends rules with generative AI. SageMaker Catalog suggests glossary terms and PII classifications. And with the AWS MCP Server, generally available since May 2026, a coding agent can call any AWS API through one tool.
What AWS deliberately doesn't ship is an agent that owns your data operations from the alert to the verified fix. AWS sells building blocks, and its AI features help a person inside one service. Of the four big platforms we cover in this series (Databricks, Snowflake, Google Cloud and AWS), AWS has the thinnest packaged agent layer for data operations, so the space between "something broke" and "it's fixed and proven" is widest here. That space is the job Data Workers is built for.
Keep every AWS service exactly where it is. Ask the Data Agent to write your next query. Hand Data Workers everything that has to be traced, fixed, approved and proven across services.
Key takeaways
- •AWS is where your data lives and runs. Data Workers runs the operations on top of it. Incidents, failed loads, schema changes, access requests, cost cleanup and catalog upkeep run at whatever autonomy level you choose per domain.
- •AWS's AI helps a person inside one service. The SageMaker Data Agent generates code and SQL, Amazon Q generative SQL says to review before running, Glue Data Quality recommendations are reviewed and saved by a person, and the EMR troubleshooting agent recommends a fix.
- •AWS ships agent building blocks. Bedrock AgentCore, the AWS MCP Server and the Agent Toolkit let you build your own agents. Data Workers ships finished ones, and they're callable from AgentCore Gateway over MCP.
- •Data Workers owns the loop across services: detect, diagnose, fix, review, verify and remember, from Aurora through DMS, Glue, S3 Tables and Redshift to dbt and the dashboard, with a tamper-evident receipt on every change.
- •IAM and Lake Formation stay the lock. Data Workers acts with the roles and grants you give it, and its calls show up in CloudTrail like any other principal's.
- •Zero migration. Nothing moves out of your accounts. Data Workers stores metadata and scrubbed facts about your data, and can run on the models you already use in Amazon Bedrock.
This guide compares what AWS covers with what Data Workers adds, service by service. For the step-by-step climb from a traditional AWS data team to an autonomous data platform, read the playbook From AWS to an autonomous data platform. For the executive version, read the data leader's guide for AWS teams. If you run Databricks or Snowflake next to AWS, the sibling guides Data Workers on Databricks and Data Workers on Snowflake cover those platforms. Our earlier reference stack for data engineering on AWS and Redshift MCP server setup cover adjacent ground.
Six things Data Workers adds on top of AWS
1. Incidents fixed and closed across services. A failed Glue Data Quality ruleset, a CloudWatch alarm on a Glue job or a Security Hub finding becomes a resolution loop instead of a ticket. The Autonomous Data-Conductor traces the cause in whichever system it lives in, has the right agent fix it there and confirms the downstream number is right again. An incident isn't closed until the checks pass.
2. Schema changes caught before a job writes nulls. Most broken AWS pipelines start upstream: a column renamed in Aurora, carried by DMS, read by a Glue job whose mapping still expects the old name. The job succeeds and the data is wrong. The Schema Evolution agent sees the source change and maps its blast radius through Glue jobs, S3 Tables, Redshift and dbt before the next run.
3. Access requests that clear the queue. Lake Formation is a precise lock, and the request queue in front of it is a ticket backlog on the platform team. Data Workers reads Lake Formation permissions and LF-Tags, proposes a least-privilege, time-boxed grant with its reasoning, and records who approved it and how to revoke it.
4. One context graph across AWS and everything next to it. Data Context Wizard builds one governed graph from the Glue Data Catalog, Lake Formation tags, Redshift, DMS replication tasks, your dbt project, Airflow and BI, plus Snowflake, Databricks and BigQuery if you run them. The agents that fix things know how your whole estate fits together, including the parts AWS doesn't host.
5. Spend cut where it starts. The Cost Savings & Cleanup agent finds tables nobody has queried in months, idle compute and duplicate pipelines, and it archives only after checking dependencies. On AWS it reads usage over AWS's MCP servers today, next to the Snowflake, Databricks and BigQuery bills it covers directly. Our design target is a 25 to 40% cut in warehouse spend. That's a target we're engineering toward, not a billed average.
6. Receipts your auditors can read. CloudTrail records every API call. A Data Workers receipt adds what CloudTrail can't know: why the change was made, what it touched downstream, who approved it and how to undo it. Evidence builds up as a side effect of the work, in Spellbook Data Catalog (in preview).
Behind all six are 20+ specialist agents in the Data-Agents Swarm, one context graph, and the coding agent your team already uses (Claude Code, Codex or Cursor) as the way in. Every agent is an MCP server.
One incident, six systems
Here is a scenario most AWS data teams will recognize. It's an illustration, not a customer case.
- •01:40. A product team renames
regiontosales_regionin the Aurora PostgreSQL orders database. - •01:42. The AWS DMS change-data-capture task carries the change into the S3 landing zone. New files have the new column.
- •03:00. The nightly Glue job merges the landing files into the
ordersIceberg table in S3 Tables. Its mapping still readsregion, so every new row gets a null region. The job succeeds. - •04:30. An Airflow DAG on MWAA runs dbt on Redshift, and
fct_revenue_by_regionputs the night's orders under "unknown". - •06:10. A Glue Data Quality completeness rule on
ordersfails and publishes the result. - •09:00. Regional sales leads open the Amazon Quick revenue dashboard for their weekly review.
| Step | What AWS-native tooling sees | What Data Workers does |
|---|---|---|
| Rename in Aurora, carried by DMS | DMS replicates the change as designed. A Glue crawler registers the new schema on its next run. | The Schema Evolution agent flags the contract change and maps its blast radius: the Glue job, the orders table, the dbt model and the dashboard. |
| Glue job writes null regions | A successful job run in CloudWatch. If an engineer asks, the Data Agent can help debug the script in a notebook. | The Conductor ties the nulls to the stale mapping. The Pipeline agent prepares a fix to the job script as an approval-gated pull request, with the blast radius attached. |
| Glue Data Quality fails | The completeness rule fails, and the result is available to alert on. Resolving it is up to your team. | Data Workers reads the result over the AWS API or MCP server, confirms the rename as the cause and routes the proposal to the on-call engineer in Spellbook. |
| Backfill and rebuild | Nothing happens until a person reruns the job and the DAG. | After approval, Data Workers reruns the Glue job for the affected partitions (over the AWS API or MCP server today) and triggers the dbt rebuild through the DAG. |
| Verify the result | The ruleset passes on its next run. | Data Workers reruns the ruleset, diffs row counts and values to confirm no null regions remain, checks the dbt tests and records a tamper-evident receipt. The pattern is remembered, so the next rename is held before the Glue job runs. |

AWS raised a clear signal at 06:10, and Glue Data Quality deserves the credit for it. The difference is what happens between 06:10 and 09:00. With AWS alone, an engineer spends that window in CloudWatch, the Glue console, the dbt project and the Airflow UI. With Data Workers, one engineer spends it on one approval.
What AWS covers, as of October 2026
AWS shipped a lot this year, and some of the names changed. The "next generation of Amazon SageMaker" brings SageMaker Unified Studio and SageMaker Catalog together on top of Amazon DataZone. Agent SOPs in the AWS MCP Server became agent skills. Here is what AWS's own documentation and What's New posts say it ships for data teams.
| Area | What AWS ships | Status (Oct 2026) |
|---|---|---|
| Data and AI workspace | SageMaker Unified Studio: notebooks, Query Editor, Visual ETL, Workflows (Airflow), lineage, data quality authoring, data profiling and anomaly detection (Aug 18, 2026) | GA |
| Notebook and SQL agent | SageMaker Data Agent: generates code, SQL and execution plans, diagnoses errors, uses SageMaker Catalog business context (Jun 4, 2026); billed per credit | GA (since Nov 2025) |
| Catalog | SageMaker Catalog (built on Amazon DataZone): glossary, metadata forms, subscriptions; AI suggests glossary terms and PII and PHI classifications that producers accept or modify; metadata sync with Collibra (both ways), Atlan and Alation | GA |
| Technical catalog | Glue Data Catalog: Iceberg v3 optimization, statistics and crawlers (Oct 1, 2026); federation to remote Iceberg catalogs; metadata exports to S3 Tables | GA; exports in preview |
| Data quality | Glue Data Quality: DQDL rules, anomaly detection, ADVANCED rule recommendations through Amazon Bedrock that you review and save (Sep 2026) | GA |
| Processing | Glue 6.0 (Spark 4.1, Iceberg v3), EMR, Athena, DMS, zero-ETL | GA |
| Spark agents | Spark troubleshooting agent (finds the root cause, recommends code changes) and upgrade agent for EMR, as Kiro powers | Available (Apr 2026) |
| Warehouse | Redshift: Serverless AI-driven scaling by default (Apr 27, 2026), Iceberg v3 tables, cross-Region data lake queries (Oct 1, 2026) | GA |
| Warehouse assistant | Amazon Q generative SQL in Redshift query editor v2: generates SQL from prompts; "Review the generated SQL before running it" | GA in listed Regions |
| Table storage | S3 Tables: managed Iceberg, Iceberg V3 data types (Sep 30, 2026), Kinesis streaming tables | GA |
| Access | Lake Formation: fine-grained permissions and LF-Tags; table permissions extended to the underlying S3 files (Jun 11, 2026) | GA |
| Business-user AI | Amazon Quick: AI assistant with dashboards (Amazon Quick), always-on agents (Sep 9, 2026) | GA |
| Agent building | Bedrock AgentCore: Runtime, Harness, Gateway, Memory, Identity, Policy, Evaluations, Registry and more | GA (Policy Mar 2026, Harness Jun 2026, new Runtime Sep 2026) |
| MCP | AWS MCP Server: managed, any AWS API through one tool, IAM-gated, CloudTrail-logged, no added charge; Agent Toolkit skills and plugins | GA (May 6, 2026) |
| MCP (open source) | awslabs MCP servers: Redshift (read-only), Data Processing for Glue, EMR and Athena (read-only unless --allow-write), S3 Tables (read-only unless --allow-write) | Available |
| AI-assisted migration | Redshift with Agent Toolkit: build, query, troubleshoot and migrate to Redshift from Claude Code, Kiro or Cursor | Available (Aug 27, 2026) |
That is a broad, well-built platform. If your whole estate is AWS and your team is happy building its own agents, much of it is enough, and we'll say so in the section on when AWS alone is enough.
Why doesn't AWS just do this itself?
Because AWS's business is building blocks, and its AI is designed to stay on the right side of the shared responsibility model.
AWS sells services to every kind of customer, from a two-person startup to a bank. Each service team builds the best primitive for its job and an assistant that helps a person use it: the Data Agent writes the notebook cell, Q writes the SQL, Glue Data Quality drafts the rules, SageMaker Catalog suggests the classification. In every case a person reviews and runs the result. When customers want agents that act, AWS gives them AgentCore and the AWS MCP Server, and the customer builds the agent, writes the policies and owns the outcome. That design is right for AWS. It keeps each service's blast radius inside the service, and it keeps AWS out of decisions about your business logic.
Owning a data incident across services is a different product category. It needs blast-radius scoping across Glue jobs, Iceberg tables, Redshift, dbt and dashboards; approvals that vary by domain; rollback; a receipt that records why a change was made; context about systems AWS doesn't host, like your dbt repo, a Snowflake account or Tableau; and the liability for changes made in your code and in tools AWS doesn't run. Taking that on would put AWS inside your business logic, across every customer's different stack. It's a sensible line for AWS not to cross, and it's the product Data Workers is.
AWS's answer so far is to give your engineers' coding agents the keys: the AWS MCP Server can call any AWS API, gated by IAM. That's useful. It runs with one engineer's role, in one session, and the approval policy, the cross-system verification and the shared record are left to you to build.
One platform across the data lifecycle
A data team's work runs across ten stages: keeping context current, answering business questions, data quality, observability and incidents, pipelines, schema changes and migrations, governance and access, security, cost, and the data under models. AWS offers a service or a feature for almost every one of them, and each adds another console, another set of permissions and another handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, on top of the services you already run.
We score the same ten stages on every comparison page, so you can compare across pages. AWS leads on its home stages, Pipelines & Ingestion and MLOps & Models, where its services are deep and mature. Data Workers covers all ten.

| Stage | Data Workers | AWS | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 7 | The Glue Data Catalog and SageMaker Catalog hold schemas, glossary terms, lineage and AI-suggested metadata for AWS assets. Data Workers keeps one context graph across AWS, dbt, Airflow, BI and other warehouses. |
| Analytics & Insights | 8 | 7.5 | Redshift, Athena, Amazon Quick and the SageMaker Data Agent answer questions and write SQL for AWS data. Data Workers answers across every platform from the same governed graph. |
| Data Quality | 8 | 7 | Glue Data Quality writes DQDL rules, recommends them with generative AI and detects anomalies; a person reviews and saves each ruleset. Data Workers writes, runs and repairs checks and dbt tests. |
| Observability & Incidents | 8.5 | 5 | CloudWatch, Glue Data Quality and SageMaker anomaly detection raise the signal, and the EMR troubleshooting agent recommends a fix. Data Workers owns the loop after the alert: diagnose, fix, verify, record. |
| Pipelines & Ingestion | 8.5 | 9 | AWS's home stage: Glue 6.0, EMR, DMS, zero-ETL and Kinesis streaming tables move and process data at any scale. Data Workers writes the pipeline change and triggers the rerun behind approval. |
| Schema & Migration | 8 | 6 | Glue crawlers pick up new schemas and the Agent Toolkit has Redshift migration skills for one engineer's session. Data Workers catches upstream schema changes before jobs run and plans moves in parity-checked waves. |
| Governance & Access | 8.5 | 7.5 | Lake Formation enforces fine-grained access and SageMaker Catalog runs subscription approvals. Data Workers works the request queue and proposes least-privilege grants through Lake Formation. |
| Security & Privacy | 8 | 8 | IAM, KMS, Lake Formation and CloudTrail secure and log every AWS call. Data Workers acts with the roles you give it and leaves a tamper-evident receipt on every change. |
| Cost / FinOps | 8 | 6 | Cost Explorer reports the bill and Redshift Serverless AI-driven scaling tunes Redshift compute. Data Workers removes waste after a dependency check, across AWS and the platforms next to it. |
| MLOps & Models | 7.5 | 8.5 | AWS's home stage: SageMaker AI, managed MLflow, HyperPod and Bedrock build, train and serve models. Data Workers keeps the data under models healthy and connects to MLflow and W&B. |
The outcomes a data leader buys
The lifecycle view shows breadth. This view scores eight outcomes a data leader pays for. AWS leads on three, all on its own ground: building blocks for your own agents, help inside AWS notebooks and editors, and access enforcement on AWS data. Data Workers leads on the five outcomes that cross a service boundary or happen after the alert.

| Outcome | Data Workers | AWS | Why we scored it this way |
|---|---|---|---|
| Building blocks for your own agents | 4 | 9 | Bedrock AgentCore (Runtime, Gateway, Memory, Policy, Evaluations, Registry), the AWS MCP Server and the Agent Toolkit are a full kit for building agents. Data Workers ships finished agents, callable from AgentCore Gateway over MCP. |
| Help inside AWS notebooks and editors | 6 | 8 | The SageMaker Data Agent writes code and SQL in notebooks and the Query Editor, and Amazon Q generative SQL works in Redshift query editor v2. Data Workers works through the coding agent your team already uses and through Spellbook. |
| Access enforced on AWS data | 6 | 9 | Lake Formation and IAM enforce access, now down to the S3 files under a table. Data Workers works through Lake Formation: it reads permissions and LF-Tags and proposes grants for Lake Formation to apply. |
| Incidents fixed across services and verified | 9 | 3 | AWS's AI detects, explains and recommends inside one service. Data Workers traces the cause from Glue to DMS to dbt, fixes it where it starts behind approval and checks the result downstream. |
| Schema changes caught before jobs break | 8 | 4 | Glue crawlers register a new schema after it lands. Data Workers sees the source change, maps its blast radius through Glue jobs, Redshift and dbt, and proposes the fix before the next run. |
| One context across AWS and other platforms | 9 | 5 | SageMaker Catalog syncs both ways with Collibra and feeds Atlan and Alation, and Glue federates remote Iceberg catalogs. Data Workers keeps one governed graph across AWS, Snowflake, Databricks, dbt and BI. |
| Spend cut across every platform you pay for | 7 | 5 | Redshift Serverless AI-driven scaling tunes Redshift, and Cost Explorer reports the bill. Data Workers reads AWS usage over MCP today, next to Snowflake, Databricks and BigQuery, and archives waste after a dependency check. |
| Audit evidence for every change | 9 | 6 | CloudTrail logs every API call, which is strong evidence of what was called. Every Data Workers change also records why, what it touched downstream, who approved it and how to undo it. |
These are directional scores of scope, not benchmarks. They measure what each side covers, not answer quality, and we've shown the reasoning so you can check every line.
Where AWS stops
Each limit below comes from AWS's own documentation, and each follows from its choice to build primitives and assistants rather than own your outcomes.
The assistants hand the change back to a person. The Redshift documentation for Amazon Q generative SQL says to "review the generated SQL before running it," and Q's questions "cannot reference an external schema." Glue Data Quality's new AI recommendations are reviewed, adjusted and saved by a person. SageMaker Catalog's AI classifications are suggestions that "producers can accept or modify." The Spark troubleshooting agent "identifies the root cause" and "provides specific code recommendations." Each is the right design for an assistant.
Help lives inside one service. The SageMaker Data Agent works in Unified Studio notebooks and the Query Editor. Q generative SQL works on the connected Redshift database. The EMR agents work on Spark jobs. An incident that starts in Aurora, travels through DMS and Glue and lands in a dbt model on Redshift crosses four of those boundaries.
Signals stop at the alert. CloudWatch, Glue Data Quality and SageMaker Unified Studio anomaly detection tell you a check failed or a value fell outside its predicted range. Tracing the cause, fixing it, backfilling, rebuilding downstream and proving the dashboard is right is your team's work.
Agents that act are yours to build and govern. AgentCore gives you a runtime, a gateway, memory and Cedar-compatible policies. The AWS MCP Server lets an agent call any API under IAM. What the agent should do with a failed ruleset, which domains it may touch, who approves and how the outcome is verified are left to you.
Context ends where AWS ends. SageMaker Catalog syncs with Collibra in both directions and feeds Atlan and Alation, and Glue can federate remote Iceberg catalogs. Your dbt project, your Airflow DAGs, a Snowflake account and your BI semantics still live outside one governed graph.

Where the two overlap
"Both" means Data Workers works through the AWS service rather than competing with it.
| Job to be done | AWS-native | Data Workers | What we recommend |
|---|---|---|---|
| Storage, compute and engines | S3, S3 Tables, Redshift, Glue, EMR, Athena | Not our job | AWS |
| Access enforcement | IAM, Lake Formation, LF-Tags | Reads permissions and tags, proposes grants for Lake Formation to apply | Both: Lake Formation enforces, Data Workers works the queue |
| Writing a query or notebook cell | SageMaker Data Agent, Q generative SQL | Your coding agent, with Data Workers context over MCP | AWS inside Unified Studio; Data Workers when the question spans platforms |
| Technical and business catalog | Glue Data Catalog, SageMaker Catalog | Reads both into one governed graph; Spellbook as the control plane for agent work | Both: AWS catalogs stay the source for AWS assets |
| Quality rules | Glue Data Quality, Unified Studio data quality | Writes and repairs checks and dbt tests, turns recurring incidents into rules | Both: keep your rulesets; Data Workers acts when they fail |
| Incident resolution | Alarms and recommendations | Conductor: diagnose, fix, verify, remember across services | Data Workers |
| Schema changes before they break | Crawlers register the new schema | Schema Evolution and Data Change Review agents | Data Workers |
| Replication | AWS DMS, zero-ETL | Reads and manages DMS replication tasks; maps downstream impact | Both |
| Cost cleanup | Cost Explorer, Redshift AI-driven scaling | Removes waste after a dependency check, across platforms | Both: AWS tunes each service, Data Workers removes the waste |
| Security findings | Security Hub | Reads findings and routes them to the governance and incident agents | Both |
| Building your own agents | AgentCore, AWS MCP Server, Agent Toolkit | Finished agents, callable as MCP tools | Both |
| Models | SageMaker AI, Bedrock | Keeps the data under models healthy; runs on Bedrock models if you choose | AWS for models; Data Workers for the data under them |
What it costs as agents do more of the work
On AWS, each AI feature carries its own meter. The SageMaker Data Agent is billed per credit, at $0.04 a credit. SageMaker Catalog charges for requests, metadata storage, compute units and recommendation tokens beyond its free allowances. Glue Data Quality runs on Glue compute. The AWS MCP Server and Agent Toolkit are free, but every API call they make runs on resources you pay for, and agents you build on AgentCore are billed by consumption.
None of that is unreasonable, and none of it goes away with Data Workers: your AWS services keep running on your AWS bill. Data Workers is priced the other way round. The Apache 2.0 core is free. The Pilot Program is $7,500 one-time, credited in full against your first year. Scale starts at $1,000 a month and Enterprise at $3,000 a month, both at the annual rate. Seats are unlimited, there's no usage meter, and there's no markup on model spend because you bring your own model, including models you already use in Amazon Bedrock. See pricing.
The useful question isn't which line is smaller this month. It's how many engineer hours sit between an AWS alarm and a verified fix today, and whether the cost of closing that gap grows with every task an agent takes on. Ours doesn't.
The case for your CFO
The outcome. Today AWS tells your team when a job or a check fails, and engineers turn that into a fix. Data Workers closes that second step for the domains you allow: failed loads are traced, fixed and verified, access requests arrive as proposed grants instead of tickets, and the dashboards finance and the board read are right before anyone opens them.
The risk story. Data Workers starts observe-only (L0 and L1): it reads and reports. At propose (L2) every change is a draft a person approves. At act reversibly (L3) it runs only changes it can undo, and only in the domains you've moved up. Autonomous (L4) is a per-domain choice, never a default. It acts with the IAM roles and Lake Formation grants you give it, every write is scoped before it runs, no agent approves its own work, and every receipt records who or what acted, why, what it touched and how to undo it. Zero migration: your data stays in your accounts.
Why now. Since May, any coding agent can call any AWS API through the AWS MCP Server. Agents are already acting on your estate one session at a time. The choice is whether that work runs under one approval flow and one record.
The first win. Failed Glue loads and Data Quality failures in one domain, handled in propose mode.
What stays the same. Every AWS service, IAM, Lake Formation, CloudTrail, SageMaker Catalog, your dbt project and Airflow DAGs, and the coding agent your engineers already use.
The path. Start with a pilot. See pricing; the pilot is credited in full against the first year.
The sentence to repeat upstairs: "AWS runs our data; Data Workers runs the operations on top of it, fixing what breaks across services behind our approvals and proving every fix with a receipt."
The fastest first win: failed loads and data-quality failures
Start with the failures that end the same way every time. A Glue job that failed or wrote bad rows usually needs a mapping fix or a rerun, a backfill and a check. Connect Data Workers read-only to the Glue Data Catalog, Lake Formation, your DMS tasks, Redshift and your dbt project, and put that domain in propose mode. Each failure arrives in the Spellbook inbox with the root cause, the proposed fix or rerun and its blast radius. When your team has approved those proposals as-is for a few weeks, move the domain up to reversible actions. Access requests through Lake Formation make a good second domain, because they're high volume and the rules are clear. The playbook covers the full sequence in From AWS to an autonomous data platform.
What each Data Workers product does on AWS
Data Context Wizard. One governed graph across your AWS services and the tools around them, with 50+ connectors. Its Glue connector reads databases, tables, partitions and table versions from the Glue Data Catalog, and Lake Formation permissions, LF-Tags and data lake settings. Its AWS DMS connector reads replication tasks and their status. Redshift connects over the Redshift Data API or the Redshift MCP server today, and S3 Tables over their Iceberg REST endpoint or MCP today. Every fact carries its source, author, confidence and the time it was observed, and an authority score orders a human promotion queue: nothing becomes authoritative without a named person's approval.
Data-Agents Swarm. 20+ specialist agents that do the work AWS's assistants recommend. On AWS the ones that matter most are the Schema Evolution agent, the Pipeline Building agent, the Incident Debugging agent, the Data Change Review agent, the Access & Governance agent working through Lake Formation, and the Data Security agent, which reads AWS Security Hub findings and routes them to triage. The swarm is MCP-native, so a team building on Bedrock AgentCore can register Data Workers agents as Gateway targets and call them as tools.
Autonomous Data-Conductor. The orchestrator that owns an outcome rather than a step. It runs detect, diagnose, fix, review, verify and remember across the estate, directs the right agents, scopes the blast radius before acting and leaves a tamper-evident receipt on every change. A Glue Data Quality failure, a CloudWatch alarm or a Security Hub finding is one input to the detect step.
Spellbook Data Catalog. The control plane for agent work and the business-user app. Every proposed change lands in one inbox (approve, steer, send back or roll back), asset pages cover the whole estate from Aurora to the dashboard, and an authority guard enforced in code stops any agent from approving its own work. Spellbook doesn't replace SageMaker Catalog or Lake Formation; approved access changes go through Lake Formation. It's in preview, and Slack and Teams access is coming.
Autonomy guardrails and security
AWS governs agents at the level of the API call: IAM decides what a role may call, AgentCore Policy intercepts tool calls before they run, and CloudTrail records them. Those are the right controls, and Data Workers runs inside them. It adds governance at the level of the work across services.
- •Autonomy is set per domain. Failed-load fixes can run at L3 (act, reversibly) while schema changes stay at L2 (propose).
- •New deployments start observe-only, with read-only roles. You extend autonomy one domain at a time as the receipts earn trust.
- •Every write is scoped before it runs, with blast radius computed across Glue, S3 Tables, Redshift, dbt and BI.
- •Every action is approved or reversible, and leaves a tamper-evident receipt.
- •No agent can approve or promote its own work. This is enforced in code, not left to a prompt.
- •Least privilege, inside your accounts. Data Workers acts with the IAM roles and Lake Formation grants you create. Its calls appear in CloudTrail like any other principal's, and Enterprise can run in your own VPC.

"We can build this with AgentCore and the AWS MCP Server. Why buy it?"
You can build agents with them, and AWS has made that easier every quarter. AgentCore gives you a runtime, memory, a gateway, identity, Cedar-compatible policy and evaluations. The AWS MCP Server gives an agent every AWS API, and the Agent Toolkit gives it tested skills. A strong platform team could assemble an incident agent on top.
What you'd be building is the part AWS leaves to you: the context graph across AWS and non-AWS systems, the specialist agents for schema changes, pipelines, access and cost, the loop that verifies a fix downstream and remembers it, per-domain autonomy levels, the approval inbox and the receipt format your auditors accept. That's a product, and it's the one Data Workers ships. The two fit together. Data Workers agents are MCP servers, so they can sit behind AgentCore Gateway next to the agents you build, and Data Workers can run on Bedrock models.
MCP isn't the difference on either side. Both sides speak it. The difference is whether your data operations run under one operating model across the estate, or as many sessions under many roles.
How it fits together

Getting started takes no migration. You create read-only IAM roles and connect Data Workers to the Glue Data Catalog and Lake Formation, Redshift, DMS, your dbt project and your orchestrator, plus any other warehouse you run. The Context Wizard builds the graph, every agent starts in observe-only mode, and the first things you see are lineage from Aurora through DMS, Glue and Redshift to the dashboard, and what each recent failure's fix would have been, with its blast radius.
When AWS alone is enough
AWS alone can be enough when:
- •Your estate is small and AWS-only, failures are rare, and your team fixes them fast.
- •Your need is help inside one service, like writing queries in Unified Studio or drafting Glue Data Quality rules.
- •You want to build and run your own data agents on AgentCore, and you have the platform team to own their policies, verification and audit trail.
Data Workers earns its place when any of these is true: incidents cross services, so a failure in Aurora or DMS shows up as a wrong number in Redshift; dbt, Airflow or a second warehouse sits in the critical path; the access and cleanup queues compete with incident work for the same engineers; or you need graded, reversible autonomy with receipts your auditors accept.
FAQ
Does Data Workers replace any AWS service? No. Redshift, Glue, S3 Tables, EMR, Athena, DMS and SageMaker keep doing their jobs. IAM and Lake Formation stay the enforcement point. Data Workers runs the operations across them.
How is Data Workers different from the SageMaker Data Agent? The Data Agent helps a person write code and SQL inside SageMaker Unified Studio, and the person runs it. Data Workers owns outcomes across services and tools: it traces a failure, proposes or applies the fix at the autonomy level you set, verifies the result downstream and leaves a receipt.
Do agents get admin access to our AWS accounts? No. You create the roles. Data Workers starts read-only, and at higher levels it acts only inside the permissions and blast radius you define for each domain. Every call appears in CloudTrail.
Does data leave our AWS accounts? Data Workers stores metadata and scrubbed facts about your data (definitions, lineage, owners, incident history), not copies of your tables. PII is scrubbed before anything is stored, tenants are isolated, and Enterprise can run in your own VPC or on-premise.
Can we use Amazon Bedrock models? Yes. Data Workers supports Bedrock as a model provider, so you can run it on models you already use there, with no markup on model spend.
Does it work with Glue Data Quality and SageMaker Catalog? Yes. Data Workers reads Glue Data Quality results as signals and keeps your rulesets in place. SageMaker Catalog and the Glue Data Catalog stay the source for AWS assets, and Data Workers reads them into one graph with the rest of your estate.
We also run Snowflake or Databricks. Does that change anything? It makes the case stronger. Data Workers keeps one context graph and one approval flow across AWS and the other platforms. See Data Workers on Snowflake and Data Workers on Databricks.
Is Data Workers open source? The core is Apache 2.0 and free to run. See pricing for what the platform adds.
Sources
AWS capabilities, statuses and pricing are current as of October 2, 2026, from AWS's own documentation and What's New posts:
- •SageMaker Data Agent and its business-context update (June 4, 2026), checked Oct 2, 2026
- •What is SageMaker Unified Studio, data profiling and anomaly detection (Aug 18, 2026) and metadata sync with third-party catalogs (Mar 3, 2026), checked Oct 2, 2026
- •SageMaker Catalog automatic data classification (Nov 30, 2025), checked Oct 2, 2026
- •What is Amazon DataZone, checked Oct 2, 2026
- •AWS Glue documentation history and Glue Data Quality rule recommendations (Sep 22, 2026), checked Oct 2, 2026
- •Glue Data Catalog Iceberg V3 support (Oct 1, 2026), checked Oct 2, 2026
- •Apache Spark troubleshooting and upgrade agents (Apr 3, 2026), checked Oct 2, 2026
- •Amazon Q generative SQL in Redshift query editor v2, checked Oct 2, 2026
- •Redshift Serverless AI-driven scaling (Apr 27, 2026), Redshift with Agent Toolkit for AWS (Aug 27, 2026) and Redshift cross-Region data lake queries (Oct 1, 2026), checked Oct 2, 2026
- •S3 Tables Iceberg V3 data types (Sep 30, 2026), checked Oct 2, 2026
- •Lake Formation access to underlying S3 data (Jun 11, 2026), checked Oct 2, 2026
- •AWS MCP Server general availability and Agent Toolkit for AWS (May 6, 2026), checked Oct 2, 2026
- •awslabs MCP servers, including the Redshift, Data Processing and S3 Tables server READMEs, checked Oct 2, 2026
- •What is Amazon Bedrock AgentCore, Policy GA (Mar 3, 2026) and the new AgentCore Runtime (Sep 18, 2026), checked Oct 2, 2026
- •Amazon Quick always-on agents (Sep 9, 2026), checked Oct 2, 2026
- •Amazon SageMaker pricing, checked Oct 2, 2026
Product names and statuses change quickly. If we've got something wrong, tell us and we'll fix it.