Comparison
Comparison17 min readBy The Data Workers Team

Data Workers on AWS: The Agentic Data Platform on Top of Redshift, Glue, S3 Tables and SageMaker

AWS is where your data lives and runs. Data Workers runs the operations across Redshift, Glue, S3 Tables, Lake Formation and SageMaker, with approvals and receipts.

AWS is where your data lives and runs. Data Workers is the agentic data platform that runs the operations on top of it: it finds the cause of a broken number, fixes it where it started, checks the result downstream and leaves a receipt, across every AWS service you use and the tools around them.

Here is the estate most AWS data teams run. Aurora and other operational databases replicate through AWS DMS or zero-ETL into an S3 landing zone. Glue jobs on Spark turn that into Apache Iceberg tables, more and more of them in S3 Tables, registered in the Glue Data Catalog and locked down with Lake Formation. Redshift serves the warehouse, with dbt models built on it and Airflow DAGs on MWAA or Step Functions running the schedule. Athena and EMR handle the ad hoc and heavy work. Amazon Quick (formerly QuickSight) dashboards sit on top, and the newer teams work in SageMaker Unified Studio. Often there's a Snowflake account or a Databricks workspace next to all of it.

AWS gives that estate excellent services and a growing set of AI features. The SageMaker Data Agent writes code and SQL in notebooks. Glue Data Quality now recommends rules with generative AI. SageMaker Catalog suggests glossary terms and PII classifications. And with the AWS MCP Server, generally available since May 2026, a coding agent can call any AWS API through one tool.

What AWS deliberately doesn't ship is an agent that owns your data operations from the alert to the verified fix. AWS sells building blocks, and its AI features help a person inside one service. Of the four big platforms we cover in this series (Databricks, Snowflake, Google Cloud and AWS), AWS has the thinnest packaged agent layer for data operations, so the space between "something broke" and "it's fixed and proven" is widest here. That space is the job Data Workers is built for.

Keep every AWS service exactly where it is. Ask the Data Agent to write your next query. Hand Data Workers everything that has to be traced, fixed, approved and proven across services.

Key takeaways

  • •AWS is where your data lives and runs. Data Workers runs the operations on top of it. Incidents, failed loads, schema changes, access requests, cost cleanup and catalog upkeep run at whatever autonomy level you choose per domain.
  • •AWS's AI helps a person inside one service. The SageMaker Data Agent generates code and SQL, Amazon Q generative SQL says to review before running, Glue Data Quality recommendations are reviewed and saved by a person, and the EMR troubleshooting agent recommends a fix.
  • •AWS ships agent building blocks. Bedrock AgentCore, the AWS MCP Server and the Agent Toolkit let you build your own agents. Data Workers ships finished ones, and they're callable from AgentCore Gateway over MCP.
  • •Data Workers owns the loop across services: detect, diagnose, fix, review, verify and remember, from Aurora through DMS, Glue, S3 Tables and Redshift to dbt and the dashboard, with a tamper-evident receipt on every change.
  • •IAM and Lake Formation stay the lock. Data Workers acts with the roles and grants you give it, and its calls show up in CloudTrail like any other principal's.
  • •Zero migration. Nothing moves out of your accounts. Data Workers stores metadata and scrubbed facts about your data, and can run on the models you already use in Amazon Bedrock.

This guide compares what AWS covers with what Data Workers adds, service by service. For the step-by-step climb from a traditional AWS data team to an autonomous data platform, read the playbook From AWS to an autonomous data platform. For the executive version, read the data leader's guide for AWS teams. If you run Databricks or Snowflake next to AWS, the sibling guides Data Workers on Databricks and Data Workers on Snowflake cover those platforms. Our earlier reference stack for data engineering on AWS and Redshift MCP server setup cover adjacent ground.

Six things Data Workers adds on top of AWS

1. Incidents fixed and closed across services. A failed Glue Data Quality ruleset, a CloudWatch alarm on a Glue job or a Security Hub finding becomes a resolution loop instead of a ticket. The Autonomous Data-Conductor traces the cause in whichever system it lives in, has the right agent fix it there and confirms the downstream number is right again. An incident isn't closed until the checks pass.

2. Schema changes caught before a job writes nulls. Most broken AWS pipelines start upstream: a column renamed in Aurora, carried by DMS, read by a Glue job whose mapping still expects the old name. The job succeeds and the data is wrong. The Schema Evolution agent sees the source change and maps its blast radius through Glue jobs, S3 Tables, Redshift and dbt before the next run.

3. Access requests that clear the queue. Lake Formation is a precise lock, and the request queue in front of it is a ticket backlog on the platform team. Data Workers reads Lake Formation permissions and LF-Tags, proposes a least-privilege, time-boxed grant with its reasoning, and records who approved it and how to revoke it.

4. One context graph across AWS and everything next to it. Data Context Wizard builds one governed graph from the Glue Data Catalog, Lake Formation tags, Redshift, DMS replication tasks, your dbt project, Airflow and BI, plus Snowflake, Databricks and BigQuery if you run them. The agents that fix things know how your whole estate fits together, including the parts AWS doesn't host.

5. Spend cut where it starts. The Cost Savings & Cleanup agent finds tables nobody has queried in months, idle compute and duplicate pipelines, and it archives only after checking dependencies. On AWS it reads usage over AWS's MCP servers today, next to the Snowflake, Databricks and BigQuery bills it covers directly. Our design target is a 25 to 40% cut in warehouse spend. That's a target we're engineering toward, not a billed average.

6. Receipts your auditors can read. CloudTrail records every API call. A Data Workers receipt adds what CloudTrail can't know: why the change was made, what it touched downstream, who approved it and how to undo it. Evidence builds up as a side effect of the work, in Spellbook Data Catalog (in preview).

Behind all six are 20+ specialist agents in the Data-Agents Swarm, one context graph, and the coding agent your team already uses (Claude Code, Codex or Cursor) as the way in. Every agent is an MCP server.

One incident, six systems

Here is a scenario most AWS data teams will recognize. It's an illustration, not a customer case.

  • •01:40. A product team renames region to sales_region in the Aurora PostgreSQL orders database.
  • •01:42. The AWS DMS change-data-capture task carries the change into the S3 landing zone. New files have the new column.
  • •03:00. The nightly Glue job merges the landing files into the orders Iceberg table in S3 Tables. Its mapping still reads region, so every new row gets a null region. The job succeeds.
  • •04:30. An Airflow DAG on MWAA runs dbt on Redshift, and fct_revenue_by_region puts the night's orders under "unknown".
  • •06:10. A Glue Data Quality completeness rule on orders fails and publishes the result.
  • •09:00. Regional sales leads open the Amazon Quick revenue dashboard for their weekly review.
StepWhat AWS-native tooling seesWhat Data Workers does
Rename in Aurora, carried by DMSDMS replicates the change as designed. A Glue crawler registers the new schema on its next run.The Schema Evolution agent flags the contract change and maps its blast radius: the Glue job, the orders table, the dbt model and the dashboard.
Glue job writes null regionsA successful job run in CloudWatch. If an engineer asks, the Data Agent can help debug the script in a notebook.The Conductor ties the nulls to the stale mapping. The Pipeline agent prepares a fix to the job script as an approval-gated pull request, with the blast radius attached.
Glue Data Quality failsThe completeness rule fails, and the result is available to alert on. Resolving it is up to your team.Data Workers reads the result over the AWS API or MCP server, confirms the rename as the cause and routes the proposal to the on-call engineer in Spellbook.
Backfill and rebuildNothing happens until a person reruns the job and the DAG.After approval, Data Workers reruns the Glue job for the affected partitions (over the AWS API or MCP server today) and triggers the dbt rebuild through the DAG.
Verify the resultThe ruleset passes on its next run.Data Workers reruns the ruleset, diffs row counts and values to confirm no null regions remain, checks the dbt tests and records a tamper-evident receipt. The pattern is remembered, so the next rename is held before the Glue job runs.
Incident timeline across the stack: what AWS, your team and Data Workers each do, step by step

AWS raised a clear signal at 06:10, and Glue Data Quality deserves the credit for it. The difference is what happens between 06:10 and 09:00. With AWS alone, an engineer spends that window in CloudWatch, the Glue console, the dbt project and the Airflow UI. With Data Workers, one engineer spends it on one approval.

What AWS covers, as of October 2026

AWS shipped a lot this year, and some of the names changed. The "next generation of Amazon SageMaker" brings SageMaker Unified Studio and SageMaker Catalog together on top of Amazon DataZone. Agent SOPs in the AWS MCP Server became agent skills. Here is what AWS's own documentation and What's New posts say it ships for data teams.

AreaWhat AWS shipsStatus (Oct 2026)
Data and AI workspaceSageMaker Unified Studio: notebooks, Query Editor, Visual ETL, Workflows (Airflow), lineage, data quality authoring, data profiling and anomaly detection (Aug 18, 2026)GA
Notebook and SQL agentSageMaker Data Agent: generates code, SQL and execution plans, diagnoses errors, uses SageMaker Catalog business context (Jun 4, 2026); billed per creditGA (since Nov 2025)
CatalogSageMaker Catalog (built on Amazon DataZone): glossary, metadata forms, subscriptions; AI suggests glossary terms and PII and PHI classifications that producers accept or modify; metadata sync with Collibra (both ways), Atlan and AlationGA
Technical catalogGlue Data Catalog: Iceberg v3 optimization, statistics and crawlers (Oct 1, 2026); federation to remote Iceberg catalogs; metadata exports to S3 TablesGA; exports in preview
Data qualityGlue Data Quality: DQDL rules, anomaly detection, ADVANCED rule recommendations through Amazon Bedrock that you review and save (Sep 2026)GA
ProcessingGlue 6.0 (Spark 4.1, Iceberg v3), EMR, Athena, DMS, zero-ETLGA
Spark agentsSpark troubleshooting agent (finds the root cause, recommends code changes) and upgrade agent for EMR, as Kiro powersAvailable (Apr 2026)
WarehouseRedshift: Serverless AI-driven scaling by default (Apr 27, 2026), Iceberg v3 tables, cross-Region data lake queries (Oct 1, 2026)GA
Warehouse assistantAmazon Q generative SQL in Redshift query editor v2: generates SQL from prompts; "Review the generated SQL before running it"GA in listed Regions
Table storageS3 Tables: managed Iceberg, Iceberg V3 data types (Sep 30, 2026), Kinesis streaming tablesGA
AccessLake Formation: fine-grained permissions and LF-Tags; table permissions extended to the underlying S3 files (Jun 11, 2026)GA
Business-user AIAmazon Quick: AI assistant with dashboards (Amazon Quick), always-on agents (Sep 9, 2026)GA
Agent buildingBedrock AgentCore: Runtime, Harness, Gateway, Memory, Identity, Policy, Evaluations, Registry and moreGA (Policy Mar 2026, Harness Jun 2026, new Runtime Sep 2026)
MCPAWS MCP Server: managed, any AWS API through one tool, IAM-gated, CloudTrail-logged, no added charge; Agent Toolkit skills and pluginsGA (May 6, 2026)
MCP (open source)awslabs MCP servers: Redshift (read-only), Data Processing for Glue, EMR and Athena (read-only unless --allow-write), S3 Tables (read-only unless --allow-write)Available
AI-assisted migrationRedshift with Agent Toolkit: build, query, troubleshoot and migrate to Redshift from Claude Code, Kiro or CursorAvailable (Aug 27, 2026)

That is a broad, well-built platform. If your whole estate is AWS and your team is happy building its own agents, much of it is enough, and we'll say so in the section on when AWS alone is enough.

Why doesn't AWS just do this itself?

Because AWS's business is building blocks, and its AI is designed to stay on the right side of the shared responsibility model.

AWS sells services to every kind of customer, from a two-person startup to a bank. Each service team builds the best primitive for its job and an assistant that helps a person use it: the Data Agent writes the notebook cell, Q writes the SQL, Glue Data Quality drafts the rules, SageMaker Catalog suggests the classification. In every case a person reviews and runs the result. When customers want agents that act, AWS gives them AgentCore and the AWS MCP Server, and the customer builds the agent, writes the policies and owns the outcome. That design is right for AWS. It keeps each service's blast radius inside the service, and it keeps AWS out of decisions about your business logic.

Owning a data incident across services is a different product category. It needs blast-radius scoping across Glue jobs, Iceberg tables, Redshift, dbt and dashboards; approvals that vary by domain; rollback; a receipt that records why a change was made; context about systems AWS doesn't host, like your dbt repo, a Snowflake account or Tableau; and the liability for changes made in your code and in tools AWS doesn't run. Taking that on would put AWS inside your business logic, across every customer's different stack. It's a sensible line for AWS not to cross, and it's the product Data Workers is.

AWS's answer so far is to give your engineers' coding agents the keys: the AWS MCP Server can call any AWS API, gated by IAM. That's useful. It runs with one engineer's role, in one session, and the approval policy, the cross-system verification and the shared record are left to you to build.

One platform across the data lifecycle

A data team's work runs across ten stages: keeping context current, answering business questions, data quality, observability and incidents, pipelines, schema changes and migrations, governance and access, security, cost, and the data under models. AWS offers a service or a feature for almost every one of them, and each adds another console, another set of permissions and another handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, on top of the services you already run.

We score the same ten stages on every comparison page, so you can compare across pages. AWS leads on its home stages, Pipelines & Ingestion and MLOps & Models, where its services are deep and mature. Data Workers covers all ten.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, AWS goes deep on its own area
StageData WorkersAWSWhy we scored it this way
Catalog & Context97The Glue Data Catalog and SageMaker Catalog hold schemas, glossary terms, lineage and AI-suggested metadata for AWS assets. Data Workers keeps one context graph across AWS, dbt, Airflow, BI and other warehouses.
Analytics & Insights87.5Redshift, Athena, Amazon Quick and the SageMaker Data Agent answer questions and write SQL for AWS data. Data Workers answers across every platform from the same governed graph.
Data Quality87Glue Data Quality writes DQDL rules, recommends them with generative AI and detects anomalies; a person reviews and saves each ruleset. Data Workers writes, runs and repairs checks and dbt tests.
Observability & Incidents8.55CloudWatch, Glue Data Quality and SageMaker anomaly detection raise the signal, and the EMR troubleshooting agent recommends a fix. Data Workers owns the loop after the alert: diagnose, fix, verify, record.
Pipelines & Ingestion8.59AWS's home stage: Glue 6.0, EMR, DMS, zero-ETL and Kinesis streaming tables move and process data at any scale. Data Workers writes the pipeline change and triggers the rerun behind approval.
Schema & Migration86Glue crawlers pick up new schemas and the Agent Toolkit has Redshift migration skills for one engineer's session. Data Workers catches upstream schema changes before jobs run and plans moves in parity-checked waves.
Governance & Access8.57.5Lake Formation enforces fine-grained access and SageMaker Catalog runs subscription approvals. Data Workers works the request queue and proposes least-privilege grants through Lake Formation.
Security & Privacy88IAM, KMS, Lake Formation and CloudTrail secure and log every AWS call. Data Workers acts with the roles you give it and leaves a tamper-evident receipt on every change.
Cost / FinOps86Cost Explorer reports the bill and Redshift Serverless AI-driven scaling tunes Redshift compute. Data Workers removes waste after a dependency check, across AWS and the platforms next to it.
MLOps & Models7.58.5AWS's home stage: SageMaker AI, managed MLflow, HyperPod and Bedrock build, train and serve models. Data Workers keeps the data under models healthy and connects to MLflow and W&B.

The outcomes a data leader buys

The lifecycle view shows breadth. This view scores eight outcomes a data leader pays for. AWS leads on three, all on its own ground: building blocks for your own agents, help inside AWS notebooks and editors, and access enforcement on AWS data. Data Workers leads on the five outcomes that cross a service boundary or happen after the alert.

Spider chart comparing Data Workers and AWS on the outcomes a data leader buys
OutcomeData WorkersAWSWhy we scored it this way
Building blocks for your own agents49Bedrock AgentCore (Runtime, Gateway, Memory, Policy, Evaluations, Registry), the AWS MCP Server and the Agent Toolkit are a full kit for building agents. Data Workers ships finished agents, callable from AgentCore Gateway over MCP.
Help inside AWS notebooks and editors68The SageMaker Data Agent writes code and SQL in notebooks and the Query Editor, and Amazon Q generative SQL works in Redshift query editor v2. Data Workers works through the coding agent your team already uses and through Spellbook.
Access enforced on AWS data69Lake Formation and IAM enforce access, now down to the S3 files under a table. Data Workers works through Lake Formation: it reads permissions and LF-Tags and proposes grants for Lake Formation to apply.
Incidents fixed across services and verified93AWS's AI detects, explains and recommends inside one service. Data Workers traces the cause from Glue to DMS to dbt, fixes it where it starts behind approval and checks the result downstream.
Schema changes caught before jobs break84Glue crawlers register a new schema after it lands. Data Workers sees the source change, maps its blast radius through Glue jobs, Redshift and dbt, and proposes the fix before the next run.
One context across AWS and other platforms95SageMaker Catalog syncs both ways with Collibra and feeds Atlan and Alation, and Glue federates remote Iceberg catalogs. Data Workers keeps one governed graph across AWS, Snowflake, Databricks, dbt and BI.
Spend cut across every platform you pay for75Redshift Serverless AI-driven scaling tunes Redshift, and Cost Explorer reports the bill. Data Workers reads AWS usage over MCP today, next to Snowflake, Databricks and BigQuery, and archives waste after a dependency check.
Audit evidence for every change96CloudTrail logs every API call, which is strong evidence of what was called. Every Data Workers change also records why, what it touched downstream, who approved it and how to undo it.

These are directional scores of scope, not benchmarks. They measure what each side covers, not answer quality, and we've shown the reasoning so you can check every line.

Where AWS stops

Each limit below comes from AWS's own documentation, and each follows from its choice to build primitives and assistants rather than own your outcomes.

The assistants hand the change back to a person. The Redshift documentation for Amazon Q generative SQL says to "review the generated SQL before running it," and Q's questions "cannot reference an external schema." Glue Data Quality's new AI recommendations are reviewed, adjusted and saved by a person. SageMaker Catalog's AI classifications are suggestions that "producers can accept or modify." The Spark troubleshooting agent "identifies the root cause" and "provides specific code recommendations." Each is the right design for an assistant.

Help lives inside one service. The SageMaker Data Agent works in Unified Studio notebooks and the Query Editor. Q generative SQL works on the connected Redshift database. The EMR agents work on Spark jobs. An incident that starts in Aurora, travels through DMS and Glue and lands in a dbt model on Redshift crosses four of those boundaries.

Signals stop at the alert. CloudWatch, Glue Data Quality and SageMaker Unified Studio anomaly detection tell you a check failed or a value fell outside its predicted range. Tracing the cause, fixing it, backfilling, rebuilding downstream and proving the dashboard is right is your team's work.

Agents that act are yours to build and govern. AgentCore gives you a runtime, a gateway, memory and Cedar-compatible policies. The AWS MCP Server lets an agent call any API under IAM. What the agent should do with a failed ruleset, which domains it may touch, who approves and how the outcome is verified are left to you.

Context ends where AWS ends. SageMaker Catalog syncs with Collibra in both directions and feeds Atlan and Alation, and Glue can federate remote Iceberg catalogs. Your dbt project, your Airflow DAGs, a Snowflake account and your BI semantics still live outside one governed graph.

Matrix of where Data Workers and AWS can read, fix and verify across every system in the estate

Where the two overlap

"Both" means Data Workers works through the AWS service rather than competing with it.

Job to be doneAWS-nativeData WorkersWhat we recommend
Storage, compute and enginesS3, S3 Tables, Redshift, Glue, EMR, AthenaNot our jobAWS
Access enforcementIAM, Lake Formation, LF-TagsReads permissions and tags, proposes grants for Lake Formation to applyBoth: Lake Formation enforces, Data Workers works the queue
Writing a query or notebook cellSageMaker Data Agent, Q generative SQLYour coding agent, with Data Workers context over MCPAWS inside Unified Studio; Data Workers when the question spans platforms
Technical and business catalogGlue Data Catalog, SageMaker CatalogReads both into one governed graph; Spellbook as the control plane for agent workBoth: AWS catalogs stay the source for AWS assets
Quality rulesGlue Data Quality, Unified Studio data qualityWrites and repairs checks and dbt tests, turns recurring incidents into rulesBoth: keep your rulesets; Data Workers acts when they fail
Incident resolutionAlarms and recommendationsConductor: diagnose, fix, verify, remember across servicesData Workers
Schema changes before they breakCrawlers register the new schemaSchema Evolution and Data Change Review agentsData Workers
ReplicationAWS DMS, zero-ETLReads and manages DMS replication tasks; maps downstream impactBoth
Cost cleanupCost Explorer, Redshift AI-driven scalingRemoves waste after a dependency check, across platformsBoth: AWS tunes each service, Data Workers removes the waste
Security findingsSecurity HubReads findings and routes them to the governance and incident agentsBoth
Building your own agentsAgentCore, AWS MCP Server, Agent ToolkitFinished agents, callable as MCP toolsBoth
ModelsSageMaker AI, BedrockKeeps the data under models healthy; runs on Bedrock models if you chooseAWS for models; Data Workers for the data under them

What it costs as agents do more of the work

On AWS, each AI feature carries its own meter. The SageMaker Data Agent is billed per credit, at $0.04 a credit. SageMaker Catalog charges for requests, metadata storage, compute units and recommendation tokens beyond its free allowances. Glue Data Quality runs on Glue compute. The AWS MCP Server and Agent Toolkit are free, but every API call they make runs on resources you pay for, and agents you build on AgentCore are billed by consumption.

None of that is unreasonable, and none of it goes away with Data Workers: your AWS services keep running on your AWS bill. Data Workers is priced the other way round. The Apache 2.0 core is free. The Pilot Program is $7,500 one-time, credited in full against your first year. Scale starts at $1,000 a month and Enterprise at $3,000 a month, both at the annual rate. Seats are unlimited, there's no usage meter, and there's no markup on model spend because you bring your own model, including models you already use in Amazon Bedrock. See pricing.

The useful question isn't which line is smaller this month. It's how many engineer hours sit between an AWS alarm and a verified fix today, and whether the cost of closing that gap grows with every task an agent takes on. Ours doesn't.

The case for your CFO

The outcome. Today AWS tells your team when a job or a check fails, and engineers turn that into a fix. Data Workers closes that second step for the domains you allow: failed loads are traced, fixed and verified, access requests arrive as proposed grants instead of tickets, and the dashboards finance and the board read are right before anyone opens them.

The risk story. Data Workers starts observe-only (L0 and L1): it reads and reports. At propose (L2) every change is a draft a person approves. At act reversibly (L3) it runs only changes it can undo, and only in the domains you've moved up. Autonomous (L4) is a per-domain choice, never a default. It acts with the IAM roles and Lake Formation grants you give it, every write is scoped before it runs, no agent approves its own work, and every receipt records who or what acted, why, what it touched and how to undo it. Zero migration: your data stays in your accounts.

Why now. Since May, any coding agent can call any AWS API through the AWS MCP Server. Agents are already acting on your estate one session at a time. The choice is whether that work runs under one approval flow and one record.

The first win. Failed Glue loads and Data Quality failures in one domain, handled in propose mode.

What stays the same. Every AWS service, IAM, Lake Formation, CloudTrail, SageMaker Catalog, your dbt project and Airflow DAGs, and the coding agent your engineers already use.

The path. Start with a pilot. See pricing; the pilot is credited in full against the first year.

The sentence to repeat upstairs: "AWS runs our data; Data Workers runs the operations on top of it, fixing what breaks across services behind our approvals and proving every fix with a receipt."

The fastest first win: failed loads and data-quality failures

Start with the failures that end the same way every time. A Glue job that failed or wrote bad rows usually needs a mapping fix or a rerun, a backfill and a check. Connect Data Workers read-only to the Glue Data Catalog, Lake Formation, your DMS tasks, Redshift and your dbt project, and put that domain in propose mode. Each failure arrives in the Spellbook inbox with the root cause, the proposed fix or rerun and its blast radius. When your team has approved those proposals as-is for a few weeks, move the domain up to reversible actions. Access requests through Lake Formation make a good second domain, because they're high volume and the rules are clear. The playbook covers the full sequence in From AWS to an autonomous data platform.

What each Data Workers product does on AWS

Data Context Wizard. One governed graph across your AWS services and the tools around them, with 50+ connectors. Its Glue connector reads databases, tables, partitions and table versions from the Glue Data Catalog, and Lake Formation permissions, LF-Tags and data lake settings. Its AWS DMS connector reads replication tasks and their status. Redshift connects over the Redshift Data API or the Redshift MCP server today, and S3 Tables over their Iceberg REST endpoint or MCP today. Every fact carries its source, author, confidence and the time it was observed, and an authority score orders a human promotion queue: nothing becomes authoritative without a named person's approval.

Data-Agents Swarm. 20+ specialist agents that do the work AWS's assistants recommend. On AWS the ones that matter most are the Schema Evolution agent, the Pipeline Building agent, the Incident Debugging agent, the Data Change Review agent, the Access & Governance agent working through Lake Formation, and the Data Security agent, which reads AWS Security Hub findings and routes them to triage. The swarm is MCP-native, so a team building on Bedrock AgentCore can register Data Workers agents as Gateway targets and call them as tools.

Autonomous Data-Conductor. The orchestrator that owns an outcome rather than a step. It runs detect, diagnose, fix, review, verify and remember across the estate, directs the right agents, scopes the blast radius before acting and leaves a tamper-evident receipt on every change. A Glue Data Quality failure, a CloudWatch alarm or a Security Hub finding is one input to the detect step.

Spellbook Data Catalog. The control plane for agent work and the business-user app. Every proposed change lands in one inbox (approve, steer, send back or roll back), asset pages cover the whole estate from Aurora to the dashboard, and an authority guard enforced in code stops any agent from approving its own work. Spellbook doesn't replace SageMaker Catalog or Lake Formation; approved access changes go through Lake Formation. It's in preview, and Slack and Teams access is coming.

Autonomy guardrails and security

AWS governs agents at the level of the API call: IAM decides what a role may call, AgentCore Policy intercepts tool calls before they run, and CloudTrail records them. Those are the right controls, and Data Workers runs inside them. It adds governance at the level of the work across services.

  • •Autonomy is set per domain. Failed-load fixes can run at L3 (act, reversibly) while schema changes stay at L2 (propose).
  • •New deployments start observe-only, with read-only roles. You extend autonomy one domain at a time as the receipts earn trust.
  • •Every write is scoped before it runs, with blast radius computed across Glue, S3 Tables, Redshift, dbt and BI.
  • •Every action is approved or reversible, and leaves a tamper-evident receipt.
  • •No agent can approve or promote its own work. This is enforced in code, not left to a prompt.
  • •Least privilege, inside your accounts. Data Workers acts with the IAM roles and Lake Formation grants you create. Its calls appear in CloudTrail like any other principal's, and Enterprise can run in your own VPC.
The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous

"We can build this with AgentCore and the AWS MCP Server. Why buy it?"

You can build agents with them, and AWS has made that easier every quarter. AgentCore gives you a runtime, memory, a gateway, identity, Cedar-compatible policy and evaluations. The AWS MCP Server gives an agent every AWS API, and the Agent Toolkit gives it tested skills. A strong platform team could assemble an incident agent on top.

What you'd be building is the part AWS leaves to you: the context graph across AWS and non-AWS systems, the specialist agents for schema changes, pipelines, access and cost, the loop that verifies a fix downstream and remembers it, per-domain autonomy levels, the approval inbox and the receipt format your auditors accept. That's a product, and it's the one Data Workers ships. The two fit together. Data Workers agents are MCP servers, so they can sit behind AgentCore Gateway next to the agents you build, and Data Workers can run on Bedrock models.

MCP isn't the difference on either side. Both sides speak it. The difference is whether your data operations run under one operating model across the estate, or as many sessions under many roles.

How it fits together

How Data Workers fits with AWS: your coding agent on top, Data Workers in the middle, your estate underneath

Getting started takes no migration. You create read-only IAM roles and connect Data Workers to the Glue Data Catalog and Lake Formation, Redshift, DMS, your dbt project and your orchestrator, plus any other warehouse you run. The Context Wizard builds the graph, every agent starts in observe-only mode, and the first things you see are lineage from Aurora through DMS, Glue and Redshift to the dashboard, and what each recent failure's fix would have been, with its blast radius.

When AWS alone is enough

AWS alone can be enough when:

  • •Your estate is small and AWS-only, failures are rare, and your team fixes them fast.
  • •Your need is help inside one service, like writing queries in Unified Studio or drafting Glue Data Quality rules.
  • •You want to build and run your own data agents on AgentCore, and you have the platform team to own their policies, verification and audit trail.

Data Workers earns its place when any of these is true: incidents cross services, so a failure in Aurora or DMS shows up as a wrong number in Redshift; dbt, Airflow or a second warehouse sits in the critical path; the access and cleanup queues compete with incident work for the same engineers; or you need graded, reversible autonomy with receipts your auditors accept.

FAQ

Does Data Workers replace any AWS service? No. Redshift, Glue, S3 Tables, EMR, Athena, DMS and SageMaker keep doing their jobs. IAM and Lake Formation stay the enforcement point. Data Workers runs the operations across them.

How is Data Workers different from the SageMaker Data Agent? The Data Agent helps a person write code and SQL inside SageMaker Unified Studio, and the person runs it. Data Workers owns outcomes across services and tools: it traces a failure, proposes or applies the fix at the autonomy level you set, verifies the result downstream and leaves a receipt.

Do agents get admin access to our AWS accounts? No. You create the roles. Data Workers starts read-only, and at higher levels it acts only inside the permissions and blast radius you define for each domain. Every call appears in CloudTrail.

Does data leave our AWS accounts? Data Workers stores metadata and scrubbed facts about your data (definitions, lineage, owners, incident history), not copies of your tables. PII is scrubbed before anything is stored, tenants are isolated, and Enterprise can run in your own VPC or on-premise.

Can we use Amazon Bedrock models? Yes. Data Workers supports Bedrock as a model provider, so you can run it on models you already use there, with no markup on model spend.

Does it work with Glue Data Quality and SageMaker Catalog? Yes. Data Workers reads Glue Data Quality results as signals and keeps your rulesets in place. SageMaker Catalog and the Glue Data Catalog stay the source for AWS assets, and Data Workers reads them into one graph with the rest of your estate.

We also run Snowflake or Databricks. Does that change anything? It makes the case stronger. Data Workers keeps one context graph and one approval flow across AWS and the other platforms. See Data Workers on Snowflake and Data Workers on Databricks.

Is Data Workers open source? The core is Apache 2.0 and free to run. See pricing for what the platform adds.

Sources

AWS capabilities, statuses and pricing are current as of October 2, 2026, from AWS's own documentation and What's New posts:

Product names and statuses change quickly. If we've got something wrong, tell us and we'll fix it.

Ready to go autonomous and agentic?

We’re building the future of data infrastructure right now. See how your enterprise data stack can operate fully agentic today.