Product
Product11 min readBy The Data Workers Team

You're on AWS Glue Data Catalog and SageMaker Catalog: AWS Registers and Shares Your Data. Data Workers Keeps It True

On AWS Glue Data Catalog and SageMaker Catalog? Keep them where your lake is registered, governed and shared. Data Workers works beside them and does the operations work that keeps them true, behind approvals.

Your AWS lake is registered in the Glue Data Catalog. Crawlers keep schemas and partitions current, S3 Tables register their Iceberg tables there, and Athena, EMR, Redshift Spectrum and Glue jobs read the same definitions. Lake Formation decides who sees which table, column and row, with LF-Tags. On top, SageMaker Catalog, built on Amazon DataZone inside SageMaker Unified Studio, turns tables into assets with business names, glossary terms, metadata forms and data products: producers publish, consumers subscribe, the owner approves and access is granted. Glue tables now take glossary terms and semantic search themselves (preview since June), the SageMaker Data Agent writes SQL from that business context, and Glue Data Quality drafts rules with generative AI. Glue and SageMaker Catalog are where your AWS estate registers, governs and shares data. Data Workers keeps it true: it catches the change the night it lands, fixes the cause through its owner and leaves a receipt.

Key takeaways

  • •Glue and SageMaker Catalog keep their jobs: the metastore, crawlers, Lake Formation grants, LF-Tags, assets and subscriptions stay where they are.
  • •Glue, Lake Formation and SageMaker Catalog connect over AWS's APIs or MCP servers today, in your team's client next to Data Workers. Data Workers reads the dbt manifest natively and traces what a table change breaks.
  • •Every catalog and grant change goes to its owner as a proposal with evidence.
  • •A crawler's schema change arrives with its blast radius and a fix; after a named person approves, the rerun is queued and a receipt records it.
  • •Start with a pilot on one lake domain, read-only first, on the ladder from L0 manual to L4 autonomous.

Glue and SageMaker Catalog are where your lake is registered, governed and shared. Data Workers is the control plane.

Here is a night at a grocery delivery marketplace. The mobile app sends order events to Amazon Data Firehose, which writes JSON to S3 by day. A Glue crawler, raw_app_orders, runs at 01:00 with its default schema change policy, which updates the table to match the data. dbt Cloud builds stg_app_orders, fct_orders_daily and fct_basket_margin on Athena at 02:00. The Commerce project publishes fct_orders_daily to SageMaker Catalog as "Daily orders"; Marketing, Finance and Supply subscribe, and SageMaker Unified Studio manages their Lake Formation grants. The Daily GMV dashboard in Amazon Quick reads it through Athena. This is an illustration, not a customer case.

TimeSystemWhat happens
Tue 18:40Mobile appRelease 6.4 ships. The iOS client now serializes price as a quoted string, "12.99", instead of a number
18:41Amazon Data FirehoseDelivers the new events to s3://orders-raw/app_orders/dt=2026-09-29/, as designed
Wed 01:00Glue crawlerraw_app_orders reads the new files and, under its default policy, changes price from double to string in raw.app_orders. The Glue Data Catalog records table version 18
02:00dbt Cloud on AthenaThe nightly job fails on fct_orders_daily: Athena can't sum a varchar
02:01Data WorkersThe failure alert pages the on-call, who hands it to Data Workers; it opens an incident and classifies the cause as a type change on price that breaks three dbt models that sum or multiply it
02:03Data Workersblast_radius_analysis maps the readers from the dbt manifest: three dbt models and the Daily GMV dashboard. The "Daily orders" asset in SageMaker Catalog and its three subscriber projects go on the plan for the Commerce owner. Without a fix, the Daily GMV dashboard misses Tuesday
02:06Spellbook and SlackData Workers drafts one change set: a dbt diff that casts price safely in stg_app_orders and fails the run if any value won't cast; for the platform owner, the crawler set to add new columns only, with partitions inheriting from the table; for the Commerce data owner, a republish of the asset with a column note; and a Slack note to the mobile team naming release 6.4. Approval requests go to the named owners in Slack
07:15SpellbookThe analytics engineering owner reviews the diff, blast radius and undo, approves and merges it
07:18dbt CloudThe owner reruns the nightly job after approval; it succeeds
07:41Athena and Amazon QuickData Workers verifies: all 41,200 Tuesday-evening orders carry a price, Tuesday GMV sits inside its weekly band and the dashboard's Athena query returns Tuesday. It writes the receipt
09:30Glue and SageMaker CatalogThe platform owner applies the crawler setting. The Commerce data owner republishes "Daily orders" with the column note. Lake Formation grants for the three subscribers stay exactly as they were
Incident timeline across the stack: what Glue and SageMaker Catalog, your team and Data Workers each do, step by step

The crawler did its job, and the catalog kept version 17, so the platform owner confirmed the change in one look. The cause sat in an app release and the break in dbt, neither of which a catalog changes. Four owners each approved their own part, and the 09:00 trading meeting saw the right GMV.

JobWhat Glue and SageMaker Catalog doWhat Data Workers does
The registerThe Glue Data Catalog holds databases, tables, partitions and every table versionJoins the tables to dbt, runs, usage and quality results, with the catalog as the record
The discoveryCrawlers infer schemas from S3 and update tables on a scheduleCatches what a changed table breaks in dbt the night it lands
The meaningSageMaker Catalog adds business names, glossary terms, metadata forms and data productsLeaves that context where it is and names the asset each proposal touches
The sharingProducers publish, consumers subscribe, the owner approves and access is grantedMaps which subscribers a change reaches and drafts the republish for the owner
The permissionLake Formation enforces grants down to column, row and cell, with LF-TagsLeaves grants and tags in Lake Formation; any grant change is proposed for the owner, by design
The fixRecords the new version and the quality resultProposes the change where the cause lives, such as a dbt diff, and queues the rerun after approval
The proofGlue Data Quality scores and anomaly detectionVerifies downstream in Athena and the dashboard's query and writes a receipt

Why doesn't Glue or SageMaker Catalog just do this itself?

Because the catalogs are the shared register and permission layer for every AWS engine, and their writes follow that job. A crawler writes table definitions; SageMaker Catalog writes business metadata and, after an owner approves, the grants that fulfil a subscription. Glue Data Quality drafts rules that, in AWS's words, you review, adjust and save, and the SageMaker Data Agent writes code for you to run. AWS's MCP servers for the Glue Data Catalog and S3 Tables run read-only by default, and the data processing server's README asks for caution when --allow-write is combined with sensitive data access. That is the right design for a service every account depends on.

Changing a dbt model, rerunning a dbt Cloud job or asking a mobile team to fix a release is a different product with a different liability: blast radius across every reader, scoped credentials, a named owner's approval, rollback and downstream proof. A metastore that rewrote those systems would be the shared dependency changing its own dependents. Data Workers is the product on the other side of that line.

Every tool owns a slice. Data Workers covers the whole lifecycle

Glue and SageMaker Catalog own registering, governing and sharing data on AWS. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the register and grants you already run.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Glue and SageMaker Catalog goes deep on its own area
StageData WorkersGlue and SageMaker CatalogWhy we scored it this way
Catalog & Context99A shared home stage. The Glue Data Catalog is the metastore every AWS engine reads, and SageMaker Catalog adds business names, glossaries, metadata forms, data products and OpenLineage-compatible lineage. Data Workers joins the tables around it to dbt, runs, usage and quality results.
Analytics & Insights85SageMaker Unified Studio adds semantic search, the query editor and the SageMaker Data Agent, which writes SQL and Python from business context; dashboards live in Amazon Quick. Data Workers answers questions from governed context and checks the numbers behind them.
Data Quality86.5Glue Data Quality runs rulesets on catalog tables, detects anomalies, writes results back to the catalog and drafts rules with generative AI (Sep 2026); Unified Studio shows profiles and anomalies (Aug 2026). Data Workers fixes what fails in the system that caused it and proves the fix.
Observability & Incidents8.53The catalogs show quality results and lineage; incident response happens elsewhere. Data Workers detects, diagnoses, fixes and verifies across systems, with a receipt for every incident.
Pipelines & Ingestion8.56Crawlers discover S3 data and register tables, Iceberg V3 included (Oct 2026); the ETL runs in Glue jobs and workflows next door. Data Workers catches source changes the night they land and queues reruns through your orchestrator.
Schema & Migration85Crawlers change table definitions by default and the catalog keeps every version, with settings to add new columns only or ignore changes. Data Workers scores a type change against everything downstream in dbt and proposes the fix with its undo.
Governance & Access8.59The home stage: Lake Formation grants down to column, row and cell, LF-Tags, cross-account sharing, and SageMaker Catalog subscriptions an owner approves before access is granted. Data Workers runs the request queue around them and proposes each grant for the owner.
Security & Privacy87Lake Formation data filters protect PII at row and cell level, CloudTrail audits access and Bedrock Guardrails sit inside SageMaker Catalog. Data Workers flags new sensitive column names in pull request review and proposes masking for the owner.
Cost / FinOps82The catalogs do not manage spend. Data Workers reads AWS Cost Explorer by service with a forecast and drafts the fix for the owner of the workload behind it.
MLOps & Models7.56.5SageMaker Catalog catalogs models alongside data, and SageMaker AI trains and serves them next door. Data Workers keeps the data under those models healthy.

How Glue, SageMaker Catalog and Data Workers work together

Your people stay in Kiro, Claude Code or Amazon Q Developer, SageMaker Unified Studio and Amazon Quick. Spellbook Data Catalog (in preview) is where the data team looks: each finding, proposed change, blast radius, approver and rollback, with approval requests reaching people in Slack or email. Underneath, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm's 20+ specialist agents do the work, and the Autonomous Data-Conductor runs each fix end to end.

How Data Workers fits with Glue and SageMaker Catalog: your coding agent on top, Data Workers in the middle, your estate underneath

What Data Workers reads. dbt natively: the manifest, sources, models and tests, which is how a crawler's change is traced to the models it breaks. Context Wizard keeps each fact with its provenance, next to AWS Cost Explorer and Slack, among 50+ connectors. The Glue Data Catalog, Lake Formation, SageMaker Catalog, Athena, Firehose and Amazon Quick connect over their APIs or MCP servers today: AWS's data processing MCP server runs in your team's client next to Data Workers, so tables, versions and grants come into the same conversation. Agents get one governed view through explain_table, trace_cross_platform_lineage and blast_radius_analysis.

What Data Workers writes, and where. To the catalogs, nothing directly. Crawler settings, table fixes and republishing go to the platform or data owner as proposals with evidence; approved facts land in the Context Wizard graph. Lake Formation grants are proposed after a dry run that shows effective privileges, the sensitive columns reached and a least-privilege grant with an expiry. dbt fixes go to the owner as a diff to merge; reruns are queued through your orchestrator, or run by the owner in dbt Cloud.

Setup today. Give Data Workers read access to the dbt project in scope and its manifest. For your assistant, clone the open-source repository and add start-agent.sh entries to your client config, as the client setup docs show. AWS's data processing MCP server can sit alongside, read-only by default.

// Example: .mcp.json for Claude Code
{
  "mcpServers": {
    "dw-connectors": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-connectors"]
    },
    "dw-context-catalog": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-context-catalog"]
    },
    "aws-dataprocessing": {
      "command": "uvx",
      "args": ["awslabs.aws-dataprocessing-mcp-server@latest"],
      "env": { "AWS_REGION": "us-east-1" }
    }
  }
}

The AWS server stays read-only because the config leaves out --allow-write. List the tools with /mcp in Claude Code. A platform lead can then ask "which Glue tables changed type this week, and who reads them?" and get versions, lineage and grants in one answer.

One request, L0 to L4. The ladder is set per domain, so Commerce can sit at a different level from Finance.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. Connected, not acting. The team works crawler changes by hand.
  • •L1 observe. Data Workers reports each breaking change with the dbt models it reaches, their subscribers and owners, and logs what each action would have needed.
  • •L2 propose. It drafts the dbt diff, crawler setting and republish with their blast radius; owners approve before anything runs.
  • •L3 act reversibly. For proven classes, such as queuing the rerun of an affected dbt job after a fix merges, it acts, verifies and records. Code still arrives as diffs.
  • •L4 autonomous. For a scoped, trusted class in one domain, it runs the loop end to end and posts the receipt. Catalog and Lake Formation changes still go to their owners.

What changes for your team

Six jobs that run on autopilot with Data Workers next to Glue and SageMaker Catalog, with a concrete example of each

AWS lake teams spend much of their week after the catalog changes: why Athena fails after a crawl, which models read a table, who to tell, which grant to add. With Data Workers, six jobs run on autopilot at the level you set: incidents caught the night they land, failing Glue Data Quality rules arriving with a cause and a fix, Cost Explorer spend routed to the workload's owner, access requests arriving dry-run, audits that link asset, table version, approver, diff and undo, and migrations planned in waves with parity checks the owner signs off.

Leads get their week back for what only people decide: what an asset means and who should subscribe.

Keep Glue and SageMaker Catalog, or consolidate?

Keep Glue and SageMaker Catalog if you love them; Data Workers works with them from day one. Many teams consolidate once Data Workers runs that slice too.

The Glue Data Catalog is the metastore your engines need and Lake Formation is the permission system; they stay, by design. Teams consolidate the tooling around them: Lambda functions that watch crawlers, scripts that compare table versions, spreadsheets of which dbt model reads which table. If you are weighing building this layer on the Glue APIs and AWS's MCP servers, read build it ourselves with Claude Code and MCP servers: reading the catalog is the easy part; cross-system context, approvals and rollback are the work.

The case for your CFO

The outcome: what your catalog publishes is what your dashboards show. A crawler change becomes a fixed model before the meeting it would have spoiled.

The risk story is plain. Data Workers changes nothing in Glue or Lake Formation and proposes every catalog and grant change to its owner. Every other change shows its blast radius, goes to a named approver, is verified downstream and leaves a receipt with the undo. Unanswered requests expire and escalate; they never auto-grant. No agent can promote its own work. Autonomy is set per domain, and an org-wide stop halts all autonomous dispatch. The agents run in your infrastructure; your data stays in your systems, and the hosted Conductor sees workflow metadata only. Zero migration.

Why now: the SageMaker Data Agent, Amazon Quick's agents and every MCP client answer from the same Glue tables, so a wrong column spreads faster than a weekly review can catch. The first win is one lake domain, read-only, with a report of every breaking schema change and its readers. What stays the same: your AWS accounts, the metastore, Lake Formation, your assets and subscriptions. Start with a pilot; the path is on the pricing page, and the pilot is credited in full against the first year. See the ROI of agentic data operations.

The sentence to repeat upstairs: "Glue and SageMaker Catalog are where we register and share our data; Data Workers makes sure the tables behind them stay right every night, and every fix comes with an approval and a receipt."

Getting started

Start with a pilot. Pick the lake domain your SageMaker Catalog program cares about most, connect Data Workers read-only to the dbt project behind it, with AWS's MCP server alongside, and let it report schema changes, failures and their readers for a few weeks before turning on the first fix class. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.

FAQ

Does Data Workers write to the Glue Data Catalog or Lake Formation? No. They connect over AWS's APIs or MCP servers today, and Data Workers changes nothing in them. Crawler settings, table fixes and Lake Formation grants go to the owner as proposals, and the owner applies them.

How does this relate to SageMaker Catalog subscriptions? They stay the way consumers get access, and the owner still approves each one. Data Workers shows every subscriber a schema change reaches and drafts the republish for the asset's owner.

How is this different from the SageMaker Data Agent and Amazon Q Developer? Both work on your request, in your session. Data Workers runs the operations underneath: it watches the dbt runs and models built on every table, fixes the cause through its owner and proves the result.

Can we stop crawlers from changing schemas instead? Yes, for curated tables: a crawler can add new columns only or ignore changes, and Data Workers proposes that setting where it fits. Upstream changes still arrive, and adapting models, rerunning jobs and telling subscribers is the work Data Workers runs.

Does Data Workers replace Glue Data Quality? No. Keep your rulesets and anomaly detection in Glue. Data Workers traces what fails, proposes the fix in the system that caused it and verifies it, so the next Glue Data Quality run passes.

What about personal data in the lake? Lake Formation's grants and data filters stay in charge. When a pull request adds a column, Data Workers' review flags names and annotations that look sensitive (it reads names, not values), and Data Workers proposes any LF-Tag or masking change to the owner as a dry run.

Sources

  • •AWS, Amazon SageMaker Catalog product page, https://aws.amazon.com/sagemaker/catalog/ (checked Oct 3, 2026)
  • •AWS docs, What is Amazon SageMaker Unified Studio?, https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/what-is-sagemaker-unified-studio.html (checked Oct 3, 2026)
  • •AWS docs, Catalog in IDC-based domains (inventory and publishing), https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/working-with-business-catalog.html (checked Oct 3, 2026)
  • •AWS docs, Data discovery, subscription, and consumption, https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/discover-data.html (checked Oct 3, 2026)
  • •AWS docs, Preventing a crawler from changing an existing schema, https://docs.aws.amazon.com/glue/latest/dg/crawler-schema-changes-prevent.html (checked Oct 3, 2026)
  • •AWS docs, What is AWS Lake Formation?, https://docs.aws.amazon.com/lake-formation/latest/dg/what-is-lake-formation.html (checked Oct 3, 2026)
  • •AWS What's New, Amazon SageMaker Data Agent integrates business context into conversations (Jun 4, 2026), https://aws.amazon.com/about-aws/whats-new/2026/06/amazon-sagemaker-data-agent-bdc/ (checked Oct 3, 2026)
  • •AWS What's New, AWS Glue Data Catalog now supports business context and semantic search, Preview (Jun 17, 2026), https://aws.amazon.com/about-aws/whats-new/2026/06/aws-glue-data-catalog/ (checked Oct 3, 2026)
  • •AWS What's New, AWS Glue Data Quality now supports anomaly detection and writing results to the AWS Glue Data Catalog (Jul 27, 2026), https://aws.amazon.com/about-aws/whats-new/2026/07/aws-glue-data-quality-catalog-anomaly-detection-write-results (checked Oct 3, 2026)
  • •AWS What's New, Amazon SageMaker Unified Studio now supports data profiling and anomaly detection (Aug 18, 2026), https://aws.amazon.com/about-aws/whats-new/2026/05/smus-data-profiling (checked Oct 3, 2026)
  • •AWS What's New, AWS Glue Data Quality delivers context-specific rule recommendations in seconds (Sep 22, 2026), https://aws.amazon.com/about-aws/whats-new/2026/09/glue-data-quality-rule-recommendations/ (checked Oct 3, 2026)
  • •AWS What's New, AWS Glue Data Catalog now supports table optimization, statistics, and crawlers for Apache Iceberg V3 (Oct 1, 2026), https://aws.amazon.com/about-aws/whats-new/2026/10/aws-glue-iceberg-v3-optimization/ (checked Oct 3, 2026)
  • •AWS What's New, Amazon Quick announces autonomous agents (Jun 17, 2026), https://aws.amazon.com/about-aws/whats-new/2026/06/amazon-quick/ (checked Oct 3, 2026)
  • •awslabs/mcp, AWS Data Processing MCP Server README (release Sep 22, 2026), https://github.com/awslabs/mcp/tree/main/src/aws-dataprocessing-mcp-server (checked Oct 3, 2026)
  • •awslabs/mcp, Amazon S3 Tables MCP Server README (release Sep 8, 2026), https://github.com/awslabs/mcp/tree/main/src/s3-tables-mcp-server (checked Oct 3, 2026)
  • •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)
  • •Data Workers open-source repository, agent tools and start-agent.sh, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
  • •Data Workers, Spellbook Data Catalog product page, https://dataworkers.io/product/spellbook-data-catalog/ (checked Oct 3, 2026)