Hive and Glue to Unity Catalog Migration With Parity Checks: A How-To Over MCP
Move Hive metastore and AWS Glue tables to Unity Catalog in approved waves: Data Workers plans the parity checks, holds the completion gate, repoints jobs and dashboards, and leaves a receipt.
Databricks moves your tables into Unity Catalog with SYNC, CLONE and UCX. Data Workers plans each wave with its parity checks, holds the completion gate for the owner's sign-off, moves the jobs and dashboards that read it behind a named approval, and leaves a receipt you can roll back from.
If you still run tables in the workspace hive_metastore or an AWS Glue Data Catalog, you know the plan on paper. Databricks calls the per-workspace Hive metastore a legacy feature and recommends migrating the tables and the workloads that reference them, then disabling direct access. The copy is the easy part: SYNC SCHEMA ... DRY RUN, a deep clone for managed Delta tables, or the UCX table migration workflow. The hard part is the week after. Is main.finance.ap_invoices really the same as the Glue table? Which Airflow DAGs, dbt sources and dashboards still read the old name? Who approved the switch, and how do you go back? This guide shows how Data Workers, the agentic data platform, answers those questions wave by wave over MCP today.
Key takeaways
- •Databricks does the copy. SYNC, CREATE TABLE CLONE, CTAS and UCX stay your tools for moving tables.
- •Data Workers plans the checks for every table in a wave: row counts, schemas and aggregates as comparison queries your team runs, and grants against the Hive or Glue original.
- •A wave isn't done until the gate says so.
gate_migration_completereturns the tables still blocking until validation and parity are signed off. - •Repoints ship behind approval. Jobs, dbt sources and dashboards move only after a named owner signs off, and each wave keeps its rollback path.
- •Design target: 4 to 8 weeks for a migration that often takes 6 to 12 months. It's a target, not a measured result.
What it connects
Data Workers works across both sides of the move and the readers in between. Hive metastore federation puts the Glue and Hive tables in reach of your SQL warehouse, so your coding agent lists their tables, partitions and schemas through Databricks' managed DBSQL MCP server and hands them to Data Workers; Lake Formation grants come in the same way through AWS's MCP server. Data Workers' Databricks connector reads Unity Catalog grants and applies approved ones through the permissions API, and dbt and Airflow show who reads each table today. The agents run as MCP servers, so your coding agent calls them from the session it already uses.
| Databricks surface | What it does in the move | Status, October 2026 |
|---|---|---|
| UCX | Databricks Labs toolkit: assessment, group migration and table migration workflows (SYNC for external tables, deep clone for managed) | Labs project, provided as-is outside Databricks support SLAs; its code migration workflow is listed as under development; latest release v0.60.1 (Oct 2025) |
| Upgrade wizard | Catalog Explorer bulk copy of schemas and tables into Unity Catalog as external tables | Documented UI path |
SYNC | Upgrades external tables and managed tables outside workspace storage; DRY RUN checks first; writes an upgraded_to property on the source | Documented SQL command |
CREATE TABLE CLONE | Deep clones managed Delta tables into managed Unity Catalog tables, including managed tables in workspace storage that SYNC can't upgrade | Documented; simpler than CTAS |
| Hive metastore federation | Mirrors an internal, external or Glue metastore as a foreign catalog; read and write for the internal metastore, read-only for external and Glue metastores | GA; shallow clone in Public Preview |
What moves in each direction
| Data Workers reads | Data Workers sends back |
|---|---|
| Glue tables, partitions and Lake Formation grants, handed over by your coding agent | Proposed parity queries for your SQL warehouse: counts, schemas, aggregates |
| Hive metastore tables and partitions, through federation and the same agent | A wave gate result: pass, or the tables still blocking |
| Unity Catalog grants, read by the Databricks connector | Repoint changes for jobs, dbt sources and dashboards, after approval |
| dbt and Airflow lineage: who reads each table today | Unity Catalog grants matched to the old ones, after approval |
| The copy results from SYNC, CLONE or UCX | A receipt per wave, with the way back |

Three agents share the work. The Migration agent plans waves with map_migration_dependencies, which builds the blast radius from lineage and returns a wave-ordered plan, then plans each wave's validation and parity checks, proposes the comparison queries for your SQL warehouse and tracks the results at gate_migration_complete. The Governance agent compares grants and applies new ones through approvals. The Autonomous Data-Conductor sequences each wave and holds every write at the autonomy level you set. For the wider picture, read Data Workers on Databricks and, if your Glue estate is large, Data Workers on AWS.
Prerequisites
- •A Unity Catalog target with a storage credential and external locations for the data you're moving, and the target catalog and schemas created.
- •Glue federated as a foreign catalog (optional, recommended). With a service credential, a connection and a foreign catalog in place, one SQL warehouse can query the Glue original and the Unity Catalog copy side by side. Workspace Hive tables are already readable as
hive_metastore. - •A service principal for Data Workers with
USE CATALOG,USE SCHEMAandSELECTon both sides,CAN USEon one SQL warehouse, and read access tosystem.queryandsystem.access. - •AWS read access for the AWS MCP server in your coding agent:
glue:Get*andglue:SearchTableson the databases in scope, pluslakeformation:ListPermissionsandlakeformation:GetResourceLFTags. - •Your agreed tolerance. The Migration agent's wave plans default to row counts within 0.01% and aggregates within 0.1%. Set your own per wave.
No write grants yet. Data Workers acts with the grants you give it, and the first waves run read-only.
Setup
Add the three agents to your MCP client. The connectors read the same variables in any client.
Example: MCP client config
{
"mcpServers": {
"dw-migration": {
"command": "./start-agent.sh",
"args": ["dw-migration"],
"env": {
"DATABRICKS_HOST": "https://<workspace>.cloud.databricks.com",
"DATABRICKS_TOKEN": "<OAuth token for the dw-migrator service principal>",
"DATABRICKS_HTTP_PATH": "/sql/1.0/warehouses/<warehouse-id>",
"DATABRICKS_CATALOG": "main"
}
},
"dw-governance": { "command": "./start-agent.sh", "args": ["dw-governance"] },
"dw-conductor": { "command": "./start-agent.sh", "args": ["dw-conductor"] }
}
}Then ask for the plan. Your coding agent calls map_migration_dependencies for the source, reviews the wave order with you, and you approve the waves you want in the first month. Each wave is a named list of tables and the readers that move with them.
One wave, end to end
Here is wave 2 of a typical move. It's an illustration, not a customer case. The finance wave covers eight tables: three in Glue (finance_raw) and five in the workspace Hive metastore (finance). The platform team has already run SYNC for the external tables and a deep clone for the managed ones.
| Time | Step | What happens | Who decides |
|---|---|---|---|
| 09:00 | Copy | The platform team runs SYNC SCHEMA and CREATE TABLE ... DEEP CLONE into main.finance. | Your team, in Databricks |
| 09:40 | Read | The coding agent hands Data Workers the three Glue tables, their partitions and the Lake Formation grants, then the five Hive tables. | Data Workers, read-only |
| 09:52 | Validate | The platform team runs the validation queries Data Workers proposed (row count, schema match, column stats) on its SQL warehouse: 8 of 8 tables match on count and schema. Data Workers logs the results against the wave. | Your team, in Databricks |
| 10:05 | Parity | The team's aggregate check shows SUM(amount) on ap_invoices differs. The owner compares the Glue partition list with the copy, using the check Data Workers proposes, and finds two that landed after the sync. | Your team runs it; Data Workers proposes the next check |
| 10:20 | Gate | gate_migration_complete returns "not complete" with ap_invoices blocking, and Data Workers proposes a re-sync. | Data Workers proposes |
| 11:10 | Fix | An engineer approves; the re-sync runs; the team reruns the proposed parity queries and ap_invoices matches. | A named engineer |
| 11:15 | Grants | One Lake Formation grant, the AP analysts' read access, has no Unity Catalog match. Data Workers drafts the grant. | Data Workers proposes |
| 11:30 | Repoint plan | Lineage and query history show three Airflow DAGs, two dbt sources and four dashboard tiles still reading the old names. Data Workers drafts each change with its blast radius. | Data Workers proposes |
| 14:00 | Approve | The finance data owner approves the grant and the repoint in Spellbook. | A named person |
| 14:25 | Repoint and verify | The DAG and dbt changes deploy as reviewed changes through the Pipelines agent, the tiles move, and the finance team checks the September close totals against the source, with the result logged at the wave gate. | Data Workers, after approval |

The receipt it leaves. One record per wave in Spellbook Data Catalog (in preview): the eight tables and how each was copied, the parity results per table with counts, aggregates and the tolerance used, the re-sync and why, the grants compared and the one added, every repointed reader with a link to its change, the approvers and times, and the rollback steps. The Hive and Glue tables are untouched apart from the upgraded_to property SYNC writes, which the receipt records.
Rollback per wave. Because the originals stay in place, going back means reverting the repoints in reverse dependency order: dashboards first, then dbt sources, then DAGs. Data Workers checks it is reverting the version it recorded before it starts, and the readers return to the Hive and Glue names. The Unity Catalog copies stay for the next attempt. One wave rolls back without touching the waves already done.
Why doesn't Databricks just do this itself?
Databricks built SYNC, CLONE, UCX and federation for one job: getting tables and permissions into Unity Catalog inside Databricks. That focus is the right design, and it's why the copy is reliable.
Gating a wave on its parity checks and moving everything that reads it is a different product. The readers live in Airflow, dbt, BI tools and reverse ETL jobs Databricks doesn't run. Deciding a wave is done takes business tolerances, an approver who owns the data, and a record an auditor can read. Owning a rollback across five tools means taking responsibility for changes in systems another vendor controls. Databricks keeping its tools focused on Databricks is sensible. The cross-system gate, approval and rollback layer is the product Data Workers is. If you're weighing catalogs too, Unity Catalog vs Spellbook covers where each fits, and metric views and lineage in Context Wizard covers the definitions that ride along with the tables.
The next autonomy step

The wave above runs at L1 and L2: Data Workers observes, plans each wave's parity checks and tracks the results on its own, and proposes every re-sync, grant and repoint for a person to approve. After a few waves with clean receipts, most teams move parity re-syncs for tables with no downstream readers to L3, act reversibly, because the original is still there. Repoints and grants stay at L2, propose, until the domain owner is ready. Autonomy is set per domain, an agent can't approve its own work, and every step keeps its rollback path.
The case for your CFO
The outcome. The Hive and Glue estate retires on a schedule, and every table that moves arrives with parity results the owner signed off. Our design target is 4 to 8 weeks for a migration that often takes 6 to 12 months, and it's a target, not a promise.
The risk story. Data Workers works read-only on both sides. Anything that changes a reader or a grant waits for a named owner. Each wave has a receipt with the parity results, approvers and rollback steps, and the originals stay in place until you retire them. Data Workers works with your Databricks and AWS accounts as they are.
Why now. Databricks calls the workspace Hive metastore legacy and recommends disabling direct access once tables move. Every quarter on two catalogs is a quarter of two permission models.
The first win. One wave, read-only: a parity plan and comparison queries for tables you've already synced, run by your team, with the readers still pointing at the old names.
What stays the same. SYNC, CLONE, UCX, Catalog Explorer, your jobs and dashboards.
The pilot path. Start with a pilot on one wave. The pilot is credited in full against the first year.
One sentence for upstairs: "Databricks copies the tables; Data Workers plans every wave's parity checks and holds the gate until we sign off, moves the jobs and dashboards behind an approval, and leaves a receipt we can roll back from."
FAQ
Does Data Workers replace UCX or SYNC? No. Those stay your copy tools. Data Workers plans the checks on what they produced, holds the gate and moves the readers.
Can it plan checks for Glue tables that Unity Catalog can only read? Yes. Federation makes Glue a read-only foreign catalog, which is all a parity query needs: your team runs the comparison side by side on one SQL warehouse, and the Lake Formation grants come in through AWS's MCP server in your coding agent.
What counts as parity? Row counts, schema match, column stats, sampled hashes and aggregates within your tolerance, plus a checksum bisection for tables that need row-level certainty. Data Workers puts each in the wave plan as a query your team runs, and the gate holds until the owner signs off on the results.
What if a wave fails after repointing? Revert that wave's repoints in reverse order. The originals and the other waves are untouched.
Sources
Databricks capabilities and statuses, checked October 2, 2026: migrate to Unity Catalog (paths, CLONE vs CTAS; updated Sep 11, 2026), UCX (workflows, Labs support status, code migration workflow under development; Sep 11, 2026), UCX releases (v0.60.1, Oct 3, 2025), SYNC (DRY RUN, upgraded_to; Sep 11, 2026), Hive metastore (legacy status, disabling direct access; Sep 11, 2026), Hive metastore federation (GA, read and write for internal, read-only for Glue and external; Sep 11, 2026) and federation for AWS Glue (service credential, connection, foreign catalog; Sep 29, 2026). Databricks pages differ on Glue write support; we follow the federation overview and treat Glue foreign catalogs as read-only. Product names and statuses change quickly; if we've got something wrong, tell us and we'll fix it.