Product
Product12 min readBy The Data Workers Team

You're on DataHub: Keep the Open Map, Add the Control Plane That Keeps It True

DataHub is the metadata graph your team and its agents trust. Data Workers works beside it and does the operations work that keeps it true: it traces a failing assertion to its cause, proposes the fix and verifies it, behind approvals.

Your platform team runs DataHub because it wanted a metadata graph it could own. Ingestion recipes pull entities, schemas and column-level lineage from Kafka, the lake, Airflow, dbt and the dashboards. Owners, domains, glossary terms and data products sit on every dataset that matters. On DataHub Cloud, freshness and volume assertions watch the tables people depend on and raise a DataHub incident when one fails. DataHub now calls itself "The Context Platform for AI Agents", and agents read the same graph: the DataHub MCP server serves search, entities, lineage and the real SQL behind a dataset to Claude, Cursor and the rest; Ask DataHub answers questions in the UI, Slack and Teams; and on DataHub Cloud, custom agents run tasks that keep the context current. DataHub is the open map your team trusts. Data Workers works beside it and does the operations work that keeps it true: it finds the cause behind a failing assertion, reruns the right job and proves the number, behind your approvals.

Key takeaways

  • •DataHub keeps its job. The graph, ingestion, lineage, glossary, assertions, incidents, Ask DataHub, its agents and the MCP server stay exactly where they are.
  • •DataHub's graph stays your team's map. DataHub connects over its API or MCP server today: its MCP server runs in your team's client next to Data Workers, on DataHub Core or DataHub Cloud, while Data Workers builds its own lineage and run history from dbt and Airflow.
  • •A failing assertion turns into a fix. Data Workers ties the failure to its cause in Kafka Connect, Airflow or dbt, proposes the change with its blast radius and routes it to the domain's named owner.
  • •DataHub stays the record of its own metadata. Data Workers writes nothing to DataHub. Description changes are proposed for the owner and reach DataHub through dbt docs on its next ingestion.
  • •Start with a pilot. One domain, read-only, then one fix class, on the ladder from L0 manual to L4 autonomous.

DataHub is the open map. Data Workers is the control plane.

DataHub describes the estate from a Kafka topic to a dashboard chart. Core v1.7.0, released August 4, 2026, made metrics and semantic models first-class entities and put data products directly into the lineage graph; the LTS line got v1.6.0.3 on September 25, 2026. Around that graph sits everything that decides whether it is accurate: the connectors, the Spark jobs, the Airflow DAGs, the dbt models and the dashboards. Data Workers runs that side with one context, one approval flow and one audit trail, next to DataHub's graph.

Here is a Thursday at a delivery marketplace that runs an open lakehouse. This is an illustration, not a customer case.

TimeSystemWhat happens
Wed 23:05Kafka ConnectA broker restart fails one task of the S3 sink for the orders.v2 topic. Hour 23 stops landing in the raw bucket
Wed 23:41Kafka ConnectThe task restarts and replays the gap into S3 from its committed offsets
Thu 00:30Airflow and SparkThe orders_daily DAG run starts on schedule. Its sensor only checks that Wednesday's folder exists, so the Spark task writes Wednesday's partition to the Iceberg table lake.orders in Nessie while the replay is still landing: 19% fewer rows than a normal Wednesday
Thu 01:10dbt on Trinodbt builds fct_orders_daily and mart_courier_payouts on the short partition. Unique and not-null tests pass
Thu 02:00DataHubThe DataHub Cloud volume assertion on lake.orders fails (row count change of minus 19% against a 10% bound). DataHub raises a Volume incident and posts it to the team's Slack channel
Thu 02:03Data WorkersThe row-count baseline the team records with monitor_metrics shows the same drop. Data Workers reads the Airflow run, joins it to the sink task's failure window, and get_root_cause names the cause: the DAG ran before the replay finished
Thu 02:08Data Workersblast_radius_analysis lists two dbt models, the Superset "Courier payouts" dashboard (from a team note in the context graph) and the payout export due at 06:00
Thu 02:12Data WorkersIt proposes two changes: a backfill of Wednesday's partition through Airflow, then the dbt rebuild, with the undo step written into the plan; and a sensor change that waits for the sink's task status and expected file count, as a diff for the owner to merge
Thu 02:31Slack and SpellbookThe request reaches the domain's on-call owner; he reviews the plan, the blast radius and the undo path in Spellbook and approves the backfill
Thu 02:32Airflow, Spark, dbtData Workers queues the task rerun through Airflow and reads its status, then queues the dbt rebuild. The monitor_metrics baseline confirms the row count sits within 1% of the trailing four Wednesdays and the payout-total baseline holds. It writes the receipt
Thu 04:00DataHubThe assertion's next evaluation passes. The owner resolves the DataHub incident and links the receipt
Thu 06:00SupersetThe payout export and dashboard run on the full partition
Thu 10:15GitHubThe owner merges the sensor diff after review
Incident timeline across the stack: what DataHub, your team and Data Workers each do, step by step

DataHub did its part exactly: the assertion caught a drop every dbt test missed and the lineage showed the payout dashboard two hops downstream. The cause lived in a connector, a sensor and a schedule, and the fix needed the owner's approval. Without the trace and the backfill, couriers would have been paid on 81% of Wednesday's orders.

JobWhat DataHub doesWhat Data Workers does
The mapIngests entities, schemas, owners, glossary terms and column-level lineage from every sourceBuilds its own lineage and run history from the dbt manifest and Airflow runs; DataHub's graph reaches your team's client over its MCP server
The signalRuns freshness and volume assertions on DataHub Cloud and raises an incident when one failsRuns its own checks on the same tables and opens an incident with the cause attached
The causeShows lineage and the history of the dataset; Ask DataHub traces impactTraces from the Iceberg table through the Airflow DAG to the Kafka Connect task that left the gap
The fixRecords owners, terms and incident stateProposes the backfill and the DAG change with their blast radius and routes them to the named owner
The proofRe-evaluates the assertion on its scheduleQueues the rerun, checks the result, and keeps a receipt with the cause, plan, approver, checks and undo path
The recordThe owner resolves the incident; stewards update descriptionsDrafts the description change; the dbt docs change reaches DataHub on its next dbt ingestion

Why doesn't DataHub just do this itself?

Because DataHub built a great product for one job: a context layer that describes every system you run. That job needs read access to nearly everything and changes to nothing outside the graph, which is why platform teams run it with confidence.

DataHub's AI follows the same line. The MCP server's mutation tools change DataHub metadata (tags, glossary terms, owners, domains, descriptions, structured properties and documents); they stay off until you set TOOLS_IS_MUTATION_ENABLED=true, and each is annotated readOnlyHint: false so a client can ask for confirmation. Since DataHub Cloud v2.3.0 (September 30, 2026), Ask DataHub can propose changes to descriptions, tags, domains, structured properties, glossary terms and documents. Custom agents, in public beta in the DataHub Cloud Context add-on, run tasks against the metadata graph on a schedule or on events, such as filling in missing descriptions, with tool review and human Decisions checkpoints. All of it keeps the context layer healthy.

Rerunning a Spark job, editing an Airflow sensor and rebuilding dbt models is a different product. It needs the code, the blast radius across systems DataHub describes but doesn't run, scoped credentials, a named-owner approval, a verified result, a rollback path and the liability for the change. That product is Data Workers.

Every tool owns a slice. Data Workers covers the whole lifecycle

DataHub owns one slice outright: the metadata graph people and agents trust. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail. We use the same DataHub scores as our data catalog comparison of Collibra, Alation, DataHub and OpenMetadata. On Catalog & Context, DataHub ties Data Workers: that is its home stage.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, DataHub goes deep on its own area
StageData WorkersDataHubWhy we scored it this way
Catalog & Context99DataHub's home stage, tied: an open metadata graph of entities, schemas, owners, glossary terms, domains and column lineage, served to agents over its MCP server. Data Workers works next to it, joining dbt, run results and tests.
Analytics & Insights83Ask DataHub (public beta) drafts SQL, the open-source Analytics Agent runs it, and the Metrics Catalog defines metrics once. Data Workers answers questions across platforms from approved definitions.
Data Quality86.5DataHub Cloud runs freshness and volume assertions, and Core takes assertion results from dbt and Great Expectations. Data Workers writes, runs and repairs checks and dbt tests.
Observability & Incidents8.56Assertions can raise a DataHub incident and notify owners in Slack. Data Workers owns the loop after the alert: diagnose, fix, verify, record.
Pipelines & Ingestion8.53DataHub ingests pipeline metadata; it doesn't change pipelines. Data Workers proposes the DAG or model change and queues the rerun behind approval.
Schema & Migration84Schema history and lineage show a change's reach. Data Workers detects the change at the source and drafts the fix with rollback SQL.
Governance & Access8.57Policies, ownership, domains and data products, with mutation tools off by default. Data Workers routes approvals to each domain's named owners.
Security & Privacy85Classification through tags and glossary terms. Data Workers flags sensitive column names in pull request review and dry-runs grants against sensitive columns.
Cost / FinOps82Usage and popularity, not warehouse spend. Data Workers attributes Snowflake spend to the dbt model and drafts the fix for its owner.
MLOps & Models7.54ML models and features as catalog entities with lineage. Data Workers keeps the data under models healthy.

These are directional scores of scope, not benchmarks, and the reasoning is shown so you can check every line.

How DataHub and Data Workers work together

Your people stay where they are: analysts in Ask DataHub, engineers in Claude Code or Cursor with DataHub's MCP server attached. Spellbook Data Catalog (in preview) is where the data team looks: each proposed change, the owner it touches, its blast radius, the checks and the rollback. Between them, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm does the work with more than 20 specialist agents, and the Autonomous Data-Conductor runs each fix end to end (detect, diagnose, fix, review, verify, remember) under per-domain guardrails.

How Data Workers fits with DataHub: your coding agent on top, Data Workers in the middle, your estate underneath

What Data Workers reads, and where DataHub fits. DataHub connects over its API or MCP server today: the MCP server runs in your engineers' client next to Data Workers' agents, so the people using that client bring DataHub's datasets, owners, terms and lineage into the same conversation. Data Workers' own context comes from what it reads natively, each fact kept with its source and time:

  • •dbt. The manifest and run results: models, sources and tests; change review's lineage diff also reads the Superset dashboards declared as exposures when a fix goes through a pull request.
  • •Orchestration. Airflow DAG runs and task instances, which is how a run that started too early shows up.
  • •Lineage. dbt lineage and Airflow runs, plus the team's notes for hops dbt can't see, walked with trace_cross_platform_lineage.
  • •Owners and domains. Each domain has named owners in Data Workers, usually the people DataHub lists. They approve changes, and domains are where you set autonomy levels.

DataHub's assertions stay DataHub's signal to your team. Data Workers confirms a failure independently with its own checks (run_quality_check on Snowflake and BigQuery, a monitor_metrics baseline elsewhere) before anything is proposed. explain_table returns a table's definition, lineage, documentation and trust score in one call; quality comes from get_quality_score and open incidents from get_incident_history.

When a fix should change what DataHub says, Data Workers proposes it for the owner. Approved descriptions go back as dbt docs changes (local files by default, or a pull request when your team turns on the GitHub pull-request target), and DataHub picks them up on its next dbt ingestion. Approved facts land in the Context Wizard graph, where a named person promotes them to authoritative; no agent can promote its own work. Tags, owners and terms in DataHub stay with their owners, by design.

Setup over MCP today. Add DataHub's MCP server and Data Workers' agents to the same client. For Data Workers, clone the open-source repository and add start-agent.sh entries, as the client setup docs show, next to DataHub's MCP server, so both serve the same client.

// Example: .mcp.json for Claude Code
{
  "mcpServers": {
    "datahub": {
      "command": "uvx",
      "args": ["mcp-server-datahub"],
      "env": {
        "DATAHUB_GMS_URL": "https://<your-datahub>",
        "DATAHUB_GMS_TOKEN": "<read-only personal access token>"
      }
    },
    "dw-context-catalog": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-context-catalog"],
    },
    "dw-quality": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-quality"]
    },
    "dw-incidents": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-incidents"]
    }
  }
}

Leave DataHub's mutation tools off; Data Workers needs read access only. List the tools with /mcp in Claude Code. Then "why did the orders volume assertion fail?" gets the owner and downstream dashboards from DataHub's get_entities and get_lineage, the cause from get_root_cause and the fix's reach from blast_radius_analysis, in one answer.

One request, L0 to L4. The autonomy ladder is set per domain.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. Your engineer chases the assertion by hand.
  • •L1 observe. Data Workers explains each failing assertion with its cause, owner and blast radius, and changes nothing.
  • •L2 propose. Data Workers drafts the backfill plan or the DAG diff; the owner approves in Spellbook before anything runs.
  • •L3 act reversibly. For proven classes, such as backfilling a short partition after an upstream replay, Data Workers queues the rerun, verifies the result and keeps the undo step ready for the owner.
  • •L4 autonomous. For a scoped, trusted class in one domain, such as late loads into orders, Data Workers fixes and verifies on its own and posts the receipt for review.

The safety model is in is it safe to let AI agents change production data; where data and credentials live is in where does our data go. The agents run in your infrastructure: your data stays in your systems, and the hosted Conductor sees workflow metadata only.

Teams on other catalogs follow the same pattern: see You're on OpenMetadata, You're on Alation and You're on Atlan. For choosing and wiring, see Data Workers vs DataHub and Claude Code with DataHub.

What changes for your team

Six jobs that run on autopilot with Data Workers next to DataHub, with a concrete example of each

DataHub teams spend much of the week turning signals into work: reading a failing assertion at 2 a.m., finding which DAG left the gap and whether a backfill is safe. With Data Workers on top, those jobs run on autopilot at the level you set.

  • •Incidents. A failing DataHub assertion arrives traced to its cause, with the fix drafted for the owner.
  • •Data quality. Failed checks turn into fixes, reruns and a verified result the owner signs off.
  • •Cloud spend. Snowflake spend is traced to the dbt model behind it, with the change drafted for its owner.
  • •Access. A request for a table found in DataHub is dry-run for effective privileges, sensitive columns and policy conflicts, with a least-privilege recommendation and an expiry. On Databricks, Data Workers applies the approved Unity Catalog grant; elsewhere the grant is proposed for the owner to apply.
  • •Audits. Every change carries the dataset, owner, diff, checks and rollback path (who owns the agents).
  • •Migrations. A legacy warehouse move runs in planned waves, with parity checks tracked and the completion gate held for the owner's sign-off.

Keep DataHub, or consolidate?

Keep DataHub if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

For most teams the answer is to keep it: your recipes, glossary, ownership and lineage live there, and your agents already read it over MCP. What teams consolidate is the tooling around it: the runbook wiki for assertion failures and the spreadsheet that maps datasets to DAGs and on-call owners. If you are weighing building this layer yourself on DataHub's MCP server and the Agent Context Kit, read build it ourselves with Claude Code and MCP servers: the connection is the easy part; cross-system context, approvals, verification and rollback are the work.

The case for your CFO

The outcome: the metadata graph your platform team already runs pays back in correct numbers. Every owner, lineage edge and assertion in DataHub sits next to the context Data Workers uses to fix breaks before payouts or reports go out wrong.

The risk story is plain. Data Workers writes nothing to DataHub; DataHub stays the record of what the data is and who owns it. At L1 the agents observe and change nothing. At L2 every change is proposed with its blast radius and goes to the domain's named owner. At L3 only proven, reversible change classes run without a fresh approval. Every action leaves a receipt: the trigger, the diff, the blast radius, the checks with before and after values, the approver and the undo path. Requests no one answers expire and escalate; they never auto-grant. There is zero migration: DataHub, Kafka, Iceberg, Airflow, dbt and Superset stay where they are.

Why now: agents already read DataHub over MCP every day; with Data Workers in the same client, fixes follow, with a person approving each one. The first win is one domain, read-only, where every failing assertion arrives explained and owned. Your catalog, ingestion, review process and permissions stay the same. For the numbers, see the ROI of agentic data operations.

The sentence to repeat upstairs: "DataHub tells us what our data is and who owns it; Data Workers keeps it right, with the owner approving every fix and a receipt for each one."

Getting started

Start with a pilot. Pick one domain your teams lean on, such as orders, add DataHub's MCP server next to Data Workers in one team's client, connect Data Workers read-only to the dbt project and Airflow behind that domain, and let it explain every failing assertion in that domain before you turn on the first fix class. Plans and the pilot path are on the pricing page, and the pilot is credited in full against the first year.

FAQ

Does Data Workers write to DataHub? No. DataHub connects over its API or MCP server today, and Data Workers changes nothing in it. When a fix should change a description, Data Workers proposes it for the owner, and the dbt docs change reaches DataHub on its next dbt ingestion.

We already have Ask DataHub, custom agents and the MCP server. Why add Data Workers? DataHub's AI works on the graph: it finds data, traces impact, drafts SQL and changes metadata. Data Workers works on the systems behind the graph: the connector, the Spark job, the Airflow DAG and the dbt model, with an approval and a receipt for each change.

Does Data Workers work with DataHub Core, or only DataHub Cloud? Both: DataHub's MCP server works with either. Volume and freshness assertions are a DataHub Cloud feature; on Core, Data Workers' own checks provide the signal.

How does Data Workers know who approves a change? From the domain owners you name in Data Workers, usually the people DataHub lists on the dataset. Requests reach that person in Slack or email, and the decision happens in Spellbook. Unanswered requests expire and escalate; they never auto-grant.

What does Data Workers store, and where? The agents run in your infrastructure and keep the context graph, receipts and audit log there, not copies of your tables. The hosted Conductor sees workflow metadata only.

Sources

  • •DataHub, homepage ("The Context Platform for AI Agents"; DataHub Core and DataHub Cloud; © Acryl Data, Inc.), https://datahub.com/ (checked Oct 3, 2026)
  • •DataHub, release notes (v1.7.0 Aug 4, 2026; v1.7.0.1 Sep 3, 2026; v1.6.0.3 Sep 25, 2026), https://docs.datahub.com/docs/releases (checked Oct 3, 2026)
  • •DataHub, GitHub releases (v1.6.0.3 published Sep 25, 2026), https://github.com/datahub-project/datahub/releases (checked Oct 3, 2026)
  • •DataHub Cloud, v2.3.0 release notes (Sep 30, 2026: Context module public beta; Ask DataHub proposal support; Ask DataHub and custom agents in the MCP server), https://docs.datahub.com/docs/managed-datahub/release-notes/v_2_3_0 (checked Oct 3, 2026)
  • •DataHub, Custom Agents (public beta, DataHub Cloud Context add-on; tasks, Decisions, tool review, AI Plugins), https://docs.datahub.com/docs/features/feature-guides/agents (checked Oct 3, 2026)
  • •DataHub, MCP server feature guide (managed and self-hosted; read and mutation tools; readOnlyHint: false; authentication), https://docs.datahub.com/docs/features/feature-guides/mcp (checked Oct 3, 2026)
  • •Acryl Data, mcp-server-datahub repository and releases (v0.7.1, Sep 16, 2026; TOOLS_IS_MUTATION_ENABLED), https://github.com/acryldata/mcp-server-datahub (checked Oct 3, 2026)
  • •DataHub, Ask DataHub (public beta; UI, Slack, Teams, Chrome extension; first-draft SQL), https://docs.datahub.com/docs/features/feature-guides/ask-datahub (checked Oct 3, 2026)
  • •DataHub, Analytics Agent (open source, Apr 30, 2026), https://docs.datahub.com/docs/features/feature-guides/analytics-agent (checked Oct 3, 2026)
  • •DataHub, Agent Context Kit, https://docs.datahub.com/docs/dev-guides/agent-context/agent-context (checked Oct 3, 2026)
  • •DataHub, Volume Assertions (DataHub Cloud Observe; incidents and Slack notifications), https://docs.datahub.com/docs/managed-datahub/observe/volume-assertions (checked Oct 3, 2026)
  • •DataHub, Incidents (Core and Cloud; raise and resolve), https://docs.datahub.com/docs/incidents/incidents (checked Oct 3, 2026)
  • •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)
  • •Data Workers open-source repository (agent setup, start-agent.sh), https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)