Product
Product12 min readBy The Data Workers Team

You're on OpenMetadata: Keep the Open Map, Add the Control Plane That Fixes What Its Tests Find

OpenMetadata is the open metadata platform your team runs: entities, lineage, test cases and an MCP server on by default. Data Workers works beside it and does the operations work that keeps it true, behind approvals.

Your platform team runs OpenMetadata. Ingestion workflows crawl BigQuery, dbt, Airflow and Superset into one knowledge graph of entities, so every table has its columns, owner, tier, glossary terms and lineage from the source database to the dashboard. When a test case fails, the Incident Manager opens an incident for someone to acknowledge and assign. Since OpenMetadata 2.0.0 (August 24, 2026; the current release is 2.0.3, September 30), the Context Center holds your articles and documents next to the assets, and the MCP server ships "installed and enabled by default", so Claude, Cursor or Claude Code can search the catalog, walk lineage and run root_cause_analysis on a failed test. Some of you self-host it on Apache 2.0; others run Collate, the managed edition with Collate AI. A failed test is a precise signal, but someone still has to find the cause behind the table, change the code, rebuild the data and prove the number. Data Workers works beside OpenMetadata and turns that signal into a verified fix, behind your approvals.

Key takeaways

  • •OpenMetadata keeps its job. Entities, glossary, lineage, test cases, the Incident Manager, the Context Center and its MCP server stay exactly where they are.
  • •OpenMetadata and Data Workers sit side by side. OpenMetadata connects over its API or MCP server today, in your team's client next to Data Workers' agents, and Data Context Wizard builds its own context from schema changes, run history and checks in PostgreSQL, BigQuery, dbt and Airflow.
  • •A failed test turns into a fix. Data Workers traces the failure to its cause in the systems OpenMetadata describes, proposes the change with its blast radius and routes it to the named owner.
  • •OpenMetadata stays the record of its own metadata. Data Workers writes nothing to it; approved descriptions arrive as dbt docs changes on OpenMetadata's next dbt ingestion.
  • •Start with a pilot. One domain's test cases, read-only, then one fix class, on the ladder from L0 manual to L4 autonomous.

OpenMetadata is the open map with the inspector's checklist. Data Workers is the control plane that turns a failed check into a verified fix.

OpenMetadata is the open map of your estate, with more than 130 connectors by its own count, and its test cases are the inspector's checklist: row counts, value ranges, custom SQL and table diffs. Its MCP server gives agents read tools for search, entities, lineage and test definitions, plus root_cause_analysis, which walks quality lineage back to a failure's origin. On Collate, the Documentation Agent writes descriptions, the Tier Agent classifies assets by usage and impact, and the Quality Agent creates tests.

Around the map sit the systems that decide whether the roads are sound: the app database, the load DAG, the warehouse, the dbt models. Data Workers works there with one context, one approval flow and one audit trail, next to OpenMetadata's graph and test results.

Here is a Tuesday at a grocery delivery business in Chicago that runs on Google Cloud. This is an illustration, not a customer case.

TimeSystemWhat happens
Mon 22:40PostgreSQL (Cloud SQL)An app migration changes orders.placed_at from timestamp to timestamptz. The app used to store Chicago wall-clock time; it now stores UTC
Tue 02:00Managed Service for Apache AirflowThe orders_to_bigquery DAG run loads the night's orders into raw.orders in BigQuery, now with UTC times
Tue 02:05Data WorkersReviewing the app migration merged in GitHub, Data Workers records the type change on orders.placed_at. Nothing has failed yet, so it logs the change with its downstream tables
Tue 02:30BigQuery and dbtdbt builds fct_daily_orders with date(placed_at). The 2,734 orders placed after 19:00 Chicago time on Monday now carry Tuesday's date. Every dbt test passes
Tue 06:10OpenMetadataThe scheduled test suite runs. columnValuesToBeBetween on fct_daily_orders.order_count fails: Monday shows 9,412 orders against a floor of 11,000. The Incident Manager opens an incident and marks it New
Tue 06:12Data WorkersIts own run_quality_check on fct_daily_orders in BigQuery confirms the shortfall against the same floor. It joins that to the 02:05 type change and finds the cause with get_root_cause: dates are taken from UTC timestamps. blast_radius_analysis lists four dbt models, the "Daily Orders" and "Courier Staffing" Superset dashboards and the weekly finance export (all three from the team's context-graph notes)
Tue 06:25Data WorkersIt proposes a dbt diff in stg_orders that derives order_date as date(placed_at, 'America/Chicago'), plus a dbt test that checks daily counts in fct_daily_orders against raw.orders by Chicago day, as a diff for the owner to merge. The plan rebuilds the Monday and Tuesday partitions; the undo step (revert the commit, rebuild the same two partitions) is written down before anything runs
Tue 08:05OpenMetadataThe data platform lead acknowledges the incident and assigns it to the analytics engineer who owns fct_daily_orders, which opens a task for her
Tue 08:12Slack and SpellbookThe approval request is already in her Slack. She reviews the diff, the blast radius and the undo step in Spellbook, merges the diff and approves the rebuild. dbt CI passes
Tue 08:30Airflow and BigQueryData Workers queues the rebuild of both partitions through Airflow and reads the task status. Then run_quality_check passes on fct_daily_orders in BigQuery: Monday now holds 12,146 orders, above the 11,000 floor, and the new dbt test passes. It writes the receipt
Tue 09:00SupersetThe dashboards' cache refreshes on schedule (Superset connects over its API today); the staffing lead sees Monday evening's orders back before the courier plan is set
Tue 09:20OpenMetadataThe engineer reruns the test suite; the test case passes. She resolves the incident with the reason and a comment that links the receipt
Incident timeline across the stack: what OpenMetadata, your team and Data Workers each do, step by step

OpenMetadata did its part exactly: the right test on the right table, an incident with an owner and a status, and lineage to the dashboards. The cause lived in an app migration, a load DAG and a dbt model. Without the trace, Tuesday's courier staffing would have been planned on a Monday that looked 22% quieter than it was.

JobWhat OpenMetadata doesWhat Data Workers does
The mapHolds entities, columns, owners, tiers, glossary terms and lineage from source to dashboardBuilds its own lineage and run history from PostgreSQL, BigQuery, dbt and Airflow; OpenMetadata's graph reaches your team's client over its MCP server
The checkRuns test cases and opens an incident when one fails, with acknowledge, assign and resolvePicks up the failed test result and opens its own investigation with the cause
The causeroot_cause_analysis traces the failure back along quality lineageTraces from the BigQuery table through dbt and the Airflow DAG to the PostgreSQL type change
The fixRecords what owners and stewards change in the catalogProposes the dbt change with its blast radius and routes it to the named owner for approval
The proofReruns the test case and records the resolutionQueues the rebuild, runs quality checks on the rebuilt BigQuery table and keeps a receipt with the cause, diff, approver, checks and undo path
The catalog recordOwners update descriptions and resolve the incidentDrafts the column description for the owner; the dbt doc change reaches OpenMetadata on its next dbt ingestion

Why doesn't OpenMetadata just do this itself?

Because OpenMetadata built a great product for one job: the open, shared map of what your data is, where it flows and whether it passes its tests. That job needs read access to nearly everything and write access to almost nothing, which is why teams can open it to everyone.

Its write surface follows from that job. The MCP server's write tools, create_entity, patch_entity, create_lineage and create_test_case, change OpenMetadata's own graph, and the reference for create_entity is frank: "Writes take effect immediately, so confirm the target with the user before calling." root_cause_analysis is read-only. Collate's Documentation Agent "will either update or suggest descriptions, keeping your teams the ability to accept or deny the agent's requests." Every one of those writes lands in the catalog.

Changing the dbt model, rebuilding BigQuery partitions through the Airflow DAG and proving the rebuilt table passes its checks is a different product. It needs the code, blast radius across systems OpenMetadata doesn't own, scoped credentials, a named-owner approval, a checked result, a rollback path and liability for changes in someone else's tool. That product is Data Workers.

Every tool owns a slice. Data Workers covers the whole lifecycle

OpenMetadata owns one slice outright: the open map of your data and the tests that inspect it. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the catalog you already run. The scores match our data catalog comparison of Collibra, Alation, DataHub and OpenMetadata, where OpenMetadata ties Data Workers on Catalog & Context.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, OpenMetadata goes deep on its own area
StageData WorkersOpenMetadataWhy we scored it this way
Catalog & Context99OpenMetadata's home stage, tied: one knowledge graph of entities, glossary, lineage and the Context Center, served to agents by an MCP server that is on by default in 2.0. Data Workers works next to it, joining run results, schema changes and grants.
Analytics & Insights84AskCollate and AI Analytics answer questions on Collate, and the catalog feeds data insights. Data Workers answers questions across platforms from approved definitions.
Data Quality87A strong stage: test cases at table and column level, a profiler and Collate's Quality Agent. Data Workers writes, runs and repairs checks and dbt tests across platforms.
Observability & Incidents8.55A failed test opens an incident with an assignee, and root_cause_analysis walks quality lineage to the origin. Data Workers owns the loop after it: diagnose across systems, fix, verify, record.
Pipelines & Ingestion8.53OpenMetadata ingests pipeline metadata and status; it doesn't change pipelines. Data Workers proposes the model or DAG change and queues the rerun behind approval.
Schema & Migration83Lineage and column types show what a schema change touches. Data Workers catches the change in its migration pull request, scopes its blast radius and drafts the fix with rollback.
Governance & Access8.57Glossary, domains, roles and policies, tasks and custom intake forms for governance workflows. Data Workers routes approvals to named owners and applies approved Unity Catalog grants.
Security & Privacy85PII classification and tags on catalog assets. Data Workers dry-runs grants against sensitive columns and flags sensitive column names in pull request review.
Cost / FinOps82Usage and tiering, not warehouse spend. Data Workers attributes Snowflake spend to the dbt model and drafts the fix for its owner.
MLOps & Models7.53ML model entities in the catalog, and AI Governance Studio on Collate. Data Workers keeps the data under models healthy.

These are directional scores of scope, not benchmarks, and the reasoning is shown so you can check every line.

How OpenMetadata and Data Workers work together

Engineers keep asking Claude Code or Cursor with OpenMetadata's MCP server connected; Collate users keep AskCollate. Spellbook Data Catalog (in preview) is where the data team looks: each proposed change, its owner, blast radius, checks and rollback. Underneath, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm's 20+ specialist agents do the work, and the Autonomous Data-Conductor runs each fix end to end (detect, diagnose, fix, review, verify, remember) under per-domain guardrails.

How Data Workers fits with OpenMetadata: your coding agent on top, Data Workers in the middle, your estate underneath

What Data Workers reads, and where OpenMetadata fits. OpenMetadata connects over its API or MCP server today: the MCP server runs in your engineers' client next to Data Workers' agents, so the people using that client bring OpenMetadata's entities, lineage and test results into the same conversation.

  • •Schemas and changes. PostgreSQL and BigQuery schemas, with each change recorded from its migration pull request and the dbt manifest. A description becomes authoritative only when a named person promotes it; no agent can promote its own work.
  • •Lineage. dbt lineage, plus the Superset dashboards your team records in the context graph, joined across PostgreSQL, BigQuery, dbt and Airflow with trace_cross_platform_lineage; explain_table returns a table's definition, lineage, documentation and trust score.
  • •Checks. Data Workers' own run_quality_check on BigQuery confirms what a failed OpenMetadata test reports; its incidents come from get_incident_history, quality from get_quality_score.
  • •Writes. None to OpenMetadata. Catalog changes are proposed for the owner; approved descriptions go back as dbt docs changes (local files by default, or a pull request when your team turns on the GitHub pull-request target) that OpenMetadata picks up on its next dbt ingestion, and approved facts land in the Context Wizard graph. The Incident Manager stays your team's.

Setup today. For your engineers' clients, add OpenMetadata's MCP server and Data Workers' agents side by side: clone the open-source repository and add start-agent.sh entries, as the client setup docs show. For OpenMetadata, its docs give the /mcp endpoint and recommend OAuth through your OpenMetadata sign-in, with a personal access token or, for unattended agents, a bot token as alternatives.

// Example: .mcp.json for Claude Code
{
  "mcpServers": {
    "openmetadata": {
      "type": "http",
      "url": "https://<your-openmetadata-server>/mcp"
    },
    "dw-context-catalog": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-context-catalog"]
    },
    "dw-schema": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-schema"]
    },
    "dw-quality": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-quality"]
    },
    "dw-incidents": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-incidents"]
    }
  }
}

List the tools with /mcp. Then "why did the order_count test fail on fct_daily_orders?" gets the test, owner and downstream dashboards from OpenMetadata's get_entity_details and get_entity_lineage, the upstream change from the migration pull request, the cause from get_root_cause, the reach from blast_radius_analysis and the table's state from run_quality_check, in one answer.

One request, L0 to L4. The autonomy ladder is set per domain.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. Connected, not acting. Your engineer works the incident by hand.
  • •L1 observe. Each failed test arrives explained with cause, owner and blast radius; nothing changes. The permission ladder logs what each action would have needed.
  • •L2 propose. Data Workers drafts the dbt diff or rebuild plan; the owner approves in Spellbook before anything runs.
  • •L3 act reversibly. For proven classes, such as rerunning a failed load task after an upstream fix, Data Workers acts, verifies and keeps the undo path ready.
  • •L4 autonomous. For a scoped, trusted class, such as late loads into one domain, Data Workers fixes, verifies and posts the receipt for review.

More on how approvals work, autonomy levels L0 to L4, whether it is safe to let AI agents change production data and where your data goes. The agents run in your infrastructure: your data stays in your systems, and the hosted Conductor sees workflow metadata only.

What changes for your team

Six jobs that run on autopilot with Data Workers next to OpenMetadata, with a concrete example of each

OpenMetadata teams spend much of the week between a red test and a fix: finding the upstream change, the dbt model behind it and what the app team shipped, then rerunning the load. With Data Workers on top, those jobs run on autopilot at the level you set.

  • •Incidents. A failed OpenMetadata test arrives traced to its cause, with the fix drafted for the owner before the assignee opens the task.
  • •Data quality. Failing test cases turn into fixes and rebuilds the owner signs off, plus a dbt test so the same break fails in CI next time.
  • •Cloud spend. Snowflake spend is traced to the dbt model behind it and BigQuery spend is read from the Jobs API, with changes drafted for their owners.
  • •Access. When someone requests a table they found in OpenMetadata, Data Workers dry-runs the grant: effective privileges, the PII, PCI or PHI columns it would reach, policy conflicts and a least-privilege recommendation with an expiry. On Databricks it applies the approved Unity Catalog grant; elsewhere the grant is proposed for the owner to apply.
  • •Audits. Every change carries the table, the owner, the diff, the checks and a rollback path; see who owns the agents.
  • •Migrations. A legacy warehouse move runs in planned waves, with parity checks tracked and the completion gate held for the owner's sign-off.

The platform team gets its week back for new connectors, better tests and a catalog people trust.

Keep OpenMetadata, or consolidate?

Keep OpenMetadata if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

Most teams keep it. Your entities, glossary, tests and Context Center articles live in OpenMetadata, it is Apache 2.0 like the Data Workers core, and your AI tools already read it over MCP. What teams consolidate is the tooling around it: a separate observability tool, a script that turns failed tests into tickets, a runbook wiki mapping tables to dbt models and owners. If you are weighing building this layer yourself on OpenMetadata's MCP server, read build it ourselves with Claude Code and MCP servers: the connection is the easy part; cross-system context, approvals and rollback are the work. For the wider category, see what is an agentic data platform and Data Workers vs data observability.

The case for your CFO

The outcome: the open-source catalog your team already runs starts paying back in correct numbers. Every test case and lineage edge in OpenMetadata sits next to the context Data Workers uses to fix breaks before a staffing plan, a forecast or a finance export goes out wrong.

The risk story is plain. Data Workers writes nothing to OpenMetadata. At L1 the agents observe and change nothing. At L2 every change is proposed with its blast radius and goes to a named owner. At L3 only proven, reversible change classes run without a fresh approval. Every action leaves a receipt: trigger, diff, blast radius, checks, approver and undo path. Unanswered requests expire and escalate; they never auto-grant. Zero migration: OpenMetadata, PostgreSQL, BigQuery, dbt, Airflow and Superset stay put.

Why now: OpenMetadata 2.0 turned its MCP server on by default, so agents already read your catalog; the same context can drive approved fixes. The first win is one domain of tested tables, read-only, where every failed test arrives explained, owned and scoped. What stays the same: your catalog, tests, incident process and permissions. For the numbers, see the ROI of agentic data operations. Start with a pilot (pricing); the pilot is credited in full against the first year.

The sentence to repeat upstairs: "OpenMetadata tells us what broke; Data Workers fixes it, with the owner approving every change and a receipt for each one."

Getting started

Start with a pilot. Pick one domain where OpenMetadata's test cases guard tables people rely on, such as orders, connect Data Workers read-only to the PostgreSQL, BigQuery, dbt and Airflow behind it, and let it explain every failed test before you turn on the first fix class. Plans are on the pricing page, and the pilot is credited in full against the first year.

FAQ

Does Data Workers write to OpenMetadata? No. OpenMetadata connects over its API or MCP server today, and Data Workers changes nothing in it. When a fix should change a description, Data Workers drafts it for the owner, and the dbt docs change reaches OpenMetadata on its next dbt ingestion.

OpenMetadata already has `root_cause_analysis`. What does Data Workers add? root_cause_analysis traces a failed test back along quality lineage inside the catalog. Data Workers carries the investigation into the systems behind it, such as the PostgreSQL type change and the dbt model, then proposes the fix, queues the rebuild after approval and runs the checks.

We use Collate AI. Does it overlap with Data Workers? They meet at the catalog. Collate's Documentation, Tier and Quality agents improve the metadata and tests inside Collate; your team can accept or deny the Documentation Agent's suggestions. Data Workers works on the dbt models, DAGs and warehouse tables those tests watch.

Does Data Workers resolve incidents in OpenMetadata's Incident Manager? No. Your assignee keeps acknowledging and resolving them. Data Workers brings the cause, the approved fix and the checked result, with the receipt linked in the resolution comment.

How does Data Workers decide who approves a change? Each domain has named owners in Data Workers, usually the people OpenMetadata lists as owners. Requests reach them in Slack or email; they decide in Spellbook.

What does Data Workers store, and where? The agents run in your infrastructure and keep the context graph, receipts and audit log there, not copies of your tables. The hosted Conductor sees workflow metadata only.

Sources

  • •OpenMetadata, homepage ("The #1 open context layer for humans, AI assistants, and agents"), https://open-metadata.org/ (checked Oct 3, 2026)
  • •OpenMetadata, GitHub releases (2.0.0 Aug 24, 2026; 2.0.3 Sep 30, 2026), https://github.com/open-metadata/OpenMetadata/releases (checked Oct 3, 2026)
  • •OpenMetadata, MCP server setup and authentication, https://docs.open-metadata.org/latest/how-to-guides/mcp (checked Oct 3, 2026)
  • •OpenMetadata, MCP tools reference, https://docs.open-metadata.org/latest/how-to-guides/mcp/reference (checked Oct 3, 2026)
  • •OpenMetadata, Connect Claude Code, https://docs.open-metadata.org/latest/how-to-guides/mcp/claude-code (checked Oct 3, 2026)
  • •OpenMetadata, Incident Manager, https://docs.open-metadata.org/latest/how-to-guides/data-quality-observability/incident-manager (checked Oct 3, 2026)
  • •OpenMetadata, data quality tests in YAML, https://docs.open-metadata.org/latest/how-to-guides/data-quality-observability/quality/tests-yaml (checked Oct 3, 2026)
  • •Collate, homepage (AI Governance Studio, Aug 25, 2026), https://www.getcollate.io/ (checked Oct 3, 2026)
  • •Collate, pricing, https://www.getcollate.io/pricing (checked Oct 3, 2026)
  • •Collate, Collate AI, https://docs.getcollate.io/collateai (checked Oct 3, 2026)
  • •Collate, Documentation Agent, https://docs.getcollate.io/collateai/documentation-agent (checked Oct 3, 2026)
  • •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)
  • •Data Workers open-source repository, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)