You're on Google Knowledge Catalog: Turn the Context You Curate Into Fixes Your Stewards Approve
Already curating Knowledge Catalog? Data Workers takes in the glossary, aspects and lookupContext bundles your team's assistant hands it and acts on what they reveal, with the steward on the term approving.
Your team has done the curation work. Knowledge Catalog, called Dataplex Universal Catalog until April 10, 2026, holds entries for every BigQuery dataset, Cloud Storage bucket and Pub/Sub topic you run. Each finance glossary term has a steward in its contacts and entry links to the columns it defines. Aspects carry overviews, guidelines and the Gemini data insights you chose to publish. Auto data quality scans run set checks, null checks and SQL assertions on the tables that matter and publish a data-quality-scorecard aspect. And since lookupContext arrived in preview in June, agents call it through the Knowledge Catalog MCP server and get all of it in one bundle: schema, joins, terms, descriptions, profile and quality results, filtered by IAM. Knowledge Catalog is where your Google Cloud meaning lives. Data Context Wizard is where every agent reads it, next to lineage, quality and usage from every other system, and Data Workers acts on what that context reveals, with the steward named on the term approving.
That last part is what this guide covers. A failed set check at 3 a.m. is a perfect signal; someone still has to find the cause, change the code, rerun the job and prove the number before the revenue review.
Key takeaways
- •Knowledge Catalog keeps its job. The glossary, aspects, quality scans, Gemini data insights, data products and IAM stay as they are.
- •Your curation becomes first-class context. Terms, stewards, scorecards and
lookupContextbundles come into Data Context Wizard with their source, joined to lineage across Postgres, Kafka, dbt, Dagster and Looker. - •The context leads to action. When a scan fails or a term and a model disagree, Data Workers traces the cause, proposes the fix with its blast radius and routes it to the steward on the term.
- •Writes stay where they belong. Your team's assistant reads Knowledge Catalog through its read-only discovery MCP tools, side by side with Data Workers. Fixes land in dbt and your scheduler; stewards keep editing their own terms.
- •Start with a pilot. One glossary domain and its linked tables, read-only, then one fix class, on the ladder from L0 manual to L4 autonomous.
Knowledge Catalog is where your meaning lives. Data Workers acts on it.
Knowledge Catalog is very good at curating and serving context on Google Cloud. A glossary holds up to 5,000 terms, each with a description, an overview, contacts, synonyms, related terms and links to columns. Auto data quality runs row-level rules (range, null, set, regex, uniqueness), aggregate checks and custom SQL, and scores at job, column and dimension level. lookupContext packs up to ten entries into a YAML, XML or JSON bundle within a context_budget character limit.
Around that layer sits everything that decides whether the data is right: the application database, the CDC stream, the dbt models, the scheduler and the dashboards. Data Workers covers that side with one context, one approval flow and one audit trail, and uses your curation as its map.
Here is a Thursday on a BigQuery-centred revenue team. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| 14:05 | Postgres | The orders service ships a release that adds the status partially_refunded |
| 14:06 | Debezium and Kafka | Debezium streams the change events; the BigQuery sink lands them in raw.orders_cdc |
| 02:00 | dbt and Dagster | Dagster runs the nightly dbt build. fct_bookings maps statuses with a CASE that has no branch for the new value, so 1,184 orders fall to other and out of gross bookings |
| 03:10 | Knowledge Catalog | The quality scan's set check on fct_bookings.status_group fails and the scorecard aspect is published; the steward gets the email |
| 07:40 | Claude Code | An analyst asks why gross bookings fell 3.8%. The agent calls lookup_context and gets the glossary term "Gross bookings" (includes partially refunded orders at the gross amount), its steward and the failed check |
| 07:42 | Data Workers | It traces fct_bookings back through the CDC table to the Postgres column, finds the new status value, and lists the blast radius: three dbt models, two Looker Explores and the "Revenue" data product |
| 07:55 | GitHub and dbt | Data Workers proposes a diff that maps partially_refunded into booked revenue as the term defines it, with a test for unmapped statuses |
| 08:30 | Spellbook | The steward named on the term reviews the diff, the term and the blast radius, and approves; the model owner merges it and dbt CI passes |
| 08:50 | Dagster | The affected partition reruns; Data Workers checks the order count and gross bookings against Postgres and writes the receipt |
| 09:05 | Knowledge Catalog | The steward reruns the quality scan on demand and it passes |
| 10:00 | Looker | The revenue review opens on $2.51M, the right number |

Knowledge Catalog did its part exactly: the scan caught the break and the glossary said what the number should mean. The fix lived in a dbt model and a Dagster job, approved by the steward first. Without the trace, the review would have opened about $95,000 short.
| Job | What Knowledge Catalog does | What Data Workers does |
|---|---|---|
| The meaning | Holds glossary terms, stewards, overviews and entry links to columns | Brings terms in as sourced context and checks models and metrics against them |
| The signal | Runs quality scans and publishes the scorecard aspect, with email and logging alerts | Picks up the failure with the term and the steward attached and opens an incident |
| The bundle | Serves schema, joins, terms, profile and quality results through lookupContext, filtered by IAM | Adds what sits outside Google's catalog: upstream CDC, dbt code, scheduler runs and BI usage |
| The cause | Shows lineage for BigQuery jobs | Traces across Postgres, Debezium, Kafka, dbt and Dagster to the change that broke it |
| The fix | Records what the steward edits | Proposes the code change with its blast radius and routes it to the steward on the term |
| The proof | Rescans and updates the scorecard | Verifies counts after the rerun and keeps a receipt with the cause, diff, approver and rollback |
Why doesn't Knowledge Catalog just do this itself?
Because Knowledge Catalog is built to describe, check and serve Google Cloud metadata at scale, under IAM, for every team and agent in the organization. A quality scan that reports and alerts is the right design for that job: fast, never changing data, safe across thousands of tables.
The cause of most failures sits in systems Google doesn't run. In the incident above it was a Postgres release, a Debezium stream and a dbt CASE statement. Changing those means knowing the code, scoping the blast radius, getting the right approval, rerunning the job, proving the result and owning the rollback. Google draws its own lines sensibly here: the discovery MCP tools are read-only, and governance workflows for metadata change requests, in private preview, cover catalog metadata. Writing to production pipelines across vendors is a different product with a different liability. That product is Data Workers.
Every tool owns a slice. Data Workers covers the whole lifecycle
Knowledge Catalog owns one slice of the lifecycle outright on Google Cloud: curating and serving context, and governing who sees it. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the catalog you already curate. We use the same Knowledge Catalog scores as our Knowledge Catalog context engine comparison, with the reasons written for teams building on it.

| Stage | Data Workers | Knowledge Catalog | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 9.5 | Knowledge Catalog's home stage: glossaries, aspects, Gemini data insights and lookupContext give agents rich, IAM-filtered context on Google Cloud data. Data Workers brings that curation in as a first-class source and joins it to the rest of the estate. |
| Analytics & Insights | 8 | 6 | Gemini data insights drafts descriptions, sample queries and relationship graphs (GA on BigQuery). Data Workers answers questions across platforms from approved definitions. |
| Data Quality | 8 | 7.5 | Auto data quality runs row, aggregate and SQL rules on BigQuery and Iceberg tables and publishes a scorecard aspect. Data Workers runs checks on BigQuery, Snowflake and Postgres and drafts the fix for what fails. |
| Observability & Incidents | 8.5 | 4 | Quality alerts by email and Cloud Logging, plus lineage for impact. Data Workers detects, diagnoses, fixes and verifies across systems. |
| Pipelines & Ingestion | 8.5 | 2 | Knowledge Catalog gathers metadata, not data. Data Workers repairs and reruns the pipelines behind a failing table, behind an approval. |
| Schema & Migration | 8 | 3 | Discovery picks up schema as it changes, with no migration tooling. Data Workers traces upstream schema changes and plans migrations in approved waves. |
| Governance & Access | 8.5 | 9 | A home stage: IAM on every entry, Catalog Admin, Editor and Viewer roles, data products (GA) and governance workflows for data product access (preview). Data Workers routes approvals for definitions and changes, with enforcement left in IAM. |
| Security & Privacy | 8 | 6 | IAM-filtered lookupContext and ACL-aware search on Google Cloud. Data Workers runs least-privilege connectors and tenant-scoped tools across every platform. |
| Cost / FinOps | 8 | 1 | No cost features; Gemini features bill under Gemini in BigQuery. Data Workers works from query history and spend to remove the cause. |
| MLOps & Models | 7.5 | 4 | Ingests Vertex AI models, datasets and feature groups as entries. Data Workers keeps the data under the models healthy. |
How Knowledge Catalog and Data Workers work together
Your engineers and analysts stay where they are: Claude Code or Gemini CLI for SQL and dbt, Gemini Enterprise for plain-language questions. Spellbook Data Catalog (in preview) is where the data team looks: each proposed change, the term and steward it touches, its blast radius and its rollback. Between them, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm does the work with more than 20 specialist agents, and the Autonomous Data-Conductor runs each fix end to end (detect, diagnose, fix, review, verify, remember) under per-domain guardrails.

What Context Wizard does with your curation. Three things come in, each with its source and time.
- •Glossary terms. A term's definition, steward and linked columns become a governed business rule owned by that steward. When a dbt model or metric disagrees, the conflict goes to the steward. The walkthrough is our guide to bringing glossary terms and aspects into Data Context Wizard.
- •Aspects. Overviews, guidelines and published Gemini insights arrive as proposals at the "derived" level. A named person promotes them; no agent can (see the context engine comparison).
- •lookupContext bundles. Your agent hands Data Workers the bundle it used, quality results included, over MCP, and it sits next to Data Workers' own view of the table:
explain_tablefor definition, lineage and trust score,trace_cross_platform_lineagebeyond BigQuery, andget_incident_historyfor open incidents.
Data Workers also imports and exports Open Knowledge Format (OKF) bundles, the Apache 2.0 markdown format (spec v0.2) published in Google Cloud's GitHub organization, so curated knowledge moves between catalogs and git without being retyped.
Setup over MCP today. Add Google's discovery MCP endpoint and Data Workers' agents to the same client. For Data Workers, clone the open-source repository and add start-agent.sh entries, as the client setup docs show. Google's server authenticates with OAuth 2.0 and IAM: grant the MCP Tool User role (roles/mcp.toolUser) and the Dataplex catalog role Google's remote MCP page lists, and use the dataplex.readonly scope, since search_entries, lookup_context and lookup_entry only read.
// Example: .mcp.json for Claude Code
{
"mcpServers": {
"knowledge-catalog": {
"type": "http",
"url": "https://dataplex.googleapis.com/mcp",
"headers": { "Authorization": "Bearer ${KC_ACCESS_TOKEN}" }
},
"dw-context-catalog": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-context-catalog"]
},
"dw-schema": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-schema"]
},
"dw-quality": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-quality"]
}
}
}Set KC_ACCESS_TOKEN from gcloud auth print-access-token for a session, or use a service account for anything longer. The discovery endpoint is all this pattern needs. List the tools with your client's own command (/mcp in Claude Code). Then "why did the gross bookings check fail?" gets the term and scorecard from lookup_context, lineage and the upstream model change from trace_cross_platform_lineage and the dbt manifest, the fix's reach from blast_radius_analysis and the table's state from run_quality_check, in one answer.
Where writes go. Fixes land where your team already reviews change: a dbt diff for the owner to merge, a Dagster rerun, a sink connector setting drafted for its owner. Knowledge Catalog stays the system of record for its terms and aspects, by design: Data Workers drafts a clearer definition, the approved fact lands in the Context Wizard graph, and the steward makes the edit in Knowledge Catalog (or your team's MCP client does, where Google's MCP server allows it).
One request, L0 to L4. The autonomy ladder is set per domain.

- •L0 manual. Data Workers is connected but not acting. Your engineer chases the scan by hand.
- •L1 observe. Data Workers explains each failure with its cause and steward, and changes nothing.
- •L2 propose. Data Workers drafts the dbt diff or the rerun plan with its blast radius. The steward on the term approves in Spellbook before anything runs.
- •L3 act reversibly. For proven change classes, such as rerunning a failed partition after an upstream fix, Data Workers applies, verifies and can roll back.
- •L4 autonomous. For a scoped, trusted class like late CDC loads into one dataset, Data Workers fixes and verifies on its own and posts the receipt for review.
The safety model behind each step is in is it safe to let AI agents change production data; where data and credentials live is in where does our data go. For the rest of the Google Cloud estate, including BigQuery agents and Managed Service for Apache Airflow (formerly Cloud Composer), read Data Workers on Google Cloud.
The same pattern holds for every source of meaning. See the hub, bring your own context, and the guides for teams on Databricks Genie Ontology, Snowflake Horizon Context and the dbt Semantic Layer. For upstream schema changes, see BigQuery schema change impact on downstream dashboards.
What changes for your team

Governance teams on Knowledge Catalog spend much of the week turning signals into tickets: forwarding a scan email, finding who owns the dbt model, asking the steward whether a new status counts. With Data Workers on top, those jobs run on autopilot at the level you set.
- •Incidents. A failed scan arrives traced to its cause, with the term, steward and proposed fix attached.
- •Data quality. Scorecard failures become fixes and reruns that the next scan confirms.
- •Cloud spend. BigQuery spend is read from the Jobs API, and each fix goes to its owner drafted.
- •Access. A request for a table behind a glossary term reaches that term's steward, and IAM stays the enforcement point.
- •Audits. Every change carries the term, the approving steward, the diff, the verification and a rollback path.
- •Migrations. A move off a legacy warehouse into BigQuery is planned in waves, each with its parity checks, and the stewards get a drafted mapping to relink their glossary terms to the new tables.
Stewards get their week back for curation only they can do.
Keep Knowledge Catalog, or consolidate?
Keep Knowledge Catalog if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
For most Google Cloud teams the answer is to keep it. Knowledge Catalog is built into how BigQuery, IAM and Gemini see your data, and your stewards' curation belongs there. What teams consolidate is the tooling around it: a separate observability tool, a script that turns scan results into tickets, a spreadsheet mapping glossary terms to dbt models. Data Workers runs those jobs with one context, one approval flow and one audit trail. If you are weighing building this layer yourself on the MCP server, read build it ourselves with Claude Code and MCP servers: the connection is the easy part; cross-system context, approvals and rollback are the work.
The case for your CFO
The outcome: the catalog the company already invested in pays back in correct numbers. Terms, stewards and quality scans become the map Data Workers uses to fix breaks before a revenue review, a board deck or a customer report.
The risk story is plain. Agents read Knowledge Catalog through read-only discovery tools under IAM. Every change Data Workers proposes shows its blast radius, goes to the steward named on the term, lands through your existing dbt and scheduler workflows, is verified and leaves a receipt: who approved it, what it touched and how to undo it. Autonomy is set per domain from L0 manual to L4 autonomous and can be dialled back any time. There is zero migration: Knowledge Catalog, BigQuery, dbt and Dagster stay where they are.
Why now: since June, lookupContext hands your curated context straight to agents, so the same curation can drive fixes as well as answers. The first win is one glossary domain, read-only, where every failed scan arrives explained and owned. What stays the same: your glossary, stewards, IAM, Gemini settings and review process. For the numbers, see the ROI of agentic data operations. Start with a pilot (pricing); the pilot is credited in full against the first year.
The sentence to repeat upstairs: "We already curate Knowledge Catalog; Data Workers turns that curation into fixes, approved by the steward named on the term, with a receipt."
Getting started
Start with a pilot. Pick one glossary domain your agents lean on, such as finance, connect the discovery MCP server read-only next to Data Workers, and let Data Workers explain every quality failure on the linked tables before you turn on the first fix class. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.
FAQ
Does Data Workers write to Knowledge Catalog? No. Agents use the read-only discovery tools (search_entries, lookup_context, lookup_entry) and hand what they find to Data Workers over MCP. Stewards keep editing their own terms and aspects; when a fix suggests a clearer definition, Data Workers drafts it for the steward.
Do we need the data products MCP server? Not for this pattern. Its create and update tools manage data products and assets, which stays with your data product owners. The read-only discovery endpoint covers everything this page describes.
How does Data Workers know who to ask for approval? From your curation. The steward in a glossary term's contacts becomes the owner of that definition in Data Context Wizard, and changes that touch the term's columns route to that person in Spellbook. Terms don't inherit contacts from their categories in Knowledge Catalog, so set them on each term you want routed.
Will Data Workers replace our auto data quality scans? No. Keep the scans and the scorecard aspect; they are the signal. Data Workers adds the trace, fix, rerun and receipt, and runs its own checks outside Google Cloud with run_quality_check.
Can our glossary leave Google Cloud? Yes. Data Context Wizard imports and exports Open Knowledge Format bundles, the open markdown format from Google Cloud's GitHub organization, carrying each term's source and owner in its frontmatter, so Snowflake, Databricks and dbt work can use the same definitions.
What does Data Workers store? Metadata and scrubbed facts about your data, such as definitions, lineage, owners and incident history, not copies of your tables.
Sources
- •Google Cloud, Retrieve data context with lookupContext (updated Sept 30, 2026), https://docs.cloud.google.com/knowledge-catalog/docs/retrieve-data-context (checked Oct 2, 2026)
- •Google Cloud, Knowledge Catalog MCP reference (updated Sept 21, 2026), https://docs.cloud.google.com/dataplex/docs/reference/mcp (checked Oct 2, 2026)
- •Google Cloud, Use the Knowledge Catalog remote MCP server (OAuth 2.0, roles, scopes), https://docs.cloud.google.com/dataplex/docs/use-remote-mcp (checked Oct 2, 2026)
- •Google Cloud, Create a business glossary, https://docs.cloud.google.com/knowledge-catalog/docs/create-glossary (checked Oct 2, 2026)
- •Google Cloud, Auto data quality overview (updated Sept 30, 2026), https://docs.cloud.google.com/dataplex/docs/auto-data-quality-overview (checked Oct 2, 2026)
- •Google Cloud, Data insights for structured data, https://docs.cloud.google.com/knowledge-catalog/docs/use-data-insights-structured-data (checked Oct 2, 2026)
- •Google Cloud, Knowledge Catalog release notes (rename Apr 10, MCP servers preview May 27, lookupContext preview Jun 4, governance workflows for data product access Jul 24, dbt import GA Sept 24, 2026), https://docs.cloud.google.com/dataplex/docs/release-notes (checked Oct 2, 2026)
- •GoogleCloudPlatform, Open Knowledge Format (spec v0.2, Apache 2.0), https://github.com/GoogleCloudPlatform/open-knowledge-format (checked Oct 2, 2026)
- •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)
- •Data Workers open-source repository, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)