Knowledge Catalog Business Glossary and Aspects in Data Context Wizard: A How-To
Connect your Knowledge Catalog business glossary, contacts and overview aspects to Data Context Wizard over MCP, so every agent uses governed, sourced definitions.
Your stewards already write the meaning of your business in Knowledge Catalog. The finance glossary says what "net revenue" and "active customer" mean, the contacts aspect says who owns each term, and entry links tie each term to the BigQuery tables and columns that carry it. Knowledge Catalog holds what your business terms mean on Google Cloud. Data Context Wizard turns those terms into governed definitions, with their source attached, that every agent uses on every platform, and it flags the places where dbt disagrees so the term's owner can settle it on the record.
This guide shows how to connect the two over MCP today: what moves in each direction, the roles you need, a setup example, and one run end to end with the receipt it leaves. Knowledge Catalog is the new name for Dataplex Universal Catalog (renamed on April 10, 2026). The API, CLI and IAM names still say dataplex, so you'll see both names below.
Key takeaways
- •The glossary stays where it is. Stewards keep writing terms, categories and contacts in Knowledge Catalog, under its IAM roles.
- •Your coding agent is the bridge. It reads glossary terms through the Knowledge Catalog API and entries through Google's remote MCP server, then hands them to Data Context Wizard over MCP.
- •Every term lands as a proposal with its source. It arrives at authority "derived", pending review, carrying the term's resource name and update time. Only a named human can make it authoritative.
- •Conflicts go to the term's owner. When dbt's MetricFlow metric and the glossary disagree, the steward named in the
contactsaspect decides in Spellbook's Approvals inbox. - •The approved fact goes back as a docs change. Data Workers proposes the dbt description update; the steward edits the glossary term in Knowledge Catalog.
- •Start read-only with one glossary and the viewer role.
What this connects
Three parts of Knowledge Catalog carry meaning, and all three are readable today.
- •The business glossary. A glossary holds categories (nested up to three levels) and up to 5,000 terms. Each term has a description, an overview, contacts, and links to synonyms, related terms and data assets. You list terms with
glossaries.terms.listin the REST API. - •Aspects. Aspects are typed metadata attached to an entry or to one of its columns. The system
overview,contactsandschemaaspect types are the ones that matter here: what a table says it is, who owns it, and what each column holds. You read them withlookupEntry, choosing a view and, optionally, the aspect types you want. - •Entry links. A glossary term attaches to a table or column through an entry link of type
definition(orsynonymbetween terms). This is how you know that the "Net revenue" term governsfct_orders.net_amount.
On the Data Workers side, the receiving end is the Data Context & Catalog agent, which runs Data Context Wizard. It is an MCP server, so the coding agent your team already uses (Claude Code, Codex or Cursor) can call Knowledge Catalog and Context Wizard in the same session. The Google Cloud integration guide covers the rest of the estate. Glossary terms and aspects come through the coding agent and Google's own API and MCP server, which is what this guide sets up.
| From Knowledge Catalog and dbt into Data Workers | From Data Workers back to your team |
|---|---|
| Glossary terms: display name, description, overview, synonyms | A governed definition per term, sourced to the glossary, served to every agent over MCP |
The contacts aspect: the steward who owns a term | A conflict in Spellbook's Approvals inbox for that steward when dbt and the glossary disagree |
| Entry links: which tables and columns each term defines | A receipt for each decision: proposer, both sources, approver, time |
overview and schema aspects on BigQuery entries | A dbt docs change for the approved wording, in schema YAML, after approval |
| dbt MetricFlow metrics: expression, grain, owner | A suggested glossary edit the steward makes in Knowledge Catalog |

Prerequisites
- •Knowledge Catalog read access. Grant the identity your coding agent uses
roles/dataplex.catalogVieweron the project that holds the glossary. That covers listing glossary terms and looking up entries. Editing terms needsroles/dataplex.catalogEditor, and that role stays with your stewards. - •MCP access. Google's remote MCP server needs
roles/mcp.toolUserfor tool calls, plus the Dataplex API enabled. It uses Streamable HTTP with OAuth; thedataplex.readonlyscope is enough for this guide. Google's setup page names Catalog Admin for full access; the read-only path here doesn't need it. - •Your dbt project. Context Wizard reads MetricFlow metrics from your dbt manifest and imports them in bulk. The dbt integration guide covers the token and artifacts.
- •Spellbook Data Catalog for approvals. It's in preview, and it's where stewards approve or send back each proposal.
- •A named steward per glossary. The
contactsaspect is how the run below finds the owner, so fill it in for the terms you start with.
Setup
Register both MCP servers with your coding agent, then give it one prompt that says how to read the glossary. Data Workers' agents come from the open-source repository: clone it, run npm install, set DW_HOME to the clone and add start-agent.sh entries, as the client setup docs show. The configuration below is for Claude Code; Codex and Cursor take the same two entries in their own config files.
Example
{
"mcpServers": {
"knowledge-catalog": {
"type": "http",
"url": "https://dataplex.googleapis.com/mcp",
"headers": { "Authorization": "Bearer ${KC_ACCESS_TOKEN}" }
},
"dw-catalog": {
"command": "${DW_HOME}/start-agent.sh",
"args": ["dw-context-catalog"]
}
}
}# Example: refresh the token, then list the finance glossary's terms
export KC_ACCESS_TOKEN="$(gcloud auth print-access-token)"
curl -s -H "Authorization: Bearer ${KC_ACCESS_TOKEN}" \
"https://dataplex.googleapis.com/v1/projects/acme-data/locations/us/glossaries/finance/terms?pageSize=200"The project, location and glossary ids are placeholders. Access tokens from gcloud expire, so refresh before a long session or use a service account. Google's discovery MCP endpoint offers search_entries, lookup_context and lookup_entry; glossary terms come through the REST call above.
Then tell the agent the rules once, for example in your project's agent instructions: read terms from the finance glossary; for each term, read the steward from its contacts and the tables it defines through entry links; hand each term to Context Wizard with define_business_rule, putting the term's resource name and update time in the rule text; never call a write tool in Knowledge Catalog.
One run, end to end
This is an illustration, not a customer case. It's the shape of the first conflict most teams find.
At 09:02 an analyst asks Claude Code for Q3 net revenue by region. The finance glossary's "Net revenue" term, last edited in March, says revenue "net of discounts, before refunds". In August the finance team changed the policy, and the dbt MetricFlow metric net_revenue now subtracts refunds. Both are in production. The contacts aspect names Priya, the finance data steward, as the term's owner.
| Time | System | What happens | Who decides |
|---|---|---|---|
| 09:03 | Knowledge Catalog | The agent lists the term, reads its contacts, and follows the entry link to fct_orders.net_amount with lookupEntry | Read-only |
| 09:04 | dbt | Context Wizard reads the net_revenue metric from the dbt manifest | Read-only |
| 09:05 | Context Wizard | resolve_metric for "net revenue" returns the dbt definition, which subtracts refunds; the agent compares it with the glossary wording and flags the mismatch | Read-only |
| 09:06 | Context Wizard | The agent records the glossary term with define_business_rule, naming the conflict and its source, at a confidence below the certified definition so it stays pending the steward's review | Data Workers proposes |
| 09:07 | Spellbook | The proposal waits in Approvals for Priya, with both definitions side by side; the agent gives the analyst the dbt number and says the definition is under review | Data Workers |
| 11:40 | Spellbook | Priya approves the dbt logic as the current policy and adds a note on the March wording | A named steward |
| 11:41 | Context Wizard | The definition becomes authoritative and the receipt is recorded | Approved at 11:40 |
| 11:45 | GitHub | Data Workers proposes the dbt docs change for net_revenue with propose_dbt_doc_writeback | A reviewer merges |
| 14:10 | Knowledge Catalog | Priya updates the glossary term's description herself | The steward |

The receipt is the record a reviewer or auditor reads later. It holds the proposal (the glossary term's resource name and its March update time, and the dbt metric with its manifest version), who proposed it and when, Priya as the approver with her note, the authority change from "derived" to authoritative, and the follow-ups: the docs proposal for dbt and the glossary edit. Any agent can fetch it later with get_change_receipt.
Two things never happened in this run. No agent wrote to Knowledge Catalog, and no agent promoted its own proposal. Context Wizard rejects an agent's attempt to set a definition authoritative without a named human approver (promotion_requires_human_approval), and the dbt docs change waits in the same Approvals queue before it's written. It opens as a pull request when you turn on the GitHub target; otherwise it writes to your local project.
Why doesn't Knowledge Catalog just do this itself?
Knowledge Catalog built a great product for one job: harvesting and governing metadata across Google Cloud, with a glossary, aspects, search and lineage in one place. Its design follows from that job. Its writes target its own terms, entries and aspects. Its dbt import (GA since September 24, 2026) brings dbt metadata into Knowledge Catalog, one way. Review of metadata change requests to glossaries and entries is in private preview. That's a sensible scope for a catalog.
Settling which definition is right across dbt, the warehouse and every agent that answers questions is a different product. It needs both sources side by side, the owner on record, approvals that cover a glossary term and a dbt metric in the same decision, rollback, receipts, and changes written into tools Google doesn't own. That's the product Data Workers is. If you're choosing a catalog control plane, Knowledge Catalog vs Spellbook Data Catalog covers that decision. This page is the wiring.
The next autonomy step

The run above sits at L1: Data Workers observes and proposes, and people decide everything. Once a domain's receipts look right, move up one step at a time.
- •L2 for docs on approved terms. Turn on the GitHub target so every docs change for a term the steward already approved opens as a pull request on its own. A reviewer still merges, and the steward still owns the glossary.
- •A standing check. Have the coding agent re-read the finance glossary on a schedule and submit only terms whose update time changed, so a new edit in Knowledge Catalog reaches every agent through review.
- •The next glossary. Add product or marketing terms once finance is clean. Keep model changes at "propose" while docs move faster; autonomy is set per domain.
The case for your CFO
The outcome. One definition per business term, used by every agent and every dashboard, with the owner on record. When the board deck says net revenue, everyone means the same thing.
The risk story. At the observe level, agents read the glossary, aspects and dbt and change nothing. At the propose level, every definition lands as a proposal; only a named human makes it authoritative, and no agent approves its own work. Each receipt holds the sources, the approver, the before and after, and the follow-ups. Agents get no write access to Knowledge Catalog, and nothing migrates.
Why now. Google ships an MCP server and lookupContext for agents, and dbt ships its own. Agents are already reading your definitions; you want them reading reviewed ones.
The first win. One finance glossary, read-only, and a list of every term where dbt and the glossary disagree, each routed to its owner.
What stays the same. Knowledge Catalog, its glossary and IAM; dbt and your reviewers; the coding agent your team already uses.
The pilot path. Start with a pilot on one glossary and one dbt project. The pilot is credited in full against the first year.
One sentence for upstairs: "Our stewards keep writing definitions in Knowledge Catalog; Data Workers makes every agent use them, catches where dbt disagrees, and gets the owner to settle it on the record."
Sources
Checked October 2, 2026: Knowledge Catalog business glossary (structure, limits, roles, private preview of metadata change requests), the glossary terms list API, aspects and entry links (updated April 10, 2026), the lookupEntry API, the remote MCP reference (updated September 21, 2026), using the remote MCP server (roles and scopes, updated September 30, 2026), the release notes (rename on April 10, 2026; dbt import GA on September 24, 2026) and the dbt metadata import. Product names and statuses change quickly; if we've got something wrong, tell us and we'll fix it.