Knowledge Catalog Context for AI Agents vs Data Workers: Gemini Drafts It, a Named Owner Approves It
Knowledge Catalog and Gemini draft context for agents on Google Cloud. Data Workers scores that context, gates it on a named owner and covers every other platform.
Google now calls Knowledge Catalog "the universal context engine for your enterprise", and on Google Cloud it earns the line. Entries for BigQuery, Spanner, Pub/Sub and Cloud Storage appear on their own. Gemini data insights drafts table and column descriptions, sample queries and relationship graphs, and publishes them as aspects. The lookupContext method hands an agent a ready bundle of schema, joins, descriptions, profile and quality results, and the discovery MCP server lets Claude Code, Cursor or Gemini CLI call it directly. If your team is wiring agents to that context, this page is for you.
Your estate probably looks like this. BigQuery holds product and marketing data, and Knowledge Catalog is how people and agents find it. Finance runs a Snowflake mart. The ML team trains churn models in Databricks on Unity Catalog. dbt builds the models, Airflow schedules them, and Looker sits on top. Gemini can describe the BigQuery half of that estate in an afternoon. Nobody has decided which of those descriptions and joins an agent should treat as the truth.
Knowledge Catalog is the context engine for Google Cloud: it gathers context and Gemini drafts it. Data Workers is the context authority for the whole estate: it records where each fact came from, scores it, and keeps it a proposal until a named owner approves it. Data Workers is the agentic data platform: it runs the whole data lifecycle, and its Data Context Wizard keeps one governed context graph across every platform. Gemini enrichment stays the fastest way to describe Google Cloud data. Data Workers decides what agents can trust, on Google Cloud and everywhere else, and runs the fixes when that context breaks.
A note on names: Google renamed Dataplex Universal Catalog to Knowledge Catalog on April 10, 2026. API, IAM and CLI names are unchanged, so the endpoints still read dataplex.googleapis.com.
Key takeaways
- •Knowledge Catalog leads on gathering context for Google Cloud. Auto-ingest and semantic search are GA, Gemini data insights is GA on BigQuery, and
lookupContextpacks up to ten assets into one agent-ready bundle. - •Gemini enrichment is a proposal until someone approves it. In Knowledge Catalog, publishing insights is a per-run choice and review of metadata changes is in private preview. Data Workers holds AI-written context as a proposal, scores it and needs a named owner to promote it.
- •Provenance and scoring on every fact. Google labels insights Agent or User. Data Workers records source, actor, confidence, authority and time on every fact, and scores trust on origin, usage, closeness to approved facts, freshness and correction history.
- •Context from every platform. Federation with Unity Catalog, Glue and Snowflake Horizon is preview. Data Workers reads Snowflake, dbt and Airflow today, and takes Knowledge Catalog entries and bundles from your coding agent over MCP.
- •Context that runs the fixes. The same graph drives the Data-Agents Swarm and the Autonomous Data-Conductor, with a receipt on every change.
- •Nothing migrates. Knowledge Catalog stays the inventory and Gemini keeps drafting.
This page is the context-for-agents view. For choosing a catalog and control plane, read our Dataplex alternative comparison, Knowledge Catalog vs Spellbook Data Catalog. For the wiring, read how to bring Knowledge Catalog glossary terms and aspects into Data Context Wizard. The Snowflake twin of this page is Snowflake Horizon Context vs Data Workers.
Six things you get with Data Workers on top of Knowledge Catalog
1. AI-written context stays a proposal until an owner approves it. Descriptions and joins from Gemini, from your coding agent or from our own agents enter the Context Wizard at the "derived" authority level, marked pending review. Promotion to authoritative needs a named human, and the check is enforced in code.
2. Every proposal is checked before anyone reviews it. A documentation change is validated against the live schema: every column it describes must exist, generated text can't overwrite a human-written description as authoritative, and downstream lineage impact is listed. Proposed joins are checked for cardinality, with fan-trap and chasm-trap warnings.
3. A trust score on every fact. Each fact gets a score from 0 to 1 built from five signals: where it came from, how often it's used, how close it sits to facts a human approved, how fresh it is and how often it has been corrected. When a fact drifts well below a competing definition of the same thing, the drift detector flags it.
4. Provenance an auditor can follow. Every fact records its source, the actor that wrote it (agent or human), confidence, authority tier and when it was observed and last confirmed, stored tenant-isolated.
5. One graph across every platform. The Context Wizard reads BigQuery next to Snowflake, dbt and Airflow, from a library of 50+ connectors, and imports MetricFlow, Databricks metric views, Cube and Wren MDL definitions. An agent asking about a BigQuery table also sees the Databricks model that reads it.
6. One review queue, with receipts. Proposed definitions and conflicts land in Spellbook Data Catalog, in preview today, next to every other change agents propose. Approving one leaves a receipt with the approver, the diff and the rollback path.
One join, five systems
Here is a Tuesday morning on a BigQuery-centred team. It's an illustration, not a customer case.
- •08:10, BigQuery. A product release starts writing one row per user per day to
app_events. Until today it held one row per account per day. - •08:40, Knowledge Catalog. An analyst regenerates data insights on
subscriptionsandapp_eventsand picks "Generate and publish". The published description still says one row per account per day, and the suggested join onaccount_idassumes it. Both carry the source label Agent. - •09:15, Knowledge Catalog data product. The "Customer 360" data product's Gemini sample query is regenerated on the same join.
- •10:05, coding agent. An engineer asks Claude Code for weekly active subscribers. The agent calls
lookup_contexton the Knowledge Catalog MCP server, gets the join and the description, and writes SQL that returns 212,400. The approved figure is 71,900. - •10:20, Databricks and Looker. The churn model in Databricks, built on account-day grain in Unity Catalog, disagrees with a Looker tile someone built from the agent's SQL.
| Where the context lives | What Knowledge Catalog sees | What Data Workers does |
|---|---|---|
BigQuery app_events | The schema change, ingested automatically | Reads the new schema through its BigQuery connector and records the change with its time |
| Gemini description and join | Published as aspects, source Agent, visible to everyone with access | The coding agent hands both to the Context Wizard as proposals over MCP; the schema check flags the grain and the join check flags a fan trap |
| "Customer 360" data product | A sample query built on the join | Lists the data product as a consumer of the disputed join |
| Coding agent's answer | lookup_context returns the bundle it was given | Returns the approved account-grain definition with its approver; the agent answers 71,900 |
| Databricks churn model | Readable through federation in preview | Reads the Unity Catalog definition, which matches the approved one and raises its trust score |
| The owner | No review step between publish and use | Gets one item in Spellbook: approve or reject the join, with sources side by side |

The owner rejects the join at 11:30, and a steward marks the correct join as User in Knowledge Catalog, so the next Gemini run won't override it. Gemini drafted fast and mostly well. The miss came from a schema change that landed after the description was written, and the answer reached a person before anyone looked. Data Workers put the check between the draft and the answer.
What Knowledge Catalog gives agents, as of October 2026
Google labels its parts one by one. Where the release notes and the doc page disagree, we use the more conservative status and say so.
| Capability | What it does for agents | Status (Oct 2026) |
|---|---|---|
lookupContext | One bundle for up to ten assets: technical metadata, joins from query logs and data insights, descriptions (source-system and auto-generated), terms, profile and quality results; YAML, XML or JSON, with a character budget; filtered by IAM | Launched in preview June 4, 2026; the doc page (updated September 30) shows no label |
| Discovery MCP server | search_entries, lookup_context and lookup_entry, all read-only | No preview label on the doc page |
| Data products MCP server | Reads plus create and update for data products, assets and their aspects | Preview, announced May 2026 |
| Lineage MCP server | Queries lineage graphs and impact | Preview since May 27, 2026 |
| Data insights | Gemini drafts descriptions, sample queries and relationship graphs; "Generate and publish" makes them searchable and visible to others | GA on BigQuery; Iceberg tables and the Relationships tab are preview |
| Source labels | Insights are labelled Agent or User; for joins, User beats table constraints, which beat query history, which beat Agent | Documented |
| Data products | Curated asset bundles with contracts and access-request approvals; Gemini drafts sample queries and documentation templates | GA since May 25, 2026 |
| Verified queries and semantic guardrails; automated context curation | Verified SQL patterns and pre-generated questions; continuous enrichment | Preview, announced April 22, 2026 |
| Metadata change review | Governance workflows for change requests to glossaries and entries | Private preview |
| Reach beyond Google Cloud | Federation with Unity Catalog, Glue and Snowflake Horizon; database connectors | Preview |
| dbt metadata import | dbt Core, Fusion and dbt Cloud runs into Knowledge Catalog, lineage for BigQuery targets | GA since September 24, 2026 |
lookupContext is the strongest piece of this. It solves a real agent problem, packing the right metadata into a context window, and it respects IAM, so an agent sees only what its caller can see.
One platform, not one more tool
Context is one stage of a data team's lifecycle. The same team writes quality checks, handles incidents, changes pipelines, runs migrations, manages access, protects sensitive data, cuts spend and keeps models fed. Each point tool adds another console, another contract and another handoff.
Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail. We score the same ten lifecycle stages on every comparison page. Here we've scored Knowledge Catalog with data insights, data quality, lineage and its MCP servers. Google leads on catalog and context, and on governance and access through data products and IAM.

| Stage | Data Workers | Knowledge Catalog | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 9.5 | Knowledge Catalog's home stage. Auto-ingest across Google Cloud, semantic search, Gemini data insights and lookupContext give agents rich context on Google Cloud data. Data Workers keeps one scored, approved graph across every platform. |
| Analytics & Insights | 8 | 6 | Data insights drafts descriptions, sample queries and relationship graphs (GA on BigQuery). Data Workers answers questions across platforms from approved context. |
| Data Quality | 8 | 7.5 | Auto data quality and profiling are GA for BigQuery and Lakehouse tables, and their results ride along in lookupContext. Data Workers runs checks on BigQuery, Snowflake and Postgres and holds other tables to baselines the team records. |
| Observability & Incidents | 8.5 | 4 | Quality scan alerts, change feeds to Pub/Sub and lineage impact, with no incident workflow. Data Workers detects, diagnoses, fixes and verifies. |
| Pipelines & Ingestion | 8.5 | 2 | Knowledge Catalog gathers metadata, not data. Data Workers proposes and repairs pipelines behind an approval. |
| Schema & Migration | 8 | 3 | Change feeds flag schema updates, and there's no migration tooling. Data Workers plans schema changes and migrations in approved waves. |
| Governance & Access | 8.5 | 9 | A home stage. Data products carry access-request approvals that grant IAM (GA); review of metadata changes is private preview. Data Workers adds human-only promotion of definitions, with enforcement left in IAM. |
| Security & Privacy | 8 | 6 | IAM-filtered lookupContext, ACL-aware search and IAM on Google Cloud. Data Workers runs least-privilege connectors and tenant-scoped tools everywhere. |
| Cost / FinOps | 8 | 1 | No cost features; Gemini enrichment bills under Gemini in BigQuery. Data Workers works from query history and spend to remove the cause. |
| MLOps & Models | 7.5 | 4 | Ingests Vertex AI models, datasets and feature groups. Data Workers keeps the data under the models healthy. |
Knowledge Catalog vs Data Workers on the context outcomes you buy
This view scores eight outcomes a data leader pays for when the goal is context agents can trust. Google leads on three, all on its own surface: context bundles, drafting at scale and search. Data Workers leads on every outcome that decides what's trusted or spans platforms.

| Outcome | Data Workers | Knowledge Catalog | Why we scored it this way |
|---|---|---|---|
| Context bundles for Google Cloud assets | 6 | 9 | Google leads on its own surface. lookupContext returns schema, joins, descriptions, profile and quality results for up to ten assets in one call, filtered by IAM. |
| Descriptions and joins drafted at scale | 6 | 9 | Gemini data insights drafts descriptions, sample queries and relationship graphs across BigQuery (GA). Data Workers drafts docs too and enters them as proposals; Gemini does this across Google Cloud out of the box. |
| Search across Google Cloud metadata | 7 | 9 | Semantic search is GA and ACL-aware. Data Workers takes the results your coding agent finds into its graph. |
| Where each fact came from | 9 | 6 | Google labels insights Agent or User and ranks join sources. Every Data Workers fact records its source, actor, confidence, authority and time. |
| Trust scored per fact | 9 | 4 | Google ranks joins by source type. Data Workers scores every fact on origin, usage, closeness to approved facts, freshness and correction history. |
| Nothing authoritative without a named approver | 9 | 3 | Publishing insights is a per-run choice, and metadata change review is private preview. Data Workers' authority guard needs a named person to promote a fact. |
| Context from Snowflake, Databricks, dbt and BI | 9 | 4 | Federation and database connectors are preview, and dbt import is one-way. Data Workers reads Snowflake, dbt, Airflow, Tableau and Looker today; other BI tools connect over their API or MCP server. |
| Context used to fix and verify | 9 | 2 | Knowledge Catalog informs agents. Data Workers uses the approved graph to fix across systems and writes the outcome back. |
The scores measure scope, not answer quality. They're directional judgments, not benchmarks, and the reasoning is shown so you can argue with any line.
Where Knowledge Catalog stops
Each limit below comes from Google's own pages. None is a flaw. Knowledge Catalog exists to make Google Cloud data easy to find and describe, and fast enrichment with light friction is the right design for that job.
Publishing is the review step. Data insights offers "Generate and publish" or "Generate without publishing". Once published, insights are "indexable, searchable, and visible to other users in your organization", and that includes agents calling lookup_context. Governance workflows for metadata change requests are in private preview.
Labels record origin, not approval. Google marks insights Agent or User and ranks join sources: User, then table constraints, then query history, then Agent. That's useful provenance. It says who wrote a fact, not whether its owner signed off, and an Agent-sourced query "may be replaced during a regeneration".
The bundle mixes sources. lookupContext returns descriptions "captured in the source system and auto-generated in Knowledge Catalog" side by side. An agent reading the bundle has to work out which is which and what changed since it was written.
Google Cloud first. Federation with Unity Catalog, Glue and Snowflake Horizon is preview, the database connectors are preview, and dbt import brings metadata in one way.
Gemini has its own terms. Google notes that Gemini in BigQuery "doesn't support the same compliance and security offerings as BigQuery", and Gemini features bill under Gemini in BigQuery or Gemini Code Assist. Check both before you enrich regulated data.
Context informs. It doesn't fix. Nothing in the context stack owns a broken pipeline, checks that a fix held, or records the outcome for next time.

Why doesn't Knowledge Catalog just do this itself?
Because Knowledge Catalog is built for coverage and speed on Google Cloud, and it does that job well. Its value is that Gemini can describe thousands of tables without a queue in front of every one, and that any agent can pull the result under the caller's IAM. Putting an approval on every Gemini draft would slow the thing that makes enrichment worth turning on, so a per-run publish choice and Agent or User labels are a sensible default.
Deciding what's authoritative across the estate is a different product. It means holding definitions from Snowflake, Databricks, dbt and Looker that Google doesn't govern, scoring them against each other, keeping a record of who approved each one, and stopping agents from promoting their own output. Acting on that context means writing to production systems across clouds: scoping the blast radius, getting approval, rolling back, leaving a receipt, and carrying liability for changes in tools Google doesn't own. A cloud provider has good reasons to keep that risk out of its catalog. That product is Data Workers.
Where the two overlap
"Both" means Data Workers works alongside the Knowledge Catalog capability.
| Job to be done | Knowledge Catalog | Data Workers | What we recommend |
|---|---|---|---|
| Finding Google Cloud assets | Auto-ingest and semantic search (GA) | Takes entries and search results from your coding agent over MCP | Knowledge Catalog |
| Drafting descriptions and joins | Gemini data insights (GA on BigQuery) | Drafts docs too, as proposals | Knowledge Catalog for Google Cloud |
| Packing context for an agent | lookupContext, discovery MCP | Approved definitions with approver and sources, over MCP | Both |
| Lineage for Google Cloud jobs | Automatic for BigQuery, Dataflow and Managed Airflow | Builds lineage from the dbt manifest and the team's context-graph notes and joins other platforms | Both |
| Deciding which draft is authoritative | Per-run publish; change review in private preview | Proposal, checks, trust score, named approver | Data Workers |
| Context from Snowflake, Databricks, dbt | Federation and connectors in preview; dbt import one-way | Read today and lined up with the approved definition | Data Workers |
| Fixing what bad context broke | Not part of the context stack | Conductor and Swarm, with a receipt on every change | Data Workers |
Where Data Workers wins in a mixed estate
Keep Gemini drafting on Google Cloud, and keep Knowledge Catalog as the inventory people search and IAM enforces.
Run trust on Data Workers. Every AI-written description and join enters as a proposal with its source, gets checked against the live schema and scored against definitions from Snowflake, Databricks and dbt, and becomes authoritative only when its owner approves it. Every agent, including your coding agent and the Data Workers agents that act, reads the approved version. Nothing moves: Data Workers stores metadata and scrubbed facts about your data, not copies of your tables.
What it costs
Knowledge Catalog bills by data compute unit hour, with SKUs still named Dataplex. Standard processing covers discovery at $0.06 per DCU-hour, with 100 free DCU-hours a month. Premium processing covers lineage, data quality and profiling at $0.089 per DCU-hour in us-central1. Metadata storage bills per GiB. Gemini features, including data insights and automated metadata generation, bill as part of Gemini in BigQuery or Gemini Code Assist.
Data Workers is priced the other way round. The Apache 2.0 core is free. A pilot is $7,500 one-time. Scale starts at $1,000 a month and Enterprise at $3,000 a month (billed annually). Seats are unlimited, there's no usage meter, and there's no markup on model spend because you bring your own model. See pricing.
The fastest first win: one data product your agents already use
Pick one data product, or the ten BigQuery tables your agents query most. Connect Data Workers to BigQuery with a read-only role, then add your dbt project and the one other platform where the same metrics live. Give your coding agent both MCP servers and one instruction: before using a Gemini description or join, hand it to the Context Wizard. Within a pilot you have every AI-written fact your agents rely on listed with its source and score, the conflicts with Databricks or dbt routed to their owners, and an approved definition for each, with a record of who approved it.
What each Data Workers product adds next to Knowledge Catalog
Data Context Wizard. One governed graph across warehouses, lakehouses, dbt, orchestration and BI, with provenance, trust scores and the authority guard. It scrubs PII before storage, isolates each tenant and keeps a tamper-evident audit. It also imports and exports Open Knowledge Format bundles, the format behind Google's open knowledge-catalog project, through the same governed write gate. We make the full case in Why the Data Layer Needs Its Own Context Layer.
Spellbook Data Catalog. Where people review what agents propose: approve, steer, send back or roll back, with asset pages that cover the whole estate. Knowledge Catalog and IAM stay the enforcement point for Google Cloud.
Data-Agents Swarm and Autonomous Data-Conductor. 20+ specialist agents and the conductor that runs them read the approved graph before they act. Each agent is an MCP server your coding agent can call. Fixes are verified downstream, and an approved definition fix can ship back to dbt as a doc change.
Guardrails for context and for change
Google governs Google Cloud context well: lookupContext and search respect IAM, publishing needs Catalog Editor, and data products gate access behind an owner's approval. Data Workers adds controls for context and change across the estate:
- •AI-written facts enter as proposals at the "derived" level, pending review.
- •Promotion needs a named approver. The authority guard is enforced in code, in the write path, and an agent can't approve its own output.
- •Checks run before review: live schema, human-written text protected, lineage impact, join cardinality.
- •Autonomy is set per domain. Context upkeep can run at L3 (act, reversibly) while definition changes stay at L2 (propose).
- •New deployments start observe-only.
- •Every action is approved or reversible, with a receipt.

"Knowledge Catalog already gives our agents context over MCP. Isn't that enough?"
For finding and describing Google Cloud data, it often is. The discovery MCP server and lookupContext are a clean way to give Claude Code or Gemini CLI the schema, joins and descriptions they need, and your coding agent can call them next to Data Workers.
What MCP carries is whatever the catalog holds. If a Gemini draft went stale an hour ago, lookup_context serves it, correctly filtered by IAM, to every agent that asks. Access control decides who may see a fact. It doesn't decide whether the fact is right, who approved it, or whether Databricks defines it differently. MCP is the transport. The scoring, the approval and the cross-platform graph are the product, and that's what Data Workers ships.
How it fits together

There's no migration. Knowledge Catalog entries, Gemini descriptions, aspects and lookupContext bundles reach it through your coding agent over MCP today. Add your dbt project, your orchestrator and your other platforms. Every agent starts observe-only, and the first thing you see is the list of AI-written facts your agents use, each with its source, its score and any definition elsewhere that disagrees.
The case for your CFO
The outcome. Every agent answer about your data rests on a definition its owner approved, on Google Cloud and every other platform. Fewer wrong numbers reach a deck, and less time goes to tracing where an agent got an answer.
The risk story. At L1 Data Workers only observes. At L2 it flags and proposes; a named owner approves. At L3 it acts only on reversible changes, and L4 autonomy is something you grant per domain, never by default. Every change leaves a receipt with the diff, the approver, the blast radius and the rollback path, and fixes are verified downstream. Nothing migrates: Knowledge Catalog stays the inventory and Gemini keeps drafting.
Why now. In 2026 Google shipped lookupContext and MCP servers for Knowledge Catalog, so Gemini drafts now flow straight into agent answers. The faster enrichment runs, the more it matters who approves what agents treat as true.
The first win. One data product: every AI-written fact your agents use, scored, with conflicts routed to owners and approved definitions on record.
What stays the same. Knowledge Catalog, Gemini data insights, IAM, BigQuery, dbt and the coding agent your engineers already use.
The path. Start with a pilot. It's $7,500 one-time and credited in full against the first year; see pricing.
The sentence for upstairs: "Gemini can draft our metadata; Data Workers makes sure a named owner approves it before any agent treats it as the truth, on Google Cloud and everywhere else."
When Knowledge Catalog alone is enough
Knowledge Catalog can be enough if your data lives in Google Cloud, a small team curates every published insight, and agents mainly need to find and read assets. Once definitions also live in Snowflake, Databricks or dbt, or agents start answering questions that reach a board, you need context that's scored, approved and consistent across platforms, and that's the job Data Workers does.
FAQ
What does Knowledge Catalog give AI agents? Knowledge Catalog gives agents Google Cloud metadata through the lookupContext method and a discovery MCP server with search_entries, lookup_context and lookup_entry. A bundle covers schema, joins, descriptions, glossary terms, profile and quality results for up to ten assets, filtered by IAM. lookupContext launched in preview on June 4, 2026.
What is Gemini metadata enrichment in Knowledge Catalog? It's data insights: Gemini drafts table and column descriptions, sample queries and relationship graphs, GA on BigQuery. You choose per run whether to publish them as aspects, which makes them searchable and visible to others, or to view them for your session only.
Does anyone review Gemini-generated descriptions before agents use them? In Knowledge Catalog, the person who publishes decides, and the output is labelled Agent until someone edits it and sets the source to User. Review of metadata change requests through governance workflows is in private preview. Data Workers holds AI-written facts as proposals until a named owner approves them.
Knowledge Catalog context for AI agents vs Data Workers: which should I choose? Use both. Keep Knowledge Catalog and Gemini for finding and describing Google Cloud data. Add Data Workers when agents need context that's scored, approved by an owner and consistent across Snowflake, Databricks, dbt and BI.
Does Data Workers read lookupContext and data insights? Knowledge Catalog entries, descriptions, aspects and lookupContext bundles reach Data Workers over MCP today, through your coding agent calling both servers.
Will Data Workers overwrite our Gemini or steward descriptions? Knowledge Catalog stays the system of record for its own aspects, by design. Approvals live in the Context Wizard and Spellbook, where every decision gets an owner and a receipt, and the steward carries the approved fact into the Knowledge Catalog entry, for example by setting a join's source to User so Gemini won't replace it.
What does Data Workers store? Metadata and scrubbed facts about your data, such as definitions, lineage, owners and incident history, not copies of your tables. Enterprise can run in your own VPC.
Sources
Sources for Knowledge Catalog capabilities and statuses, checked October 2, 2026: retrieve data context with lookupContext (updated September 30, 2026), data insights for structured data (updated September 30, 2026), Knowledge Catalog MCP reference, use the remote MCP server, release notes (rename April 10, data products May 25, lineage MCP May 27, lookupContext June 4, dbt import GA September 24, 2026), data products overview, create a business glossary (metadata change review in private preview), catalog overview (federation preview), Introducing the Google Cloud Knowledge Catalog (April 22, 2026) and Knowledge Catalog pricing. Where Google's pages show different statuses, we've used the more conservative one and named both. If we've got something wrong, tell us and we'll fix it.