Comparison
Comparison15 min readBy The Data Workers Team

Knowledge Catalog Context for AI Agents vs Data Workers: Gemini Drafts It, a Named Owner Approves It

Knowledge Catalog and Gemini draft context for agents on Google Cloud. Data Workers scores that context, gates it on a named owner and covers every other platform.

Google now calls Knowledge Catalog "the universal context engine for your enterprise", and on Google Cloud it earns the line. Entries for BigQuery, Spanner, Pub/Sub and Cloud Storage appear on their own. Gemini data insights drafts table and column descriptions, sample queries and relationship graphs, and publishes them as aspects. The lookupContext method hands an agent a ready bundle of schema, joins, descriptions, profile and quality results, and the discovery MCP server lets Claude Code, Cursor or Gemini CLI call it directly. If your team is wiring agents to that context, this page is for you.

Your estate probably looks like this. BigQuery holds product and marketing data, and Knowledge Catalog is how people and agents find it. Finance runs a Snowflake mart. The ML team trains churn models in Databricks on Unity Catalog. dbt builds the models, Airflow schedules them, and Looker sits on top. Gemini can describe the BigQuery half of that estate in an afternoon. Nobody has decided which of those descriptions and joins an agent should treat as the truth.

Knowledge Catalog is the context engine for Google Cloud: it gathers context and Gemini drafts it. Data Workers is the context authority for the whole estate: it records where each fact came from, scores it, and keeps it a proposal until a named owner approves it. Data Workers is the agentic data platform: it runs the whole data lifecycle, and its Data Context Wizard keeps one governed context graph across every platform. Gemini enrichment stays the fastest way to describe Google Cloud data. Data Workers decides what agents can trust, on Google Cloud and everywhere else, and runs the fixes when that context breaks.

A note on names: Google renamed Dataplex Universal Catalog to Knowledge Catalog on April 10, 2026. API, IAM and CLI names are unchanged, so the endpoints still read dataplex.googleapis.com.

Key takeaways

  • •Knowledge Catalog leads on gathering context for Google Cloud. Auto-ingest and semantic search are GA, Gemini data insights is GA on BigQuery, and lookupContext packs up to ten assets into one agent-ready bundle.
  • •Gemini enrichment is a proposal until someone approves it. In Knowledge Catalog, publishing insights is a per-run choice and review of metadata changes is in private preview. Data Workers holds AI-written context as a proposal, scores it and needs a named owner to promote it.
  • •Provenance and scoring on every fact. Google labels insights Agent or User. Data Workers records source, actor, confidence, authority and time on every fact, and scores trust on origin, usage, closeness to approved facts, freshness and correction history.
  • •Context from every platform. Federation with Unity Catalog, Glue and Snowflake Horizon is preview. Data Workers reads Snowflake, dbt and Airflow today, and takes Knowledge Catalog entries and bundles from your coding agent over MCP.
  • •Context that runs the fixes. The same graph drives the Data-Agents Swarm and the Autonomous Data-Conductor, with a receipt on every change.
  • •Nothing migrates. Knowledge Catalog stays the inventory and Gemini keeps drafting.

This page is the context-for-agents view. For choosing a catalog and control plane, read our Dataplex alternative comparison, Knowledge Catalog vs Spellbook Data Catalog. For the wiring, read how to bring Knowledge Catalog glossary terms and aspects into Data Context Wizard. The Snowflake twin of this page is Snowflake Horizon Context vs Data Workers.

Six things you get with Data Workers on top of Knowledge Catalog

1. AI-written context stays a proposal until an owner approves it. Descriptions and joins from Gemini, from your coding agent or from our own agents enter the Context Wizard at the "derived" authority level, marked pending review. Promotion to authoritative needs a named human, and the check is enforced in code.

2. Every proposal is checked before anyone reviews it. A documentation change is validated against the live schema: every column it describes must exist, generated text can't overwrite a human-written description as authoritative, and downstream lineage impact is listed. Proposed joins are checked for cardinality, with fan-trap and chasm-trap warnings.

3. A trust score on every fact. Each fact gets a score from 0 to 1 built from five signals: where it came from, how often it's used, how close it sits to facts a human approved, how fresh it is and how often it has been corrected. When a fact drifts well below a competing definition of the same thing, the drift detector flags it.

4. Provenance an auditor can follow. Every fact records its source, the actor that wrote it (agent or human), confidence, authority tier and when it was observed and last confirmed, stored tenant-isolated.

5. One graph across every platform. The Context Wizard reads BigQuery next to Snowflake, dbt and Airflow, from a library of 50+ connectors, and imports MetricFlow, Databricks metric views, Cube and Wren MDL definitions. An agent asking about a BigQuery table also sees the Databricks model that reads it.

6. One review queue, with receipts. Proposed definitions and conflicts land in Spellbook Data Catalog, in preview today, next to every other change agents propose. Approving one leaves a receipt with the approver, the diff and the rollback path.

One join, five systems

Here is a Tuesday morning on a BigQuery-centred team. It's an illustration, not a customer case.

  • •08:10, BigQuery. A product release starts writing one row per user per day to app_events. Until today it held one row per account per day.
  • •08:40, Knowledge Catalog. An analyst regenerates data insights on subscriptions and app_events and picks "Generate and publish". The published description still says one row per account per day, and the suggested join on account_id assumes it. Both carry the source label Agent.
  • •09:15, Knowledge Catalog data product. The "Customer 360" data product's Gemini sample query is regenerated on the same join.
  • •10:05, coding agent. An engineer asks Claude Code for weekly active subscribers. The agent calls lookup_context on the Knowledge Catalog MCP server, gets the join and the description, and writes SQL that returns 212,400. The approved figure is 71,900.
  • •10:20, Databricks and Looker. The churn model in Databricks, built on account-day grain in Unity Catalog, disagrees with a Looker tile someone built from the agent's SQL.
Where the context livesWhat Knowledge Catalog seesWhat Data Workers does
BigQuery app_eventsThe schema change, ingested automaticallyReads the new schema through its BigQuery connector and records the change with its time
Gemini description and joinPublished as aspects, source Agent, visible to everyone with accessThe coding agent hands both to the Context Wizard as proposals over MCP; the schema check flags the grain and the join check flags a fan trap
"Customer 360" data productA sample query built on the joinLists the data product as a consumer of the disputed join
Coding agent's answerlookup_context returns the bundle it was givenReturns the approved account-grain definition with its approver; the agent answers 71,900
Databricks churn modelReadable through federation in previewReads the Unity Catalog definition, which matches the approved one and raises its trust score
The ownerNo review step between publish and useGets one item in Spellbook: approve or reject the join, with sources side by side
Incident timeline across the stack: what Knowledge Catalog, your team and Data Workers each do, step by step

The owner rejects the join at 11:30, and a steward marks the correct join as User in Knowledge Catalog, so the next Gemini run won't override it. Gemini drafted fast and mostly well. The miss came from a schema change that landed after the description was written, and the answer reached a person before anyone looked. Data Workers put the check between the draft and the answer.

What Knowledge Catalog gives agents, as of October 2026

Google labels its parts one by one. Where the release notes and the doc page disagree, we use the more conservative status and say so.

CapabilityWhat it does for agentsStatus (Oct 2026)
lookupContextOne bundle for up to ten assets: technical metadata, joins from query logs and data insights, descriptions (source-system and auto-generated), terms, profile and quality results; YAML, XML or JSON, with a character budget; filtered by IAMLaunched in preview June 4, 2026; the doc page (updated September 30) shows no label
Discovery MCP serversearch_entries, lookup_context and lookup_entry, all read-onlyNo preview label on the doc page
Data products MCP serverReads plus create and update for data products, assets and their aspectsPreview, announced May 2026
Lineage MCP serverQueries lineage graphs and impactPreview since May 27, 2026
Data insightsGemini drafts descriptions, sample queries and relationship graphs; "Generate and publish" makes them searchable and visible to othersGA on BigQuery; Iceberg tables and the Relationships tab are preview
Source labelsInsights are labelled Agent or User; for joins, User beats table constraints, which beat query history, which beat AgentDocumented
Data productsCurated asset bundles with contracts and access-request approvals; Gemini drafts sample queries and documentation templatesGA since May 25, 2026
Verified queries and semantic guardrails; automated context curationVerified SQL patterns and pre-generated questions; continuous enrichmentPreview, announced April 22, 2026
Metadata change reviewGovernance workflows for change requests to glossaries and entriesPrivate preview
Reach beyond Google CloudFederation with Unity Catalog, Glue and Snowflake Horizon; database connectorsPreview
dbt metadata importdbt Core, Fusion and dbt Cloud runs into Knowledge Catalog, lineage for BigQuery targetsGA since September 24, 2026

lookupContext is the strongest piece of this. It solves a real agent problem, packing the right metadata into a context window, and it respects IAM, so an agent sees only what its caller can see.

One platform, not one more tool

Context is one stage of a data team's lifecycle. The same team writes quality checks, handles incidents, changes pipelines, runs migrations, manages access, protects sensitive data, cuts spend and keeps models fed. Each point tool adds another console, another contract and another handoff.

Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail. We score the same ten lifecycle stages on every comparison page. Here we've scored Knowledge Catalog with data insights, data quality, lineage and its MCP servers. Google leads on catalog and context, and on governance and access through data products and IAM.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Knowledge Catalog goes deep on its own area
StageData WorkersKnowledge CatalogWhy we scored it this way
Catalog & Context99.5Knowledge Catalog's home stage. Auto-ingest across Google Cloud, semantic search, Gemini data insights and lookupContext give agents rich context on Google Cloud data. Data Workers keeps one scored, approved graph across every platform.
Analytics & Insights86Data insights drafts descriptions, sample queries and relationship graphs (GA on BigQuery). Data Workers answers questions across platforms from approved context.
Data Quality87.5Auto data quality and profiling are GA for BigQuery and Lakehouse tables, and their results ride along in lookupContext. Data Workers runs checks on BigQuery, Snowflake and Postgres and holds other tables to baselines the team records.
Observability & Incidents8.54Quality scan alerts, change feeds to Pub/Sub and lineage impact, with no incident workflow. Data Workers detects, diagnoses, fixes and verifies.
Pipelines & Ingestion8.52Knowledge Catalog gathers metadata, not data. Data Workers proposes and repairs pipelines behind an approval.
Schema & Migration83Change feeds flag schema updates, and there's no migration tooling. Data Workers plans schema changes and migrations in approved waves.
Governance & Access8.59A home stage. Data products carry access-request approvals that grant IAM (GA); review of metadata changes is private preview. Data Workers adds human-only promotion of definitions, with enforcement left in IAM.
Security & Privacy86IAM-filtered lookupContext, ACL-aware search and IAM on Google Cloud. Data Workers runs least-privilege connectors and tenant-scoped tools everywhere.
Cost / FinOps81No cost features; Gemini enrichment bills under Gemini in BigQuery. Data Workers works from query history and spend to remove the cause.
MLOps & Models7.54Ingests Vertex AI models, datasets and feature groups. Data Workers keeps the data under the models healthy.

Knowledge Catalog vs Data Workers on the context outcomes you buy

This view scores eight outcomes a data leader pays for when the goal is context agents can trust. Google leads on three, all on its own surface: context bundles, drafting at scale and search. Data Workers leads on every outcome that decides what's trusted or spans platforms.

Spider chart comparing Data Workers and Knowledge Catalog on the outcomes a data leader buys
OutcomeData WorkersKnowledge CatalogWhy we scored it this way
Context bundles for Google Cloud assets69Google leads on its own surface. lookupContext returns schema, joins, descriptions, profile and quality results for up to ten assets in one call, filtered by IAM.
Descriptions and joins drafted at scale69Gemini data insights drafts descriptions, sample queries and relationship graphs across BigQuery (GA). Data Workers drafts docs too and enters them as proposals; Gemini does this across Google Cloud out of the box.
Search across Google Cloud metadata79Semantic search is GA and ACL-aware. Data Workers takes the results your coding agent finds into its graph.
Where each fact came from96Google labels insights Agent or User and ranks join sources. Every Data Workers fact records its source, actor, confidence, authority and time.
Trust scored per fact94Google ranks joins by source type. Data Workers scores every fact on origin, usage, closeness to approved facts, freshness and correction history.
Nothing authoritative without a named approver93Publishing insights is a per-run choice, and metadata change review is private preview. Data Workers' authority guard needs a named person to promote a fact.
Context from Snowflake, Databricks, dbt and BI94Federation and database connectors are preview, and dbt import is one-way. Data Workers reads Snowflake, dbt, Airflow, Tableau and Looker today; other BI tools connect over their API or MCP server.
Context used to fix and verify92Knowledge Catalog informs agents. Data Workers uses the approved graph to fix across systems and writes the outcome back.

The scores measure scope, not answer quality. They're directional judgments, not benchmarks, and the reasoning is shown so you can argue with any line.

Where Knowledge Catalog stops

Each limit below comes from Google's own pages. None is a flaw. Knowledge Catalog exists to make Google Cloud data easy to find and describe, and fast enrichment with light friction is the right design for that job.

Publishing is the review step. Data insights offers "Generate and publish" or "Generate without publishing". Once published, insights are "indexable, searchable, and visible to other users in your organization", and that includes agents calling lookup_context. Governance workflows for metadata change requests are in private preview.

Labels record origin, not approval. Google marks insights Agent or User and ranks join sources: User, then table constraints, then query history, then Agent. That's useful provenance. It says who wrote a fact, not whether its owner signed off, and an Agent-sourced query "may be replaced during a regeneration".

The bundle mixes sources. lookupContext returns descriptions "captured in the source system and auto-generated in Knowledge Catalog" side by side. An agent reading the bundle has to work out which is which and what changed since it was written.

Google Cloud first. Federation with Unity Catalog, Glue and Snowflake Horizon is preview, the database connectors are preview, and dbt import brings metadata in one way.

Gemini has its own terms. Google notes that Gemini in BigQuery "doesn't support the same compliance and security offerings as BigQuery", and Gemini features bill under Gemini in BigQuery or Gemini Code Assist. Check both before you enrich regulated data.

Context informs. It doesn't fix. Nothing in the context stack owns a broken pipeline, checks that a fix held, or records the outcome for next time.

Matrix of where Data Workers and Knowledge Catalog can read, fix and verify across every system in the estate

Why doesn't Knowledge Catalog just do this itself?

Because Knowledge Catalog is built for coverage and speed on Google Cloud, and it does that job well. Its value is that Gemini can describe thousands of tables without a queue in front of every one, and that any agent can pull the result under the caller's IAM. Putting an approval on every Gemini draft would slow the thing that makes enrichment worth turning on, so a per-run publish choice and Agent or User labels are a sensible default.

Deciding what's authoritative across the estate is a different product. It means holding definitions from Snowflake, Databricks, dbt and Looker that Google doesn't govern, scoring them against each other, keeping a record of who approved each one, and stopping agents from promoting their own output. Acting on that context means writing to production systems across clouds: scoping the blast radius, getting approval, rolling back, leaving a receipt, and carrying liability for changes in tools Google doesn't own. A cloud provider has good reasons to keep that risk out of its catalog. That product is Data Workers.

Where the two overlap

"Both" means Data Workers works alongside the Knowledge Catalog capability.

Job to be doneKnowledge CatalogData WorkersWhat we recommend
Finding Google Cloud assetsAuto-ingest and semantic search (GA)Takes entries and search results from your coding agent over MCPKnowledge Catalog
Drafting descriptions and joinsGemini data insights (GA on BigQuery)Drafts docs too, as proposalsKnowledge Catalog for Google Cloud
Packing context for an agentlookupContext, discovery MCPApproved definitions with approver and sources, over MCPBoth
Lineage for Google Cloud jobsAutomatic for BigQuery, Dataflow and Managed AirflowBuilds lineage from the dbt manifest and the team's context-graph notes and joins other platformsBoth
Deciding which draft is authoritativePer-run publish; change review in private previewProposal, checks, trust score, named approverData Workers
Context from Snowflake, Databricks, dbtFederation and connectors in preview; dbt import one-wayRead today and lined up with the approved definitionData Workers
Fixing what bad context brokeNot part of the context stackConductor and Swarm, with a receipt on every changeData Workers

Where Data Workers wins in a mixed estate

Keep Gemini drafting on Google Cloud, and keep Knowledge Catalog as the inventory people search and IAM enforces.

Run trust on Data Workers. Every AI-written description and join enters as a proposal with its source, gets checked against the live schema and scored against definitions from Snowflake, Databricks and dbt, and becomes authoritative only when its owner approves it. Every agent, including your coding agent and the Data Workers agents that act, reads the approved version. Nothing moves: Data Workers stores metadata and scrubbed facts about your data, not copies of your tables.

What it costs

Knowledge Catalog bills by data compute unit hour, with SKUs still named Dataplex. Standard processing covers discovery at $0.06 per DCU-hour, with 100 free DCU-hours a month. Premium processing covers lineage, data quality and profiling at $0.089 per DCU-hour in us-central1. Metadata storage bills per GiB. Gemini features, including data insights and automated metadata generation, bill as part of Gemini in BigQuery or Gemini Code Assist.

Data Workers is priced the other way round. The Apache 2.0 core is free. A pilot is $7,500 one-time. Scale starts at $1,000 a month and Enterprise at $3,000 a month (billed annually). Seats are unlimited, there's no usage meter, and there's no markup on model spend because you bring your own model. See pricing.

The fastest first win: one data product your agents already use

Pick one data product, or the ten BigQuery tables your agents query most. Connect Data Workers to BigQuery with a read-only role, then add your dbt project and the one other platform where the same metrics live. Give your coding agent both MCP servers and one instruction: before using a Gemini description or join, hand it to the Context Wizard. Within a pilot you have every AI-written fact your agents rely on listed with its source and score, the conflicts with Databricks or dbt routed to their owners, and an approved definition for each, with a record of who approved it.

What each Data Workers product adds next to Knowledge Catalog

Data Context Wizard. One governed graph across warehouses, lakehouses, dbt, orchestration and BI, with provenance, trust scores and the authority guard. It scrubs PII before storage, isolates each tenant and keeps a tamper-evident audit. It also imports and exports Open Knowledge Format bundles, the format behind Google's open knowledge-catalog project, through the same governed write gate. We make the full case in Why the Data Layer Needs Its Own Context Layer.

Spellbook Data Catalog. Where people review what agents propose: approve, steer, send back or roll back, with asset pages that cover the whole estate. Knowledge Catalog and IAM stay the enforcement point for Google Cloud.

Data-Agents Swarm and Autonomous Data-Conductor. 20+ specialist agents and the conductor that runs them read the approved graph before they act. Each agent is an MCP server your coding agent can call. Fixes are verified downstream, and an approved definition fix can ship back to dbt as a doc change.

Guardrails for context and for change

Google governs Google Cloud context well: lookupContext and search respect IAM, publishing needs Catalog Editor, and data products gate access behind an owner's approval. Data Workers adds controls for context and change across the estate:

  • •AI-written facts enter as proposals at the "derived" level, pending review.
  • •Promotion needs a named approver. The authority guard is enforced in code, in the write path, and an agent can't approve its own output.
  • •Checks run before review: live schema, human-written text protected, lineage impact, join cardinality.
  • •Autonomy is set per domain. Context upkeep can run at L3 (act, reversibly) while definition changes stay at L2 (propose).
  • •New deployments start observe-only.
  • •Every action is approved or reversible, with a receipt.
The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous

"Knowledge Catalog already gives our agents context over MCP. Isn't that enough?"

For finding and describing Google Cloud data, it often is. The discovery MCP server and lookupContext are a clean way to give Claude Code or Gemini CLI the schema, joins and descriptions they need, and your coding agent can call them next to Data Workers.

What MCP carries is whatever the catalog holds. If a Gemini draft went stale an hour ago, lookup_context serves it, correctly filtered by IAM, to every agent that asks. Access control decides who may see a fact. It doesn't decide whether the fact is right, who approved it, or whether Databricks defines it differently. MCP is the transport. The scoring, the approval and the cross-platform graph are the product, and that's what Data Workers ships.

How it fits together

How Data Workers fits with Knowledge Catalog: your coding agent on top, Data Workers in the middle, your estate underneath

There's no migration. Knowledge Catalog entries, Gemini descriptions, aspects and lookupContext bundles reach it through your coding agent over MCP today. Add your dbt project, your orchestrator and your other platforms. Every agent starts observe-only, and the first thing you see is the list of AI-written facts your agents use, each with its source, its score and any definition elsewhere that disagrees.

The case for your CFO

The outcome. Every agent answer about your data rests on a definition its owner approved, on Google Cloud and every other platform. Fewer wrong numbers reach a deck, and less time goes to tracing where an agent got an answer.

The risk story. At L1 Data Workers only observes. At L2 it flags and proposes; a named owner approves. At L3 it acts only on reversible changes, and L4 autonomy is something you grant per domain, never by default. Every change leaves a receipt with the diff, the approver, the blast radius and the rollback path, and fixes are verified downstream. Nothing migrates: Knowledge Catalog stays the inventory and Gemini keeps drafting.

Why now. In 2026 Google shipped lookupContext and MCP servers for Knowledge Catalog, so Gemini drafts now flow straight into agent answers. The faster enrichment runs, the more it matters who approves what agents treat as true.

The first win. One data product: every AI-written fact your agents use, scored, with conflicts routed to owners and approved definitions on record.

What stays the same. Knowledge Catalog, Gemini data insights, IAM, BigQuery, dbt and the coding agent your engineers already use.

The path. Start with a pilot. It's $7,500 one-time and credited in full against the first year; see pricing.

The sentence for upstairs: "Gemini can draft our metadata; Data Workers makes sure a named owner approves it before any agent treats it as the truth, on Google Cloud and everywhere else."

When Knowledge Catalog alone is enough

Knowledge Catalog can be enough if your data lives in Google Cloud, a small team curates every published insight, and agents mainly need to find and read assets. Once definitions also live in Snowflake, Databricks or dbt, or agents start answering questions that reach a board, you need context that's scored, approved and consistent across platforms, and that's the job Data Workers does.

FAQ

What does Knowledge Catalog give AI agents? Knowledge Catalog gives agents Google Cloud metadata through the lookupContext method and a discovery MCP server with search_entries, lookup_context and lookup_entry. A bundle covers schema, joins, descriptions, glossary terms, profile and quality results for up to ten assets, filtered by IAM. lookupContext launched in preview on June 4, 2026.

What is Gemini metadata enrichment in Knowledge Catalog? It's data insights: Gemini drafts table and column descriptions, sample queries and relationship graphs, GA on BigQuery. You choose per run whether to publish them as aspects, which makes them searchable and visible to others, or to view them for your session only.

Does anyone review Gemini-generated descriptions before agents use them? In Knowledge Catalog, the person who publishes decides, and the output is labelled Agent until someone edits it and sets the source to User. Review of metadata change requests through governance workflows is in private preview. Data Workers holds AI-written facts as proposals until a named owner approves them.

Knowledge Catalog context for AI agents vs Data Workers: which should I choose? Use both. Keep Knowledge Catalog and Gemini for finding and describing Google Cloud data. Add Data Workers when agents need context that's scored, approved by an owner and consistent across Snowflake, Databricks, dbt and BI.

Does Data Workers read lookupContext and data insights? Knowledge Catalog entries, descriptions, aspects and lookupContext bundles reach Data Workers over MCP today, through your coding agent calling both servers.

Will Data Workers overwrite our Gemini or steward descriptions? Knowledge Catalog stays the system of record for its own aspects, by design. Approvals live in the Context Wizard and Spellbook, where every decision gets an owner and a receipt, and the steward carries the approved fact into the Knowledge Catalog entry, for example by setting a join's source to User so Gemini won't replace it.

What does Data Workers store? Metadata and scrubbed facts about your data, such as definitions, lineage, owners and incident history, not copies of your tables. Enterprise can run in your own VPC.

Sources

Sources for Knowledge Catalog capabilities and statuses, checked October 2, 2026: retrieve data context with lookupContext (updated September 30, 2026), data insights for structured data (updated September 30, 2026), Knowledge Catalog MCP reference, use the remote MCP server, release notes (rename April 10, data products May 25, lineage MCP May 27, lookupContext June 4, dbt import GA September 24, 2026), data products overview, create a business glossary (metadata change review in private preview), catalog overview (federation preview), Introducing the Google Cloud Knowledge Catalog (April 22, 2026) and Knowledge Catalog pricing. Where Google's pages show different statuses, we've used the more conservative one and named both. If we've got something wrong, tell us and we'll fix it.