Comparison
Comparison13 min readBy The Data Workers Team

On Google Cloud? What Data Workers Adds on Top of Knowledge Catalog (Dataplex) and the BigQuery Agents

Google Cloud now ships Knowledge Catalog (formerly Dataplex), a BigQuery Data Engineering Agent and managed MCP servers. Here is where they stop, and what Data Workers adds.

Your analytics run on BigQuery and Looker. Finance probably still has a Snowflake warehouse, the ML team has a Databricks workspace, and half your transformations are dbt models that never moved to Dataform. That's where your catalog goes blind and your incidents stall.

At Google Cloud Next '26, Google repackaged its data portfolio as the Agentic Data Cloud. Dataplex Universal Catalog was renamed Knowledge Catalog and repositioned as a "universal context engine" for agents. The BigQuery Data Engineering Agent went GA. BigQuery, Spanner, Cloud SQL and a dozen other services got managed MCP servers.

So when a GCP customer asks us "Google now has a context engine, a data engineering agent and MCP everywhere, so what's left for Data Workers?", it's a reasonable question. The short answer: Google has built a context and agent layer that is read-wide and write-narrow. It can increasingly see data outside GCP, but it acts only on GCP services, and most of its reach beyond Google is still in preview and read-only. Data Workers is the autonomous back office for the whole estate, and it treats BigQuery as one important system among several.

Use Knowledge Catalog and Conversational Analytics for your GCP data. Use Data Workers for everything that lives or breaks outside GCP, and for anything an AI writes that a person should approve before it's trusted.

Key takeaways

  • •Data Workers turns your Google Cloud estate into an autonomous data platform. Cost optimization, migrations, incidents, governance and catalog upkeep run on their own, across every cloud you use, at whatever autonomy level you choose per domain.
  • •IAM and VPC Service Controls stay the lock, Knowledge Catalog stays the governance plane for GCP assets, and LookML stays the metric source. Data Workers is the back office for everything Google can't see, and the review queue for anything that writes back into GCP.
  • •Google's agents act only on GCP. The Data Engineering Agent works in Dataform and BigQuery Pipelines. Catalog federation to Unity Catalog, Glue and Snowflake is read-only and in preview.
  • •Knowledge Catalog is a read-side context plane that other agents call. It is not an agent workforce that fixes things. Data Workers' Autonomous Data-Conductor owns detect → diagnose → fix → review → verify across GCP and everything else.
  • •Third-party coverage is thin where it matters. The Snowflake metadata connector is community-built and "not officially supported by Google", there's no Databricks connector, and dbt import is in preview.
  • •Data Workers vs. Dataplex, in one line: Dataplex tells an agent what your GCP data means. Data Workers runs the work across all of your data, with a governed write path and receipts.

This guide compares what Google Cloud covers with what Data Workers adds. For the step-by-step path from a traditional data team to an autonomous data platform, read From Google Cloud to an autonomous data platform. For the leadership view, read You built on Google Cloud. Now make it agentic & autonomous.

Six things you get on top of Google Cloud

1. A data platform that runs itself, across more than GCP. Google has added agents to GCP, and they act on GCP services only. Data Workers is the agentic data platform for everything. The back office (governance, compliance, pipelines, catalog upkeep, cost and incidents) runs on its own across every platform you use, and your team moves up an autonomy ladder from observe-only to fully autonomous one domain at a time. Google's agents act on GCP services only.

2. One control plane across every cloud, with no migration. Connect BigQuery, Knowledge Catalog and Looker alongside Snowflake, Databricks, dbt and Airflow. Nothing moves. You get one context graph, one review inbox and one autonomy policy across all of them. Knowledge Catalog can read some of these, in preview. Data Workers lets you run them.

3. Cost optimization that never stops. The Cost Savings & Cleanup agent right-sizes BigQuery slots and breaks spend down by table, query, team and pipeline, across BigQuery, Snowflake and Databricks. It finds tables nobody has queried in months and duplicate pipelines, and archives only after checking dependencies. Our design target is a 25–40% cut in warehouse spend. That's a target we're engineering toward, not a billed average.

4. Migrations you approve instead of staffing. Teradata, Redshift, Oracle or Snowflake workloads into BigQuery, or out of it. Google's transfer and translation services move data and SQL into BigQuery. The Data Migration agent adds what makes a cutover safe: it re-translates anything that fails validation, proves parity with row counts, checksums, statistical profiles and referential-integrity checks, and cuts over incrementally while the legacy system stays live. Each wave is a single approval in Spellbook. Our target is 4–8 weeks for work that usually takes consultants 6–12 months.

5. Incidents closed across every hop. The Conductor traces a wrong Looker number through BigQuery to its root cause, even when that's a dbt model or a Snowflake share. It has the right agent fix it there and confirms the tile is right again, with a receipt on every change.

6. Governance and compliance by default. Access requests go from ticket to IAM grant or policy tag, with a receipt. AI-generated metadata stays a proposal until someone approves it. Audit evidence is assembled from receipts instead of being hunted down every cycle.

Behind all six are 20+ specialist agents, one context graph across every platform, and the coding agent your team already uses (Claude Code, Codex or Cursor) as the way in. Spellbook is where you look: every asset, every change and every receipt, with one-click rollback.

One incident, five systems

Here's a scenario most GCP shops with a mixed estate will recognize (an illustration, not a customer case):

  • •Monday. Finance's Snowflake team renames a column in a table they share with analytics.
  • •Monday night. A dbt model, run by your self-managed Airflow, joins that share into BigQuery and silently drops the unmatched rows.
  • •Tuesday. Knowledge Catalog's lineage shows the BigQuery table, but upstream it stops at the edge of GCP.
  • •Tuesday, 09:00. The Looker revenue tile is down, and someone asks Conversational Analytics why.
StepWhat Google-native tooling seesWhat Data Workers does
Rename in SnowflakeNothing. The Snowflake connector is community-built, and federation is read-only and in preview.The Schema Evolution agent flags the change and every downstream consumer, including the BigQuery table.
dbt model drops rowsdbt import is in preview. The Data Engineering Agent works in Dataform, not dbt.The Pipeline agent opens a PR on the dbt model, and the Change Review agent attaches the blast-radius report.
BigQuery table is incompleteData quality scans may catch the row drop.The Conductor ties the BigQuery symptom to the Snowflake root cause.
Looker tile is wrongConversational Analytics explains the drop from LookML.The Conductor confirms the tile recovered, closes the receipt, and records the pattern.

Google gives you the best explanation of the symptom inside GCP. Data Workers closes the ticket wherever it started. The rest of this guide is about that difference.

What Google Cloud covers, as of September 2026

Google renamed a lot in 2026: Dataplex Universal Catalog became Knowledge Catalog, BigLake became Lakehouse for Apache Iceberg, Vertex AI became Gemini Enterprise Agent Platform, and Composer became Managed Service for Apache Airflow.

AreaWhat Google shipsStatus (Sept 2026)
Context engineKnowledge Catalog (formerly Dataplex): metadata harvesting, Gemini enrichment, ACL-aware semantic search, data productsGA (several features in preview)
Data engineering agentBigQuery Data Engineering Agent: builds, modifies, troubleshoots and migrates pipelinesGA; runs in Dataform, BigQuery Pipelines and BigQuery Studio
Data science agentData Science Agent in BigQuery and Colab EnterpriseGA
Business Q&AConversational Analytics in BigQuery and Looker; Conversational Analytics APIGA (proactive "agentic workflows" in preview)
SemanticsLookML; BigQuery Graph and MeasuresGA / Preview
MCPManaged MCP servers for BigQuery, Cloud SQL, Spanner, Dataform and more; Knowledge Catalog and Data Lineage MCPGA / Preview
Agent toolingADK, A2A, Agent Runtime with Memory Bank, Data Agent Kit, open-source MCP Toolbox for DatabasesMixed; Data Agent Kit in preview
LineageColumn-level lineage across BigQuery, Spark, Airflow and LookerGA
Cross-cloudCross-cloud Lakehouse (BigQuery and Spark over AWS and Azure data); catalog federation to Unity Catalog, Glue and Snowflake/PolarisPreview, read-only
GovernanceIAM (including deny policies), VPC Service Controls, Model Armor on MCP traffic, Agent Identity; Agent GatewayGA; Agent Gateway in preview

Knowledge Catalog's hybrid semantic search and verified-query approach are well designed. For a GCP-only estate, it's a good context plane.

Here's Data Workers vs. Dataplex (Knowledge Catalog) and the rest of Google's stack, side by side. Google leads on its home ground: automatic metadata enrichment, self-serve answers inside BigQuery and Looker, and the security perimeter. The five outcomes where we lead are all about data and work that sit outside GCP, or that need a person's sign-off before an AI result is trusted.

Spider chart comparing Data Workers and Google Cloud on eight outcomes a data leader buys
OutcomeData WorkersGoogle CloudWhy we scored it this way
Metadata enriched automatically79Gemini writes descriptions, glossaries and entity extractions at scale. We enrich too, but Google's reach inside GCP is wider.
Self-serve answers in BigQuery and Looker58Conversational Analytics is GA in BigQuery and Looker. Our business-user app is Spellbook Data Catalog (preview): one chat box across every cloud, with Spellbook for Slack and Teams coming.
Security perimeter for agents79IAM deny policies, VPC Service Controls and Model Armor are best in class. We operate inside that perimeter rather than replace it.
Catalog coverage of Snowflake, Databricks, dbt93The Snowflake connector is community-built and unsupported, there's no Databricks connector, and dbt import is in preview.
AI metadata reviewed before it's trusted94Gemini's output lands as catalog content, and glossary approval workflows are in preview. In Spellbook, AI-written metadata stays a proposal until a person approves it.
Pipelines built and fixed outside Dataform82The Data Engineering Agent works in Dataform and BigQuery Pipelines only. Our agents work in dbt, self-managed Airflow and non-GCP targets.
Incidents fixed across GCP and non-GCP hops83The Database Observability Agent and proactive workflows are in preview and GCP-only. We trace and fix through every hop.
One autonomy policy across every cloud83Google's controls decide whether an agent may call a GCP tool. We set how much an agent may do on its own, per domain, on every platform.

The scores measure scope (what each side covers), not answer quality, and they're our directional judgments, not benchmarks. We've shown the reasoning so you can argue with any line.

Google's agents stop at the GCP edge

Every limit below comes from Google's own documentation. Each follows from a consistent design choice: GCP is the center, and everything else is a source to read.

1. The agents act only on GCP services

The BigQuery Data Engineering Agent is capable, but it works in Dataform and BigQuery Pipelines. Google's materials say nothing about it working on dbt, on self-managed Airflow, or on any target outside GCP. Conversational Analytics runs on BigQuery, Looker and GCP databases, and reaches other clouds only through the preview Lakehouse federation. We found no documentation of any Google agent writing to Snowflake, Databricks, external dbt or external Airflow.

2. Reach beyond GCP is read-only and mostly preview

  • •Catalog federation is read-only and in preview. Google's docs say creating, updating or deleting resources in the remote catalog isn't supported. The Unity Catalog integration covers only external locations on AWS or GCP; UC default storage isn't supported.
  • •The Snowflake metadata connector is a community connector, explicitly "not officially supported by Google."
  • •There's no Databricks metadata connector.
  • •dbt Core and MetricFlow metadata import is in preview.
  • •The only multi-vendor connectivity Google offers is open-source DIY (MCP Toolbox, ADK, A2A), which means you build and govern those agents yourself.

3. Knowledge Catalog serves context; it doesn't do the work

This is the core of the Data Workers vs. Dataplex comparison. Knowledge Catalog is a read-side context and governance plane that other agents call. It harvests, enriches, searches and serves. Its MCP server can create and update data products and aspects, but it doesn't detect that a pipeline broke, trace the cause across systems, fix it, and verify the fix. The Database Observability Agent and the proactive analytics workflows are in preview and scoped to GCP databases and BigQuery metrics.

4. Enrichment is AI-generated, and promotion isn't gated

Gemini auto-generates descriptions, glossaries and entity extractions at scale, which is useful. But an auto-generated description that nobody has checked is still a guess. Knowledge Catalog's governance workflows for glossary approval are in preview. In Data Workers, nothing an agent writes becomes authoritative until a named human approves it, and that rule is enforced in code.

Matrix of where Data Workers and Google Cloud can read, fix and verify across every system in a Google Cloud estate

The overlap, honestly

Job to be doneGoogle-nativeData WorkersWhat we recommend
Warehouse and lakehouse computeBigQuery, Managed Spark, IcebergNot our jobGoogle Cloud
Access enforcement on GCP assetsIAM, VPC-SC, policy tagsProposes and applies IAM and policy-tag changes through GCPBoth: GCP enforces, Data Workers operates
Business Q&A over BigQuery and LookerConversational Analytics, Looker agentsSpellbook chat (preview), backed by the Search & Research and Data Science & Insights agentsBoth: Conversational Analytics for BigQuery- and Looker-only questions, Spellbook when the answer spans clouds or turns into a fix
Business Q&A across cloudsCross-cloud Lakehouse (preview)Context-grounded Q&A across BigQuery, Snowflake, Databricks and PostgresData Workers
SemanticsLookML, BigQuery Measures (preview)Ingests LookML, dbt semantics, Snowflake semantic views and UC metric views into one governed graphBoth
Context layerKnowledge CatalogData Context Wizard: cross-platform, provenance on every fact, human-gated promotionData Workers for anything beyond GCP
CatalogKnowledge CatalogSpellbook Data Catalog: an estate-wide agentic control planeBoth: Knowledge Catalog stays the GCP governance plane
Pipeline authoringData Engineering Agent (Dataform, BigQuery Pipelines)Pipeline Building agent across dbt, Airflow, Dagster and Prefect (Dataform over its API or MCP today)Google for Dataform-only; Data Workers for mixed stacks
Data qualityKnowledge Catalog data quality scansQuality Monitoring agent that routes each issue into a resolution loopBoth: scans are a signal we consume
Incident detection and resolutionDatabase Observability Agent (preview, GCP DBs)Autonomous Data-Conductor running detect → diagnose → fix → review → verify across systemsData Workers
MigrationData Engineering Agent (into BigQuery); Snowflake transfer serviceData Migration agent in any direction, with parity verificationBoth
Cost and cleanupSlot recommendations, FinOps hubCost Savings & Cleanup agent across every warehouse you pay forData Workers
Building your own agentsADK, Agent Runtime, Data Agent KitOpen-source Apache 2.0 core; our agents are callable from ADK agents over MCPBoth

What it costs

On GCP we don't argue on price. Google bundles its core Gemini in BigQuery features into BigQuery editions at no extra charge, and Conversational Analytics is billed as the BigQuery queries it runs. Data Workers is priced as a flat platform fee: the Apache 2.0 core is free, Scale starts at $1,000 a month and Enterprise at $3,000 a month (billed annually), with unlimited seats, no usage meter and no markup on model spend. The comparison that matters is coverage. Google's agents can't fix a dbt model, a Snowflake table or a Databricks job at any price.

The fastest first win: one catalog for Snowflake, Databricks and dbt

The quickest way to see the gap is to put the systems Knowledge Catalog can't see next to the ones it can. Data Context Wizard ingests Knowledge Catalog's glossary, aspects and lineage alongside Snowflake, Databricks, dbt and Airflow, and Spellbook shows end-to-end lineage in one place. Gemini-generated descriptions come in as proposals, so your team approves what becomes the official definition.

What each Data Workers product adds on Google Cloud

Spellbook Data Catalog vs. Knowledge Catalog (Dataplex): inventory versus control plane

Knowledge Catalog is the most ambitious hyperscaler catalog on the market, and Google is right that the catalog is where context for agents should live. We go one step further. Once agents are doing the work, the catalog has to become the control plane for that work. It's no longer just the place agents look things up. We make the full argument in From the Traditional Data Catalog to the Agentic Data Catalog.

Spellbook Data Catalog on a GCP estate:

  • •Agents keep the metadata current, governed like every other change, so it doesn't rot the way manually curated metadata does.
  • •Asset pages cover the whole estate. A BigQuery table's deep-wiki page shows the Pub/Sub or Fivetran source, the Dataform or dbt model, the Snowflake or Databricks consumers, and the Looker explores built on it.
  • •One inbox for all agent work: approve, steer, send back or roll back, for every agent on every platform.
  • •An authority guard enforced in code. No agent, Gemini-powered or ours, can promote its own output to canonical.
  • •Knowledge Catalog stays the governance plane for GCP. Approved changes are applied through IAM, policy tags and Knowledge Catalog aspects.

Spellbook is in preview today.

Data Context Wizard vs. Knowledge Catalog's context engine

Google calls Knowledge Catalog a "universal context engine". Universal, in practice, means GCP plus a set of preview connectors. Data Context Wizard differs in three ways:

  • •It really is cross-cloud. BigQuery, Snowflake, Databricks, Postgres, dbt, Airflow, Looker and other BI tools, plus existing catalogs, all feed one graph, through 50+ connectors that we support ourselves.
  • •Provenance on every fact. Each fact records its source, author, confidence, when it was observed, and its tenant. Gemini-generated descriptions come in as proposed facts with low confidence, not as truth.
  • •Scored and gated trust. A multi-signal authority score orders the human review queue, and a named human approves what becomes authoritative.

The graph scrubs PII before storage, isolates each tenant, keeps a tamper-evident audit, and supports GDPR right-to-be-forgotten.

The Data-Agents Swarm vs. the BigQuery Data Engineering Agent

Google shipped one capable data engineering agent that works inside one set of tools. The Data-Agents Swarm is 20 specialized agents across 50+ integrations, each built for one kind of work: pipelines, incidents, quality, schema evolution, change review, access and governance, security, identity, cost, migration, observability, streaming, ingestion, MLOps and search. Specialization matters because the checks that make a schema migration safe aren't the checks that make an access grant safe.

On a GCP estate, these matter most:

  • •Pipeline Building agent for teams that run dbt next to (or instead of) Dataform, or Airflow outside Managed Airflow.
  • •Streaming agent for Pub/Sub and Kafka pipelines that feed BigQuery and other destinations.
  • •Cost Savings & Cleanup agent for BigQuery slot right-sizing and the Snowflake or Databricks bills sitting next to it.
  • •Data Access & Governance and Identity agents that take access requests from ticket to IAM grant to receipt.

Autonomous Data-Conductor: owning the outcome

Nothing in Google's portfolio owns an incident end to end across systems. Autonomous Data-Conductor does. It detects the problem, traces the cause through lineage (including the non-GCP hops), directs the right agents to fix it where it lives, scopes the blast radius, routes to approval at whatever level the domain requires, and verifies that the Looker number is right again. Every change leaves a signed receipt and can be reversed in one click, and each outcome is written back to the context graph.

Autonomy guardrails and security

Google's security primitives are excellent. IAM deny policies can block MCP use, VPC Service Controls wrap MCP servers, Model Armor scans MCP traffic, and every call is audit-logged. Those controls decide whether an agent may call a GCP tool.

Data Workers adds the layer above that: how much an agent may do on its own, per domain, across every platform.

  • •Autonomy runs from L0 (manual) to L4 (fully autonomous), set per domain.
  • •New deployments start observe-only, and teams move up the ladder as the receipts earn trust.
  • •Blast radius is scoped across platforms before any write.
  • •Every change gets a signed receipt and one-click reversal.
  • •No agent can approve or promote its own work, enforced in code.
  • •Least privilege. Data Workers acts with the IAM roles you grant it, inside your VPC-SC perimeter, so Google's controls still apply to everything we touch there.
The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous

"Google has MCP servers for everything now. Doesn't that make Data Workers redundant?"

No, though it makes connectors redundant, which is fine by us. Google's managed MCP servers are a great way for any agent, including ours, to reach BigQuery or Spanner with IAM enforced. But an MCP server is a door, not an operator. The value is in what's on the other side of the door: one operating context across the estate, specialists that know how to do each kind of work safely, a loop that verifies outcomes, and memory that makes the next occurrence cheaper. MCP is how we connect. The operating model is what we sell.

How it fits together

How Data Workers fits with Google Cloud: your coding agent on top, Data Workers in the middle, your Google Cloud estate underneath

There's no migration. You connect Data Workers with a scoped service account, add Knowledge Catalog, Looker, your dbt or Dataform project, your orchestrator and your non-GCP platforms, and let Data Context Wizard build the graph. Every agent starts observe-only. The first things you'll see are cross-platform lineage and incident history, including the hops Knowledge Catalog can't see today.

When Google Cloud alone is enough

  • •Your estate is GCP-only. BigQuery, Dataform, Managed Airflow and Looker, with no Snowflake or Databricks in the critical path.
  • •Your only need is Q&A on data that never leaves BigQuery and LookML. Conversational Analytics is strong there.
  • •Your pipeline work lives in Dataform, and the Data Engineering Agent covers it.

Data Workers earns its place when more than one platform sits in your critical path; when your transformations run in dbt or your orchestration runs outside GCP; when you want governance, compliance, pipeline management and catalog upkeep to run autonomously; or when you need graded, reversible autonomy with receipts an auditor will accept.

FAQ

Is Data Workers a Dataplex alternative? It's a complement on GCP-only estates and the better foundation on multi-cloud ones. Knowledge Catalog stays the governance plane for GCP assets. Spellbook and Data Context Wizard sit across the estate and apply changes back through GCP's own controls.

Can Google ADK agents or Gemini Enterprise call Data Workers? Yes. The swarm is MCP-native, and Gemini Enterprise supports bring-your-own MCP.

Does Data Workers work with the BigQuery Data Engineering Agent? Yes. Dataform changes made by Google's agent are picked up by our Schema Evolution and Change Review agents, like any other change, and assessed for blast radius across the estate.

Does Data Workers bypass IAM or VPC Service Controls? No. It acts with the IAM roles you grant, inside your VPC-SC perimeter, and applies changes to GCP assets through GCP's own controls.

What happens if an agent gets something wrong? Every write is scoped before it runs and goes to review at whatever autonomy level you've set for that domain. It leaves a signed receipt and can be reversed in one click. No agent can approve its own work.

What does Data Workers store? Metadata and scrubbed facts about your data (definitions, lineage, owners, incident history), not copies of your tables. PII is scrubbed before anything is stored, tenants are isolated, and Enterprise can run in your own VPC or on-premise.

Is Data Workers open source? The core is Apache 2.0 and free to run. See pricing for what the enterprise platform adds.

Sources for Google Cloud capabilities and statuses: Google Cloud documentation, release notes and blog posts current as of September 29, 2026, including What's new in the Agentic Data Cloud, the BigQuery Data Engineering Agent, Knowledge Catalog release notes, managed connectivity, catalog federation with Databricks, supported MCP products and MCP IAM controls. If we've got something wrong, tell us and we'll fix it.

Ready to go autonomous and agentic?

We’re building the future of data infrastructure right now. See how your enterprise data stack can operate fully agentic today.