Product
Product9 min readBy The Data Workers Team

What is an agentic data platform?

An agentic data platform is one where AI agents do the operational work of a data estate across every engine, under governance. The five layers that make one, and how catalogs, observability and platform-native agents compare.

An agentic data platform is one where AI agents do the operational work of a data estate (detect, diagnose, fix, verify and document) across every engine you run, under governance: shared context, scoped access, approvals, autonomy set per domain and a receipt on every change. Data Workers is the agentic data platform: all five layers that definition needs, across every engine you run, built on the warehouses, pipelines, catalogs and assistants you already have.

The definition matters because "agent" is now on every data product page. This page gives buyers a test they can apply to any vendor, including us.

Key takeaways

  • •Agentic means the agents do the work. Answering questions about data is useful. Running the estate, which means finding the break, fixing it at the source and proving the fix held, is the job that defines the category.
  • •Five layers make a platform agentic. Cross-engine context, specialist agents, a governed write path with approvals and rollback, autonomy set per domain from L0 to L4, and receipts with an audit trail.
  • •Every adjacent category is strong at its own job. Catalogs hold the inventory, observability raises the alarm, platform-native agents answer on their own platform, coding agents write code and semantic layers define the metrics.
  • •Data Workers brings all five layers and works with all of those tools. It reads them as context, acts across them through one approval flow and leaves one audit trail.
  • •Zero migration. Your warehouse, dbt project, orchestrator, catalog and assistants stay where they are.

The definition, in one paragraph

A data platform stores and moves data. An agentic data platform also operates it. When a column is renamed upstream, a test starts failing or a dashboard drifts, agents notice, trace the cause across systems, propose or apply a fix within the limits you set, check that the fix held downstream and record what happened. People stay in charge through approvals and autonomy levels. The platform earns more autonomy, domain by domain, as its record shows it was right.

Our pillar pages go deeper on the category and on agentic data engineering, the practice these platforms carry out. Our resource guides give an overview of agentic data platforms and the options, the 2026 landscape and the agentic data stack. This page is the buyer's test: the five layers, and how to check any vendor against them.

The five layers that make a platform agentic

How Data Workers fits with Your engines: your coding agent on top, Data Workers in the middle, your estate underneath

1. Cross-engine context. Agents need one governed view of lineage, owners, freshness, quality, usage and business meaning across every engine, with provenance on each fact. Without it, an agent rediscovers the estate on every run. In Data Workers, the Data Context Wizard holds that graph across Snowflake, Databricks, BigQuery, dbt, orchestrators and BI, and brings your existing catalogs, glossaries and semantic layers in as first-class sources (bring your own context). Public tools such as trace_cross_platform_lineage, explain_table and search_across_platforms answer from it.

2. Specialist agents. Data operations span incidents, quality, schema, pipelines, governance, security, cost, migration and models. Data Workers runs the Data-Agents Swarm: more than 20 specialist agents, each an MCP server that owns one domain. The incident agent runs diagnose_incident and get_root_cause, the schema agent runs assess_impact, the quality agent runs run_quality_check, and the governance agent runs check_policy, which evaluates each action against the policies your team sets. The Autonomous Data-Conductor routes each piece of work to the right specialist and holds state across every hop. Why a swarm of specialist agents explains the design.

3. A governed write path. Reading is safe; writing to production is where the category is decided. Before a change, Data Workers computes its blast radius (column-level, through blast_radius_analysis). Credentials are scoped per agent. Your team sets the rules for which actions need review; those changes, and anything irreversible, go to a named owner in an approval workflow with timeouts and escalation. Every change carries a rollback path, and a global kill switch halts every agent at once. The agents run in your infrastructure on every tier and hold the warehouse credentials and model key; the hosted Conductor sees workflow metadata only, never rows or credentials. The safety page covers what an agent can and can't do at each level.

4. Autonomy per domain, L0 to L4. One switch for the whole estate is the wrong control. Data Workers sets autonomy per domain on one ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous. Your team might set freshness fixes to L3 while schema changes and access grants stay at L2, where every change is proposed and approved. At L1 observe the permission ladder logs what approval each action would have needed, and your team reviews the agents' records against what it did, before any write is turned on. The autonomy levels explained walks through each step.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous

5. Receipts and audit. Every action leaves a tamper-evident receipt: what triggered it, the diff, the blast radius, the checks that ran with before and after values, who approved it and how to undo it. Receipts feed the audit trail auditors ask for, and they are the evidence that moves a domain up the ladder. Spellbook Data Catalog (in preview) is where people review, approve, roll back and audit, and the same approvals reach Claude Code, Cursor, ChatGPT and other MCP clients.

A product that is missing the governed write path, per-domain autonomy or receipts can still be excellent. It is a different category.

One incident on an agentic data platform

An illustration, not a customer case. The estate runs Postgres, Fivetran, Snowflake, dbt, Airflow and Looker.

TimeSystemWhat happens
01:50PostgresA product team renames plan_tier to plan_code
02:05FivetranThe sync lands the renamed column in Snowflake
02:30dbtfct_subscriptions builds green and writes null plans
02:40Data WorkersThe quality agent flags the null spike; the Conductor opens an incident
02:44Data WorkersLineage traces it through Fivetran to the Postgres rename; blast radius lists two dbt models, one Airflow DAG and three Looker tiles
02:50Data WorkersThe schema agent proposes a dbt diff mapping plan_code; revenue is at L2, so it routes to the domain owner
07:30SpellbookThe owner reviews the diff and blast radius and approves
07:42AirflowThe backfill DAG run completes
07:50Data WorkersThe verify step re-runs the checks on the tables behind all three tiles and records before and after results; the receipt records diff, approver, checks and rollback path
08:10LookerFinance opens a correct revenue tile

Each of the five layers did one job here: context found the cause across four systems, specialists did the work, the write path gated the change, the revenue domain's autonomy level decided who approved, and the receipt proved it.

How the adjacent categories compare

Buyers weigh five neighbours. Each is strong at its home job, and each made sensible design choices for that job.

Comparison matrix of Your engines and Data Workers on the outcomes a data leader buys

Data catalogs. Catalogs hold the inventory: assets, glossary, owners, lineage and stewardship workflows. Collibra's MCP Server brings governed data, glossary terms, lineage and impact analysis into Claude, ChatGPT, Databricks, Snowflake Cortex and other MCP tools, and its AI can propose catalog assets, glossary terms and quality rules for human review under Collibra's permissions. A catalog describes the estate; an agentic data platform also operates it. Data Workers reads your catalog as context and keeps it current. More in Data Workers vs a data catalog.

Data observability. Observability is the smoke alarm. Monte Carlo's Agentic Operations section runs an Operations Agent that routes work to troubleshooting, triage, cost and PII agents; the troubleshooting agent traverses lineage to find the root cause of an alert, and the MCP server and Agent Toolkit bring alerts, monitors and triage into coding agents. Those agents act inside the observability workflow (alerts, monitors, investigations) with the user's own permissions, which is right for a monitoring product. Data Workers takes the alert as a signal and runs the fix, the verification and the receipt. More in Data Workers vs data observability.

Platform-native agents. Each data platform now ships agents for its own estate. Snowflake Cortex Agents build and run agents inside Snowflake's governed environment, with access governed by Snowflake privileges, and people reach them in Snowflake CoWork and CoCo. Databricks' Genie family (Genie One, Genie Agents, Genie Code) is governed through Unity Catalog, and Agent Bricks adds a Supervisor Agent that orchestrates Genie Agents, functions and MCP servers. Google's Data Agent Kit brings BigQuery, Dataflow, Spark and Knowledge Catalog into the IDE and into coding agents. Microsoft Fabric data agents answer questions over OneLake and semantic models and enforce read-only access by design. These agents are deep on their own platform. Data Workers works across all of them, so a change that spans Snowflake and Databricks, or BigQuery and Looker, has one context, one approval and one receipt.

Coding agents. Claude Code, Cursor and Codex are where engineers work, and they write code well. They run with the session and keep a person in the loop who can deny each tool call, as the MCP specification recommends. Data Workers is what they call over MCP when the work touches production data, and it keeps running when the editor closes. See build it ourselves with Claude Code and working with the coding agents you use.

Semantic and context layers. The dbt Semantic Layer, Cube and similar tools define metrics once on top of your models. That meaning is valuable context, and Data Workers plugs it in as a first-class source with provenance. A semantic layer serves definitions by design; it doesn't operate the pipelines under them.

Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and keeps every one of those tools as a source or a surface. If your company already bought AI assistants, the assistants hub shows how they connect.

How to test any vendor against the definition

Ask five questions, one per layer:

  • •Does it hold context across every engine you run, with provenance, or only its own?
  • •Which agents do the work, and which domains do they own?
  • •What happens before a write: blast radius, approval by a named owner, rollback?
  • •Can you set autonomy per domain and lower it at any time?
  • •What record does each change leave, and can an auditor read it?

A vendor that answers all five concretely is offering an agentic data platform. The security and deployment page answers the follow-up question, where the agents run.

The case for your CFO

The outcome is a data estate that runs itself within limits you set: incidents traced and fixed overnight, quality checks written and kept current, idle spend removed after a dependency check, access requests answered with a receipt. The team spends its hours on the data products the business asked for.

The risk story is built into the definition. Agents propose before they act, a named owner approves anything irreversible, every change carries its blast radius, checks and rollback path, and autonomy rises per domain only when the receipts show the agents were right. Data Workers runs in your infrastructure on every tier, your data stays in your systems, and the hosted Conductor sees workflow metadata only. Zero migration: the warehouse, dbt, the orchestrator, the catalog and the assistants stay.

Why now: your people already use assistants and coding agents, and every platform ships its own agents. Agents will touch your data either way. The choice is whether they do it through one governed write path across every engine.

The first win is cross-system incident triage at L2, approved by the domain owner, inside the pilot. Start with a pilot: $7,500 one-time, and the pilot is credited in full against the first year. After it, Scale starts from $1,000/month and Enterprise from $3,000/month, billed annually, with unlimited seats, no usage meter, no markup on model spend, bring your own model, and an Apache 2.0 core. See pricing, the ROI calculator and the ROI page.

The sentence to repeat upstairs: "We're adding an agentic data platform that operates every engine we already run, under our approvals, with a receipt on every change."

FAQ

Is an agentic data platform a new warehouse? No. It sits on top of the warehouses, lakehouses, pipelines and BI you run and operates them. Data Workers needs zero migration.

Our warehouse vendor ships agents. Isn't that enough? For questions and work inside that platform, those agents are excellent. Most estates span more than one engine plus ingestion, dbt, orchestration and BI. Data Workers works across all of them with one context and one approval flow, and connects to platform agents over MCP.

What's the difference between agentic and autonomous? Agentic describes who does the work: agents. Autonomy is how much they do without asking, set per domain from L0 manual to L4 autonomous. Most teams start at L1 or L2 and move up as receipts accumulate.

Do we have to replace our catalog or observability tool? No. Data Workers reads them as context and signals. Many teams consolidate once Data Workers runs that slice too, and that choice stays yours.

How do approvals reach people? In Spellbook, and in the MCP clients your team already uses, such as Claude Code, Cursor and ChatGPT. See how approvals work.

Can we try the agents before buying? Yes. The Apache 2.0 core runs in your coding agent on your own model key; see the open-source docs. The platform adds your estate, governed writes, the context graph, the Conductor and Spellbook.

Sources

  • •Snowflake, Cortex Agents: fully managed service for building and running agents inside Snowflake's governed environment; data access governed by Snowflake privileges; users interact with agents in Snowflake CoWork and Cortex Code (CoCo). Checked October 2, 2026.
  • •Databricks, Genie: Genie One, Genie Agents and Genie Code, governed through Unity Catalog. Checked October 2, 2026.
  • •Databricks, Agent Bricks: Supervisor Agent orchestrating Genie Agents, Unity Catalog functions, MCP servers and custom agents; Agent Bricks CLI (Beta). Checked October 2, 2026.
  • •Google Cloud, Data Agent Kit overview: IDE extension and coding-agent plugin across BigQuery, Dataflow, Spark, Airflow and Knowledge Catalog. Last updated September 30, 2026; checked October 2, 2026.
  • •Microsoft, Fabric data agent concepts: read-only access and read-only queries by design; Purview policies apply. Checked October 2, 2026.
  • •Monte Carlo, Operations agent (routes to troubleshooting, triage, cost and PII agents; runs with the user's permissions), Troubleshooting agent (traverses lineage to root cause) and MCP server (investigate alerts, explore lineage, create and manage monitors; Agent Toolkit; read-only mode). Checked October 2, 2026.
  • •Collibra, MCP Server: governed data, glossary terms, lineage and impact analysis in Claude, ChatGPT, Databricks, Snowflake Cortex and other MCP tools; AI proposes assets, glossary terms and quality rules for human review; writes use Collibra's permissions. Checked October 2, 2026.
  • •dbt Labs, About the dbt Semantic Layer: metrics defined on top of models with joins handled. Checked October 2, 2026.
  • •Model Context Protocol, Specification 2025-11-25: Tools: there should always be a human in the loop able to deny tool invocations. Checked October 2, 2026.
  • •Data Workers, Autonomous Data-Conductor (autonomy per domain, L0 to L4), open-source core (Apache 2.0, public tool registrations), pricing and what is an autonomous agentic data platform. Checked October 2, 2026.