Index · Q4 2026
Every vendor in the agentic data engineering market, including Data Workers, graded on one six-level autonomy scale, one cell at a time, every cell cited.
Published 2026-09-09 · 21 vendors · 6 lanes · 64 graded cells · CC BY 4.0
Why this exists
Buyers keep asking the same question in different words: when this vendor says agent, does the agent alert, suggest, draft, apply with approval, apply on its own with a receipt, or close the loop and verify? The word agent now covers all six. This index grades what each vendor publicly evidences, lane by lane, so the answer can be checked rather than believed.
Six levels. The question for every lane is the same: between the alert and the verified fix, what does the agent do on its own, and what does it leave to a person? Alerts are not fixes.
L0 · Alert
The system detects a condition and notifies a human. The human diagnoses and does everything else.
L1 · Suggest
The agent explains what happened or recommends what to do. It does not produce the change itself.
L2 · Propose
The agent produces the actual artefact (SQL, model, test, description, pipeline code, review verdict). A human applies it.
L3 · Apply with approval
The agent applies the change itself, but only after a named human approves. A dry-run or preview precedes the approval and the change carries a record of what was done.
L4 · Autonomous with receipts
Within a declared policy the agent applies routine changes without waiting for a human. Every change is logged with a receipt (what changed, why, blast radius) and can be rolled back.
L5 · Self-verifying operation
The agent detects, fixes, then verifies that the fix produced the intended outcome, learns from the result, operates across systems, and escalates to a human under a governed policy when it cannot verify.
Every cell is the highest level a public page evidences. A dashed cell marked with an asterisk is a claim: the level is stated by the vendor but the mechanism is not documented. n.e. means no public page evidences the lane at all, which is different from L0. Open a vendor below for the note and the citation behind every cell.
| Vendor | Incident resolution | Data quality | Schema and change review | Pipeline build | Catalog and documentation | Cross-cloud scope |
|---|---|---|---|---|---|---|
| Direct peers | ||||||
| Upriver | L3 | L3 | L3 | L3 | L2* | n.e. |
| Alkera | L3 | L3 | L3 | L3 | L2* | L2 |
| Altimate AI | n.e. | L2 | L2 | L3 | n.e. | L2 |
| Datus AI | L3 | L1 | L1 | L3 | L2 | L2 |
| Recce | n.e. | n.e. | L2 | n.e. | n.e. | n.e. |
| Publisher of this index | ||||||
| Data Workers | L3 | L3 | L3 | L2 | L2 | L2 |
| Platform vendors | ||||||
| Databricks (Lakeflow, Genie) | L2 | L0 | n.e. | L2 | L2 | n.e. |
| Snowflake (Cortex Code, TensorStax) | L2 | n.e. | n.e. | L3 | L1 | n.e. |
| Google (BigQuery data engineering agent, Dataplex) | L2 | L1* | n.e. | L2 | L2 | n.e. |
| Incumbents and catalogs | ||||||
| Validio | L1 | L1 | n.e. | n.e. | L0 | L1 |
| Ataccama | n.e. | L4* | n.e. | n.e. | L2 | L2 |
| Atlan | n.e. | n.e. | n.e. | n.e. | L3 | L2 |
| OpenMetadata (Collate) | L0 | L0 | n.e. | n.e. | L3 | L1 |
| Monte Carlo | L1 | L1 | n.e. | n.e. | n.e. | L1 |
| Adjacent | ||||||
| Actioneer | L1 | n.e. | n.e. | n.e. | n.e. | L1 |
| Nao Labs | n.e. | n.e. | n.e. | n.e. | n.e. | L1 |
| Wren AI | n.e. | n.e. | n.e. | n.e. | n.e. | L1 |
| Datafold | n.e. | L0 | L1 | L2 | L1 | L1 |
| Pivoted or repositioned | ||||||
| Ardent AI | n.e. | n.e. | n.e. | n.e. | n.e. | n.e. |
| Velum Labs | n.e. | n.e. | n.e. | n.e. | n.e. | n.e. |
| Mica AI | n.e. | n.e. | n.e. | n.e. | n.e. | n.e. |
Nobody is graded L4 or L5 on documented evidence in this edition. One L4 cell exists and it is a claim.
Data Workers was graded last, from dataworkers.io, llms.txt and the public GitHub repository only, by the same two rules that capped every other vendor: an approval gate has to be documented for L3, and unattended application has to come with a documented receipt and rollback for L4. We are not L4 anywhere. The autonomy dial on the Conductor describes an unattended mode, but no public page states the policy under which it applies changes, so it does not count. Pipeline build is graded propose because the community edition gates pipeline generation behind a paid tier a stranger cannot inspect. Spellbook is in preview and is graded propose. Cross-cloud scope is graded propose under the same cap applied to Alkera, Altimate and Atlan: connectors to three warehouses are evidenced, a gated change spanning two of them is not shown.
No customers, revenue, awards or third-party references are claimed. Data Workers is not SOC 2 certified. The site's benchmark figure is not cited here because no public run record accompanies it.
Autonomous data engineering platform with a context engine (living map) and a plan, execute, validate loop under human sign-off.
Publishes a price: no. SOC 2: yes. trust.upriverdata.com lists SOC 2 Type 2, GDPR and HIPAA with a SOC 2 Type 2 document available on request. Licence: Proprietary (nimble_udf repo public; platform closed). MCP: No first-party MCP server found; integrates via Claude and Cursor.
Funding, as publicly reported: $14M seed, June 2026 (Valley Capital Partners, Hetz Ventures).
Incident resolution: L3 Apply with approval
Homepage: issues fixed before the business Slacks you; late pipelines, logical errors and slow queries surfaced and fixed. Teardown records a mandatory human sign-off and a per-change validation report (queries run, rows sampled, before and after metrics, verdict). Approval gate plus receipt is L3.
https://www.upriverdata.com/ · ci:teardowns/features/upriver--wave1.md
Data quality: L3 Apply with approval
Flags violations of data quality standards and fixes them under the same sign-off and validation-report loop.
Schema and change review: L3 Apply with approval
Every change ships with a validation report linked to lineage and waits for sign-off. The report is the receipt; the sign-off is the gate.
ci:teardowns/features/upriver--wave1.md
Pipeline build: L3 Apply with approval
Turn any request into a validated pipeline, with the same plan, execute, validate loop and human sign-off before production.
Catalog and documentation: L2 Propose (claim)
A living map of the full data environment, continuously updated in the background. Documentation is agent-maintained, but no page shows an approval or audit mechanism for doc writes, so the cell is capped at propose and marked claim.
Cross-cloud scope: not evidenced
No page lists the supported warehouses. The teardown describes one warehouse plus one orchestrator plus one repo per deployment. Not evidenced.
Fetched 2026-09-09: The repository row recorded the trust center as positioning only. The live trust center now lists SOC 2 Type 2 as an available document, so the SOC 2 field moves to yes. The /product path returned 404; product detail is taken from the homepage and the repository teardown.
Pages fetched: https://www.upriverdata.com/, https://trust.upriverdata.com/, https://www.upriverdata.com/product (404)
Data engineering agent in the CLI and IDE with column-level lineage and a living knowledge base.
Publishes a price: yes. SOC 2: no. Homepage states pending SOC 2 Type II, ISO 27001, GDPR and HIPAA. trust.alkera.ai resolves (HTTP 200). Licence: Proprietary (public repo holds DataAgentBench traces only). MCP: No MCP server found; distribution is its own CLI and IDE extension with first-party warehouse plugins.
Funding, as publicly reported: Y Combinator Summer 2026; no priced round disclosed.
Incident resolution: L3 Apply with approval
Correctness and data quality issues are triaged and fixed before they reach a report; destructive work waits for approval. Docs: each action stays within the access and change rules you set. Approval gate evidenced; no receipt or rollback documented.
Data quality: L3 Apply with approval
Same triage-and-fix loop with approval on destructive work.
Schema and change review: L3 Apply with approval
Global column lineage so migrations replicate dependencies and no change ships with unknown downstream impact, under the same approval rule. Impact analysis plus gated apply.
Pipeline build: L3 Apply with approval
New pipelines ship in hours on the tools you already run, built to your conventions, with dbt and Airflow orchestration and the same approval rule.
Catalog and documentation: L2 Propose (claim)
Living knowledge base kept current as your pipeline changes, inside the docs you already use. No approval or audit mechanism for doc writes is documented. Capped at propose, marked claim.
Cross-cloud scope: L2 Propose
One agent with plugins for Snowflake, Databricks, BigQuery, Redshift, ClickHouse and eight more, plus dbt and Airflow. Multi-warehouse drafting is evidenced; a gated apply spanning two warehouses is not.
Fetched 2026-09-09: The homepage still states Ranked first on DataAgentBench at 83.28 percent; the repository's July reading of the live board placed Alkera second behind Actioneer's Sentinel. The pricing subdomain the homepage links did not resolve on 2026-09-09; the tiers recorded in July ($0, $50, $250, custom) are retained from the repository row and alkera.ai/pricing.
Pages fetched: https://alkera.ai/, https://docs.alkera.ai/, https://pricing.alkera.ai/ (DNS did not resolve; alkera.ai/pricing returns 200)
Agentic data engineering platform: dbt pipeline build, migration, warehouse optimisation; open-source altimate-code harness.
Publishes a price: yes. SOC 2: yes. SOC 2 listed on the Enterprise tier; the repository's July pass verified a report and pen-test summary are available under NDA via a named process. Licence: MIT (altimate-code); platform proprietary. MCP: Yes; the local extension exposes an MCP endpoint with dbt, SQL, lineage and FinOps tools.
Funding, as publicly reported: $2M (SEC Form D, 2022); no later round found.
Incident resolution: not evidenced
Neither the homepage nor the altimate-code README names incident response or on-call resolution as a capability. Not evidenced.
Data quality: L2 Propose
README: data quality validation and PII detection, SQL anti-pattern detection. The harness drafts findings and tests; applying fixes to production data is not documented.
Schema and change review: L2 Propose
Column-level lineage extraction across dialects and schema-aware impact assessment. Produces the assessment; a human acts on it.
Pipeline build: L3 Apply with approval
Build, test and deploy production-quality dbt pipelines. Builder mode requires approval before dangerous operations, with DROP DATABASE, DROP SCHEMA and TRUNCATE hard-blocked; Analyst mode is read-only. Gate evidenced in the README.
Catalog and documentation: not evidenced
Documentation generation exists in the older dbt Power User extension but was not on any page fetched this pass. Not evidenced; contest with a URL.
Cross-cloud scope: L2 Propose
Thirteen warehouses in one harness including Snowflake, BigQuery and Databricks; optimisation across all three on the homepage. Drafting across platforms is evidenced; applying across two is not.
Fetched 2026-09-09: Consistent. Live pricing adds outcome-based pricing wording on the Enterprise tier. altimate-code shows 809 stars.
Pages fetched: https://altimate.ai/, https://altimate.ai/pricing, https://github.com/AltimateAI/altimate-code
Open-source agentic data engineering agent (CLI, VS Code, MCP server and client) with context learning and the Dosi semantic compiler.
Publishes a price: partial. SOC 2: no. No SOC 2 claim anywhere on the site or pricing page. Enterprise lists governance, sandboxing and approvals. Licence: Apache-2.0. MCP: Yes; both MCP server and MCP client.
Funding, as publicly reported: Undisclosed; no filing found.
Incident resolution: L3 Apply with approval
Proposed fixes for broken components inside shared workflow threads; writes go through execute_sql with explicit confirmation and statement-level SQL authorisation with AI pre-review (README). Per-statement approval is a gate, so L3, though narrower than a change-level approval with a receipt.
Data quality: L1 Suggest
Watches freshness, schema drift and metric anomalies across pipelines and explains them in the thread. Remediation of the data itself is not documented.
Schema and change review: L1 Suggest
Schema drift detection and an agent that understands warehouse, lineage and past failures. Impact explanation; no review verdict or gated apply documented.
Pipeline build: L3 Apply with approval
Plan, write, run, validate, deploy, monitor; SQL authoring and validation, pipeline, report and dashboard generation; Airflow scheduling built in; write and DDL statements stop for confirmation.
Catalog and documentation: L2 Propose
Reads schema and SQL history, then generates OSI semantic models; every run and correction settles into context. Drafts semantic definitions; humans ship them.
Cross-cloud scope: L2 Propose
Nineteen-plus dialects via adapters; Dosi compiles one semantic model into SQL for 13-plus dialects. Cross-platform drafting evidenced.
Fetched 2026-09-09: Stars moved from about 1,407 in July to 1.7k. The pricing page now lists governance, sandboxing and approvals and long-running agents on the Enterprise tier, neither of which the July row recorded.
Pages fetched: https://datus.ai/, https://github.com/Datus-ai/Datus-agent, https://datus.ai/pricing/
Formerly an AI data engineer. Now Postgres database branching for coding agents (clone any Postgres in about six seconds).
Publishes a price: yes. SOC 2: unknown. No security or compliance statement on the homepage or pricing page. Licence: Proprietary (DE-Bench benchmark repo is AGPL-3.0). MCP: No MCP server; coding agents integrate via ardent-cli.
Funding, as publicly reported: $2.15M pre-seed, September 2025 (Crane Venture Partners).
Incident resolution: not evidenced
The product is a database sandbox for other agents. It does not resolve incidents.
Data quality: not evidenced
Let agents clean, deduplicate and standardize data on an exact copy of production describes infrastructure for a third-party agent, not an agent of its own.
Schema and change review: not evidenced
Migration testing on a branch is a capability of the sandbox, not an agentic review.
Pipeline build: not evidenced
Not evidenced.
Catalog and documentation: not evidenced
Not evidenced.
Cross-cloud scope: not evidenced
Postgres only.
Fetched 2026-09-09: Consistent with the repository's pivot note. Live pricing: Starter $0 plus compute and storage, Scale $250 per month, Enterprise custom, compute $0.40 per CU-hour, storage $0.70 per GB-month.
Pages fetched: https://tryardent.com/, https://tryardent.com/pricing
AI data review agent for dbt pull requests: schema diff, data diff, lineage impact, review summary posted to the PR. Open-source core plus Recce Cloud.
Publishes a price: yes. SOC 2: yes. SOC 2 badge on the homepage and pricing page; trust.reccehq.com returns 200. Licence: Apache-2.0 (DataRecce/recce). MCP: Yes; MCP server and a Claude Code plugin in the docs setup guides.
Funding, as publicly reported: $4M, April 2025 (Heavybit).
Incident resolution: not evidenced
Review-time only. No production incident workflow.
Data quality: not evidenced
One-off questions become automated checks, but they run at PR time. Graded under change review, not production data quality.
Schema and change review: L2 Propose
On every PR the agent runs data diffs, identifies impact to the column level, writes a data review summary explaining what matters and what is safe to ignore, and posts it as a PR comment. It produces the review verdict; the human merges. It never applies changes.
Pipeline build: not evidenced
Not evidenced.
Catalog and documentation: not evidenced
Not evidenced.
Cross-cloud scope: not evidenced
Works with Snowflake, BigQuery, Redshift and Databricks, one warehouse per project. No cross-platform review evidenced.
Fetched 2026-09-09: Consistent. Live pricing: Free $0 with 100 agent reviews per month, Team $250 per month annual or $300 monthly with 1,000 reviews, Enterprise custom with SSO, BYOC and RBAC.
Pages fetched: https://reccehq.com/, https://reccehq.com/pricing/, https://docs.reccehq.com/
Agentic data quality, observability and lineage: anomaly detection, agentic root-cause analysis, catalog.
Publishes a price: no. SOC 2: yes. ISO 27001 and SOC 2 Type II stated on the homepage and pricing page; GDPR and HIPAA also stated. Licence: Proprietary. MCP: Not stated.
Funding, as publicly reported: $30M Series A, March 2026 (Plural); $47M total.
Incident resolution: L1 Suggest
Lineage tracking and incident grouping identify where issues originate, so you can debug across pipelines rather than individual tables. Agentic root-cause analysis; no fix applied.
Data quality: L1 Suggest
ML anomaly detection with adaptive thresholds, AI-assisted setup with automatic recommendations. Recommends; does not remediate.
Schema and change review: not evidenced
Schema change detection was not on any page fetched. Not evidenced; contest with a URL.
Pipeline build: not evidenced
Not evidenced.
Catalog and documentation: L0 Alert
Cataloging is listed as a platform capability alongside observability, quality and lineage. Nothing shows agent-written documentation. Graded L0 for surfacing assets and alerts.
Cross-cloud scope: L1 Suggest
Fifteen-plus platforms including Snowflake, BigQuery, Redshift, Databricks, Kafka, S3 and ADLS; RCA groups incidents across pipelines. Explains across platforms; changes nothing.
Fetched 2026-09-09: The repository row carries the vendor claim detect and autonomously resolve. Neither the homepage nor the docs fetched this pass documents any remediation mechanism; the evidenced product is detection, agentic root-cause analysis and lineage. The row's claim is therefore not graded.
Pages fetched: https://validio.io/, https://docs.validio.io/, https://validio.io/pricing
YC W26 profile: the OS for data quality across any stack. Live homepage on 2026-09-09: RouteKit, a local router that sends each coding task to the model that should do it.
Publishes a price: no. SOC 2: unknown. No security statement on the homepage. Licence: Unknown. MCP: Not stated.
Funding, as publicly reported: Y Combinator Winter 2026 (YC API); no priced round found.
Incident resolution: not evidenced
No data engineering product on the live site.
Data quality: not evidenced
The YC profile describes automated data-quality monitoring and enforcement; no product page documents it. Not evidenced.
Schema and change review: not evidenced
Not evidenced.
Pipeline build: not evidenced
Not evidenced.
Catalog and documentation: not evidenced
Not evidenced.
Cross-cloud scope: not evidenced
Not evidenced.
Fetched 2026-09-09: Disagreement. The repository row describes data-quality resolution with governed contract write-back, and the YC profile still says data quality. The live homepage sells RouteKit, a model router for coding agents, and does not mention data quality, lineage or contracts. Graded on the live page.
Pages fetched: https://velum-labs.com/, https://www.ycombinator.com/companies/velum-labs
YC S24 profile: replace the humans fixing bad data. Live homepage on 2026-09-09: a platform for publishing, reviewing and installing employee-built AI agents across ChatGPT, Claude and Cursor.
Publishes a price: no. SOC 2: unknown. No security statement on the homepage. Licence: Proprietary. MCP: Not stated.
Funding, as publicly reported: Undisclosed; Y Combinator Summer 2024.
Incident resolution: not evidenced
The YC profile's autonomous resolution of pipeline errors is not on any product page. Not evidenced.
Data quality: not evidenced
Not evidenced.
Schema and change review: not evidenced
Not evidenced.
Pipeline build: not evidenced
Not evidenced.
Catalog and documentation: not evidenced
Not evidenced.
Cross-cloud scope: not evidenced
Not evidenced.
Fetched 2026-09-09: Disagreement. The repository row describes an autonomous data-fix agent at about 30 percent of manual cost. The live homepage describes agent sharing and governance (publish, admin review, install, run, one-click rollback of agent versions) with no data pipeline capability. The YC profile still carries the data-fix description. Graded on the live page.
Pages fetched: https://usemica.com/, https://www.ycombinator.com/companies/mica-ai
Enterprise AI agent platform for operational work: conversational analytics over approved definitions, monitoring agents on business metrics, operational playbooks with approvers.
Publishes a price: no. SOC 2: unknown. Footer badges assert SOC 2, ISO/IEC 27001:2022, ISO/IEC 27701:2019 and GDPR. trust.actioneer.com has no DNS record and actioneer.com/security returns 404 (checked 2026-09-09). Asserted, not verifiable from public pages. Licence: Proprietary (can run on open-source models on-premise). MCP: Not mentioned.
Funding, as publicly reported: $5.4M pre-seed, August 2025, raised as GameRamp (BITKRAFT Ventures).
Incident resolution: L1 Suggest
Monitoring agents watch metrics and thresholds; the moment something moves, they detect it, trace it to the source, and flag what changed. Business-metric incidents, not pipeline incidents. Playbooks carry a decision to an outcome with a named approver and a recorded reasoning trail, which would be L3 for operational actions, but no page shows a playbook repairing a data pipeline.
Data quality: not evidenced
Not evidenced.
Schema and change review: not evidenced
Not evidenced.
Pipeline build: not evidenced
Not evidenced.
Catalog and documentation: not evidenced
Answers are grounded in approved data definitions in a context store; the store is human-approved. Not evidenced as agent-written documentation.
Cross-cloud scope: L1 Suggest
Snowflake, BigQuery and SaaS sources in one platform; read-only by default with write access scoped per capability. Explains across platforms.
Fetched 2026-09-09: The July row recorded a bare SOC 2 certified claim; the live footer adds ISO 27001 and ISO 27701 badges, still with no reachable trust page. Benchmark line now reads #1 on DABstep, SOTA on KramaBench and DataAgentBench.
Pages fetched: https://actioneer.com/, https://actioneer.com/resources
Open-source analytics agent built for context engineering: text-to-SQL chat over a file-system context, nao MCP, self-host with your own key.
Publishes a price: partial. SOC 2: yes. SOC 2 and AICPA badges; SOC 2 Type II reports on the Enterprise tier; compliance.getnao.io returns 200. Licence: Apache-2.0 with an enterprise carve-out (SSO, RLS, impersonation behind a licence key). MCP: Yes; nao MCP reachable from Claude, Codex, Cursor, Slack, Teams, WhatsApp, Telegram, and external MCPs can be added to the agent.
Funding, as publicly reported: About $500K, July 2025 (Kima Ventures, Sapienta VC, Y Combinator).
Incident resolution: not evidenced
Analytics only. No incident workflow.
Data quality: not evidenced
Not evidenced.
Schema and change review: not evidenced
Not evidenced.
Pipeline build: not evidenced
Not evidenced.
Catalog and documentation: not evidenced
Context is human-authored files in git (metadata, docs, tools, MCPs) that the agent reads. The agent does not write documentation. Not evidenced.
Cross-cloud scope: L1 Suggest
BigQuery, Snowflake, Postgres, Databricks, DuckDB, MotherDuck and Redshift; read-only answers with transparent reasoning.
Fetched 2026-09-09: Consistent. GitHub shows 1.6k stars. Free self-host with your own key; Cloud and Enterprise are custom and unpriced.
Pages fetched: https://getnao.io/, https://getnao.io/pricing/, https://github.com/getnao/nao
Ataccama ONE: data quality, observability, catalog, lineage, governance and MDM, now with a ONE AI Agent (digital data steward) and an MCP server.
Publishes a price: no. SOC 2: yes. Public trust center; the repository's July pass recorded ISO 27001:2022, ISO 9001 and an annual SOC 2 Type II. Licence: Proprietary. MCP: Yes; a named platform component distributing trust signals and metadata to agents, copilots and LLMs.
Funding, as publicly reported: $150M minority growth, 2023 (Bain Capital Tech Opportunities).
Incident resolution: not evidenced
No incident or on-call workflow on the pages fetched. Data-quality remediation is graded in the quality lane.
Data quality: L4 Autonomous with receipts (claim)
The ONE AI Agent autonomously generates data quality rules, suggests where to apply them, finds anomalies, and executes fixes at the source. Unattended execution is stated; no page documents a per-change audit receipt, an approval policy or rollback. Graded at the stated level and marked claim.
Schema and change review: not evidenced
Not evidenced on the pages fetched.
Pipeline build: not evidenced
Not evidenced.
Catalog and documentation: L2 Propose
Auto-generated asset descriptions and table description generation and editing; auto-generated rules suggested for placement. Drafts that stewards edit.
Cross-cloud scope: L2 Propose
One platform over Snowflake, Databricks, BigQuery and Redshift; descriptions and rules are drafted across sources. Cross-source unattended fixes inherit the claim mark from the quality lane, so the cross-cloud cell stops at propose.
Fetched 2026-09-09: Consistent. Pricing page describes three dimensions (named users, data objects, active DQ configurations) with no figures. Four pages were fetched for this vendor, one over the pass budget, because the MCP and AI pages are separate.
Pages fetched: https://www.ataccama.com/, https://www.ataccama.com/ai, https://www.ataccama.com/platform/mcp, https://www.ataccama.com/pricing
The context layer for enterprise AI: catalog, governance, Context Engineering Studio with AI-bootstrapped definitions certified by humans, Atlan MCP server.
Publishes a price: no. SOC 2: unknown. No SOC 2 statement on the homepage, pricing page or Context Engineering Studio page fetched this pass. Atlan publishes a trust portal elsewhere; not verified in this pass. Licence: Proprietary (MIT-licensed client SDKs). MCP: Yes; hosted Atlan MCP server behind OAuth, and Context Repos expose context over MCP.
Funding, as publicly reported: $105M Series C, May 2024; about $206M total.
Incident resolution: not evidenced
Not evidenced.
Data quality: not evidenced
Not evidenced on the pages fetched.
Schema and change review: not evidenced
Impact analysis is a catalog feature elsewhere in the product; not on the pages fetched. Contest with a URL.
Pipeline build: not evidenced
Not evidenced.
Catalog and documentation: L3 Apply with approval
Description Generator, Term Linkage, Metrics Generator and Ontology Generator draft context; domain experts review AI-bootstrapped definitions, resolve conflicts and approve context for production; git-like versioning with full history, staged rollouts and rollbacks. Draft, named approval, apply, rollback: L3.
Cross-cloud scope: L2 Propose
Eighty-plus connectors across Snowflake, Databricks, BigQuery, Redshift, BI and business systems in one graph; context drafted across all of them. The approval flow is documented per repo, not shown spanning platforms, so the cross-cloud cell stops at propose.
Fetched 2026-09-09: Consistent. Pricing page publishes no figures and routes to sales.
Pages fetched: https://atlan.com/, https://atlan.com/pricing/, https://atlan.com/context-engineering-studio/
Lakeflow data engineering with Genie Code agents and Lakeflow Designer, Unity Catalog governance, Genie One and Genie Ontology.
Publishes a price: yes. SOC 2: yes. Trust page states certifications and attestations for regulated industries with a compliance sub-page; specific reports not quoted on the page fetched. Licence: Proprietary platform; Unity Catalog OSS is Apache-2.0. MCP: Yes; managed MCP and Genie One expose MCP per the repository row (Data+AI Summit, June 2026).
Funding, as publicly reported: Private; $134B valuation (Series L, completed February 2026).
Incident resolution: L2 Propose
Genie Code: agents that understand your data and can author, maintain and troubleshoot data pipelines. Drafts the fix in the workspace. ZeroOps unattended remediation is announced but undocumented on the pages fetched.
Data quality: L0 Alert
Full visibility into pipeline health with real-time metrics; custom alerts guarantee you know exactly when issues occur. Alerting. Declarative expectations are a platform rule, not an agent.
Schema and change review: not evidenced
Schema evolution is a pipeline feature; no agentic change review evidenced.
Pipeline build: L2 Propose
Genie Code builds pipelines from natural language; Lakeflow Designer is AI-first authoring. The user runs and deploys.
Catalog and documentation: L2 Propose
Genie Ontology auto-extracts a learned context layer (snippets with source and definition) on top of Unity Catalog; each snippet is inspectable. Learned drafts consumed by Genie; human curation model not detailed.
https://www.databricks.com/blog/introducing-genie-one-genie-ontology-and-genie-agents · ci:teardowns/features/databricks--genie-one-ontology-agents-regate-2026-07.md
Cross-cloud scope: not evidenced
Runs on AWS, Azure and GCP but operates one lakehouse. Not evidenced across warehouses.
Fetched 2026-09-09: The repository row records Lakeflow ZeroOps (autonomous remediation) from the June summit. No page fetched this pass documents ZeroOps or its approval and audit model, so incident resolution is graded on Genie Code's troubleshooting wording only. Contest with a docs URL.
Pages fetched: https://www.databricks.com/product/data-engineering, https://www.databricks.com/trust, https://www.databricks.com/product/lakeflow (404)
Cortex Code (CoCo): an agent inside Snowflake for data engineering, analytics and ML, in Snowsight, Desktop, CLI and editors. TensorStax's autonomous data engineering agent was acquired in February 2026.
Publishes a price: yes. SOC 2: yes. Public trust center at trust.snowflake.com (body not readable by the fetcher); Snowflake is a listed company with published SOC 2 and ISO reports. Licence: Proprietary; Apache Polaris is Apache-2.0. MCP: Yes; CoCo CLI implements MCP (preview) and Snowflake managed MCP is GA.
Funding, as publicly reported: Public company (NYSE: SNOW).
Incident resolution: L2 Propose
Take actions and answer questions about credit consumption, query performance, governance and user permissions; explain and optimise existing queries. Drafts fixes inside the account; no incident workflow with a gate is documented.
https://docs.snowflake.com/en/user-guide/cortex-code/cortex-code
Data quality: not evidenced
Not evidenced on the pages fetched.
Schema and change review: not evidenced
Not evidenced.
Pipeline build: L3 Apply with approval
Create data pipelines, ML models, apps and agents in plain language, with dbt, Airflow and Spark support; generated ML pipelines run in Notebooks. The CLI has a three-tier approval system and inherits Snowflake RBAC. Gate evidenced, tiers undocumented.
https://docs.snowflake.com/en/user-guide/cortex-code/cortex-code
Catalog and documentation: L1 Suggest
CoCo reads your catalog, lineage and RBAC policies so generated code references real objects; semantic catalog search built in. Consumes the catalog; writing documentation is not evidenced.
Cross-cloud scope: not evidenced
Snowflake-native. Not evidenced across warehouses.
Fetched 2026-09-09: Consistent. Neither page names TensorStax; the acquisition is recorded in the repository row (announced 2026-02-04). Docs describe a three-tier approval system for the CLI without detailing the tiers.
Pages fetched: https://www.snowflake.com/en/product/features/cortex-code/, https://docs.snowflake.com/en/user-guide/cortex-code/cortex-code, https://trust.snowflake.com/
Gemini-powered data engineering agent in BigQuery Studio and Dataform, plus Dataplex Universal Catalog with a first-party MCP server.
Publishes a price: yes. SOC 2: yes. Google Cloud publishes SOC 2 and ISO reports through its compliance centre; not re-fetched this pass. Licence: Proprietary. MCP: Yes; first-party Dataplex MCP server (read and write on data products, assets and aspects) per the repository row.
Funding, as publicly reported: Alphabet product line.
Incident resolution: L2 Propose
Failure troubleshooting reads execution logs, root-causes failures and proposes fixes; Planning Mode emits an explicit plan for human review before acting. Proposes; human applies.
https://docs.cloud.google.com/gemini/data-agents/data-engineering-agent/agent-overview · ci:teardowns/features/bigquery-data-engineering-agents--wave-2.md
Data quality: L1 Suggest (claim)
Auto-generated data quality scorecards and assertions derived from Dataplex rules are stated as an output of pipeline generation; the teardown marks this medium confidence, not verified verbatim in primary docs. Marked claim.
ci:teardowns/features/bigquery-data-engineering-agents--wave-2.md
Schema and change review: not evidenced
Not evidenced.
https://docs.cloud.google.com/gemini/data-agents/data-engineering-agent/agent-overview
Pipeline build: L2 Propose
Natural-language request to Dataform/SQL pipeline code; plans, generates, self-compiles and fixes its own compilation errors, then hands to a human. Explicitly human-gated: it does not deploy on its own.
https://docs.cloud.google.com/gemini/data-agents/data-engineering-agent/agent-overview
Catalog and documentation: L2 Propose
Dataplex Universal Catalog generates metadata and insights and exposes a read-write MCP server; the fetched page was truncated, so the mechanism for human review is not confirmed. Graded propose.
Cross-cloud scope: not evidenced
BigQuery and Dataform only for the agent. Not evidenced across warehouses.
Fetched 2026-09-09: The BigQuery docs path recorded in the repository moved under docs.cloud.google.com/gemini/data-agents/ (HTTP 200 on 2026-09-09). Evidence is taken from the repository teardown, which captured that page in June, with the live URL cited.
Pages fetched: https://cloud.google.com/bigquery/docs/data-engineering-agent (301, target 404), https://cloud.google.com/dataplex (truncated), https://docs.cloud.google.com/bigquery/docs/data-engineering-agent-intro (404)
Agentic GenBI on an open context engine (MDL semantics), governed text-to-SQL, GenBI apps, MCP.
Publishes a price: yes. SOC 2: yes. SOC 2 Type II ticked on every cloud tier of the pricing page. Licence: Apache-2.0 (Canner/WrenAI engine; open core). MCP: Yes; MCP listed as a first-class consumption pattern.
Funding, as publicly reported: Last known $3.5M pre-A, March 2022 (Taiwania Capital).
Incident resolution: not evidenced
Not evidenced.
Data quality: not evidenced
Not evidenced.
Schema and change review: not evidenced
Not evidenced.
Pipeline build: not evidenced
Not evidenced.
Catalog and documentation: not evidenced
The context layer (MDL) is human-authored; the agent learns Knowledge, Skills and Memories for answering. Agent-written catalog documentation is not evidenced.
Cross-cloud scope: L1 Suggest
Twenty-plus connectors including Snowflake, BigQuery, Databricks and Redshift; read-only governed answers with the SQL behind them.
Fetched 2026-09-09: Consistent. Live tiers: Free $0 (20 credits), Essential $179 per month, Enterprise $559 per month, Enterprise Plus custom, open source free.
Pages fetched: https://getwren.ai/, https://getwren.ai/pricing, https://docs.getwren.ai/cp/overview
Open-source catalog repositioned as the open context layer for AI agents: knowledge graph, memory, data quality, lineage, incident manager, MCP server with read and write.
Publishes a price: partial. SOC 2: unknown. No SOC 2 statement on the homepage or README; Collate's cloud may publish one separately. Licence: Apache-2.0. MCP: Yes; MCP server for direct agent-to-context integration, read and write on the same auth engine as the APIs.
Funding, as publicly reported: $10M Series A, July 2025 (Venrock) via Collate.
Incident resolution: L0 Alert
Incident management and governance workflows, freshness checks and alerts. Human-run incident workflow over alerts.
Data quality: L0 Alert
Tests, profiling, freshness checks and alerts. Collate's test-generating agents were not on the pages fetched (blog returned 403).
Schema and change review: not evidenced
Not evidenced on the pages fetched.
Pipeline build: not evidenced
Not evidenced.
Catalog and documentation: L3 Apply with approval
Agent automation and AI workflows write metadata through the MCP server (read and write); Memory records corrections, decisions, approvals, feedback loops, and an auditable record of every change. Agent writes, approvals and an audit record on catalog changes: L3.
Cross-cloud scope: L1 Suggest
One graph over 130-plus connectors with cross-platform lineage. The graph explains across platforms; agentic writes across platforms are not separately evidenced.
Fetched 2026-09-09: Consistent. GitHub shows 15.2k stars and 130-plus connectors.
Pages fetched: https://open-metadata.org/, https://github.com/open-metadata/OpenMetadata, https://blog.open-metadata.org/announcing-openmetadata-2-0-the-open-context-layer-for-ai-agents-83b8ce8b9dde (403)
Data and AI observability with a Monitoring Agent, a Troubleshooting Agent and an MCP and agent toolkit.
Publishes a price: no. SOC 2: yes. trust.montecarlodata.com is linked from the footer (returned 403 to a plain fetch); the repository row records SOC 2 as published. Licence: Proprietary (apollo-agent collector is open source). MCP: Yes; MCP and agent toolkit under Agentic Operations, reachable from Claude Code.
Funding, as publicly reported: $236M total; $1.6B valuation (2022).
Incident resolution: L1 Suggest
Works through hundreds of hypotheses across warehouse query logs, dbt, Airflow, Databricks Workflows and GitHub or GitLab PRs and highlights the most likely cause. Human triggers it and acts on it; no fix applied.
Data quality: L1 Suggest
Monitoring plan and Monitoring Agent recommend monitors; anomalies alert. Recommends; does not remediate.
Schema and change review: not evidenced
Schema-change monitors exist in the product but were not on the pages fetched this pass. Not evidenced; contest with the docs URL.
Pipeline build: not evidenced
Not evidenced.
Catalog and documentation: not evidenced
Not evidenced.
Cross-cloud scope: L1 Suggest
Root cause traced across warehouse, dbt, Airflow and code hosting in one investigation. Explains across systems.
Fetched 2026-09-09: montecarlodata.com now redirects to montecarlo.ai. Pricing is request-only. Otherwise consistent.
Pages fetched: https://montecarlo.ai/, https://docs.getmontecarlo.com/, https://docs.getmontecarlo.com/docs/troubleshooting-agent
Data Diff, CI deployment testing on PRs, monitors, a Migration Agent delivered as a service, and a Data Knowledge Graph served over MCP.
Publishes a price: no. SOC 2: yes. SOC 2 and HIPAA compliance stated on the homepage; security.datafold.com returns 200. Licence: Proprietary (data-diff OSS repo archived). MCP: Yes; Data Diff, monitors and the Data Knowledge Graph exposed via MCP for Claude Code, Cursor and Windsurf.
Funding, as publicly reported: About $26M total (NEA, Amplify, YC).
Incident resolution: not evidenced
Not evidenced.
Data quality: L0 Alert
Monitors send alerts when data diffs fall outside predefined ranges; ML anomaly detection. Alerting.
Schema and change review: L1 Suggest
Deployment testing diffs every PR against production at value level before merge. Explains what changed; no review verdict is drafted and nothing is applied.
Pipeline build: L2 Propose
Migration Agent translates legacy code to the target with automated validation and humans reviewing what falls out, delivered as a guaranteed-outcome service. Drafts; humans review.
Catalog and documentation: L1 Suggest
Data Knowledge Graph collects lineage, business logic, usage, BI connections and git history and serves it to agents over MCP. Explains; does not write documentation.
Cross-cloud scope: L1 Suggest
Cross-database diffing compares datasets across warehouses at value level. Explains across platforms.
Fetched 2026-09-09: Consistent. Pricing redirects to a contact form.
Pages fetched: https://www.datafold.com/, https://docs.datafold.com/welcome, https://www.datafold.com/pricing (redirects to contact)
The Autonomous Agentic Data Platform: Data-Agents Swarm (20 specialised agents), Data Context Wizard (cross-cloud semantic and knowledge graph, provenance-stamped), Autonomous Data-Conductor (detect, diagnose, fix, review, verify with an autonomy dial), Spellbook Data Catalog (agent-written catalog and human control plane, preview). Runs inside Claude Code, Cursor, Codex CLI, OpenCode and any MCP client.
Publishes a price: yes. SOC 2: no. Not SOC 2 certified. llms.txt states architected for SOC 2 and that data stays inside your own infrastructure. No trust center. Licence: Apache-2.0 core (11 agents, 160-plus MCP tools in the community edition). MCP: Yes; every agent is exposed as MCP tools; runs in Claude Code, Cursor, Codex CLI, OpenCode, Gemini and any MCP client.
Funding, as publicly reported: Early stage. No public customer references. Pre-revenue.
Incident resolution: L3 Apply with approval
Agents draft the change, run it in a sandboxed dry-run, and hand it to a named human for approval; every change ships with a receipt (what changed, why, blast radius across downstream models and dashboards) and one-click rollback. Gate, dry-run and receipt are all stated. The Conductor's unattended mode (autonomy dial) would be L4 but no public page documents the policy under which changes are applied unattended, so L3.
https://dataworkers.io/llms.txt · https://dataworkers.io/product/autonomous-data-conductor/
Data quality: L3 Apply with approval
Quality monitoring agent detects; fixes go through the same propose, dry-run, approve, receipt path. No public evidence of unattended quality fixes.
Schema and change review: L3 Apply with approval
Data Change Review and Schema Evolution agents; the receipt carries blast radius across downstream models and dashboards; apply only after a named human approves. The strongest publicly documented lane.
Pipeline build: L2 Propose
The community edition drafts pipelines but states that pipeline generation requires Pro, and the Pro path is not publicly documented beyond the llms.txt claim. Graded on what a stranger can verify: propose.
https://github.com/DataWorkersProject/dataworkers-claw-community
Catalog and documentation: L2 Propose
Spellbook is an agent-written catalog with a human control plane, in preview. Agents draft descriptions and the human plane approves; the approval and audit path for catalog writes is not yet documented publicly, and the product is preview. Propose.
Cross-cloud scope: L2 Propose
Fifteen catalog connectors including Snowflake, BigQuery, Databricks, Glue, Hive Metastore and Iceberg feeding one provenance-stamped graph. Cross-cloud drafting is evidenced; a gated apply spanning two warehouses is not shown publicly. Propose, the same cap applied to Alkera, Altimate and Atlan.
https://github.com/DataWorkersProject/dataworkers-claw-community
Fetched 2026-09-09: The community repository shows 12 stars and states that pipeline generation and model training require Pro; the community edition runs on in-memory stubs by default. Published pricing: free core, $7,500 pilot, from $1,000 per month, no metering, unlimited seats. The site carries a benchmark figure without a public run record; this index does not cite it (see Forthcoming).
Pages fetched: https://dataworkers.io/, https://dataworkers.io/llms.txt, https://github.com/DataWorkersProject/dataworkers-claw-community
Send a URL to hello@dataworkers.io with the vendor, the lane and the level you believe is evidenced. A cell moves when the URL documents the mechanism; a vendor statement that a feature exists is recorded as a claim until the docs show it. Upheld contests are credited in the next edition's changelog.
Contests are reviewed within ten business days. The vendor's own docs outrank third-party coverage; a dated, reachable page outranks a slide.
Data Workers intends to publish benchmark results with a public run record, and technical white papers on the receipt format and the cross-cloud context graph. None of that material is cited in this edition and no figure from it appears here. When published, it will be graded under the same rules as every other vendor's docs and the relevant cells will move only if the documents show the mechanism.
The dataset (data.json, served at https://dataworkers.io/agentic-data-engineering-index.json) is released under CC BY 4.0. Cite as: Data Workers, The Agentic Data Engineering Index, Q4 2026.
If you are choosing
Read the L3 cells first: those are the vendors whose agents apply a change behind a documented approval gate. Then ask each one to show the receipt on a change in your stack.
Data Workers will do that in a 45-minute working session: book a time, or start with the open-source core and grade us yourself.