Product
Product11 min readBy The Data Workers Team

You're on Dremio: It Accelerates Your Iceberg Lakehouse. Data Workers Keeps What It Accelerates True

Dremio, now part of SAP, makes your Iceberg lakehouse fast. Data Workers catches a Reflection serving a day-old answer, traces the cause and gets the fix approved.

Your analysts open Tableau or a notebook, and the query lands on Dremio. Underneath sit Apache Iceberg tables on object storage, registered in an Apache Polaris catalog, organized into spaces and views your team models with dbt-dremio. Reflections, Dremio's precomputed copies of tables and aggregations, keep the dashboards fast without copying data into a warehouse. You run Dremio Cloud, Dremio Enterprise on Kubernetes or the free Community Edition, and your agents reach the lakehouse through Dremio's MCP server. Dremio calls itself "The Agentic Lakehouse for AI and Analytics", and since July 6, 2026 it is part of SAP.

That ownership change is good news for the open stack you built on. SAP says that with Dremio, SAP Business Data Cloud "will become an Apache Iceberg-native enterprise lakehouse that unifies SAP and non-SAP data", and that it is "fully committed to continuing to invest in and prioritize" Iceberg, Polaris and Arrow. Dremio's founders put it plainly: "Dremio isn't going anywhere; we are doubling down on our agentic vision." What no query engine owns is whether a fast answer is still a current one after something upstream moved. That is the job Data Workers does on top of Dremio, with a named person approving every fix.

Key takeaways

  • •Dremio keeps its job. Spaces, views, Reflections, engines and access control stay with your lakehouse team. Data Workers owns whether what Dremio serves is right, and who reads it.
  • •A stale Reflection becomes an incident with an owner. When a served number falls behind the lake, Data Workers ties it to its cause and lists every report and job that reads it.
  • •It reads the systems around Dremio natively. Airflow, dbt, Tableau and Slack connect natively among 50+ connectors. Dremio connects over its API or MCP server today, and its MCP server sits next to Data Workers in your team's client.
  • •Every fix goes through a named person. The owner approves each proposal and makes the Dremio change; Data Workers verifies the result and writes a receipt.
  • •The SAP move changes nothing here. Data Workers builds on the open Iceberg stack SAP has committed to. For the SAP side, read you're on SAP Business Data Cloud.

Dremio accelerates the lakehouse. Data Workers keeps what it accelerates true.

The job on top of Dremio is to notice that a fast answer is now an old answer, find out why, size who is reading it, hold anything that acts on it, get the fix made where it lives and prove the result. Here is one Tuesday morning at a regional home-goods retailer. This is an illustration, not a customer case.

The setup: point-of-sale lines land nightly in the Iceberg table pos.sales_lines on Amazon S3, registered in Apache Polaris, through a Spark job orchestrated by Airflow 2.10. A dbt-dremio model builds the view trading.sales_7d, rolling seven-day sales by store and product. A manually managed aggregation Reflection keeps it fast, refreshed on a schedule at 06:00, with a three-day expiration. Regional managers read it in the Tableau workbook Store Trading, declared as a dbt exposure and recorded in the context graph. At 08:00 the Airflow DAG markdown_recommendations reads the view and sends price cuts for slow sellers to the stores; the pricing team recorded that reader in the Data Workers context graph. A small DAG, serving_checks, queries the view at 07:00 and records two numbers with monitor_metrics: the latest sale date it returns, and the seven-day total.

TimeSystemWhat happens
Mon 23:40Airflow + SparkThe weekly iceberg_maintenance DAG compacts pos.sales_lines and, as approved last week, evolves its partition spec from daily to hourly on sold_at
Tue 05:20Airflow + PolarisThe nightly load commits Monday's sales as a new snapshot. The lake is current
06:00DremioThe scheduled refresh starts. As Dremio documents, when the anchor table's partition scheme has changed to be incompatible with the Reflection's and data has changed, "a full refresh is performed". The full rebuild of three years of sales lines fails on the small engine that runs refresh jobs, and Dremio retries on its documented backoff
06:00 to 07:00DremioBetween retries, the last good materialization, from Monday 06:00, stays available until its expiry on Thursday. Dremio takes a Reflection out of acceleration only when it reaches Failed status after multiple attempts, so the planner keeps choosing it: trading.sales_7d answers fast, without Monday
07:00TableauStore Trading opens. Seven-day sales are down 14% against last year across every store
07:05Data Workersserving_checks records its numbers. Two anomalies at once: the latest served sale date is two days old against a baseline of one, and the seven-day total is 14% below its range. Data Workers opens an incident
07:12Data Workers + Airflowdiagnose_incident separates the lake from the serving layer. Airflow task instances show the 05:20 load succeeded, so the table holds Monday, and that iceberg_maintenance ran its partition task at 23:40, matching last week's approved change in the context graph. trace_cross_platform_lineage follows the dbt manifest from pos.sales_lines to trading.sales_7d, and the context graph adds the Store Trading workbook. blast_radius_analysis adds markdown_recommendations: with Monday missing, about 1,100 products would cross the slow-seller threshold at 08:00
07:15Slacksend_slack_alert reaches the lakehouse owner, with the pricing data owner and the reporting lead copied
07:20Claude Code + Dremio MCPThe lakehouse owner asks Dremio's MCP server, in the same client as Data Workers, about the Reflection: its refresh jobs in sys.jobs failed with out-of-memory errors, the refresh decision was full, and the Monday materialization is still in use
07:35SpellbookShe reviews three proposals: hold markdown_recommendations; match the Reflection's partition field to the table's new hourly spec, as Dremio's docs recommend, and run the full refresh on a larger engine; and a diff to the maintenance DAG's runbook that lists the Reflections on a table before its partition spec changes. She approves the first two now and the third for review
07:40AirflowThe pricing owner pauses markdown_recommendations
07:45DremioThe lakehouse owner updates the Reflection's partitioning and starts a refresh on a larger engine
08:30DremioThe refresh completes
08:35Data Workersserving_checks reruns: the served sale date is back on its baseline and the seven-day total back in range. The receipt records cause, approvals, the Reflection change, the checks and the undo (the previous partitioning, for the owner to restore)
08:45AirflowOn the pricing owner's approval, Data Workers queues the markdown_recommendations run through Airflow with trigger_airflow_dag, and she resumes the schedule. Stores get price cuts based on a full week
Incident timeline across the stack: what Dremio, your team and Data Workers each do, step by step

Every part of Dremio did its job. Knowing that this one failed refresh mattered by 08:00 took facts outside Dremio: that the partition change came from a maintenance DAG, that the view feeds a pricing job, and that the pricing job acts on the stores within the hour.

JobWhat Dremio doesWhat Data Workers does
The queryPlans each query on Iceberg and picks the Reflection that satisfies it at the lowest costChecks the numbers people read against the baselines your team records
The ReflectionRefreshes on its policy, falls back to a full refresh when partitions no longer match, retries failuresTies a stale answer to the change behind it and to the readers it reaches
The catalogReads and writes Iceberg tables through Open Catalog, powered by Apache PolarisKeeps the table, its views, readers and owners in one context graph
Agent accessServes the lakehouse to agents over its MCP server, with the user's own permissionsProposes fixes to a named owner and dates the window a stale number was served
The fixRuns the Reflection change, engine and refresh the owner makesProposes the hold and the change in order, with the blast radius
The proofShows refresh status and job historyRe-checks the served numbers and writes a receipt: what changed, who approved it, how to undo it

Why doesn't Dremio just do this itself?

Dremio is built to answer queries on open data fast and correctly from what it holds, and it does that very well. A Reflection is "a precomputed and optimized copy of source data or a query result". Whether the business can still act on an answer after a maintenance job, a late load or a model change is a question about the orchestrator, the pricing schedule and the stores, none of which a query engine has reason to own.

Dremio's AI work follows the same clear line. The AI Agent connects "any agent to your enterprise data", the AI Semantic Layer is a "source of truth for agents", and Autonomous Reflections accelerate queries from observed patterns, used "only when fully synchronized with their source data". Dremio's MCP server, hosted with OAuth or self-hosted from the open-source dremio-mcp project, lets agents inspect schemas, query system tables and run SQL, with view creation off by default, and "agents operate with the same permissions as the authenticated user." That is the right scope for an engine vendor, now inside a larger SAP plan: SAP will deliver "a universal, open catalog built on Apache Polaris and the open Apache Iceberg REST Catalog API" as the discovery and semantic layer of Business Data Cloud.

Owning whether the data is right, from the Spark job to the price tag on the shelf, is a different product with a different liability: a context graph across systems, blast-radius scoping, named approvals, a recorded undo and receipts an auditor can read. That is Data Workers. More in is it safe to let AI agents change production data.

Every tool owns a slice. Data Workers covers the whole lifecycle

Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the Dremio lakehouse already there.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Dremio goes deep on its own area
StageData WorkersDremioWhy we scored it this way
Catalog & Context97Open Catalog on Apache Polaris and the AI Semantic Layer give engines and agents shared definitions. Data Workers keeps one governed context graph from the table to the dashboard and the job that reads it.
Analytics & Insights88.5Dremio's home stage: an Arrow-based engine and Reflections answer BI and agent queries fast on Iceberg. Data Workers answers data questions from governed definitions with lineage behind every number.
Data Quality84Reflections return what their last refresh holds. Data Workers checks the numbers people read against recorded baselines and ties a change to its cause.
Observability & Incidents8.53Job history and refresh status show how Dremio runs. Data Workers opens one data incident across the lake, the engine, the orchestrator and BI, and verifies the fix.
Pipelines & Ingestion8.55Dremio queries data in place and runs Reflection refreshes. Data Workers reads the Airflow runs that write the tables and queues approved reruns.
Schema & Migration85Iceberg schema and partition evolution are supported in place. Data Workers maps a change to the dbt models, views and readers built on it and generates migrations with rollback SQL for the owner.
Governance & Access8.56Fine-grained access control travels with the data across sources. Data Workers routes every data change to a named approver.
Security & Privacy85Permissions follow the user, including agents over MCP. Data Workers leaves a receipt on every data change.
Cost / FinOps86Autonomous Reflections and engine sizing keep query compute down. Data Workers attributes Snowflake spend to the dbt model and reads BigQuery and AWS spend.
MLOps & Models7.53Dremio serves data to agents and notebooks. Data Workers keeps the data under models and forecasts healthy.

How Dremio and Data Workers work together

How Data Workers fits with Dremio: your coding agent on top, Data Workers in the middle, your estate underneath

Data Workers reads the systems around Dremio natively: the dbt-dremio manifest for lineage, Airflow DAG runs and task instances, and Tableau and Looker (read). Slack carries alerts, and approvals land in Spellbook Data Catalog (in preview), where the approver signs in through your identity provider. Dremio and its Open Catalog connect over Dremio's API or MCP server today: the MCP server runs in your team's client, next to Data Workers, and the people using that client query Dremio through it.

Spaces, views, Reflections, engines, refresh policies and grants belong to the lakehouse owner, in Dremio or in dbt-dremio. Data Workers proposes the change as a diff for the owner to merge, or as a written change for the owner to make in Dremio, scopes what it touches and verifies afterwards.

The Data Workers side follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry. Keep Dremio's view creation off (the default) and use a personal access token or OAuth identity with read access.

# Example: Dremio's self-hosted MCP server plus Data Workers agents in Claude Code
claude mcp add dremio -- uv run --directory /path/to/dremio-mcp dremio-mcp-server run

claude mcp add dw-incidents -- /path/to/dataworkers-claw-community/start-agent.sh dw-incidents
claude mcp add dw-context-catalog -- /path/to/dataworkers-claw-community/start-agent.sh dw-context-catalog
claude mcp add dw-connectors -- /path/to/dataworkers-claw-community/start-agent.sh dw-connectors

List the tools with your client's own command (/mcp in Claude Code). In this incident: monitor_metrics and diagnose_incident (dw-incidents) score the served date and total against their baselines and name the cause, and get_incident_history shows whether the view has failed this way before; trace_cross_platform_lineage and blast_radius_analysis (dw-context-catalog) map lineage and readers; send_slack_alert and trigger_airflow_dag (dw-connectors) reach the owners and queue the approved run on Airflow 2.

In production the agents run in your infrastructure and hold the credentials and model key; your data stays in your systems, and the hosted Conductor sees workflow metadata only. More in where does our data go and read-only warehouse access for LLM agents.

One incident, L0 to L4, set per domain:

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. A regional manager calls the data team at 09:30, after the price cuts went out.
  • •L1 observe. Data Workers flags the stale served date at 07:05 with the cause and the readers at risk. Nothing changes.
  • •L2 propose. Data Workers proposes the hold and the Reflection change. Nothing reaches production until the named owner approves; an unanswered request expires and escalates, never auto-grants.
  • •L3 act reversibly. For a class with a clean record, Data Workers carries out reversible steps in the domain you open, such as queuing an approved rerun.
  • •L4 autonomous. For a scoped domain, Data Workers runs the reruns in classes the team approved in advance without waiting for a click, and every step lands in the receipt. Reflections, engines and refresh policies stay with the owner at every level.

What changes for your team

Six jobs that run on autopilot with Data Workers next to Dremio, with a concrete example of each
  • •Fast and current stop being the same claim. The numbers people read are scored against recorded baselines even when every refresh job and DAG looks green, so a stale Reflection becomes an incident with an owner, through the Incident Debugging agent.
  • •Table maintenance gets a second reader. A change to the maintenance DAG gets a blast-radius review against the views, models and jobs that read the table, through the Data Change Review agent.
  • •Lakehouse, pricing and reporting teams share one record. The same incident and receipt in Spellbook, instead of three Slack threads.
  • •Downstream actions wait for good data. Jobs that act on a view, like markdowns or replenishment, are held behind an approval while the number is wrong.

Keep Dremio, or consolidate?

Keep Dremio if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

For almost every Dremio team the answer is keep it. What teams consolidate is the tooling around it: a separate observability tool, hand-written freshness queries and "check the dashboards after maintenance" runbooks. Neighbours in the same estate: you're on Starburst, you're on ClickHouse and you're on SAP BW and Datasphere; if Business Data Cloud brings Databricks into your estate, read Data Workers on Databricks. Data Workers integrations lists what connects natively, and Data Workers vs data observability explains why table freshness checks miss this class.

Building it yourself with a coding agent and Dremio's MCP server? Read build it ourselves with Claude Code and MCP servers: querying the lakehouse is easy; the context graph, approvals, undo and receipts are the work.

The case for your CFO

The outcome. When a lakehouse answer goes stale or wrong, it is caught before the business acts on it: before price cuts, replenishment orders or the board pack, with a record of what happened.

The risk story. At L1 agents only read. At L2 a named person approves each change; an unanswered request expires and escalates. L3 covers reversible changes in domains you open. No agent can promote its own work, an org-wide stop halts all autonomous dispatch, and every change carries a receipt with its blast radius and undo. Nothing migrates: Dremio, Polaris and your Iceberg tables stay where they are.

Why now. Dremio now serves the lakehouse straight to agents, and SAP is making it the Iceberg foundation of Business Data Cloud. More readers, human and automated, will act on what it serves, faster.

The first win. Pick the views that drive decisions with money attached: pricing, replenishment, revenue. A pilot at L1 shows within weeks what the agents would have caught.

What stays the same. Dremio, dbt-dremio, Airflow, Tableau and your on-call rota. For the numbers, see the ROI of agentic data operations.

The sentence for upstairs: "Dremio makes our lakehouse fast; Data Workers makes sure what it serves is right, and gets it fixed with our approval before we price, order or report on it."

Getting started

Start with a pilot. Pick the Dremio views behind pricing, replenishment or revenue, give Data Workers read access to the dbt-dremio project, Airflow and Tableau, record the served-date and total metrics that matter, and run at L1 for a few weeks. Then turn on L2 for one domain, and open L3 for a narrow class once the receipts show the agents were right. The pilot path is on the pricing page, and the pilot is credited in full against the first year.

FAQ

Does SAP's acquisition of Dremio change how Data Workers connects? No. SAP completed the acquisition on July 6, 2026 and has said it will keep investing in Apache Iceberg, Polaris and Arrow. Data Workers connects to Dremio over its API or MCP server and reads the dbt, Airflow and BI systems around it natively, so it works the same whether you run Dremio on its own or inside Business Data Cloud. For the SAP data products side, see you're on SAP Business Data Cloud.

Does Data Workers run queries in Dremio? No, by design. Dremio's MCP server runs in your team's client next to Data Workers, with view creation off by default, and the people in that client query Dremio through it with their own permissions.

Why did the Reflection keep answering after its refresh failed? Dremio retries a failed refresh on a documented backoff, and only a Reflection in Failed status, after multiple failed attempts, stops being considered for acceleration. Until then the last materialization stays in place until its "Available Until" time. Dremio uses Autonomous Reflections, and Reflections created from its usage-based recommendations, only when fully synchronized with their source, which avoids this case for those.

Will Data Workers change our Reflections, views or engines? No. Reflections, refresh policies, views, engines and grants stay with the owner. Data Workers proposes the change as a diff or a written change, and your team makes it in Dremio or ships it through dbt-dremio.

Does Data Workers run its built-in quality checks inside Dremio? Its null, uniqueness, freshness and volume checks run on Snowflake and BigQuery. On a Dremio lakehouse the signals come from metrics your team records with monitor_metrics, such as the latest served date or a daily total, and from your own tests; Data Workers scores them against their baselines and opens the incident.

How does this fit with Dremio's AI Agent and semantic layer? They do different jobs. Dremio's AI Agent and AI Semantic Layer help people and agents ask the lakehouse questions with business context. Data Workers makes sure the data behind those answers is right and gets fixes approved by a named owner.

Sources

  • •SAP News, "SAP Completes Acquisition of Dremio" (Jul 6, 2026), https://news.sap.com/?p=243771 (checked Oct 3, 2026)
  • •SAP News, "SAP to Acquire Dremio to Unify SAP and Non-SAP Data" (May 4, 2026), https://news.sap.com/2026/05/sap-to-acquire-dremio-unify-sap-and-non-sap-data-power-agentic-ai/ (checked Oct 3, 2026)
  • •Dremio blog, "SAP Intends to Acquire Dremio" (May 4, 2026), https://www.dremio.com/blog/sap-intends-to-acquire-dremio/ (checked Oct 3, 2026)
  • •Dremio homepage (tagline, capabilities, editions, "Dremio is now part of SAP"), https://www.dremio.com/ (checked Oct 3, 2026)
  • •Dremio docs, Refresh Reflections (partition changes, refresh policies, retry policy), https://docs.dremio.com/current/acceleration/manual-reflections/refreshing-reflections (checked Oct 3, 2026)
  • •Dremio docs, Manually Manage Reflections (expiration policy, usage-based recommendations), https://docs.dremio.com/current/acceleration/manual-reflections/ (checked Oct 3, 2026)
  • •Dremio docs, View Information About Reflections ("Available Until", Failed status), https://docs.dremio.com/current/acceleration/manual-reflections/viewing-info-about-reflections (checked Oct 3, 2026)
  • •Dremio docs, Autonomous Reflections, https://docs.dremio.com/current/acceleration/autonomous-reflections (checked Oct 3, 2026)
  • •Dremio docs, MCP Server (hosted and self-hosted, system tables, permissions), https://docs.dremio.com/dremio-cloud/ai-integration/mcp-server/ (checked Oct 3, 2026)
  • •GitHub, dremio/dremio-mcp README (modes, allow_dml), https://github.com/dremio/dremio-mcp (checked Oct 3, 2026)
  • •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)