You're on MotherDuck and DuckDB: They Make Analytics Fast and Small. Data Workers Owns Whether the Data in Them Is Right
DuckDB and MotherDuck make analytics fast and small. Data Workers catches a dbt model quietly skipping rows, traces it to Dives and payouts, and gets the fix approved.
Your team runs DuckDB everywhere it fits: in a notebook, in CI, inside an app, on a laptop with a database file and a folder of Parquet. Production lives on MotherDuck, the serverless DuckDB warehouse, where databases are shared across the organization, prompt_jev() classifies text in plain SQL and Dives turn a question asked in Claude, ChatGPT or Cursor into a live chart in your workspace. The dbt project runs on a local DuckDB file in CI and on MotherDuck in production, and dbt v2 ships a built-in DuckDB adapter. DuckDB 1.5.6 (Sep 28) is current, with a v2.0 preview out since Aug 17. And DuckLabs will join Amazon Web Services as a new subsidiary, while DuckDB, DuckLake, Quack and the extensions remain free and open source under the MIT license, stewarded by the non-profit DuckDB Foundation.
That speed is the point, and it has one side effect: a small team can publish a number to the whole company, or to a payout file, in an afternoon. When an upstream change makes a model count the wrong rows, DuckDB computes exactly what the SQL says, the Dive refreshes live, every dbt test passes and nobody sees an error. Data Workers watches the systems around MotherDuck, traces a wrong number to its cause and gets the fix approved before payouts or invoices run on it.
Key takeaways
- •DuckDB and MotherDuck keep their job. Databases, shares, Dives and MotherDuck's MCP servers stay with your team. Data Workers owns whether the numbers are right.
- •Fast and green can still be wrong. When a backdated import slips past an incremental model, Data Workers ties the gap to its cause and lists every Dive, export and payout it reaches.
- •It reads the systems around MotherDuck natively. The source Postgres, the dbt manifest, Dagster runs and the build metrics your team records with
monitor_metrics. MotherDuck's MCP server runs in your team's own client, side by side with Data Workers' servers. - •Every fix goes through a named person. The owner approves each proposal and merges the model change; Data Workers queues the approved rebuild and checks the result.
- •Neutral to the ownership news. DuckLabs joining AWS changes nothing about how Data Workers works with your stack.
DuckDB and MotherDuck make analytics fast and small. Data Workers owns whether the data in them is right.
The job on top of MotherDuck is to notice that a model stopped counting what it should, find the cause, size the damage, get the fix made where it lives and prove it. Here is one Monday night at an online course marketplace, an illustration, not a customer case.
The setup: the course app writes to Postgres. A dlt pipeline, scheduled by Dagster every hour, merges enrollments into raw.enrollments on MotherDuck. The dbt project (dbt-duckdb) builds fct_enrollments at 06:00 as an incremental model that adds rows where created_at is later than the latest one already in the table. fct_enrollments feeds a Dive, "September royalties", that the partnerships team opens every morning, and payouts_monthly, which finance exports on Oct 1 to pay 212 instructors a 30% royalty. Both readers are dbt exposures, and the team has recorded them in the context graph. CI runs the same project against a local DuckDB file with sample data.
| Time | System | What happens |
|---|---|---|
| Mon Sep 29 21:00 | Postgres | The app team imports 4,860 enrollments from a newly signed university partner. Students bought in September through the partner's portal, so each row keeps its original September created_at |
| 21:05 | Data Workers + Postgres | The native Postgres read shows the enrollments row estimate up by about 4,900 since the last scan, against a normal day of about 600. Data Workers logs it to the timeline |
| 22:00 | dlt + MotherDuck | The hourly dlt run merges all 4,860 rows into raw.enrollments. Nothing fails |
| Tue 06:00 | Dagster + dbt + MotherDuck | Dagster runs dbt build. The incremental filter sees September dates older than the latest created_at already loaded and skips every imported row. All tests pass: unique, not null, accepted values |
| 06:10 | Data Workers | The team's Dagster asset check records two numbers with monitor_metrics after each build: rows landed in raw.enrollments since the last build (5,431, baseline about 600) and rows added to fct_enrollments (571, on baseline). The first is flagged; the gap between them makes it an incident |
| 06:20 | Data Workers | diagnose_incident ties the spike to the Postgres import. The dbt manifest shows fct_enrollments is an incremental model. trace_cross_platform_lineage follows the manifest from the source to the model; blast_radius_analysis adds the Dive and payouts_monthly, which pays on Oct 1 |
| 06:25 | Slack | send_slack_alert reaches the analytics engineer who owns the model, finance copied, with the evidence and a count query to run |
| 08:45 | MotherDuck MCP | The owner runs that query from Claude Code through MotherDuck's MCP server, read-only: 4,860 rows in raw.enrollments have no match in fct_enrollments. At 30% that is $41,300 of September royalties missing from the Oct 1 payout, and the Dive shows the new partner with zero enrollments |
| 09:10 | Spellbook | The owner reviews two proposals: a dbt diff that switches the incremental filter to the _dlt_load_id that dlt writes on every row, plus a test that fails when landed rows are missing from the model; and a full rebuild of fct_enrollments once the diff is merged. Undo: revert the diff. She approves both |
| 09:40 | GitHub + DuckDB | She merges the diff. CI builds the project against a local DuckDB file with a replayed sample that includes backdated rows; the new test passes |
| 09:55 | Data Workers + Dagster | With both approvals on record, trigger_dagster_job queues the rebuild job, a dbt build --full-refresh of fct_enrollments |
| 10:20 | MotherDuck | The rebuild finishes. The Dive, which queries live data, shows September with the partner's 4,860 enrollments |
| 10:30 | Data Workers | Rows added now match rows landed, and the comparison query Data Workers proposed, which the owner runs, shows no landed row missing from the model. The receipt records cause, approvals, merge, rebuild, checks and the undo |
| Thu Oct 1 09:00 | Finance | payouts_monthly exports with the partner's enrollments, and 212 instructors are paid in full |

Every part of the stack did its job: dlt landed every row, DuckDB ran the filter as written, the tests checked what they were written to check. Knowing the filter no longer fit took facts outside the warehouse: the app team imported history, and a payout reads the model on Oct 1.
| Job | What DuckDB and MotherDuck do | What Data Workers does |
|---|---|---|
| Compute | DuckDB runs analytical SQL in-process, on a laptop, in CI or in an app; MotherDuck runs it serverless for the team | Keeps the models those queries read right, and traces which numbers a wrong model reached |
| Sharing | MotherDuck shares databases across the organization; Dives query live data | Names every Dive, export and job that reads a model, so a fix reaches all of them |
| Agent access | MotherDuck's MCP servers let agents query and build Dives; the local server is read-only by default | Answers "is it still right, and what's the fix" from one governed context graph |
| Tests | Runs every test the dbt project defines, fast | Scores models against the baselines the team records, so a quiet gap becomes an incident |
| The fix | Runs the model change and rebuild the owner ships | Proposes the change to a named owner, with the blast radius, and queues the approved rebuild through Dagster |
| The proof | Answers the comparison query in seconds | Re-checks the metrics and writes a receipt: what changed, who approved it, how it was checked, how to undo it |
Why doesn't MotherDuck or DuckDB just do this itself?
Because DuckDB is built to answer SQL fast with nothing to run, and MotherDuck to make that a shared, serverless warehouse. Both scope their decisions to the query you wrote. An incremental filter is a choice in your dbt code, and whether it still fits depends on the app team's import plans and finance's payout calendar, which a database has no reason to hold.
Their AI work follows the same clear line, and it is good work. MotherDuck's remote MCP server at api.motherduck.com/mcp signs in with OAuth or a Bearer token and offers a read-only query tool and a read-write query_rw tool; a client can disable query_rw to stay read-only. The local mcp-server-motherduck (v1.0.8) opens DuckDB files, in-memory databases, S3 and MotherDuck, and runs read-only by default. Dives turn a question into a live chart, prompt_jev() (Sep 21, 2026) classifies text in SQL on paid plans, and the Tower acquisition (Aug 25, 2026) brings the runtime behind Flights, which hosts and schedules LLM-written pipelines; MotherDuck says it will turn Flights into data APIs. DuckDB ships Skills for Claude Code (Sep 16, 2026). All of it helps people and agents ask, build and publish faster.
Owning whether a number is right from the app database to the payout file is a different product: a context graph across systems, blast radius, named approvals, a recorded undo and receipts. That is Data Workers. More in is it safe to let AI agents change production data.
Every tool owns a slice. Data Workers covers the whole lifecycle
Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, on top of the MotherDuck warehouse and DuckDB files already there.

| Stage | Data Workers | MotherDuck and DuckDB | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 3 | Databases, tables and columns are listed in the catalog and over MCP, and DuckLake keeps snapshots. Data Workers keeps one governed context graph from the source database through the dbt project to every Dive and job that reads it. |
| Analytics & Insights | 8 | 8.5 | The home stage: DuckDB answers analytical SQL in-process on a laptop, in CI or in an app, and MotherDuck runs it serverless for the team, with Dives and AI functions on top. Data Workers answers from governed definitions with lineage behind every number. |
| Data Quality | 8 | 3 | DuckDB runs the tests your dbt project defines, fast. Data Workers checks that what a model serves still matches what it should count, against baselines the team records. |
| Observability & Incidents | 8.5 | 2 | Query history and errors show how queries ran. Data Workers diagnoses the data incident across the source, the loader, the model and the readers, and verifies the fix. |
| Pipelines & Ingestion | 8.5 | 6 | DuckDB reads Parquet, CSV, JSON, Postgres and lakes directly, and MotherDuck Flights runs agent-written pipelines. Data Workers sees the upstream change first and asks what it means downstream. |
| Schema & Migration | 8 | 3 | Schema changes run as SQL or dbt changes the owner writes; DuckLake records them in snapshots. Data Workers checks a source change against every model that reads it and generates migrations with rollback SQL for the owner. |
| Governance & Access | 8.5 | 3 | MotherDuck shares databases across an organization, and the MCP servers run read-only or scoped per client. Data Workers routes every data change to a named approver. |
| Security & Privacy | 8 | 3 | Open source under MIT, with MotherDuck handling cloud security. Data Workers leaves a receipt on every data change. |
| Cost / FinOps | 8 | 6 | In-process compute costs nothing extra, and MotherDuck bills by use. Data Workers traces Snowflake spend to the dbt model and reads BigQuery and AWS spend across the rest of the estate. |
| MLOps & Models | 7.5 | 3 | AI functions such as prompt_jev() classify text in SQL. Data Workers keeps the data under models and agents healthy. |
How MotherDuck, DuckDB and Data Workers work together

In this setup Data Workers reads four things: the source Postgres (schemas, columns, row estimates), the dbt manifest (sources, models, materializations and tests arrive as lineage, whether the project runs on DuckDB, MotherDuck or both), Dagster runs, and the build metrics your team records with monitor_metrics. Airflow and Prefect connect natively too, among 50+ connectors, and alerts go to Slack. Loaders such as dlt connect over their APIs; the owner runs syncs. MotherDuck's own MCP servers sit in your team's client, side by side with Data Workers' servers: the team uses them to query and build Dives, Data Workers' agents never call them, and every check rests on what Data Workers' own connectors read.
Databases, shares, Dives, models and deploys stay with the owner, by design. Data Workers proposes the change as a diff for the owner to merge, queues approved reruns through the orchestrator and verifies afterwards; data cleanups are proposed for the owner to apply.
The Data Workers side follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry. Keep MotherDuck's server read-only; its README says read-only MotherDuck connections need a read-scaling token.
# Example: MotherDuck's local MCP server (read-only by default) plus Data Workers agents in Claude Code
claude mcp add --scope user motherduck --transport stdio --env motherduck_token=<read-scaling-token> \
-- uvx mcp-server-motherduck --db-path md:
# Data Workers agents, from a clone of the open-source repo
claude mcp add dw-connectors -- /path/to/dataworkers-claw-community/start-agent.sh dw-connectors
claude mcp add dw-schema -- /path/to/dataworkers-claw-community/start-agent.sh dw-schema
claude mcp add dw-context-catalog -- /path/to/dataworkers-claw-community/start-agent.sh dw-context-catalog
claude mcp add dw-incidents -- /path/to/dataworkers-claw-community/start-agent.sh dw-incidentsOn MotherDuck's remote server (https://api.motherduck.com/mcp, OAuth), disable query_rw for this client. List tools with your client's own command (/mcp in Claude Code). In this incident: monitor_metrics and diagnose_incident (dw-incidents) score the build and name the cause; trace_cross_platform_lineage and blast_radius_analysis (dw-context-catalog) map what the model reached; assess_impact (dw-schema) checks source changes against the models that read them; send_slack_alert reaches the owner; and trigger_dagster_job (dw-connectors) queues the approved rebuild.
In production the agents run in your infrastructure and hold the database credentials and model key; your data stays in your systems, and the hosted Conductor sees workflow metadata only. More in where does our data go. For hands-on setup, see Claude Code with DuckDB for local dev, an MCP server for local DuckDB, Claude Code and MotherDuck and an MCP server for MotherDuck in the cloud.
One incident, L0 to L4, set per domain:

- •L0 manual. An instructor emails on Oct 3 that her payout looks low. Two days of reconciling, then a supplementary payout.
- •L1 observe. Data Workers flags the gap at 06:10 with the cause and blast radius. Nothing changes.
- •L2 propose. Data Workers proposes the diff and the rebuild. Nothing reaches production until the named owner approves; an unanswered request expires and escalates, never auto-grants.
- •L3 act reversibly. For a class with a clean record, Data Workers takes reversible steps in the domain you open, such as queuing the approved rebuild once CI passes. Model changes and deploys stay with the owner.
- •L4 autonomous. For a scoped domain, bulk imports into source tables are checked against every incremental model that reads them before the next build, so the partner import is flagged the night it lands.
Who sets those levels: who owns the agents, how approvals work for AI data agents and autonomy levels L0 to L4 explained.
What changes for your team

- •Green builds get a second reader. Every build is scored against what the team expects it to add, by the Incident Debugging agent, so "tests passed" stops being the last word.
- •Source changes meet their readers. Imports and schema changes on source tables are checked against every model that reads them, by the Schema Evolution agent.
- •Published numbers have a paper trail. When a Dive or export was wrong, the receipt names the model, the window and the readers.
- •One record for a small team. The app team, the analytics engineer and finance see the same incident and receipt in Spellbook Data Catalog (in preview), not a lost Slack thread.
Keep MotherDuck and DuckDB, or consolidate?
Keep MotherDuck and DuckDB if you love them; Data Workers works with them from day one. Many teams consolidate once Data Workers runs that slice too.
For DuckDB and MotherDuck, keeping them is the norm. What teams consolidate is the work around them: a separate observability tool, hand-written "did every row arrive" queries, and the runbook someone checks before payouts. If your estate grows into Snowflake or BigQuery next to MotherDuck, Data Workers reads those natively too; see Data Workers on Snowflake. Neighbours: you're on dbt, you're on Dagster and you're on ClickHouse; Data Workers integrations lists what connects natively, and Data Workers vs data observability explains why query health alone misses this class.
Building it yourself with DuckDB Skills and the MotherDuck MCP server? See build it ourselves with Claude Code and MCP servers: the query is easy; the context graph, approvals, undo and receipts are the work. The bigger picture: what is an agentic data platform.
The case for your CFO
The outcome. When an upstream change quietly makes a model skip or double rows, it is caught at the next build and corrected before payouts, invoices or board numbers run on it.
The risk story. At L1 agents only read. At L2 a named person approves each change; an unanswered request expires and escalates. L3 covers reversible steps in domains you open. Models, shares and deploys stay with the owner. No agent can promote its own work, an org-wide stop halts all autonomous dispatch, and every change carries a receipt with its undo.
Why now. Agents now build Dives and pipelines on MotherDuck in an afternoon. Speed raises the cost of a wrong number, and the team hasn't grown.
The first win. The models that pay or bill: royalties, commissions, usage invoices. A pilot at L1 shows within weeks what the agents would have caught.
What stays the same. DuckDB, MotherDuck, your Dives, dbt, Dagster and your loaders. For the numbers, see the ROI of agentic data operations.
The sentence for upstairs: "DuckDB and MotherDuck keep our analytics fast and cheap; Data Workers makes sure the numbers are right, and gets them fixed with our approval when they aren't, before we pay or bill on them."
Getting started
Start with a pilot. Pick the MotherDuck models behind payouts, invoices or leadership's Dives, give Data Workers read access to the source databases, the dbt project and the orchestrator, record the build metrics that matter, and run at L1 for a few weeks. Then turn on L2 for one domain. The pilot path is on the pricing page, and the pilot is credited in full against the first year.
FAQ
Does Data Workers connect to MotherDuck directly? Data Workers works from the systems around MotherDuck, which it reads natively: the source Postgres, the dbt manifest, Dagster runs and the build metrics your team records with monitor_metrics. Your team's MCP client runs MotherDuck's MCP server side by side with Data Workers' servers, so engineers query MotherDuck and build Dives there; Data Workers' agents never call it, and their checks rest on Data Workers' own reads.
Does DuckLabs joining AWS change anything? Not for Data Workers. DuckDB's announcement (Aug 26, 2026) says DuckLabs will join Amazon Web Services as a new subsidiary, expected to be effective in early September, with "no changes for our projects' roadmap, licensing, and governance model": DuckDB, DuckLake, Quack and the extensions stay MIT under the non-profit DuckDB Foundation. MotherDuck now sells support for open-source DuckDB. Data Workers stays neutral to the engine and the cloud.
Why did dbt's tests pass? They tested what was in the model: unique keys, no nulls, accepted values. Rows that never entered the model break none of those. The gap shows only when you compare what landed with what was added.
We use DuckLake. Does that help? Yes. DuckLake (v1.0, April 2026) keeps tables as Parquet with snapshots and time travel, so the owner can compare a model before and after a fix in plain SQL. Data Workers works the same way on top: it watches the sources and the dbt project, and the receipt records the window a wrong number covered.
Which MotherDuck plan do we need? Any. MotherDuck lists Lite, Business and Enterprise plans, and Dives are available on all of them. Data Workers prices separately, with no usage meter.
Sources
- •DuckDB, News (DuckDB 1.5.6 Sep 28, 2026; "A Preview of DuckDB v2.0" Aug 17, 2026; "Try DuckDB v2.0-dev" Sep 2, 2026; "DuckDB Skills for Claude Code" Sep 16, 2026), https://duckdb.org/news/ (checked Oct 3, 2026)
- •DuckDB, GitHub releases (v1.5.6, Sep 28, 2026), https://github.com/duckdb/duckdb/releases (checked Oct 3, 2026)
- •DuckDB, "DuckLabs to Join AWS, Projects to Remain Open Source" (Aug 26, 2026), https://duckdb.org/2026/08/26/ducklabs-to-join-aws.html (checked Oct 3, 2026)
- •DuckDB, "DuckDB Now Ships inside dbt v2" (Sep 22, 2026), https://duckdb.org/2026/09/22/dbt-fusion.html (checked Oct 3, 2026)
- •DuckLake, project site (v1.0, April 2026), https://ducklake.select/ (checked Oct 3, 2026)
- •MotherDuck blog, "DuckDB outgrows its nest" (Jordan Tigani, Aug 26, 2026), https://motherduck.com/blog/duckdb-amazon/ (checked Oct 3, 2026)
- •MotherDuck blog, "MotherDuck Now Offers Open Source DuckDB Support" (Sep 14, 2026), https://motherduck.com/blog/duckdb-support/ (checked Oct 3, 2026)
- •MotherDuck blog, "Introducing prompt_jev(): bringing Jev to MotherDuck SQL" (Sep 21, 2026), https://motherduck.com/blog/motherduck-supports-jev/ (checked Oct 3, 2026)
- •MotherDuck blog, "MotherDuck Acquires Tower: Agents Can Answer. Now They Can Build." (Aug 25, 2026), https://motherduck.com/blog/motherduck-acquires-tower/ (checked Oct 3, 2026)
- •MotherDuck docs, Dives (available on all plans), https://motherduck.com/docs/key-tasks/ai-and-motherduck/dives/ (checked Oct 3, 2026)
- •MotherDuck docs, MCP server setup (remote server, OAuth or Bearer token,
queryandquery_rw), https://motherduck.com/docs/key-tasks/ai-and-motherduck/mcp-setup/ (checked Oct 3, 2026) - •MotherDuck, mcp-server-motherduck README and releases (read-only by default, read-scaling token, v1.0.8 Aug 19, 2026), https://github.com/motherduckdb/mcp-server-motherduck (checked Oct 3, 2026)
- •MotherDuck, Pricing (Lite, Business, Enterprise), https://motherduck.com/pricing/ (checked Oct 3, 2026)
- •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)