You're on dlt: Your Engineers Load Anything With It. Data Workers Owns Whether What It Loaded Is Right
dlt evolves your schema on its own: new columns, nested tables, variant columns. Data Workers catches when that change breaks a dbt model or a payout, traces it and gets the fix approved.
Your team writes pipelines. A source groups the resource functions that yield records from a REST API, a database or a bucket; pipeline.run() extracts, normalizes and loads them as a load package into BigQuery, Snowflake, Databricks, DuckDB or Postgres. Incremental cursors live in pipeline state, merge dispositions dedupe on a primary key, and every load writes a row to _dlt_loads and, when the schema moves, a new version to _dlt_version. The pipelines run wherever you put them: an Airflow DAG through dlt deploy ... airflow-composer, Dagster, Prefect, GitHub Actions, or dltHub's managed runtime. More and more of them are written by a coding agent: dltHub's AI Harness, part of the dltHub platform, gives Claude Code, Cursor or Codex skills such as /create-rest-api-pipeline and /debug-pipeline, and the open-source dlt MCP server (0.3.0, beta) lets the agent inspect pipelines and tables while it works.
The feature your engineers love most is schema evolution. A new field becomes a new column, a nested object flattens into parent__child columns, a list becomes a nested table linked by _dlt_parent_id, and a value that changes type lands in a variant column instead of failing the load. That is exactly right for a loader: nothing gets dropped. It also means a load can be green while the dbt model on top quietly reads a column that stopped being filled. Data Workers watches what dlt loaded, traces it through dbt, BI and the jobs that read it, and gets the fix approved before anything downstream builds on it.
Key takeaways
- •dlt keeps its job. Sources, resources, pipeline code, contracts, schedules and loads stay with your engineers. Data Workers works on what the loads land.
- •Green loads get checked too. New tables and columns are read from dlt's schema versions, and null and volume checks plus the metrics your team records turn a successful load that moved data out from under a model into a diagnosed incident.
- •Connected over dlt's API or MCP server today. Data Workers reads the loaded datasets,
_dlt_loadsand_dlt_versionnatively in the warehouse; the dlt MCP server sits next to the Data Workers agents in the same client. - •Every fix goes through a named person. The owner approves the hold, the dbt diff and any pipeline change, and runs the dlt side; each step leaves a receipt.
- •Autonomy is set per domain. Start at L1 observe, move to L2 propose, and open L3 act reversibly for narrow classes once the record supports it.
dlt is the loader your engineers write. Data Workers is the owner of what it loaded.
dlt's job is to get every record from the source into the destination, infer and evolve the schema as the source changes, and record exactly what changed on the way. The job after the load is different: notice that a successful load changed the shape a model depends on, size the damage, hold anything that would ship the error, get the model fixed, rebuild in order and prove the number. Here is one Monday at a language-learning subscription app that loads its own billing service with dlt. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| Sun 22:40 | Subscriptions API | Billing engineering ships v3 of the in-house charges endpoint to support stacked promo codes. The single discount object on each charge becomes a discounts list |
| Mon 02:00 | Airflow + dlt | The subscriptions_ingest DAG runs the dlt pipeline (REST API source, charges resource, merge on id) into the BigQuery dataset subscriptions_raw. The normalizer unpacks the list into a new nested table, charges__discounts, with 52,900 rows linked by _dlt_parent_id. On the 38,400 new charges, discount__code and discount__amount arrive null. The schema evolves under the default contract, _dlt_loads records status 0 and _dlt_version gains a version. The DAG run succeeds |
| 02:25 | Airflow + dbt | Airflow runs dbt build. stg_subs__charges coalesces the missing discount to zero, so fct_net_revenue overstates Monday's net revenue by $57,930. Every dbt test passes |
| 02:40 | Data Workers + BigQuery | The daily discount total the team records as a metric reads $0 against a 30-day range of $58k to $64k, and the null check on charges.discount__amount jumps for the new load. The new schema version, read over the dlt MCP server, lists charges__discounts, a table no model reads. Data Workers opens an incident |
| 02:48 | Data Workers + BigQuery | Diagnosis: the new nested table and the null columns arrived in the same load (one _dlt_load_id), and _dlt_version gained a version at 02:00. The discounts moved to charges__discounts; none were lost. Blast radius: fct_net_revenue, fct_affiliate_commissions, the Looker Explore "Net revenue by plan", and the Airflow DAG affiliate_payouts, which sends the weekly commission file to the affiliate network at 08:00 |
| 02:50 | PagerDuty + Slack | Data Workers raises a PagerDuty incident with the diagnosis and sends the approval request to the pipeline owner in Slack |
| 06:30 | Spellbook | The owner reviews three proposals: pause affiliate_payouts; a dbt diff that sums charges__discounts into stg_subs__charges on _dlt_parent_id; and a schema contract for the charges resource that freezes new tables, so the next shape change stops for review. She approves the hold and the diff, and puts the contract on the sprint after weighing the freshness trade |
| 06:40 | Airflow + GitHub | She pauses affiliate_payouts in Airflow and merges the dbt diff |
| 06:45 | Data Workers + Airflow | Data Workers queues the approved dbt run through Airflow 2's REST API; the revenue models rebuild |
| 07:05 | Data Workers + BigQuery | Data Workers verifies: the recorded discount metric reads $57,930, inside its baseline; volume holds at 38,400 charges and id is unique; the owner reports the new dbt test (net equals gross less nested discounts) passing. No reload was needed: dlt had loaded every value |
| 07:10 | PagerDuty | Data Workers resolves the PagerDuty incident with a link to the receipt: cause, approvals, runs, checks passed and the undo |
| 07:20 | Airflow | The owner unpauses affiliate_payouts; the 08:00 file goes out on true net revenue |
| 09:00 | Looker | The weekly revenue review opens "Net revenue by plan" on verified numbers |

dlt did everything it promises. It read every charge, kept every discount, created the nested table with clean parent keys and recorded the new schema version. Billing changed an API shape, and a dbt model written a year earlier still read the flattened column. Catching it takes knowledge dlt was never meant to hold: that discount__amount feeds net revenue, that zero is a plausible but wrong default, and that a job pays partners from that number at 08:00.
| Job | What dlt does | What Data Workers does |
|---|---|---|
| The load | Extracts, normalizes and loads any source from Python code; incremental cursors, merge and replace dispositions, retries | Reads what the loads land, natively in the warehouse; connects to dlt over its API or MCP server |
| The schema | Infers and evolves it: new columns, flattened objects, nested tables, variant columns; contracts can freeze or discard | Reads each new table and column from dlt's schema version and traces what it means for every model, dashboard and job downstream |
| The signal | load_info, _dlt_loads, _dlt_version and Slack hooks report each load and its schema updates | Checks the data itself: nulls, volume, freshness and the metrics your team records, so a green load with a broken model still raises an incident |
| The diagnosis | The dlt MCP server and /debug-pipeline help an engineer read pipelines, schemas and state | Joins the load, the schema history and lineage into one cause, with the blast radius across dbt models, dashboards and downstream jobs |
| The fix | Runs whatever pipeline code and contract the team commits | Proposes the dbt diff, the hold and the contract change to the owner, who approves and applies them |
| The proof | Keeps load and schema history in the dataset | Re-checks the tables and writes a receipt: what changed, who approved it, how it was checked, how to undo it |
Why doesn't dlt just do this itself?
Because dlt is a library for loading data faithfully, and its defaults say so. Under the default evolve contract, in its own docs, "New tables may always be created. New columns may always be appended to the existing table," and data that doesn't fit a column's type goes to a variant column. The alternative is a choice your engineers make per resource: freeze raises a DataValidationError and stops the load, discard_row and discard_value drop what doesn't fit. Each setting trades freshness against breakage, and dlt rightly leaves that trade to you. What it can't know is that a particular flattened column feeds net revenue and a partner payout.
dltHub's AI features point the same way. The AI Harness teaches a coding agent "how to build production-grade pipelines" through skills, rules and MCP servers, on the principle "Agents propose, humans validate, deterministic tooling enforces the boundaries", and dlthub-start (beta) scaffolds a workspace for it. The dlt MCP server is built for inspection: it lists pipelines and tables, returns schemas and schema diffs, reads pipeline state and runs SELECT queries. The dltHub platform adds a managed runtime, data quality checks and an observability dashboard for the pipelines it runs.
Owning whether loaded data is right across an internal API, a warehouse, dbt, BI and a payout job is a different product: a context graph of every resource, table and consumer, blast-radius scoping, approvals that name a person, a recorded undo and receipts an auditor can read. That is Data Workers. More in is it safe to let AI agents change production data.
Every tool owns a slice. Data Workers covers the whole lifecycle
dlt owns loading any source into your warehouse as Python code. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the dlt pipelines already there.

| Stage | Data Workers | dlt | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 4 | dlt infers a schema for every source and keeps its history in _dlt_version; dltHub adds a context graph for its own pipelines. Data Workers keeps one governed context graph of what each table means, who owns it and what reads it. |
| Analytics & Insights | 8 | 1 | Not dlt's job: it loads data, and business numbers live in BI. Data Workers answers data questions from governed definitions with lineage behind every number. |
| Data Quality | 8 | 6 | Schema contracts freeze, discard or evolve on the way in, and dltHub adds data quality checks. Data Workers checks what landed against its baselines and turns a wrong value into an incident. |
| Observability & Incidents | 8.5 | 4 | load_info, _dlt_loads and Slack hooks report each load and its schema updates. Data Workers diagnoses the data incident when the load succeeded, proposes the fix and verifies it. |
| Pipelines & Ingestion | 8.5 | 9 | dlt's home stage: REST APIs, databases and files loaded as code, incremental loading, merge dispositions and an AI Harness that builds pipelines. Data Workers plans the reruns around it for the owner. |
| Schema & Migration | 8 | 6 | Schema evolution adds columns, nested tables and variant columns on its own, and contracts govern it. Data Workers traces each change through dbt, BI and downstream jobs. |
| Governance & Access | 8.5 | 3 | Pipelines run with the credentials you give them; RBAC and audit logs sit in dltHub Enterprise. Data Workers routes every data change to a named approver. |
| Security & Privacy | 8 | 3 | Secrets providers for Google and AWS keep credentials out of code. Data Workers leaves a receipt on every data change. |
| Cost / FinOps | 8 | 3 | Open source is free to run; warehouse spend sits in the warehouse. Data Workers attributes Snowflake spend to the dbt models behind it and reads BigQuery spend from the Jobs API. |
| MLOps & Models | 7.5 | 2 | dlt loads the data models and agents train and answer on. Data Workers keeps the data under models and agents healthy. |
How dlt and Data Workers work together

Data Workers connects to dlt over its API or MCP server today. What it needs from each load sits in the dataset dlt writes: the tables, the nested tables with their _dlt_parent_id keys, _dlt_loads and _dlt_version are ordinary tables it reads natively in BigQuery, Snowflake, Databricks or PostgreSQL. Pipeline code, contracts, secrets, schedules and loads stay with your engineers: Data Workers never runs, pauses or reloads a dlt pipeline. It proposes the dbt diff or the contract change for the owner to merge, and queues approved downstream runs through the orchestrator. The rest of this incident runs on native connections: BigQuery, dbt, Airflow, Looker (read), PagerDuty and Slack, among 50+ connectors. More on the wiring in Data Workers + Airflow and Data Workers + dbt.
The dlt MCP server and the Data Workers agents run side by side in one client. Your engineer asks the dlt MCP for the schema diff on charges and asks Data Workers what that change did to revenue, in the same session. The Data Workers side follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry.
Example: the dlt MCP server plus four Data Workers agents in .mcp.json (Claude Code) or .cursor/mcp.json (Cursor).
{
"mcpServers": {
"dlt": {
"command": "uv",
"args": ["run", "--with", "dlt-mcp[search]", "python", "-m", "dlt_mcp"]
},
"dw-incidents": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-incidents"] },
"dw-quality": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-quality"] },
"dw-schema": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-schema"] },
"dw-catalog": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-context-catalog"] }
}
}List the tools with your client's own command (/mcp in Claude Code). In this incident: the dlt MCP server shows the new schema version with the nested table; run_quality_check (dw-quality) runs the null and volume checks; monitor_metrics (dw-incidents) tracks the daily discount total against its baseline; diagnose_incident names the cause; blast_radius_analysis and trace_cross_platform_lineage (dw-context-catalog) map what the change reached; trigger_airflow_dag queues the approved rebuild on Airflow 2; remediate re-checks the quality assertions after the rebuild and escalates any failure to a person. On the dlt side, get_table_schema_diff, get_pipeline_local_state and execute_sql_query cover the engineer's inspection.
In production the agents run in your infrastructure and hold the warehouse credentials and model key, just as your dlt pipelines do. Your data stays in your systems; the hosted Conductor sees workflow metadata only. More in where does our data go.
One incident, L0 to L4, set per domain:

- •L0 manual. Finance asks on Wednesday why commissions jumped, after the payout went out.
- •L1 observe. Data Workers flags the zero discount total at 02:40 with the cause and blast radius. Nothing changes.
- •L2 propose. Data Workers proposes the hold, the dbt diff and the contract; nothing moves until the named owner approves.
- •L3 act reversibly. For a class with a clean record, Data Workers queues the rebuilds itself, verifies them and sends any failed check to a person.
- •L4 autonomous. For a scoped domain, Data Workers checks every load as it lands, so the owner's fix is waiting before the first downstream job runs.
Who sets those levels: who owns the agents, how approvals work for AI data agents and autonomy levels L0 to L4 explained.
What changes for your team

- •Schema evolution stops being a gamble. Engineers keep
evolvewhere freshness matters, because a change that breaks something downstream reaches the owner with a cause attached. - •Agent-built pipelines get a second pair of eyes on the output. When a coding agent writes a new resource with the AI Harness, Data Workers checks what that resource actually lands against the models that read it.
- •Jobs that pay or push data out get a guard. When a payout, billing or activation job reads a table that just changed shape, the hold is proposed before it runs.
- •Data and platform teams share one record. Both see the same incident and receipt in Spellbook Data Catalog (in preview), linked from PagerDuty.
The Schema Evolution agent is the one that watches what lands; for the general pattern on a warehouse, see BigQuery schema change impact on downstream dashboards, and for where loading sits in the stack, data ingestion vs ETL.
Keep dlt, or consolidate?
Keep dlt if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
For most dlt teams the answer is keep it: pipelines as reviewable Python, an open-source license and any destination are hard to replace. What teams consolidate is the tooling around the loaded data: a separate observability tool, Slack hooks that post schema updates nobody reads, and runbooks that say "check _dlt_version and eyeball the dashboard". Many estates also run a managed mover next to dlt: see you're on Airbyte and you're on Fivetran, and if Airflow runs your pipelines, you're on Airflow. One incident record spans them.
Weighing a build on the dlt MCP server and a coding agent? Read build it ourselves with Claude Code and MCP servers. Reading a schema diff is the easy part; the context graph, approvals, undo and receipts are the work.
The case for your CFO
The outcome. When a dlt load changes the shape of data behind revenue, payouts or billing, it is caught the same night and corrected before money moves on it, with a record of what went wrong and how it was checked.
The risk story. At L0 and L1, agents only read. At L2 they propose and a named person approves; an unanswered request expires and escalates, never auto-grants. At L3 they act on reversible changes inside the domains you open; L4 is a later choice per domain. Pipeline code, contracts and loads stay with your engineers. No agent can promote its own work, and an org-wide stop halts all autonomous dispatch. Every change carries a receipt: what changed, who approved it, the blast radius and how to undo it. Nothing migrates: dlt, your warehouse and your dbt project stay where they are.
Why now. Coding agents now write dlt pipelines. More pipelines, written faster, against more internal APIs means more schema changes reaching models nobody re-read.
The first win. L1 on the dlt datasets that feed money numbers: every load is checked, and an API shape change becomes a diagnosed incident before the first downstream job runs.
What stays the same. dlt, your pipeline repo, your orchestrator, your dbt project, your warehouse and your on-call rota. For the numbers, see the ROI of agentic data operations.
The sentence for upstairs: "dlt lets our engineers load anything; Data Workers makes sure what it loaded is right and gets it fixed with our approval when it isn't, before we pay or report on it."
Getting started
Start with a pilot. Pick the dlt pipelines that feed the numbers leaders, partners and customers see, give Data Workers read access to their datasets, and run at L1 for a few weeks. Then turn on L2 for one domain, and open L3 for a narrow class once the receipts show the agents were right. For the broader picture, see what is an agentic data platform and what integrations Data Workers supports. The pilot path is on the pricing page, and the pilot is credited in full against the first year.
FAQ
How does Data Workers connect to dlt? Over dlt's API or MCP server today, with the dlt MCP server in the same client for schemas and pipeline state. Data Workers reads the loaded datasets, _dlt_loads and _dlt_version natively in your warehouse. It works wherever the pipelines run.
Our load succeeded. How can the data be wrong? Under the default contract, dlt evolves the schema instead of failing: a list becomes a nested table, a type change becomes a variant column, and the old column simply stops being filled. Every value is in the warehouse; the model on top just reads the wrong place. Data Workers checks the loaded data and traces what each change reached.
Wouldn't a schema contract catch this? A freeze contract on tables would have stopped the load with a DataValidationError, and for some money tables that is the right call. It also means no fresh data until someone ships a change. Many teams keep evolve for freshness; Data Workers is what makes that safe, and it can propose a contract per resource when the record shows one is worth the trade.
Will Data Workers run, pause or reload our dlt pipelines? No. Pipeline code, contracts, schedules and loads stay with your engineers. Data Workers proposes fixes as diffs for the owner to merge and queues approved downstream runs, such as a dbt rebuild, through your orchestrator.
Do we still need dltHub's AI Harness? If your engineers build pipelines with coding agents, yes: it is good at writing and debugging them. Data Workers works on what those pipelines land and everything downstream.
Sources
- •dltHub, homepage ("The agentic data layer on top of any data warehouse"; AI Harness; context graph; dlt "Apache 2.0 licensed and always free to use"), https://dlthub.com/ (checked Oct 3, 2026)
- •dltHub, pricing (dlt open source; dltHub plan with managed runtime, data quality checks, observability dashboard and AI Harness; Enterprise with RBAC and audit logs), https://dlthub.com/pricing (checked Oct 3, 2026)
- •dlt on PyPI (1.30.0, released Aug 11, 2026), https://pypi.org/project/dlt/, and release notes, https://github.com/dlt-hub/dlt/releases (checked Oct 3, 2026)
- •dlt-mcp on PyPI (0.3.0, Feb 19, 2026, Development Status Beta, Apache 2.0; tools list_pipelines, list_tables, get_table_schemas, execute_sql_query, get_load_table, get_pipeline_local_state, get_table_schema_diff), https://pypi.org/project/dlt-mcp/ (checked Oct 3, 2026)
- •dlthub-start on PyPI (0.10.8, beta, Sep 15, 2026), https://pypi.org/project/dlthub-start/ (checked Oct 3, 2026)
- •dltHub docs, AI Harness introduction ("a part of the dltHub platform"), https://dlthub.com/docs/hub/ai-harness/introduction, and REST API source with the AI Harness, https://dlthub.com/docs/dlt-ecosystem/llm-tooling/llm-native-workflow (checked Oct 3, 2026)
- •dlt docs, Schema evolution, https://dlthub.com/docs/general-usage/schema-evolution (checked Oct 3, 2026)
- •dlt docs, Schema and data contracts, https://dlthub.com/docs/general-usage/schema-contracts (checked Oct 3, 2026)
- •dlt docs, Destination tables (nested tables,
_dlt_loads,_dlt_version), https://dlthub.com/docs/general-usage/destination-tables (checked Oct 3, 2026) - •dlt docs, Deploy with Airflow and Google Composer, https://dlthub.com/docs/walkthroughs/deploy-a-pipeline/deploy-with-airflow-composer (checked Oct 3, 2026)
- •Data Workers open-source repository (tool registrations in dw-incidents, dw-quality, dw-context-catalog, dw-schema, dw-connectors), https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
- •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)