You're on Meltano: Your Team Runs Singer Pipelines as Code. Data Workers Owns Whether What They Load Is Right
Meltano runs your Singer taps and targets as code, in Git. Data Workers checks what every run loads, catches the green run that left data behind, and gets the fix approved.
Your data engineers live in meltano.yml. Each extractor is a Singer tap, each loader a Singer target, both plugins from Meltano Hub's 600+ connectors. The select extra decides which streams and properties a tap pulls from its discovered catalog, and incremental state keeps a bookmark per stream so each run picks up where the last one stopped. A pipeline is one line, meltano run tap-xero target-bigquery dbt-bigquery:build, and every change goes through a pull request: your data stack as code, with Git-based governance. When a source has no tap, someone writes one with the Meltano SDK. You run it yourselves on Meltano Open, or on Meltano Cloud, which adds schedules, monitoring and, since August, Melty AI: a Diagnose button on failed runs, diagnosis inside failure alerts, a knowledge kit called Agent Melty seeded into every workspace repo, and Melty MCP, which brings your workspaces into Claude.
The project is in steady hands. Meltano is MIT-licensed and maintained by Matatika, which acquired the open-source project in March 2026 after Arch, its previous steward, shut down. Meltano 4.4.0 shipped on Sep 29, 2026, the same day as Meltano SDK 0.54.7.
Meltano measures the run: did the tap extract what it was told to, did the target load it, did state advance. That is the right contract for an EL tool, and it means a run can be green while the table behind your revenue number stopped getting rows. Data Workers checks what every run loads, traces it through dbt and the reports on top, and gets the fix approved before anyone reports on a wrong number.
Key takeaways
- •Meltano keeps its job. Taps, targets,
meltano.yml, state, schedules and runs stay with your engineers. Data Workers works on what the runs land. - •A green run gets checked too. Freshness, volume, null and duplicate checks on loaded tables, plus the metrics your team records, turn a successful run that left data behind into a diagnosed incident.
- •Connected over Meltano's API or MCP server today. Data Workers reads the warehouse and the dbt project natively; Melty MCP sits next to the Data Workers agents in the same Claude session.
- •Every fix goes through a named person. The owner approves the hold and the dbt diff, changes
meltano.ymland reruns the pipeline; each step leaves a receipt. - •Autonomy is set per domain. Start at L1 observe, move to L2 propose, and open L3 act reversibly for narrow classes once the record supports it.
Meltano runs your Singer pipelines as code. Data Workers owns whether what they load is right.
Meltano runs the taps and targets your team declared, exactly as declared, and keeps state so nothing is extracted twice or skipped. The job after the run is different: notice that a successful run did not deliver what the business needs, find the cause, hold the report, get the project fixed, reload and prove the number. Here is one month-end at a B2B software company whose finance data comes from Xero through Meltano Cloud. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| Wed Sep 30, 11:20 | GitHub + Meltano Cloud | To trim compute hours, a data engineer replaces tap-xero's default select rule (.) with an explicit list of the streams finance uses. The pull request is reviewed and merged, and Meltano Cloud publishes the workspace. credit_notes is not on the list |
| 16:00 | Xero | Finance closes September and issues 212 credit notes for service credits and billing corrections, worth $48,300 |
| Thu Oct 1, 02:00 | Meltano Cloud | The scheduled pipeline runs meltano run tap-xero target-bigquery dbt-bigquery:build. The tap extracts every selected stream; invoices, payments and contacts land in the BigQuery dataset raw_xero. credit_notes is not extracted, which is exactly what the project asked for. The run succeeds |
| 02:14 | dbt | fct_net_revenue subtracts credit notes from invoices. With no new credits, September net revenue is overstated by $48,300. Every dbt and Elementary test passes |
| 02:20 | Data Workers + BigQuery | The load-lag metric the team records with monitor_metrics on raw_xero.credit_notes flags the last load as 24 hours old while every other Xero table loaded at 02:00, and the volume check finds no new rows. Two more metrics the team records with monitor_metrics fall outside their baselines: the daily credit-note row count (0 against 9 to 40 a day over 30 days) and the daily credit-note total ($0 at month-end). Data Workers opens an incident |
| 02:27 | Data Workers + dbt | Diagnosis: one stream is missing with no failed job and fresh loads on every other Xero table, so the stream was not extracted. Blast radius from the dbt lineage: stg_xero__credit_notes, fct_net_revenue, fct_customer_balances, and, from the team's context-graph note, the Evidence board pack, which builds at 09:30 for the 10:00 close review |
| 02:30 | Slack | Data Workers posts the incident to #data-finance and sends the Meltano project owner an approval request that links to Spellbook |
| 07:30 | Claude + Melty MCP | The owner asks Melty MCP for last night's run log in the same Claude session where the Data Workers agents run. The log lists the selected streams; credit_notes is absent since Wednesday's change |
| 07:40 | Spellbook | She opens the request and reviews three proposals: hold the 09:30 board-pack build; restore credit_notes to tap-xero's select list; a dbt diff that adds a source freshness test on credit_notes, so a missing stream fails the build next time. She approves all three |
| 07:50 | GitHub + Meltano Cloud | She holds the board-pack build, merges the select fix and the dbt diff, and asks Melty MCP to publish the workspace and trigger the pipeline. The stream's bookmark from Wednesday is still in state, so the tap resumes from it |
| 08:25 | Data Workers + BigQuery | Data Workers verifies: 212 rows landed, stg_xero__credit_notes has unique credit_note_id, the load-lag and row-count metrics are back in range, the recorded credit total reads $48,300, and the owner reports the new freshness test passing. The finance owner confirms the total against Xero's credit note report, which Data Workers does not read |
| 08:30 | Slack + Spellbook | Data Workers closes the incident with a link to the receipt: cause, approvals, the rerun, checks passed and how to undo each change |
| 09:30 | Evidence | The board pack builds on verified net revenue; the close review starts on time |

Meltano did everything it promises. It read the select list from a reviewed commit, extracted every stream on it, and kept the deselected stream's bookmark in state, so the fix needed no backfill. Catching the gap takes knowledge Meltano was never meant to hold: that credit_notes feeds net revenue, that credits spike at month-end, and that a board pack reads the number at 09:30. Xero and Evidence stay outside Data Workers' native set; both connect over their APIs today, and here the team's context-graph note is enough to name the board pack.
| Job | What Meltano does | What Data Workers does |
|---|---|---|
| The run | Runs taps, targets and dbt from meltano.yml; schedules, triggers and retries on Meltano Cloud | Reads what each run lands, natively in the warehouse; connects to Meltano over its API or MCP server |
| The catalog | Discovers streams and properties per tap; select rules decide what is extracted | Knows which tables feed which models, metrics and reports, and who owns each one |
| The signal | Run status, logs, row counts per run, an Unstable flag for pipelines with recent failures | Checks the data itself: freshness, volume, nulls, duplicates and recorded metrics, so a green run that left data behind still raises an incident |
| The diagnosis | Melty AI explains failed runs from logs, history, code and config | Explains a data incident when every run succeeded, with the blast radius across dbt models and reports |
| The fix | Runs whatever meltano.yml the team merges | Proposes the hold, the dbt diff and the project change to the owner, who approves and applies them |
| The proof | Keeps state and run history | Re-checks the tables and writes a receipt: what changed, who approved it, how it was checked, how to undo it |
Why doesn't Meltano just do this itself?
Because Meltano is built to run declared pipelines faithfully, and that is the right design. In Meltano's docs the select extra holds "an array of stream and property selection rules that are applied to the extractor's discovered catalog file". If a rule leaves a stream out, not extracting it is correct; flagging it would break every team that deselects streams on purpose. Meltano knows the run. It does not know that one deselected stream sits under a revenue number.
Meltano's AI features follow the same scope, and they are well chosen: AI diagnosis (Aug 7, 2026; in failure alert emails since Aug 27) explains failed runs from logs, recent history, pipeline code and configuration, redacted, on your own Claude API key. Agent Melty (Sep 3, 2026) seeds every new workspace repo with a knowledge base for coding agents. Melty MCP lets Claude list workspaces and pipelines, read runs, logs and row counts, diagnose failures, create pipelines and trigger runs, with exactly the permissions the user already has; it is documented and available on request while it awaits listing in Claude's connector directory. Secrets are entered in the app, never through the assistant. All of it serves the pipelines inside a workspace.
Owning whether the loaded data is right across Xero, BigQuery, dbt and a board report is a different product: a context graph of every table and consumer, blast-radius scoping, approvals that name a person, a recorded undo and receipts an auditor can read. That is Data Workers. More in is it safe to let AI agents change production data.
Every tool owns a slice. Data Workers covers the whole lifecycle
Meltano owns running Singer pipelines as code. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail.

| Stage | Data Workers | Meltano | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 3 | Each tap discovers a catalog of streams and properties, and meltano.yml records what the project extracts. Data Workers keeps one governed context graph of what each table means, who owns it and what reads it. |
| Analytics & Insights | 8 | 2 | Meltano feeds BI tools, semantic layers and notebooks; the numbers live there. Data Workers answers data questions from governed definitions with lineage behind every number. |
| Data Quality | 8 | 4 | dbt and Elementary tests run inside the Meltano project. Data Workers checks what landed for nulls, uniqueness and volume, tracks lateness against a baseline your team records, and turns a missing load into an incident. |
| Observability & Incidents | 8.5 | 4 | Meltano Cloud monitors runs, flags Unstable pipelines and explains failures with AI diagnosis. Data Workers diagnoses the data incident when every run succeeded, proposes the fix and verifies it. |
| Pipelines & Ingestion | 8.5 | 9 | Meltano's home stage: 600+ Singer taps and targets, incremental state, the Meltano SDK for custom sources and Cloud schedules. Data Workers plans the reruns around it for the owner. |
| Schema & Migration | 8 | 4 | Taps emit SCHEMA messages and targets add columns as sources change. Data Workers records each landed change against a baseline and traces it through dbt and BI. |
| Governance & Access | 8.5 | 4 | Pipeline changes go through Git review, with Entra ID sign-in and admin-only secrets on Cloud. Data Workers routes every data change to a named approver. |
| Security & Privacy | 8 | 4 | Sensitive settings are redacted in the UI, and secrets never pass through Melty AI. Data Workers leaves a receipt on every data change. |
| Cost / FinOps | 8 | 2 | Open source is free to run and Cloud bills compute hours; warehouse spend sits in the warehouse. Data Workers attributes Snowflake spend to dbt models and reads BigQuery spend from the Jobs API. |
| MLOps & Models | 7.5 | 1 | Meltano loads the data models and agents train and answer on. Data Workers keeps the data under models and agents healthy. |
How Meltano and Data Workers work together

Data Workers connects to Meltano over its API or MCP server today. What it needs sits where the target writes: the loaded tables in BigQuery, Snowflake, Databricks or PostgreSQL, read natively, and the dbt project, which gives lineage from each raw table to the models on top. Data Workers never runs, pauses or reloads a Meltano pipeline; the owner reruns through Melty MCP or Meltano Cloud. The rest of this incident runs on native connections (BigQuery, dbt, Slack) among 50+ connectors. More in Data Workers + dbt.
In Claude, your engineer asks Melty MCP for the run log and Data Workers what the missing stream did to revenue, in one conversation. Melty MCP signs in with your Meltano Cloud account. The Data Workers side follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry to the client's MCP config.
Example: four Data Workers agents in claude_desktop_config.json (Claude Desktop), with Melty MCP connected in the same app under Settings, Connectors.
{
"mcpServers": {
"dw-incidents": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-incidents"] },
"dw-quality": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-quality"] },
"dw-catalog": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-context-catalog"] },
"dw-connectors": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-connectors"] }
}
}Check the tools in your client's own tool list. Here: run_quality_check (dw-quality) runs the volume and uniqueness checks on BigQuery; monitor_metrics (dw-incidents) tracks load lag and the credit-note row count and total against their baselines; diagnose_incident names the cause; blast_radius_analysis and trace_cross_platform_lineage (dw-context-catalog) map what the missing stream reached; get_dbt_model_lineage (dw-connectors) reads the dbt graph; send_slack_alert posts the incident. If a source changes shape instead, run_quality_check reports the new columns and assess_impact (dw-schema) sizes the change downstream.
In production the agents run in your infrastructure and hold the warehouse credentials and model key. Your data stays in your systems; the hosted Conductor sees workflow metadata only. More in where does our data go.
One incident, L0 to L4, set per domain:

- •L0 manual. The CFO asks in the close review why net revenue looks high, and someone reads run logs.
- •L1 observe. Data Workers flags the empty credit-note load at 02:20 with the cause and blast radius. Nothing changes.
- •L2 propose. Data Workers proposes the hold, the
selectrestore and the dbt test; nothing moves until the named owner approves. - •L3 act reversibly. For a class with a clean record, Data Workers takes the reversible steps it is allowed, verifies the result and sends any failed check to a person.
- •L4 autonomous. For a scoped domain, Data Workers checks every run as it lands, so the owner's fix is waiting before the first report reads the table.
Who sets those levels: who owns the agents, how approvals work for AI data agents and autonomy levels L0 to L4 explained.
What changes for your team

- •Cost trims stop being risky. Engineers can narrow
selectlists and cut compute hours, because a stream that something downstream needed shows up as an incident with its consumers named. - •Code review gets a data check. A reviewed
meltano.ymlchange is checked again in what the next run actually loads. - •Month-end gets a guard. A finance table that goes quiet on its busiest day gets a proposed report hold before the build.
- •Engineers and finance share one record. One incident and receipt in Spellbook Data Catalog (in preview), linked from Slack.
The Incident Debugging agent runs the diagnosis. See also data freshness monitoring and SLAs, what is ELT and the open-source data stack guide.
Keep Meltano, or consolidate?
Keep Meltano if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
For most Meltano teams the answer is keep it: pipelines in Git, an MIT license, the Singer ecosystem and the SDK are hard to replace, and Meltano Cloud takes the hosting off your plate. What teams consolidate is the tooling around the loaded data: a separate observability tool, alert channels nobody reads, runbooks that say "check the run log". Many estates run another mover next to Meltano: see you're on dlt, you're on Airbyte and you're on Stitch, where Singer started. If Elementary already runs in your Meltano project, you're on Elementary shows how its test results feed the same incident record.
Building this yourself on Melty MCP and a coding agent? Read build it ourselves with Claude Code and MCP servers: reading a run log is the easy part; the context graph, approvals, undo and receipts are the work.
The case for your CFO
The outcome. When a Meltano run delivers less than the business relies on, it is caught the same night and corrected before anyone reports on it, with a record of how it was checked.
The risk story. At L0 and L1, agents only read. At L2 they propose and a named person approves; an unanswered request expires and escalates, never auto-grants. At L3 they act on reversible changes inside the domains you open; L4 is a later choice per domain. No agent can promote its own work, and an org-wide stop halts all autonomous dispatch. Every change carries a receipt: what changed, who approved it, the blast radius and how to undo it. Nothing migrates.
Why now. Meltano Cloud bills compute hours, so teams trim streams and schedules, and coding agents with Agent Melty and Melty MCP build pipelines in a conversation. More change, made faster, needs a check on what lands.
The first win. L1 on the Meltano datasets that feed finance and board reporting: a stream that went quiet becomes a diagnosed incident before the first report reads it.
What stays the same. Meltano, your project repo and review process, your dbt project, your warehouse and your on-call rota. For the numbers, see the ROI of agentic data operations.
The sentence for upstairs: "Meltano moves our data as code; Data Workers makes sure what it loaded is right and gets it fixed with our approval when it isn't, before we report on it."
Getting started
Start with a pilot. Pick the Meltano pipelines that feed the numbers leaders, boards and customers see, give Data Workers read access to their datasets and the dbt project, and run at L1 for a few weeks. Then turn on L2 for one domain, and open L3 for a narrow class once the receipts show the agents were right. Background: what is an agentic data platform and Data Workers vs data observability. The pilot path is on the pricing page, and the pilot is credited in full against the first year.
FAQ
How does Data Workers connect to Meltano? Over Meltano's API or MCP server today, with Melty MCP in the same Claude session for runs and logs. Data Workers reads the warehouse and the dbt project natively, so it works the same on Meltano Open and Meltano Cloud.
Every run was green. How can data be missing? A run succeeds when the tap extracts what the project selected and the target loads it. If a select change, a variant switch or a source permission change leaves a stream out, the run is still correct by its own contract. Data Workers checks the tables for volume and tracks lateness against a baseline your team records, so a quiet stream raises an incident.
Won't dbt source freshness tests catch this? They will once they exist on the right sources, which is why the fix in our example adds one. Most projects have them on a handful of tables. Data Workers checks the loaded tables after each run and proposes a test where the record shows one is worth adding.
Will Data Workers run, pause or reload our Meltano pipelines? No. meltano.yml, state, schedules and runs stay with your engineers. Data Workers proposes fixes, such as a dbt diff for the owner to merge; your team makes the project change and reruns in Meltano Cloud, through Melty MCP or from the CLI.
Do we still need Melty AI? Yes, on Meltano Cloud. It explains failed runs and builds pipelines with your coding agent. Data Workers works on what those pipelines land and everything downstream.
Is Meltano still maintained? Yes. Matatika acquired the Meltano open-source project in March 2026 and maintains it; Meltano stays MIT-licensed and community-driven, with 4.4.0 released on Sep 29, 2026.
Sources
- •Meltano, homepage ("The only EL tool built for data engineers"; 600+ connectors; "Your data stack as code" with Git-based governance; dbt and Elementary; "Meltano is an open source project by Matatika"), https://meltano.com/ (checked Oct 3, 2026)
- •Meltano blog, "Looking for Arch?" (Aaron Phethean, Mar 1, 2026: Arch shut down; Matatika acquired the Meltano open source project; Meltano remains open source and community-driven), https://meltano.com/blog/looking-for-arch (checked Oct 3, 2026)
- •Meltano, pricing (Meltano Open self-hosted; Meltano Cloud plans by compute hours), https://meltano.com/pricing/ (checked Oct 3, 2026)
- •Meltano releases (v4.4.0, Sep 29, 2026; v4.3.0, Sep 21, 2026), https://github.com/meltano/meltano/releases, and PyPI (4.4.0, MIT), https://pypi.org/project/meltano/ (checked Oct 3, 2026)
- •Meltano SDK releases (v0.54.7, Sep 29, 2026), https://github.com/meltano/sdk/releases, and https://pypi.org/project/singer-sdk/ (checked Oct 3, 2026)
- •Meltano docs, overview (Meltano Open and Meltano Cloud), https://docs.meltano.com/ (checked Oct 3, 2026)
- •Meltano docs, Melty MCP (capabilities, permissions, secrets, awaiting connector directory listing), https://docs.meltano.com/meltano-ai/melty-mcp (checked Oct 3, 2026)
- •Meltano docs, Agent Melty, https://docs.meltano.com/meltano-ai/agent-melty, and AI Diagnostics (own Claude API key; redaction before anything is sent), https://docs.meltano.com/meltano-ai/ai-diagnostics (checked Oct 3, 2026)
- •Meltano Cloud changelog: Aug 7, 2026 (AI pipeline diagnosis), https://docs.meltano.com/releases/cloud/2026-08-07-cloud-changelog; Aug 27, 2026 (diagnosis in failure alerts), https://docs.meltano.com/releases/cloud/2026-08-27-cloud-changelog; Sep 3, 2026 (Agent Melty in every workspace), https://docs.meltano.com/releases/cloud/2026-09-03-cloud-changelog; Sep 10, 2026 (Unstable pipelines), https://docs.meltano.com/releases/cloud/2026-09-10-cloud-changelog (checked Oct 3, 2026)
- •Meltano docs, Plugins (
selectextra, default.), https://docs.meltano.com/concepts/plugins, and State backends, https://docs.meltano.com/concepts/state_backends (checked Oct 3, 2026) - •Meltano Hub, tap-xero, https://hub.meltano.com/extractors/tap-xero/, and singer-io/tap-xero source (
credit_notesstream; state carried through and rewritten whole), https://github.com/singer-io/tap-xero (checked Oct 3, 2026) - •Data Workers open-source repository (tool registrations in dw-incidents, dw-quality, dw-context-catalog, dw-schema, dw-connectors), https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
- •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)