You're on Argo Workflows: Your Platform Team Runs Data and ML Workflows on Kubernetes. Data Workers Does the Operations Work on What They Produce
Argo Workflows runs your data and ML pipelines as Kubernetes resources. Data Workers catches the workflow that succeeded with the wrong output, finds the cause and hands the owner an approved fix and run plan.
Your platform team runs pipelines as Kubernetes resources. Each job is a Workflow: a DAG or a list of steps, every step a container. Reusable logic lives in WorkflowTemplates in the cluster, the nightly jobs are CronWorkflows with a schedule, a timezone and a concurrency policy, and expensive steps carry a retryStrategy and a memoize block so a rerun can skip work it already did. Argo Events starts workflows when a file lands in S3 or a message arrives on Kafka, and the templates themselves usually ship through a GitOps repo. Argo is a CNCF graduated project; v4.1.4 and v4.0.12 shipped on September 18, 2026, and the 3.7 line still takes patches.
The incidents that hurt most end with every node green: the workflow reads Succeeded, the pods exited zero, and the model scores or the warehouse table are still wrong because a step fed the next one an old input. Data Workers watches the tables your workflows write, traces a bad number back to the workflow, the node and the input behind it, and hands the owner the fix and a run plan to approve.
Key takeaways
- •Argo keeps its job. Workflows, templates, CronWorkflows, retries, the memoization cache and your GitOps flow stay where your platform team put them.
- •Green workflows get checked too. Freshness, volume, uniqueness and null checks on the tables your workflows write in Snowflake or BigQuery, plus baselines on the metrics you record, turn a
Succeededrun with stale output into a diagnosed incident the same night. - •Connected over the Argo Server API today. Argo Workflows connects over its API or MCP server today: your team's assistant reads workflows, node status and template specs, including whether a node came from a cache entry, side by side with Data Workers.
- •Every fix goes through a named person. The owner approves the template change, the hold and the run plan, and carries them out in Argo; each step leaves a receipt.
- •Autonomy is set per domain. Start at L1 observe, move to L2 propose, and open L3 act reversibly for narrow classes once the record supports it.
Argo Workflows is the engine room on your cluster. Data Workers is the operations shift for what comes out of it.
Argo schedules the pods, passes parameters and artifacts between steps, retries what you told it to, reuses cached outputs where you asked, and records the phase of every node. The work around a bad output is a different job: notice that a green night produced wrong numbers, connect them to a template change two days earlier, size the damage across the warehouse, the model registry and a customer publish, hold the publish, fix the cause, rerun and prove the result. Here is one night at a freight marketplace that prices truckload lanes with a model. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| Sun 15:20 | GitHub + Argo CD | A refactor of the lane-rates WorkflowTemplate merges in the GitOps repo and syncs to the cluster. The build-features template keeps its memoize block (cache ConfigMap lane-features-cache, maxAge 48h), but its key changes from lane-features-{{workflow.parameters.region}}-{{workflow.parameters.ds}} to lane-features-{{workflow.parameters.region}}. The date is gone |
| Mon 02:00 | Argo | The CronWorkflow lane-rates-nightly runs. The new key has no cache entry, so build-features runs Spark, writes Sunday's partition of the Iceberg table features.lane_daily and caches its output parameter snapshot_path |
| Tue 02:00 | Argo + PostgreSQL | lane-rates-nightly runs again. extract-shipments pulls Monday's 41,800 loads from the transport management system's PostgreSQL replica, plus Monday's diesel price update. build-features finds the key in the cache and returns Monday's snapshot_path without running. Spark never starts; the node shows Succeeded |
| 02:31 | Argo + MLflow | train fits the lane model on Sunday's features, logs the run to MLflow, registers version 58 and moves the champion alias to it |
| 02:48 | Argo + Snowflake | score-lanes writes 12,400 lane quotes to pricing.lane_rate_quotes in Snowflake. The workflow is Succeeded at 02:52 |
| 03:00 | Data Workers + Snowflake | The load baseline the team records with monitor_metrics on features.lane_daily (an Iceberg table queried through Snowflake) shows the newest partition two days old against a one-day norm. The day-over-day change in mean quoted rate per mile, a metric the pricing team records with Data Workers, reads 0.00% against a baseline of plus or minus 1.8%. Data Workers opens an incident |
| 03:06 | Data Workers + Argo API + Snowflake | The on-call's assistant reads the night's workflow over the Argo Server API and hands it to Data Workers: the build-features node has memoizationStatus.hit: true under key lane-features-us-central, and the template spec no longer puts ds in the key. Snowflake shows no new data in features.lane_daily since Monday 02:41. Blast radius from lineage: pricing.lane_rate_quotes, MLflow version 58 under champion, the CronWorkflow quotes-publish that pushes spot quotes to the shipper portal at 06:00, and the Metabase "Lane margin" dashboard |
| 03:10 | PagerDuty + Slack | Data Workers raises a PagerDuty incident with the diagnosis and sends the approval request in Slack to the ML platform owner for the pricing workflows |
| 05:20 | Spellbook + Argo + MLflow | The owner reviews four proposals in Spellbook: suspend quotes-publish; a template diff that puts {{workflow.parameters.ds}} back in the key; delete the stale cache entry; point champion back to version 57, then run lane-rates for Monday and resume the publish. She approves, runs argo cron suspend quotes-publish, merges the diff (Argo CD syncs it), deletes the key from the ConfigMap with kubectl and moves the alias in MLflow |
| 05:34 | Argo | She runs argo submit --from workflowtemplate/lane-rates -p ds=<Monday>. build-features misses the cache, Spark builds Monday's partition, train registers version 59 and score-lanes rewrites the quotes. Succeeded at 06:21 |
| 06:24 | Data Workers + Snowflake | Data Workers verifies: Monday's partition is present, pricing.lane_rate_quotes holds 12,400 rows with no duplicate lane_id, and the quote-change metric is back inside its range |
| 06:26 | PagerDuty | Data Workers resolves the incident and links the receipt in Spellbook: cause, approvals, the run, checks passed and the undo |
| 06:28 | Argo + shipper portal | The owner resumes quotes-publish and triggers it by hand; spot quotes reach shippers at 06:30 on current features |
| 08:00 | Metabase | The pricing review opens "Lane margin" on verified numbers |

Every part of Argo did its job: the key matched an entry younger than 48 hours, and Argo returned the cached output as designed. The fault was a one-line key change, merged two days before it mattered and visible only in the data: a feature table that stopped moving and quotes that held still on a day diesel moved. Catching it takes knowledge Argo was never meant to hold: which table carries the features, what fresh means for it, and which workflow publishes prices next.
| Job | What Argo Workflows does | What Data Workers does |
|---|---|---|
| The run | Runs DAG and steps templates as pods; CronWorkflows schedule them, Argo Events starts them on events; passes parameters and artifacts | Reads workflows, nodes and template specs over the Argo Server API |
| The signal | Node phases, retries, exit handlers, Prometheus metrics and the workflow archive | Checks the data itself: volume, uniqueness and nulls on the tables workflows write, and baselines, freshness included, on recorded metrics, so a green run with wrong output still raises an incident |
| The diagnosis | Shows which node failed, its logs, and whether a node came from the cache | Joins the workflow, the cache hit, the template change, the Iceberg snapshots and lineage into one cause, with the blast radius across tables, models, dashboards and downstream workflows |
| The fix | Runs whatever template your GitOps flow syncs | Proposes the template diff, the cache cleanup, the alias move and the hold to the owner, who approves and applies them |
| The rerun | Runs the workflow it is given, from the UI, the CLI or the API; argo retry and argo resubmit repeat a run | Proposes the run plan in order; the owner submits it in Argo, and Data Workers reads each run's status |
| The proof | Keeps the workflow record and node outputs | Re-checks the tables and writes a receipt: what changed, who approved it, how it was checked, how to undo it |
Why doesn't Argo Workflows just do this itself?
Because Argo is a general workflow engine for Kubernetes, running ML training, batch jobs, infrastructure automation and CI/CD. Its unit is the Workflow and the phase of each node. A pod that exits zero is a success, whether it wrote a fresh partition, an old one or nothing. That is the right design for an engine that stays neutral about what every container does.
Memoization shows the same choice. Argo's docs say it "was designed for 'pure' steps", whose outputs depend only on their inputs, and add: "Pure steps should not interact with the outside world, but workflows won't enforce this on you." Argo trusts the key you write. Whether it still captures everything the step depends on is a data question, outside the engine on purpose.
Argo ships no AI agent and no official MCP server. Community servers fill the gap, such as Pipekit's, built around the argo CLI. They are useful for engineers, and they stay at the level of the run.
Owning whether a workflow's output is right across a transport database, a lakehouse table, a model registry, a warehouse and a customer-facing publish is a different product: a context graph of every table, model and downstream workflow, blast-radius scoping, approvals that name a person, a recorded undo and receipts an auditor can read. That is Data Workers. For how that holds up in production, read is it safe to let AI agents change production data.
Every tool owns a slice. Data Workers covers the whole lifecycle
Argo Workflows owns running containerized workflows on Kubernetes. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the Argo workflows already there.

| Stage | Data Workers | Argo Workflows | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 1 | Not Argo's job: it knows Workflows, templates, nodes and artifacts, not what a table means or who owns it. Data Workers keeps one governed context graph of tables, models, lineage and owners. |
| Analytics & Insights | 8 | 1 | Not Argo's job: the UI shows workflow graphs and node phases, not business numbers. Data Workers answers data questions from governed definitions with lineage behind every number. |
| Data Quality | 8 | 2 | A step can run any test container, and a failed exit fails the node. Data Workers runs volume, uniqueness and null checks on the Snowflake and BigQuery tables workflows write, and flags recorded metrics, freshness included, that leave their baseline. |
| Observability & Incidents | 8.5 | 4 | Node phases, retries, exit handlers and Prometheus metrics show a broken run. Data Workers diagnoses the data incident across systems, proposes the fix and verifies it. |
| Pipelines & Ingestion | 8.5 | 9 | Argo's home stage: DAG and steps templates, WorkflowTemplates, CronWorkflows, retries, memoization and artifacts, all as Kubernetes resources. Data Workers plans the runs for the owner to submit. |
| Schema & Migration | 8 | 2 | Not Argo's job: a step writes whatever schema its container emits. Data Workers catches schema changes and assesses blast radius before downstream models break. |
| Governance & Access | 8.5 | 3 | Kubernetes RBAC, service accounts and Argo Server auth modes govern who can submit what. Data Workers routes every data change to a named approver. |
| Security & Privacy | 8 | 4 | Pods run in your cluster with Kubernetes secrets and namespace isolation; the data a step writes is classified elsewhere. Data Workers leaves a receipt on every data change. |
| Cost / FinOps | 8 | 3 | Not Argo's job: pod resources and memoization save compute, while warehouse spend sits in the warehouse. Data Workers traces Snowflake credits to the dbt model behind them through query tags. |
| MLOps & Models | 7.5 | 6 | Argo is a common engine for training and batch scoring, and Kubeflow Pipelines builds on it. Data Workers keeps the data under those models healthy. |
How Argo Workflows and Data Workers work together

Argo Workflows connects over its API or MCP server today. With an access token for a service account that can read workflows and templates in the namespaces you choose, your team's assistant reads workflows, node status (including each node's memoizationStatus), logs and template specs over the Argo Server REST API, side by side with Data Workers, whose own checks rest on the tables the workflows write. Templates, CronWorkflows, cache ConfigMaps, service accounts and runs stay with your platform team. When a fix needs a template change, Data Workers proposes the YAML as a diff, the owner merges it through your GitOps flow and submits the approved runs. The data side runs natively: Snowflake (including the Iceberg tables it reads from Polaris), PostgreSQL, PagerDuty and Slack in this incident, plus BigQuery, Databricks and 50+ connectors in all. MLflow and Metabase connect over their APIs.
Engineers keep the argo CLI, or a community Argo MCP server, next to the Data Workers agents in the same client: the Argo side lists and submits workflows; Data Workers answers "what did last night's workflows do to the data, and what's the fix". Setup follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry to your client.
# Example: Data Workers agents in Claude Code, from a clone of the open-source repo
claude mcp add --scope user dw-incidents -- "$(pwd)/start-agent.sh" dw-incidents
claude mcp add --scope user dw-quality -- "$(pwd)/start-agent.sh" dw-quality
claude mcp add --scope user dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
# Argo side stays on your own tooling, scoped to the namespace you choose
export ARGO_SERVER=argo.platform.example.com:443
export ARGO_TOKEN="Bearer $(kubectl create token argo-readonly -n pricing)"
argo list -n pricing --status SucceededList the tools with your client's own command (/mcp in Claude Code). In this incident: run_quality_check (dw-quality) runs the volume and uniqueness checks on Snowflake; monitor_metrics (dw-incidents) flags the quote-change metric against its baseline; diagnose_incident names the cause; blast_radius_analysis and trace_cross_platform_lineage (dw-context-catalog) map what the stale features reached; remediate re-checks the quality assertions after the rerun and escalates any failure to a person; get_incident_history shows whether the lane workflow has broken before. Keep the Argo token read-only.
In production the agents run in your infrastructure, often in the same cluster as Argo, and hold the warehouse credentials and model key. Your data stays in your systems; the hosted Conductor sees workflow metadata only. More in where does our data go.
One incident, L0 to L4, set per domain:

- •L0 manual. A shipper asks at 10:00 why Tuesday's quotes ignored the diesel jump. An engineer finds every node green and spots the cache hit by early afternoon; stale quotes were live for eight hours.
- •L1 observe. Data Workers flags the stale feature table at 03:00 with the cause and blast radius. Nothing changes.
- •L2 propose. Data Workers proposes the hold, the template diff, the cache cleanup, the alias move and the run plan. Nothing reaches production until the named owner approves; an unanswered request expires and escalates, never auto-grants.
- •L3 act reversibly. For a class with a clean record, Data Workers carries out reversible steps inside the domain you open, verifies them and sends any failed check to a person. Changes in Argo stay with the owner.
- •L4 autonomous. For a scoped domain, Data Workers checks the feature table right after the nightly window, so the stale step is flagged within the hour and the owner's fix is waiting before anything publishes.
Who sets those levels: who owns the agents, how approvals work for AI data agents and autonomy levels L0 to L4 explained.
What changes for your team

- •Platform engineers stop owning data outcomes by default. The Argo cluster stays theirs; the question "is the output right" gets its own owner, checks and record. See Data Workers for data platform engineers.
- •ML engineers get a guard before the registry and the publish. When a training or scoring workflow consumed stale features, the hold and the alias move are proposed before customers see the scores. See Data Workers for ML engineers and data lineage for ML features.
- •One record for platform, ML and analytics. Everyone sees the same incident and receipt in Spellbook Data Catalog (in preview), linked from the PagerDuty incident.
For the wider tool landscape, see the data orchestration tools guide and why MCP for MLOps.
Keep Argo Workflows, or consolidate?
Keep Argo Workflows if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
For most Kubernetes platform teams the answer is keep it: Argo is where the cluster's batch and ML work already runs. What teams consolidate is the tooling around the outputs: a separate observability tool, check containers copied into every template, and "resubmit and look" runbooks. Many estates also run a second orchestrator for analytics; Data Workers connects natively to Airflow, Dagster, Prefect, Azure Data Factory and Managed Service for Apache Airflow, so one incident record spans them. See you're on Airflow, you're on Dagster and you're on Kestra. How the pieces connect is in Data Workers integrations.
Weighing a build on a community Argo MCP server and a coding agent? Read build it ourselves with Claude Code and MCP servers. Listing and submitting workflows is the easy part; the context graph, approvals, undo and receipts are the work.
The case for your CFO
The outcome. When an Argo workflow produces wrong prices, model scores or finance numbers, it is caught the same night and corrected before customers see it, with a record of how it was checked.
The risk story. At L0 and L1, agents only read. At L2 they propose and a named person approves; an unanswered request expires and escalates, never auto-grants. At L3 they act on reversible changes inside the domains you open; L4 is a later choice per domain. Changes in Argo stay with the owner. No agent can promote its own work, and an org-wide stop halts all autonomous dispatch. Every change carries a receipt: what changed, who approved it, the blast radius and how to undo it.
Why now. More of the business runs on model output from Kubernetes workflows: prices, risk scores, forecasts. Those workflows report success on the run; someone has to own whether the result is right before a customer sees it.
The first win. L1 on the workflows that feed customer-facing prices or scores: a stale input becomes a diagnosed incident before the morning publish.
What stays the same. Argo Workflows, your templates, your GitOps flow, your cluster, your warehouse and your on-call rota. For the numbers, see the ROI of agentic data operations.
The sentence for upstairs: "Argo runs our data and ML workflows on Kubernetes; Data Workers checks what they produce and gets it fixed with our approval before a customer sees a wrong number."
Getting started
Start with a pilot. Pick the Argo workflows that feed the numbers customers and leaders read, give Data Workers a read-only Argo Server token and read access to the tables they write, and run at L1 for a few weeks. Then turn on L2 for one domain, and open L3 for a narrow class once the receipts show the agents were right. The pilot path is on the pricing page, and the pilot is credited in full against the first year.
FAQ
How does Data Workers connect to Argo Workflows? Over the Argo Server REST API today, with an access token for a service account you scope to the namespaces that matter. It reads workflows, node status, logs and template specs. The same API serves the 4.1, 4.0 and 3.7 lines.
Argo already retries failed steps. What does Data Workers add? A retry answers "did the pod finish". Data Workers answers "is the data right": a step that succeeded on a stale input still raises an incident, with the cause and the fix.
Will Data Workers change our templates, CronWorkflows or cache? No. In Argo it reads. Template changes arrive as a YAML diff for the owner to merge through your GitOps flow, and the owner suspends or resumes CronWorkflows, clears cache entries and submits runs.
We use memoization heavily. How do we keep it from serving stale outputs? Put every input the step depends on in the key, including the date or snapshot it reads, and set a maxAge that matches how often the input changes. Data Workers adds the backstop: lateness baselines on the tables the memoized step feeds, with the cache hit named in the diagnosis.
Does this work for Kubeflow Pipelines on Argo? The pattern is the same: the pipelines run as Argo workflows, and Data Workers checks the feature and prediction tables they read and write in your warehouse.
Should we let assistants submit Argo workflows over MCP? Start read-only, on a token scoped to one namespace, and keep anything that publishes to customers behind an approval. Data Workers adds the context and the approval flow.
Sources
- •Argo Workflows documentation, overview ("an open source container-native workflow engine for orchestrating parallel jobs on Kubernetes"; CNCF graduated; use cases; REST API), https://argo-workflows.readthedocs.io/en/latest/ (checked Oct 3, 2026)
- •Argo Workflows releases: v4.0.0 (Feb 4, 2026), v4.1.0 (Aug 11, 2026), v3.7.18 (Aug 14, 2026, latest 3.7 patch), v4.1.4 and v4.0.12 (Sept 18, 2026), https://github.com/argoproj/argo-workflows/releases (checked Oct 3, 2026)
- •CNCF, Argo project page (accepted Mar 26, 2020; graduated Dec 6, 2022), https://www.cncf.io/projects/argo/ (checked Oct 3, 2026)
- •Argo Workflows docs, Step Level Memoization (cache ConfigMaps, key, maxAge, pure steps), https://argo-workflows.readthedocs.io/en/latest/memoization/ (checked Oct 3, 2026)
- •Argo Workflows docs, Field Reference (MemoizationStatus: hit, key, cacheName), https://argo-workflows.readthedocs.io/en/latest/fields/ (checked Oct 3, 2026)
- •Argo Workflows docs, Cron Workflows (schedules, timezone, suspend, concurrency policy), https://argo-workflows.readthedocs.io/en/latest/cron-workflows/ (checked Oct 3, 2026)
- •Argo Workflows docs, Workflow Templates, https://argo-workflows.readthedocs.io/en/latest/workflow-templates/, and Retries, https://argo-workflows.readthedocs.io/en/latest/retries/ (checked Oct 3, 2026)
- •Argo Workflows docs, REST API and CLI (
argo cron suspend,argo submit --from), https://argo-workflows.readthedocs.io/en/latest/rest-api/, https://argo-workflows.readthedocs.io/en/latest/cli/argo_cron_suspend/ (checked Oct 3, 2026) - •Argo Events, https://argoproj.github.io/argo-events/ (checked Oct 3, 2026)
- •Community MCP server for Argo Workflows (pipekit/mcp-for-argo-workflows, "a golang MCP server based around the CLI capabilities of workflows", last pushed Sept 25, 2026); no MCP repository in the argoproj GitHub organization, https://github.com/pipekit/mcp-for-argo-workflows (checked Oct 3, 2026)
- •Data Workers open-source repository (tool registrations in dw-incidents, dw-quality, dw-context-catalog), https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
- •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)