Product
Product11 min readBy The Data Workers Team

You're on Control-M: Your Enterprise Schedules and Governs Batch Across Mainframe, ERP and Cloud. Data Workers Does the Operations Work on the Data Those Jobs Produce

Control-M runs and governs enterprise batch from z/OS to Snowflake. Data Workers checks the data each folder delivers, finds the cause when a job ends OK with wrong data, and hands the owner a fix and a run plan.

Your batch lives in Control-M. Every night the schedule walks through folders of jobs: z/OS jobs that run COBOL extracts, SAP jobs, File Transfer jobs that move files to Amazon S3 or a partner, and plug-in jobs for AWS Glue, Snowflake, dbt and Power BI. Calendars and dependencies decide when each job runs, and a Service with a deadline tells operations whether the business process will finish on time. Operators act on a job that ends Not OK from the Monitoring domain (hold, rerun, release, set to OK); developers keep jobs-as-code through the Automation API. BMC Software positions Control-M as "The Agentic Orchestration Layer for Enterprise Operations", on SaaS, self-hosted or hybrid.

The incidents that hurt most in a Control-M shop finish green: every job ends OK, the Service meets its deadline, and the numbers in the warehouse are still wrong. Control-M governs the batch: what runs, when, in what order, by whose authority. Data Workers runs operations on the data that batch delivers: it checks what each folder wrote, traces a bad number back to its cause, and hands the owner a fix and a run plan.

Key takeaways

  • •Control-M keeps its job. Folders, calendars, Services, job definitions and authorizations stay with your scheduling team. Data Workers works on the data the jobs produce.
  • •Ended OK gets checked too. Volume, freshness and null checks, plus metrics your team defines, run on the tables each folder writes, so a green night with wrong data becomes a diagnosed incident before the business day.
  • •Connected over the Control-M Automation API or MCP Server today. Operators run the Control-M MCP Server next to the Data Workers agents in the same assistant; Data Workers checks the tables the jobs write.
  • •Job actions stay in Control-M. Data Workers recommends holds, reruns and releases; the scheduling team carries them out in Control-M. The data owner approves the code fix in Spellbook, and every step leaves a receipt.
  • •Autonomy is set per domain. Start at L1 observe, move to L2 propose, and open L3 act reversibly for narrow classes once the record supports it.

Control-M governs the enterprise batch. Data Workers runs operations on the data it delivers.

Control-M starts each job when its conditions are met and records how it ended. Wrong data is a different job: notice it, connect it to a change on the mainframe, size the damage, stop anything that would move money on it, fix the cause, rerun the right jobs and prove the result. Here is one night at a card issuer's settlement batch. This is an illustration, not a customer case.

TimeSystemWhat happens
Mon 21:30z/OSThe cards team promotes a release that changes the copybook STLMREC. Refunds used to arrive with a negative STL-AMT; they now carry a positive amount and a new one-byte field at the end of the record, STL-DC-IND, set to C for credits. The file stays variable-length, so readers built on the old layout still parse every field they know
Tue 01:00Control-M (z/OS job)STLEXT01 runs the COBOL extract and writes the settlement file: a header, 1.42 million detail records and a trailer with the record count and net amount. Ended OK
01:22Control-M MFT + S3The File Transfer job STL_MFT_S3 moves the file in binary to the S3 landing bucket. Ended OK
01:35Control-M + AWS GlueThe Glue job stlmt_ebcdic_to_parquet decodes the EBCDIC records with Cobrix, the open-source COBOL reader for Spark, using copybook version 14. Its log notes records one byte longer than the layout and ignores the extra byte. The record count matches the trailer. Ended OK
02:05Control-M + SnowflakeSTL_SF_LOAD loads the day into RAW.CARDS.SETTLEMENT_DETAIL and the trailer into RAW.CARDS.SETTLEMENT_CONTROL. 31,840 refunds, worth $2.09M, land as positive amounts. Ended OK
02:30Control-M + dbtSTL_DBT_BUILD builds stg_settlement, fct_merchant_settlement and fct_interchange_revenue. Every not_null, unique and accepted_values test passes, because a positive amount is a valid amount. The Service "Card settlement" shows on time for its 06:30 deadline
02:52Data Workers + SnowflakeThe team's credit-share metric for the settlement table is recorded with monitor_metrics after each load. Tonight it reads 0.0% against a baseline of 2.0 to 2.4%, and the metric is flagged. run_quality_check passes volume and nulls. Data Workers opens an incident
02:58Batch on-call + Data WorkersThe batch on-call's assistant reads the folder's job runs, output and logs over the Control-M MCP Server and hands them to Data Workers: every job ended OK, and the Glue log shows the longer records. The run_quality_check profile shows no column change in Snowflake, which points at the record layout upstream
03:01Data Workersblast_radius_analysis walks lineage: the three dbt models, plus the team's context-graph notes for the Power BI "Daily settlement" report that STL_PBI_REFRESH refreshes at 05:00 and for MPAY_FILE_GEN, the job that builds the merchant payout file at 06:00. 2,140 merchants with refunds that day would be paid as if their refunds were sales
03:04PagerDuty + SlackData Workers raises a PagerDuty incident with the diagnosis and recommends holding MPAY_FILE_GEN and STL_PBI_REFRESH. The approval request for the fix goes to the settlement data owner in Slack, with the batch on-call and the cards mainframe owner copied
03:10Control-MThe batch on-call holds both jobs in the Monitoring domain
03:25SpellbookThe data owner reviews the proposal: a diff for the conversion repo that adds STL-DC-IND to the copybook as version 15 and signs credit amounts negative, plus a dbt test that fails the build when the day's net differs from the trailer, and a run plan (decode, load, dbt, then release the holds). She checks the trailer against the loaded net: $4.18M apart, twice the day's refunds. The mainframe owner confirms the release note. She approves
03:45GitHubThe data engineer on call merges the conversion diff and the dbt test
03:55Control-MThe batch on-call reruns the decode, STL_SF_LOAD (which replaces the day's rows, as it always does) and STL_DBT_BUILD. Ended OK by 04:35, and the new trailer test passes
04:40Data Workers + SnowflakeData Workers verifies: credit share 2.2%, back in its band; volume, freshness and nulls pass; remediate re-checks the quality assertions and finds none failing
04:46PagerDutyData Workers resolves the incident with the receipt: cause, approvals, jobs rerun, checks passed and the undo (revert the conversion change and rerun from decode)
04:50Control-MThe batch on-call releases both holds. Power BI refreshes at 05:00, and the payout file goes out at 06:30 with refunds netted correctly
Incident timeline across the stack: what Control-M, your team and Data Workers each do, step by step

Every part of Control-M did its job. The fault was a change in what one byte meant, made by another team for good reasons. Catching it takes knowledge a scheduler was never meant to hold: what share of settlement rows are normally credits, and which job turns those rows into money at 06:30.

JobWhat Control-M doesWhat Data Workers does
The runRuns z/OS, File Transfer, Glue, Snowflake, dbt and Power BI jobs in order by calendars and conditionsReads job runs, output and logs over the Automation API
The signalJob statuses, alerts, Services with deadlines, Execution Insights reports on failures and slow jobsChecks the data: volume, freshness, nulls and team-defined metrics on the tables each folder writes
The diagnosisAI Pilot for Monitoring analyzes a Not OK or Wait Host job from its output, logs, definition and host statusJoins job runs, warehouse checks and lineage into one cause, with the blast radius across models, reports and downstream jobs
The fixRuns whatever definitions and scripts the team deploys; AI Pilot for Planning creates and configures jobs and foldersProposes the code change as a diff for the owner to approve and merge
The rerunHolds, reruns and releases jobs from the Monitoring domain, the Automation API or the MCP ServerRecommends the holds and the run plan; the scheduling team runs them in Control-M, and Data Workers reads each run's status
The proofJob change and action audits; the Archive Service keeps run historyRe-checks the tables and writes a receipt: what changed, who approved it, how it was checked, how to undo it

Why doesn't Control-M just do this itself?

Because Control-M is built to run the enterprise's work on time and under control. Its unit is the job: when it may start, what it waits for, how it ended. It knows STL_SF_LOAD ended OK at 02:05; it does not know that 2% of settlement rows should be credits.

Control-M's AI follows that scope. Jett answers natural-language questions from the Control-M database: job status, definition changes, user actions. Execution Insights builds reports from the last four days of job history: daily snapshot, top failures, jobs slower than usual, throughput and delay. AI Pilot for Monitoring analyzes a job that is Not OK or waiting for a host; AI Pilot for Planning creates and configures jobs and folders; AI Workflow Creator, a Preview feature, builds a workflow from a prompt. The Control-M MCP Server, also in Preview, lets AI assistants trigger jobs, check workflow status and investigate failures, and supports elicitation so the person confirms an action before it runs. Each works on Control-M's own objects, inside its authorizations: the right design for a system that runs payroll, settlement and month-end close.

Owning whether the data a batch delivered is right, across a mainframe, a converter, a warehouse, dbt, BI and a payout file, is a different product: a context graph, blast-radius scoping, approvals that name a person, a recorded undo and receipts an auditor can read. A scheduler that rewrote warehouse data or dbt models would take on liability for systems other teams own. That product is Data Workers. More in is it safe to let AI agents change production data.

Every tool owns a slice. Data Workers covers the whole lifecycle

Control-M owns scheduling and governing enterprise batch. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the Control-M schedule already there.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Control-M goes deep on its own area
StageData WorkersControl-MWhy we scored it this way
Catalog & Context92Not Control-M's job: it knows jobs, folders, calendars and Services, not what a table means or who owns it. Data Workers keeps one governed context graph of tables, models, lineage and owners.
Analytics & Insights82Not Control-M's job: Jett and Execution Insights report on job health, not business numbers. Data Workers answers data questions from governed definitions with lineage behind every number.
Data Quality83A job can run any test tool, and a job that ends Not OK stops its successors. Data Workers checks the tables the batch writes after every run, including metrics the team defines.
Observability & Incidents8.57Services with deadlines, alerts, AI Pilot for Monitoring and Execution Insights track the batch closely. Data Workers diagnoses the data incident across systems, proposes the fix and verifies it.
Pipelines & Ingestion8.59Control-M's home stage: jobs across z/OS, SAP, file transfer and cloud, plug-ins for Glue, Snowflake, dbt and Power BI, and jobs-as-code. Data Workers plans the reruns for the scheduling team.
Schema & Migration82Not Control-M's job: a job moves whatever layout the source writes. Data Workers catches schema changes and assesses blast radius before downstream models break.
Governance & Access8.56Role authorizations, job change and action audits, and the Archive Service govern the schedule. Data Workers routes every data change to a named approver.
Security & Privacy85Role authorizations and encrypted file transfer (SFTP, FTPS, AS2 with PGP) protect the runs; the data itself is classified elsewhere. Data Workers leaves a receipt on every data change.
Cost / FinOps83Not Control-M's job: it shapes when work runs, while warehouse spend sits in the warehouse. Data Workers traces Snowflake credits to the dbt model behind them through query tags.
MLOps & Models7.53Plug-ins run SageMaker, Azure ML and agent frameworks as jobs. Data Workers keeps the data under those models healthy.

How Control-M and Data Workers work together

How Data Workers fits with Control-M: your coding agent on top, Data Workers in the middle, your estate underneath

Control-M connects over its API or MCP server today: with an API token scoped to a read role, your team's assistant reads job runs, statuses, output and logs, side by side with Data Workers, whose own checks rest on the tables the jobs write. Job definitions and job actions stay with your scheduling team: Data Workers recommends the actions, and a fix to a job or script arrives as a diff for the owner to merge through your promotion path. On the data side, Snowflake (with per-query credit attribution through query tags), dbt, PagerDuty and Slack connect natively; the Glue Data Catalog, S3 and Power BI connect over their APIs today; the mainframe stays with its owners and their tooling, and Data Workers sees it through the jobs and files it produces.

Operators can run the Control-M MCP Server and the Data Workers agents side by side in one assistant. Control-M's server answers "which jobs ended Not OK today" and carries out job actions with elicitation asking the person to confirm; Data Workers answers "what did last night's batch do to the data, and what's the fix". BMC labels the MCP Server Preview and recommends non-production environments, so many teams start it on a test environment while Data Workers checks the production tables the jobs write. The Data Workers side follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry to your client. The Control-M side is enabled with config systemsettings:mcp::enable and uses the endpoint and API token from BMC's MCP Server page.

// Example: Control-M MCP Server plus Data Workers agents in one client config
// (Cursor-style mcpServers block; Control-M settings follow BMC's MCP Server docs)
{
  "mcpServers": {
    "controlm": {
      "url": "https://<tenant-name>-aapi.<zone>.controlm.com/automation-api/mcp/message",
      "type": "http",
      "headers": { "x-api-key": "${env:CTM_API_TOKEN}" }
    },
    "dw-incidents": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-incidents"] },
    "dw-quality":   { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-quality"] },
    "dw-schema":    { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-schema"] },
    "dw-catalog":   { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-context-catalog"] }
  }
}

List the tools with your client's own command. In this incident: monitor_metrics and diagnose_incident (dw-incidents), run_quality_check (dw-quality), and blast_radius_analysis and trace_cross_platform_lineage (dw-context-catalog); remediate re-checks the quality assertions and escalates any failure to a person.

In production the agents run in your infrastructure and hold the warehouse credentials and model key. Your data stays in your systems; the hosted Conductor sees workflow metadata only. More in where does our data go.

One incident, L0 to L4, set per domain:

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. Merchants call on Wednesday about refunds. Every job ended OK, and two teams spend a day matching the file to the warehouse.
  • •L1 observe. Data Workers flags the credit share at 02:52 with the cause and blast radius. Nothing changes.
  • •L2 propose. Data Workers proposes the conversion diff, the trailer test and the run plan, and recommends the holds. Nothing reaches production until the named owner approves; an unanswered request expires and escalates, never auto-grants.
  • •L3 act reversibly. For a class with a clean record, Data Workers carries out reversible steps inside the domain you open, verifies them and sends any failed check to a person. Holds, reruns and releases in Control-M stay with the scheduling team.
  • •L4 autonomous. For a scoped domain, Data Workers watches the settlement metrics as each load ends, and the fix is waiting for the owner before any downstream job runs.

What changes for your team

Six jobs that run on autopilot with Data Workers next to Control-M, with a concrete example of each
  • •The batch on-call gets a data answer. When a Service is on time and the data is wrong, the operator sees the cause and the run plan instead of a ticket at 09:00.
  • •Jobs that move money get a guard. When a payout, invoicing or regulatory job reads a table that just went wrong, the hold is recommended before it runs.
  • •Mainframe, scheduling and data teams share one record. All three see the same receipt in Spellbook Data Catalog (in preview), linked from the PagerDuty incident.

Keep Control-M, or consolidate?

Keep Control-M if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

For most Control-M shops the answer is keep it: it runs work no data tool should own, from z/OS and SAP to file transfer. Teams consolidate the tooling around the data: a separate observability product, test jobs copied into every folder, runbooks that say "rerun and ask finance to check". Where Airflow or Azure Data Factory run next to Control-M, Data Workers connects to them natively, so one incident record spans them. See you're on Airflow, you're on Azure Data Factory, you're on Argo Workflows and you're on dbt.

Weighing a build on the Control-M MCP Server and a coding agent? Calling job actions is the easy part; the context graph, approvals, undo and receipts are the work. See build it ourselves with Claude Code and MCP servers.

The case for your CFO

The outcome. When a batch delivers wrong settlement, finance or regulatory numbers, the error is caught the same night and corrected before money moves or a report goes out.

The risk story. At L0 and L1, agents only read. At L2 they propose and a named person approves; an unanswered request expires and escalates, never auto-grants. At L3 they act on reversible changes inside the domains you open; L4 is a later choice per domain. Holds, reruns and releases stay with the scheduling team in Control-M. No agent can promote its own work, and an org-wide stop halts all autonomous dispatch. Every change carries a receipt: what changed, who approved it, the blast radius and how to undo it.

Why now. Control-M now invites AI agents into the schedule through its MCP Server, and mainframe data increasingly lands in cloud warehouses that feed customers and regulators. More handoffs mean more nights that end OK with wrong data.

The first win. L1 on the folders that feed payouts, the general ledger and regulatory reports: a layout change becomes a diagnosed incident before the first downstream job runs.

What stays the same. Control-M, your folders, calendars and Services, your mainframe jobs, your dbt project, your warehouse and your on-call rota. Zero migration. For the numbers, see the ROI of agentic data operations.

The sentence for upstairs: "Control-M keeps our batch on time; Data Workers checks the data that batch delivers and gets it fixed with our approval before money moves on it."

Getting started

Start with a pilot. Pick the Control-M folders that feed the numbers leaders, customers and regulators read, give Data Workers read access to the tables those folders write, with the Control-M MCP Server in your team's client, agree two or three metrics that define "right", and run at L1 for a few weeks. Then turn on L2 for one domain, and open L3 for a narrow class once the receipts show the agents were right. The pilot path is on the pricing page, and the pilot is credited in full against the first year.

FAQ

How does Data Workers connect to Control-M? Side by side. Control-M connects over its API or MCP server today, on SaaS or self-hosted: your team's assistant reads job runs, output and logs with a token whose role decides what it sees, and Data Workers checks the tables the jobs write.

Control-M has Jett, AI Pilot and an MCP Server. What does Data Workers add? Those help your team ask about the batch, analyze a failed job, build jobs and let assistants act with confirmation. Data Workers covers the data: whether what a job wrote is right, what it reached downstream, how to fix the cause, and the proof afterwards.

Will Data Workers hold, rerun or change our Control-M jobs? No. In Control-M it reads. It recommends holds, reruns and releases in a run plan, and your scheduling team carries them out in the Monitoring domain, through the Automation API or through the MCP Server with its confirmation step. Job and script changes arrive as a diff for the owner to merge.

Our jobs ended OK and the Service was on time. How would Data Workers know the data is wrong? It checks the tables each folder writes after every run: volume and nulls, plus metrics your team defines for load timing and the values that matter, such as the share of credits. A metric outside its baseline opens an incident even when every job is green. Trailer reconciliations stay in your own dbt tests.

We land mainframe files in a cloud warehouse. Does Data Workers read EBCDIC files or copybooks? Your conversion job does that, in Glue, Spark or your current tool. Data Workers reads its job output through Control-M, checks the decoded tables and proposes conversion-code changes as a diff.

Should we let assistants act on Control-M through the MCP Server? Follow BMC's guidance: it is a Preview feature, recommended for non-production environments, with authorizations from Control-M roles. Use an assistant that supports elicitation so each action asks for confirmation.

Sources

  • •BMC, Control-M product page ("The Agentic Orchestration Layer for Enterprise Operations"), https://www.bmc.com/it-solutions/control-m.html (checked Oct 3, 2026)
  • •BMC, Control-M latest release (MCP Server, Jett, AI Pilot, Archive Service, "SaaS, self-hosted, or hybrid deployment"), https://www.bmc.com/it-solutions/control-m-latest-release.html (checked Oct 3, 2026)
  • •BMC, Agentic orchestration ("the optional Control-M MCP Server"), https://www.bmc.com/it-solutions/agentic-orchestration.html (checked Oct 3, 2026)
  • •BMC, Control-M Managed File Transfer (Amazon S3, Azure, Google Cloud; SFTP, FTPS, AS2 with PGP), https://www.bmc.com/it-solutions/control-m-managed-file-transfer.html (checked Oct 3, 2026)
  • •Control-M SaaS docs, Control-M MCP Server (Preview, elicitation, authorizations, endpoint), https://documents.bmc.com/supportu/controlm-saas/en-US/Documentation/MCP_Server.htm (published Oct 3, 2026)
  • •Control-M SaaS docs, Control-M AI (Jett AI Assistant, Execution Insights, AI Workflow Creator (Preview), AI Pilot for Planning, AI Pilot for Monitoring), https://documents.bmc.com/supportu/controlm-saas/en-US/Documentation/AI_Capabilities.htm and its linked pages (checked Oct 3, 2026)
  • •Control-M SaaS docs, Control-M Integrations (plug-ins), https://documents.bmc.com/supportu/controlm-saas/en-US/Documentation/Integrations_Main.htm (checked Oct 3, 2026)
  • •Control-M Automation API quickstart and Python client (jobs-as-code), https://github.com/controlm/automation-api-quickstart, https://github.com/controlm/ctm-python-client (checked Oct 3, 2026)
  • •Cobrix, COBOL data source for Apache Spark, https://github.com/AbsaOSS/cobrix (checked Oct 3, 2026)
  • •Data Workers open-source repository (tool registrations in dw-incidents, dw-quality, dw-schema, dw-context-catalog), https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
  • •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)