Product
Product12 min readBy The Data Workers Team

You're on Factory: give Droids governed data context and approved data changes

Your engineers already delegate to Factory's Droids. Connect Data Workers over MCP and Droids work from governed data context, while every production data change goes through approvals and receipts.

Your engineers already delegate to Droids. They run droid in the terminal next to Git and their tests, review sessions and diffs in the Factory App, start Missions for large multi-feature work, and wire droid exec into CI. Some teams mention @Factory in a Slack thread and a Droid session starts from there. Your admins set a Default and Maximum Autonomy Level, write permission rules, and control which MCP servers can run through mcpPolicy. Factory is very good at the job it was built for: planning, writing, testing and shipping code across the software lifecycle. Data Workers is the agentic data platform behind it. Factory is the software factory; Data Workers is the data team its Droids call. Connected over MCP, Droids work from your governed data context, and every change to production data goes through Data Workers' approvals and leaves a receipt.

Key takeaways

  • •Factory is the software factory. Data Workers is the data team its Droids call. Droids write, test and ship the code; Data Workers brings the context, owns the change in production data and proves it worked.
  • •One mcp.json entry reaches Droid wherever it runs on that machine. The Droid CLI, droid exec and Factory App sessions share one Droid runtime and read mcp.json at user or project level, so one entry connects them all.
  • •Two locks on every write. Factory compares each tool's risk to the session's Autonomy Level and asks before anything above it; Data Workers applies its per-domain guardrail, keeps a rollback point and writes the receipt.
  • •Climb the ladder one domain at a time. Start read-only (L1), move to propose (L2), then act reversibly (L3), and let trusted domains run autonomously (L4).
  • •Nothing to migrate. Factory, your repositories, Postgres, Fivetran, Snowflake, dbt and Looker all stay. Data Workers works through each system's own API.

Factory is the software factory. Data Workers is the data team its Droids call.

A Droid reads your dbt project, your Airflow DAGs and your SQL, plans the change, writes it, runs the tests and opens the pull request. Through Factory's Snowflake and Databricks connectors it can also run a query and explain the result. What a Droid doesn't hold is the state of the data estate: which table is canonical, who owns it, what reads it downstream, what changed overnight, and who must approve a fix to finance data. That is what Data Workers holds. The Data Context Wizard keeps one governed context graph across warehouses, dbt, orchestration and BI. The Data-Agents Swarm (20+ specialist agents) does the data work. The Autonomous Data-Conductor runs each fix end to end: detect, diagnose, fix, review, verify, remember. Spellbook Data Catalog (in preview) is where people review, approve, roll back and audit.

Here is one request, end to end. It is an illustration, not a customer case.

TimeSystemWhat happens
01:45PostgresAn app release renames accounts.plan_tier to plan_code
02:20FivetranThe sync adds plan_code; raw.accounts.plan_tier in Snowflake now lands null
04:00dbtThe run succeeds; 6,212 accounts fall into the "unknown" tier in fct_arr_by_tier
08:50LookerThe ARR by tier Explore shows 23% of ARR as "unknown"
09:05SlackRevOps flags it in #data-help; a data engineer mentions @Factory in the thread and a Droid session starts
09:06FactoryThe Droid calls Data Workers over MCP (diagnose_incident, trace_cross_platform_lineage)
09:07Data WorkersTraces the nulls to the upstream rename; blast radius is 4 dbt models and 3 Looker Explores; owner is finance analytics
09:18GitHubThe Droid opens a pull request that maps plan_code in stg_accounts, with a not-null test
09:21FactoryThe Droid asks to run the Data Workers remediate tool; the engineer approves the tool call
09:34SpellbookThe finance data owner approves the backfill, as the finance domain requires
09:36SnowflakeData Workers starts the backfill of the 6,212 rows through the orchestrator, with the undo recorded first
09:52dbtAffected models rebuild; null count is zero; tests pass; receipt written
10:00LookerARR by tier is right before the forecast call
Incident timeline across the stack: what Factory, your team and Data Workers each do, step by step

Nobody left the thread. Factory did what it does best: took the request where the team already talks, planned, called the right tools, wrote the lasting code fix as a pull request your team reviews as usual, and asked before the change. Data Workers did the data work: it found the cause in Postgres, followed it through Fivetran and dbt into Looker, routed the backfill to the person who owns finance data, made the fix reversibly and proved it with null counts and tests.

What Factory doesWhat Data Workers does
Takes the request in the CLI, the Factory App, an IDE or a Slack threadAnswers it from governed context: lineage, ownership, definitions, recent changes
Plans, writes and tests code, then opens the pull requestReads and changes Snowflake, dbt, Fivetran and Looker through each system's own API
Asks before any tool whose risk is above the session's Autonomy LevelApplies per-domain guardrails and routes the change to the data owner
Runs large multi-feature work as a MissionRuns each data change in approved waves, with blast radius and parity checks
Shows the session, the diff and the review stateWrites a receipt that outlives the session: who, why, what changed, how to undo

Why doesn't Factory just do this itself?

Focus and risk. Factory built Droids for the software development lifecycle, and its design is right for that job. Its controls are deep and careful: Autonomy Levels from Off to High, risk levels on every command and MCP tool, permission rules that allow, ask or block, OS-level sandboxing, Droid Shield for secrets, and managed settings that let admins cap autonomy and allowlist MCP servers. Every one of those controls governs what a Droid session may do: which commands it runs, which servers it reaches, when it stops and asks.

Changing production data is a different product. Backfilling raw.accounts touches systems Factory doesn't run, owned by people who aren't in the session, with downstream readers nobody in the repository can see. Doing it safely takes context about every other system, a blast radius, an approver per domain, a rollback plan, verification after the fact and a receipt an auditor can read next quarter. It also means carrying the liability for changes in tools Factory doesn't own. A general software agent is right not to take that on. That is the product Data Workers is, and Droids are a great place for your engineers to ask for it.

Every tool owns a slice. Data Workers covers the whole lifecycle

Each point tool adds another console, another contract and another handoff. Data Workers covers the whole data lifecycle with one context, one approval flow and one audit trail. Factory owns a valuable slice of it: writing, testing and shipping code with Droids. It leads where that is the work, in pipeline code.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Factory goes deep on its own area
StageData WorkersFactoryWhy we scored it this way
Catalog & Context94Droids read the repository, AGENTS.md and AutoWiki pages about your code. Data Workers keeps one governed context graph across warehouses, dbt, orchestration and BI.
Analytics & Insights85A Droid can run queries through the Snowflake or Databricks connector and explain the result. Data Workers answers from governed definitions, lineage and ownership.
Data Quality85Droids write dbt tests and SQL checks, and Automated QA tests your application. Data Workers writes, runs and repairs data checks across platforms and keeps them running.
Observability & Incidents8.54Incident Response (Private Preview) starts a Droid session from a Slack alert and works toward an RCA. Data Workers detects, traces and closes data incidents across systems with a receipt.
Pipelines & Ingestion8.59Factory leads. Droids plan, write and test pipeline code, dbt models and DAGs, and ship them as pull requests, from the CLI, the app or CI. Data Workers owns the run in production behind approval.
Schema & Migration86Droids write migrations and refactors, and Missions take on large multi-feature work. Data Workers checks the blast radius across every downstream system before a schema change lands.
Governance & Access8.53Factory governs Droids: Autonomy Levels, permission rules, mcpPolicy and a Maximum Autonomy Level. Data Workers proposes least-privilege grants on your data platforms behind approvals.
Security & Privacy85Droid Shield catches secrets before a push, Security Review audits code, and the sandbox isolates Droid. Data Workers flags sensitive column names in pull request review across the estate.
Cost / FinOps82Factory reports its own LLM spend. Data Workers traces Snowflake credits to the dbt model behind them and drafts the fix for its owner.
MLOps & Models7.56Droids write training, feature and evaluation code. Data Workers keeps the data under the models healthy.

How Factory and Data Workers work together

Factory stays on top, where engineers delegate, review and approve. Data Workers sits underneath as MCP servers: the Context Wizard, the Swarm, the Conductor and the guardrails. Spellbook is where people look: data owners review and approve changes, roll them back and read the audit trail.

How Data Workers fits with Factory: your coding agent on top, Data Workers in the middle, your estate underneath

Setup over MCP today. Every Data Workers agent is a standard MCP stdio server, and Droid starts stdio servers from mcp.json. Add the incident and catalog agents from the terminal:

# Example: add two Data Workers agents to your user config (~/.factory/mcp.json)
droid mcp add dw-incidents "/path/to/dataworkers-claw-community/start-agent.sh dw-incidents"
droid mcp add dw-context-catalog "/path/to/dataworkers-claw-community/start-agent.sh dw-context-catalog"
droid mcp list

Or share them with the team in the project's .factory/mcp.json. Keep warehouse credentials out of that file, as Factory's own docs advise; Droid expands ${NAME} references in env from the shell at connect time. Example:

{
  "mcpServers": {
    "dw-incidents": {
      "type": "stdio",
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-incidents"],
      "disabledTools": ["monitor_metrics"]
    },
    "dw-context-catalog": {
      "type": "stdio",
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-context-catalog"]
    }
  }
}

Run /mcp in Droid to see the exact tool names each agent exposes, then pick what each team loads and trim the rest with disabledTools. Admins add one matcher to the org-managed mcpPolicy allowlist (stdio servers match on a substring of the command and its arguments), and on Enterprise they can set per-server and per-tool risk with mcpAutonomyOverrides. Classify the read tools low, so they run without a prompt from Low autonomy up, and classify remediate high, so Droid asks before it at every level below High. For data teams where every change must ask, cap the Maximum Autonomy Level at Medium. Example org-managed settings:

{
  "mcpPolicy": { "enabled": true, "allowlist": ["dataworkers-claw-community"] },
  "mcpAutonomyOverrides": {
    "dw-incidents": {
      "defaultLevel": "high",
      "tools": { "diagnose_incident": "low", "get_incident_history": "low", "remediate": "high" }
    }
  }
}

The same entries serve the Factory App and droid exec, which is read-only by default and only acts at the --auto level you grant. For a shared Data Workers deployment, Droid also connects to Streamable HTTP servers with --type http; mcpPolicy then matches by hostname, and mcpAutonomyUrlOverrides sets risk by URL with a safety floor: a tool that is not marked read-only never drops below high. Roll out locally first, then add the servers to the Droid Computers your Slack sessions run on. The Data Workers + Claude Code, Cursor and Codex guide covers the same wiring for other coding agents.

The same request on the autonomy ladder. Autonomy is set per domain, on the ladder L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. The engineer queries Snowflake by hand and diffs the Fivetran schema. A Droid helps with the SQL; nobody has the lineage.
  • •L1 observe. The Droid calls Data Workers and gets the cause, the blast radius and the owner. Nothing changes.
  • •L2 propose. Data Workers drafts the backfill with its blast radius and rollback plan. The finance owner approves in Spellbook; Factory asks the engineer before the tool call.
  • •L3 act reversibly. For a domain you trust, Data Workers maps renamed source columns and backfills on its own with the undo recorded first, and writes the receipt. The engineer reads it in the Droid session.
  • •L4 autonomous. The Conductor catches the rename at 02:20, before dbt runs, fixes it and verifies it. A Droid still ships the lasting code change as a pull request, and the Slack thread holds the receipt in the morning.

Factory's Autonomy Level and Data Workers' guardrails stack. Raising one never lowers the other.

What changes for your team

Engineers keep the Droids they already use. What changes is the work behind them: the jobs that used to wait for a person with the right access and the right context now run through Data Workers, at the autonomy level each domain sets.

Six jobs that run on autopilot with Data Workers next to Factory, with a concrete example of each

Data owners stop being pinged in Slack for "is this number right?" and start approving scoped changes in Spellbook. On-call stops starting from zero, because the first message to @Factory already returns the cause and the blast radius. Platform teams running Missions to refactor dozens of dbt models get each wave checked against every downstream reader before it lands. And the platform team writes one allowlist entry instead of a different MCP setup per team.

Keep Factory, or consolidate?

Keep Factory if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too. With a coding agent, keeping it is the usual answer: Factory is where your engineers build software, and Data Workers makes every data question and data change they send from a Droid governed. The same Data Workers servers also answer in Codex, Claude and Devin, so teams that run more than one agent share one context and one audit trail. For the whole category, see Your company just rolled out AI assistants. Now what?.

The case for your CFO

You already pay for Factory, and engineers already send data work to Droids. The outcome to buy now is that the questions they ask get governed answers, and the data changes they ask for get made safely: fewer wrong numbers reaching leadership, data incidents closed the same morning, and an audit trail for every change to production data.

The risk story is plain. At L1, Data Workers only reads and explains. At L2, it proposes and a named owner approves. At L3, it acts only where a change is reversible, and keeps the rollback point. L4 is reserved for domains you've trusted on evidence. Every action leaves a receipt: who asked, who approved, what changed, what it touched, and how to undo it. Factory keeps its Autonomy Levels, permission rules and sandbox, so there are two locks on every write. Nothing migrates: Factory, Postgres, Fivetran, Snowflake, dbt and Looker all stay. The cost case and how to model it are on the ROI page, and the deployment and data-handling answers are on the security page.

Why now: Droids are already the front door for engineering work, and with Missions and Slack sessions they take on bigger jobs every month. Every week without governed context is another week of agents guessing which table is canonical, and of production data changes made with personal credentials and no receipt. The safety page walks through what agents can and can't do at each level, and the build-vs-buy page covers what it takes to wire this yourself.

The first win is small and visible: read-only Data Workers incident and lineage tools in Droid for one domain, so "why is this number wrong?" gets a traced answer in the thread where it was asked. What stays the same: your repositories, your review process, your Factory admin settings and every data platform you run. Start with a pilot; pricing has the details, and the pilot is credited in full against the first year.

The sentence to repeat upstairs: "Our engineers keep Factory; Data Workers gives Droids our governed data context and routes every production data change through one approval flow with a receipt."

Getting started

Start with a pilot. Pick one domain where engineers already ask Droids about data, often finance or product analytics, add the Data Workers incident and catalog agents to mcp.json, set read tools to Low and change tools to High in your managed settings, and run that domain at L1 for the first weeks. Move it to L2 when the answers hold up, and to L3 for the fixes your team approves every time. See pricing; the pilot is credited in full against the first year.

FAQ

Does Data Workers bypass Factory's Autonomy Level or permission rules? No. Factory still decides when to ask, by comparing each tool's risk to the session's Autonomy Level, and your permission rules and Maximum Autonomy Level still apply. Data Workers adds its own per-domain guardrail on top, so a production data change needs both.

Which Factory surfaces does this work in? The Droid CLI, droid exec and Factory App sessions on your machine share one Droid runtime and read mcp.json from the user and project levels, so one entry covers them. Sessions started from Slack run on a Droid Computer, which keeps its own configuration between sessions; add the Data Workers entries to that computer once the local rollout holds.

Can our admins control which MCP servers Droids use? Yes. Enable mcpPolicy in org-managed settings and add a matcher for the Data Workers server: its hostname for Streamable HTTP, or a command or argument match for stdio. Users can't override it, and servers outside the allowlist never start.

Won't more tools crowd the Droid's context? Use disabledTools on each server entry so excluded tools never load, and scope custom droids to specific servers with mcpServers in their frontmatter. Most teams start with the incident and catalog agents and add more as domains move up the ladder.

Factory already connects to Snowflake. Why add Data Workers? Factory's Snowflake connector lets a Droid run queries, which is useful. Data Workers adds what a production change needs on top: lineage across Fivetran, dbt and Looker, ownership, a blast radius, an approver per domain, a rollback point and a receipt.

Whose credentials touch Snowflake? Data Workers' own credentials, scoped per domain, so no personal warehouse keys sit on a laptop or in a committed config. The engineer's identity is recorded on the receipt as the person who asked and approved.

Can a Mission run a whole dbt migration through Data Workers? Yes, in waves. The same holds for Factory's /migrate workflow (added September 26, 2026). The Mission plans and writes the model changes; Data Workers checks each wave's blast radius, routes it for approval, applies it reversibly and holds the next wave until its planned parity checks are signed off.

Sources

  • •Factory, "Model Context Protocol (MCP)" (transports, droid mcp commands, mcp.json levels, disabledTools, mcpPolicy, autonomy URL overrides): https://docs.factory.com/harness/mcp (checked Oct 2, 2026)
  • •Factory, "Autonomy Level" (Off, Low, Medium, High; tool risk; Default and Maximum Autonomy Level): https://docs.factory.com/autonomy-and-safety/auto-run (checked Oct 2, 2026)
  • •Factory, "Enterprise Controls & Managed Settings" (mcpAutonomyOverrides, managed settings): https://docs.factory.com/enterprise/hierarchical-settings-and-org-control (checked Oct 2, 2026)
  • •Factory, "Connectors" (Snowflake and Databricks connectors): https://docs.factory.com/harness/connectors (checked Oct 2, 2026)
  • •Factory, "Droid CLI" and "Factory App": https://docs.factory.com/droid-cli/overview and https://docs.factory.com/factory-app/overview (checked Oct 2, 2026)
  • •Factory, "Droid Exec (Headless)" (read-only by default, --auto levels): https://docs.factory.com/droid-exec/overview (checked Oct 2, 2026)
  • •Factory, "Factory Missions": https://docs.factory.com/missions/overview (checked Oct 2, 2026)
  • •Factory, "Slack" (mention @Factory in a thread): https://docs.factory.com/remote-delegations/slack (checked Oct 2, 2026)
  • •Factory, "Incident Response" and "Software Factory" (Private Preview): https://docs.factory.com/software-factory/incident-response and https://docs.factory.com/software-factory/overview (checked Oct 2, 2026)
  • •Factory, "Factory for Enterprise": https://docs.factory.com/enterprise/index (checked Oct 2, 2026)
  • •Factory, "Feature Maturity" and "Full Changelog" (/migrate workflow, CLI v0.228.0, September 26, 2026): https://docs.factory.com/changelog/feature-maturity and https://docs.factory.com/changelog/release-notes (checked Oct 2, 2026)
  • •Factory, "Droid Computers": https://docs.factory.com/droid-computers/overview (checked Oct 2, 2026)
  • •Data Workers, "Client Setup": https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)
  • •Data Workers, open-source agents and start-agent.sh (dw-incidents tools diagnose_incident, get_incident_history, remediate, monitor_metrics; dw-context-catalog trace_cross_platform_lineage): https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)