Product
Product9 min readBy The Data Workers Team

Why a swarm of specialist agents instead of one big agent?

Data Workers gives each data job its own specialist agent with narrow tools, scoped permissions and its own autonomy level, coordinated by the Conductor. Here is why that beats one generalist agent with every permission for production data work.

Because production data work is many different jobs with different risks, and each job deserves its own agent with its own tools, its own permissions and its own autonomy level. Data Workers gives incidents, quality, schema, pipelines, lineage and context, governance, cost, migrations and more each a specialist agent, and the Autonomous Data-Conductor coordinates them, so you get a smaller blast radius, least-privilege access per job, autonomy set per domain, clearer receipts and agents you can upgrade one at a time, where one generalist agent would hold every permission at once.

One big agent is a fine design for some work, and we say where below. For agents that change production data across Snowflake, dbt, Airflow, Looker and your access policies, the specialist design is what makes the change safe to approve. Data Workers is the agentic data platform built on it.

Key takeaways

  • •One job, one agent. The Data-Agents Swarm has 20+ specialist agents. Each is its own MCP server with a narrow tool set, and every tool is registered to the agent that owns it with the permission it needs: read, write or admin.
  • •The Conductor owns the outcome. The Autonomous Data-Conductor turns each problem into one proposed run with every step owned by one specialist, puts it through its governor, and holds an entity lock so two routines never work the same table at once.
  • •Autonomy is set per domain. Your team can let freshness reruns run at L3 act reversibly while access grants stay at L2 propose. One big agent has one dial for everything.
  • •Mistakes stay small. A specialist reaches only its own tools and the access your team grants that job. OWASP names excessive functionality, permissions and autonomy as the root causes of "Excessive Agency" in LLM apps; the specialist design addresses all three.
  • •The research agrees, with conditions. Anthropic and OpenAI both advise starting with one agent and splitting when tools overlap, logic branches or the work divides cleanly. Production data operations across many platforms fit those conditions, and Anthropic's warning about shared context is why every Data Workers agent reads one context graph.

How the swarm works in Data Workers

How Data Workers fits with Your estate: your coding agent on top, Data Workers in the middle, your estate underneath

Each specialist owns one job and nothing else. In the product today that includes agents for incident debugging, quality monitoring, schema evolution, pipeline building, data context and catalog, data change review, access and governance, data security, identity, cost savings and cleanup, data migration, ingestion, streaming, MLOps, observability, insights, search, usage intelligence and connectors, with the Conductor on top. Each ships as its own MCP server, so engineers can connect one agent to Claude, ChatGPT, Cursor or Codex on its own.

Four design choices in the code make the swarm work.

Narrow tools per agent. The schema agent checks compatibility with check_compatibility and does impact analysis with assess_impact. The context agent traces lineage with trace_cross_platform_lineage and sizes a change with blast_radius_analysis. The quality agent scores tables with get_quality_score. The incidents agent diagnoses and runs remediate. Each agent carries only the tools for its job, and a permissions registry maps every tool to the agent that owns it and the scope it requires.

Autonomy per agent, per operation, per customer. The autonomy controller resolves the level by operation first, then by agent, then by the default, and it can change at runtime without a restart. That is how the L0 to L4 ladder becomes a per-domain setting in practice. When an incident's root cause spans pipelines, schema, catalog and governance, the cross-domain remediation plan fans out in dependency order and each domain re-checks its own autonomy before it acts.

One coordinator that owns the outcome. The Conductor compiles goals into routines and turns each piece of work into a proposed run, with each step owned by the specialist for that job. Every candidate passes its governor first (kill switch, anomaly checks, budget, and a hard floor that keeps anything that spends money at approve-to-act); repeat triggers collapse into the run already open, and an entity lock means two routines never work the same table at the same time. Proposed runs land in the approvals inbox. The Conductor ships observe-only, and your team opens it up domain by domain.

One context, one approval flow, one audit trail. Every agent reads the same governed context graph from the Data Context Wizard, so the pipelines agent's fix accounts for the Looker Explore the context agent found. Every tool call passes through one middleware that scans responses for PII and writes a tamper-evident audit record. The permission ladder scores every call against the level your team set for that domain, and when your team switches enforcement on, calls above that level wait in one approvals queue in Spellbook Data Catalog (in preview). Approvals go to a named person: the approval guard rejects any agent, including the one that made the change.

A worked example: one schema change, five specialists

This is an illustration on a Fivetran, Snowflake, dbt, Airflow, Looker stack, not a customer record or a measured result. The team has set schema and pipelines to L2 propose, Airflow reruns to L3 act reversibly, and governance to L2 with the marketing data owner as approver.

TimeWhoWhat happensScope it used
02:10FivetranHubSpot sync adds contact_email and retypes deal_amount to textUpstream
02:13Schema agentDrift caught in SnowflakeRead-only on the schema
02:14ConductorOne run proposed; each step owned by one agentCoordination only
02:16Context agentBlast radius: 3 dbt models, 1 Airflow DAG, 2 Looker ExploresRead lineage
02:22Pipelines agentdbt cast fix and a new test proposed as a diff for the owner to mergeWrite to a branch only
02:24Governance agentcontact_email classified as PII; masking policy draftedDraft policy, no grant
07:40Analytics engineerReviews and approves the changeHuman approval
07:46Incidents and quality agentsAirflow DAG rerun at L3; quality score back in rangeRerun one DAG; read scores
08:30Data ownerApproves the masking policyHuman approval
09:00LookerPipeline review opens on correct, masked dataOne receipt for the run
Incident timeline across the stack: what Your estate, your team and Data Workers each do, step by step

Look at what never happened. The pipelines agent could not touch the masking policy. The governance agent could not edit dbt. The schema agent, which runs every night across every table, holds read access only. If any one of them had been wrong, the damage was bounded by that agent's tools and that domain's level, and the receipt shows exactly which agent did what. With one generalist agent holding all of those permissions, a misread prompt can reach every system the token can.

One big agent vs the swarm

Comparison matrix of Your estate and Data Workers on the outcomes a data leader buys
One big agentData Workers swarm
Blast radiusEverything its credentials reachOne job's tools and access
PermissionsThe union of every job's accessLeast privilege per job
AuditOne long session logA receipt per change, per agent
AutonomyOne setting for everythingSet per domain, L0 to L4
UpgradesChange one prompt, retest everythingUpgrade and evaluate one agent
SimplicityOne prompt, one loopThe Conductor coordinates for you

What the research says, and when one agent is enough

These are third-party sources, and they are fair to both designs.

Anthropic's "Building effective agents" (December 19, 2024) recommends "finding the simplest solution possible, and only increasing complexity when needed." Its multi-agent research post (June 13, 2025) reports that a lead agent with subagents outperformed a single agent by 90.2% on Anthropic's internal research eval, and also that multi-agent systems used about 15 times more tokens than chats, and that domains where every agent needs the same context or many dependencies "are not a good fit for multi-agent systems today."

OpenAI's "A practical guide to building agents" (April 2025) says to "maximize a single agent's capabilities first" and that "often a single agent with tools is sufficient," then split when prompts carry complex branching logic or tools overlap: "Some implementations successfully manage more than 15 well-defined, distinct tools while others struggle with fewer than 10 overlapping tools." It describes the manager pattern, a central agent orchestrating specialists through tool calls, and advises human oversight for high-risk, irreversible actions. Cognition's "Don't Build Multi-Agents" (June 12, 2025) argues that multi-agent collaboration in 2025 "only results in fragile systems" because context is not shared thoroughly enough between agents, and its first principle is "Share context."

OWASP's 2025 Top 10 for LLM applications lists Excessive Agency, with root causes of "excessive functionality; excessive permissions; excessive autonomy," and gives the example of an agent connected with an identity that holds UPDATE, INSERT and DELETE when it only needed SELECT.

So when does one agent suffice? A read-only analyst asking questions of one warehouse, a prototype, or an engineer pairing with a coding agent in one repo. Data Workers is designed for the other case: hundreds of overlapping tools across many platforms, writes to production, and approvers who differ by domain. The Data Workers answer to the shared-context objection is the Context Wizard: every specialist reads the same governed graph, and the Conductor holds state across every hop.

The alternatives buyers weigh

One coding agent with every MCP server attached. Claude Code, Cursor and Codex are excellent at writing code, and Claude Code's own subagents run "in its own context window with a custom system prompt, specific tool access, and independent permissions." That is the specialist idea applied inside one engineer's session. The production layer around it (per-domain approvers, rollback, receipts, cross-engine blast radius) becomes a platform your team owns; our build-vs-buy page walks through that year of work.

Multi-agent toolkits. The pattern is now mainstream. Databricks' Supervisor Agent coordinates Genie Agents, agent endpoints, Unity Catalog functions and MCP servers "across different, specialized domains," and OpenAI's Agents SDK documents orchestration through handoffs and agents as tools. These are strong kits for building your own multi-agent apps. Data Workers ships the data-operations specialists already built, scoped and governed, with the Conductor and the approvals around them.

Platform-native agents. Databricks Genie Code and Snowflake CoCo (formerly Cortex Code) are strong inside their own platform and ask a person to confirm actions by default. Databricks notes that Genie Code's auto-approve "is a productivity feature, not a security boundary," the right design for an assistant in one platform.

Observability and catalogs. Monte Carlo and its peers are the smoke alarm, and catalogs are the inventory. Their signal feeds the swarm's loop.

The case for your CFO

The outcome is operational work done safely at scale: incidents fixed, schema drift caught before dashboards break, access drafted by policy, spend traced to the model behind it. The specialist design lets you say yes, because every agent's reach is bounded by its job.

The risk story is concrete. Each agent holds only its own tools and permissions. Each domain runs at the level your team sets, from L0 manual to L4 autonomous. Changes above that level go to the named approver your team picks, no agent approves its own work, reversible changes carry a rollback, and every change leaves a receipt naming the agent, the change, the checks and how to undo it. Nothing is migrated: Snowflake, dbt, Airflow, Looker and your permission systems stay as they are. Your data stays in your systems: the agents run in your infrastructure, and the hosted Conductor sees workflow metadata only. For the full safety model read is it safe to let AI agents change production data?, and for where data lives, where does our data go?.

Why now: teams are already handing agents broad tokens to move faster. Scoping them by job is cheaper before the first incident than after it.

The first win is one domain, usually incidents or schema changes, at L2 propose. Start with a pilot: $7,500 one-time, and the pilot is credited in full against the first year. Scale is from $1,000 a month and Enterprise from $3,000 a month, billed annually, with unlimited seats, no usage meter, no markup on model spend and bring your own model. See pricing, the ROI calculator and the ROI of agentic data operations.

The sentence for upstairs: "Every agent we run does one job with one job's access, and every change it makes has an owner and a receipt."

FAQ

Isn't a swarm harder to run than one agent? For you, no. The Conductor does the coordination: it proposes the run, assigns each step to the agent that owns it, locks the table being worked on and escalates one decision. Your team sees one approvals queue and one audit trail in Spellbook.

Don't specialist agents lose context between handoffs? That is the main failure mode the research warns about. Every Data Workers agent reads the same governed context graph, and the Conductor holds the run's state across every hop, so the governance agent sees what the schema agent found. If you already keep definitions elsewhere, bring your own context shows how the Context Wizard plugs it in.

Can we turn on only some of the agents? Yes. Each agent is its own MCP server, and each domain has its own autonomy level. A strong first step is incidents and schema at L1 observe or L2 propose, with the rest left at L0.

Who owns each agent? Your team does, domain by domain: the approver for governance changes can differ from the approver for dbt fixes. Read who owns the agents for the ownership model and how approvals work.

Does the swarm cost more in model spend? Multi-agent systems can use more tokens, as Anthropic reports. Data Workers runs on your own model with no markup, and narrow agents carry smaller prompts and tool lists than one agent loaded with every tool. See which model Data Workers uses.

What makes this an agentic data platform rather than a set of agents? The shared context, the Conductor, the guardrails and Spellbook around the agents. Our explainer on what an agentic data platform is covers the whole picture, and the agent swarm vs single agent resource goes deeper on operations patterns.

Sources

  • •Anthropic, "Building effective agents," December 19, 2024: https://www.anthropic.com/engineering/building-effective-agents (checked Oct 2, 2026)
  • •Anthropic, "How we built our multi-agent research system," June 13, 2025: https://www.anthropic.com/engineering/multi-agent-research-system (checked Oct 2, 2026)
  • •OpenAI, "A practical guide to building agents," April 2025: https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf (checked Oct 2, 2026)
  • •Cognition, "Don't Build Multi-Agents," Walden Yan, June 12, 2025: https://cognition.ai/blog/dont-build-multi-agents (checked Oct 2, 2026)
  • •OWASP Gen AI Security Project, LLM06:2025 Excessive Agency: https://genai.owasp.org/llmrisk/llm062025-excessive-agency/ (checked Oct 2, 2026)
  • •Anthropic, Claude Code docs, subagents: https://code.claude.com/docs/en/sub-agents (checked Oct 2, 2026)
  • •Databricks, Genie Code agent mode (auto-approve note): https://docs.databricks.com/aws/en/genie-code/agent-mode (checked Oct 2, 2026)
  • •Databricks, Supervisor Agent (page last updated Sep 11, 2026): https://docs.databricks.com/aws/en/generative-ai/agent-bricks/multi-agent-supervisor (checked Oct 2, 2026)
  • •OpenAI Agents SDK, agent orchestration: https://openai.github.io/openai-agents-python/multi_agent/ (checked Oct 2, 2026)
  • •Snowflake, security best practices for CoCo CLI (Confirm actions is the default mode): https://docs.snowflake.com/en/user-guide/cortex-code/security (checked Oct 2, 2026)
  • •Data Workers, Spellbook Data Catalog (preview): https://dataworkers.io/product/spellbook-data-catalog/ (checked Oct 2, 2026)
  • •Data Workers, Data-Agents Swarm: https://dataworkers.io/product/data-agents-swarm/ (checked Oct 2, 2026)
  • •Data Workers, Autonomous Data-Conductor: https://dataworkers.io/product/autonomous-data-conductor/ (checked Oct 2, 2026)
  • •Data Workers pricing: https://dataworkers.io/pricing/ (checked Oct 2, 2026)
  • •Data Workers open-source agents (tool registrations): https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)