Product
Product12 min readBy The Data Workers Team

You're on Amp. Here's how Data Workers builds on it

Your engineers run Amp threads in orbs and on schedules. Connect Data Workers over MCP and every data change those threads ship is scoped, approved, verified and recorded.

Your engineers already work in Amp, built by Amp Frontier Corporation since its spin-out from Sourcegraph on December 2, 2025. Since the VS Code and Cursor extensions retired on March 5, 2026, they start threads on the web, in the macOS and iOS apps or from the CLI, turn the Dial to medium for most work and ultra for a migration, and let the agent keep going in an orb while the laptop is closed. Threads are shared with the workspace by default, the oracle gives a second opinion, the Librarian reads other repositories, and Automations wake a thread each morning to check something and report to Slack. Workspace admins set up skills, plugins and MCP servers once for everyone. For writing and shipping code, that setup is fast and very good. The open question for a data team is what happens when a thread nobody is watching ships a change to a table that a dashboard, an experiment readout and a finance export all read.

That's where Data Workers comes in. Amp is the agent that writes and ships the code. Data Workers, the agentic data platform, is the data crew that knows what each change touches and carries it into production: scoped, approved, verified and recorded. Connected over MCP, Amp threads answer data questions from Data Workers' governed context, and every change to data goes through Data Workers' approvals and receipts. Your engineers keep the agent they chose and the way they use it.

Key takeaways

  • •Amp ships the code; Data Workers carries the data change. Data Workers supplies lineage, owners and blast radius, then applies the approved change in the warehouse or lakehouse and verifies it downstream.
  • •One skill to start. Bundle the Data Workers MCP servers in an Amp skill with an mcp.json, as Amp recommends, so the tools load only when a thread touches data.
  • •Two locks on every write. Amp runs tools without asking by default, which suits code in an orb. A small Amp plugin asks a human before any Data Workers tool that changes data, and Data Workers' per-domain guardrail routes the change to its owner. No agent approves its own work.
  • •Orbs and Automations get the same context. A workspace-scoped remote MCP definition reaches every member's threads, scheduled ones included.
  • •Climb one domain at a time. Start at L1 observe (answers, lineage, impact reports on pull requests), move to L2 propose, then L3 act reversibly where the record earns it.

Amp is the agent that ships the code. Data Workers is the data crew behind it.

Amp knows your code. It reads the repository and AGENTS.md, searches other repositories through the Librarian, runs the tests in an orb and lands the change through Ship. What code alone doesn't show is the live estate: which Databricks tables a dbt model really feeds, which Looker dashboard reads a column, which experiment readout depends on yesterday's partition, and who owns each one. Data Workers keeps that in one governed context graph and does the data side of the work through 20+ specialist agents.

Here is one incident, end to end. This is an illustration, not a customer case.

TimeSystemWhat happens
22:40SegmentA mobile app release starts sending event_ts in device local time without a UTC offset.
02:00DatabricksThe nightly job partitions events by date. Late-evening events from Asia-Pacific land in the previous day.
03:00dbtfct_daily_active_users builds. Every test passes.
06:10Data WorkersA quality check flags an 11% drop in daily active users against the trend and opens an incident with the owner attached.
07:00AmpThe growth team's Automation wakes its thread in an orb. Following the Data Workers skill, the agent calls diagnose_incident.
07:03DatabricksData Workers traces the drop through lineage to the timestamp change in the silver events table and names the release that introduced it.
07:20GitHubAmp writes the fix (parse the offset, convert to UTC) and ships it as a pull request.
07:22GitHubData Workers posts the blast radius: two dbt models, the growth dashboard, one experiment readout, and a backfill plan for three partitions.
07:25SlackAmp reports in #data-growth and tags the growth data owner.
08:05GitHubCI passes and the owner approves the pull request.
08:20SpellbookThe owner approves the three-partition backfill.
08:35DatabricksData Workers reprocesses the three partitions with the previous versions kept for rollback, then reruns the affected models.
09:10LookerDaily active users are back on baseline, the experiment readout matches, and the receipt is recorded in Spellbook and posted to the thread.
Incident timeline across the stack: what Amp, your team and Data Workers each do, step by step

With Data Workers, the owner approved one pull request and one backfill, the three days of numbers were repaired, and the receipt explains the whole thing to anyone who asks later.

JobWhat Amp doesWhat Data Workers does
Understand the requestReads the prompt, the repo, AGENTS.md and earlier threads; the Librarian searches other repositoriesAdds live data context: lineage across Segment, Databricks, dbt and Looker, owners and usage
Write the changeEdits models, SQL and jobs in an orb and runs the testsSupplies the impact report and every reader the change must account for
Gate the actionRuns tools without asking by default; plugins and the MCP registry allowlist set the policyRoutes data changes to the domain owner by policy, with blast radius attached
Apply to productionShips the branch and opens the pull requestApplies the approved migration, rerun or backfill, with rollback ready
VerifyProves the code works in the orbChecks the changed tables against baselines, then dashboards and downstream runs
RecordKeeps the thread, shared with the workspaceWrites a receipt to the audit trail: who asked, who approved, what changed, how to undo it

Why doesn't Amp just do this itself?

Focus and risk. Amp built a superb product for one job: letting a frontier model write and ship software with as little friction as possible. Its docs say so plainly: "By default, Amp does not ask for approval before running tools." Agents work in orbs, threads are shared so the team can review, and git is the undo. For code, that is exactly right: a bad commit in an orb costs a revert.

Production data has no revert button. A wrong backfill overwrites the partitions a dashboard and an experiment readout read, and the damage lands in systems Amp doesn't run. Changing data safely needs a different product: a blast radius before anything runs, an approval routed to the owner of that domain rather than to whoever started the thread, a rollback path inside Databricks or Snowflake, a check that the dashboard is right afterwards, and a receipt an auditor can read. A coding-agent vendor sensibly leaves that responsibility to the server on the other end of the MCP connection, and Amp gives you plugins to put a gate in front of it. That server is Data Workers.

Every tool owns a slice. Data Workers covers the whole lifecycle

Amp goes deepest on writing, testing and shipping pipeline code. Data Workers covers every stage of the data lifecycle around that code. Each point tool adds another console, contract and handoff; Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Amp goes deep on its own area
StageData WorkersAmpWhy we scored it this way
Catalog & Context95Amp reads the repo, AGENTS.md and shared threads, and the Librarian searches your other repositories. Data Workers keeps one governed graph of tables, owners, lineage and usage across platforms.
Analytics & Insights84Amp can write a query or an analysis script on request. Data Workers answers business questions from governed metric definitions.
Data Quality85Amp writes dbt tests and checks when asked. Data Workers runs quality checks, writes the missing tests and repairs failing ones.
Observability & Incidents8.54Amp Automations can watch a job on a schedule and report to Slack. Data Workers detects data incidents, traces the cause and closes them with a receipt.
Pipelines & Ingestion8.59Amp's home stage: threads in orbs write, test and ship pipeline code, even while your laptop is closed. Data Workers builds and reruns pipelines behind approval.
Schema & Migration86Amp writes migration code and DDL. Data Workers assesses the impact, drafts each migration with its rollback SQL for the owner to apply in approved waves.
Governance & Access8.53Amp's workspace roles, MCP registry allowlist and plugins govern the agent. Data Workers proposes least-privilege grants on your data platforms behind approvals.
Security & Privacy85Amp offers isolated orbs, secrets handling and minimal data retention for its own work. Data Workers' pull request review flags new columns whose names or annotations look sensitive, and leaves a receipt on every data change.
Cost / FinOps82Amp entitlements cap spend on Amp. Data Workers traces Snowflake credits to the query and dbt model behind them and drafts the fix for the model's owner.
MLOps & Models7.54Amp writes training and feature code. Data Workers keeps the data under models healthy and connects to MLflow and W&B.

How Amp and Data Workers work together

Amp stays on top, where engineers start, steer and review threads. Data Workers sits underneath over MCP: the Data Context Wizard answers from the governed graph, the Data-Agents Swarm does the work, the Autonomous Data-Conductor runs each fix end to end, and Spellbook Data Catalog (in preview) is where people review, roll back and audit.

How Data Workers fits with Amp: your coding agent on top, Data Workers in the middle, your estate underneath

Connect the MCP server. Every Data Workers agent is a standard MCP stdio server. Clone the open-source core (dataworkers-claw-community, Apache 2.0), run npm install, and point Amp at start-agent.sh for each agent you want, following the pattern in our client setup docs. From a terminal, one line per agent adds it to your user settings. Example:

amp mcp add dw-context-catalog -- /path/to/dataworkers-claw-community/start-agent.sh dw-context-catalog
amp mcp add dw-incidents -- /path/to/dataworkers-claw-community/start-agent.sh dw-incidents
amp mcp add dw-schema -- /path/to/dataworkers-claw-community/start-agent.sh dw-schema
amp mcp add dw-quality -- /path/to/dataworkers-claw-community/start-agent.sh dw-quality

Bundle it in a skill. Amp recommends putting MCP servers in skills so their tools stay hidden until a thread needs them, which keeps the tool list clean. Create a data-changes skill under .agents/skills/ in the repository with a short SKILL.md (when to use it: any change to dbt models, SQL, DDL or pipeline jobs) and a sibling mcp.json. The includeTools field limits each server to the tools the skill needs. Example:

{
  "dw-context-catalog": {
    "command": "/path/to/dataworkers-claw-community/start-agent.sh",
    "args": ["dw-context-catalog"],
    "includeTools": ["trace_cross_platform_lineage", "blast_radius_analysis"]
  },
  "dw-incidents": {
    "command": "/path/to/dataworkers-claw-community/start-agent.sh",
    "args": ["dw-incidents"],
    "includeTools": ["diagnose_incident", "get_root_cause", "remediate"]
  },
  "dw-quality": {
    "command": "/path/to/dataworkers-claw-community/start-agent.sh",
    "args": ["dw-quality"],
    "includeTools": ["run_quality_check"]
  }
}

In the SKILL.md body, list the moments that matter: call blast_radius_analysis before editing a dbt model or DDL, monitor_metrics to see a table's load lag against its baseline before trusting a query result, run_quality_check before shipping, diagnose_incident before patching anything described as broken, and check_policy (from dw-governance) on code that touches personal data. A workspace admin can publish the skill for every member.

Reach orbs and Automations. An orb doesn't read the settings file on your laptop. For a team, run Data Workers as a remote server over Streamable HTTP and add it once as a workspace-scoped remote MCP definition; it reaches every member's threads, orbs and scheduled ones included. Example (placeholder host):

amp mcp remote --workspace acme add DataWorkers https://<your-data-workers-host>/mcp --auth bearer --bearer-token-file ./dw-token

Amp reads the token from the file and never takes it as a command argument. Servers committed to a repository's .amp/settings.json need an explicit amp mcp approve before they run, and on Enterprise an MCP registry allowlist can limit members to approved servers.

Set the gates. Amp runs tools without asking by default and leaves tool policy to plugins. We recommend a small plugin, shaped like Amp's own permissions example, that lets read tools run freely and asks a human before any Data Workers tool that changes data. Locally configured MCP tools reach the plugin as mcp__<server>__<tool>. Save it in .amp/plugins/, run plugins: reload, and a workspace admin can publish it as a global workspace plugin. Example:

import type { PluginAPI } from '@ampcode/plugin'

const WRITES = /^mcp__dw-[\w-]+__(apply_migration|rollback_migration|remediate)$/

export default function (amp: PluginAPI) {
  amp.on('tool.call', async (event, ctx) => {
    if (!WRITES.test(event.tool)) return { action: 'allow' }
    const ok = await ctx.ui.confirm({
      title: 'Data Workers change',
      message: `Amp wants to call ${event.tool}. Data Workers will still route it to the domain owner.`,
      confirmButtonText: 'Allow',
      requireHuman: true,
    })
    return ok
      ? { action: 'allow' }
      : { action: 'reject-and-continue', message: `Rejected ${event.tool}.` }
  })
}

requireHuman keeps the agent from answering its own prompt. Remote definitions are called through Amp's code_exec, so for scheduled threads and workspace servers Data Workers' guardrail is the lock that matters: at L2 the change waits for the owner in Spellbook, whatever the plugin allows.

One request, from L0 to L4. Take one request an engineer types into an Amp thread: "the events_daily freshness check failed again, fix it". Here is what happens at each level, set per domain.

LevelWhat happens when the engineer asks
L0 manualAmp helps the engineer read logs and write a patch by hand. Data Workers isn't in the loop.
L1 observeAmp calls diagnose_incident. Data Workers answers in the thread: the Databricks job timed out after a source API change, two dashboards are stale, and the owner is the growth data team.
L2 proposeData Workers drafts the fix (a retry and timeout change on the job plus a backfill plan) with its blast radius. Amp ships the pull request; the owner approves.
L3 act reversiblyFor this pre-approved class of fix, Data Workers reruns the load and the backfill itself, with rollback ready, and records the receipt. The engineer sees the result in the thread.
L4 autonomousFreshness incidents in this domain run end to end without waiting: detect, fix, verify, record. The owner reviews receipts in Spellbook and can dial the domain back at any time.

Our Claude Code, Cursor and Codex integration guide walks through the same wiring for three other coding agents side by side.

What changes for your team

Engineers keep Amp and their habits: the Dial, orbs, shared threads, Puck. What changes is the work around each data change: finding who reads a column, chasing an owner, the manual backfill, the evidence an auditor asks for later.

Six jobs that run on autopilot with Data Workers next to Amp, with a concrete example of each

The platform team stops being the lookup service for lineage and ownership. Analytics engineers ship dbt changes with the impact attached. Scheduled threads can point at production data, because every change they propose lands with an owner. The audit trail builds itself from receipts.

Keep Amp, or consolidate?

Keep Amp if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

For most data teams, Amp stays, and Data Workers sits underneath it. What teams consolidate are the extra tools around data changes: the separate lineage lookup, the impact script someone maintains, the ownership spreadsheet, the backfill runbook. Data Workers runs that work with one context, one approval flow and one audit trail, and serves every other MCP client from the same agents. If parts of your team work elsewhere, read You're on Cursor, You're on Claude or You're on OpenCode, or start from the hub, Your company just rolled out AI assistants. Now what?.

The case for your CFO

Your engineers ship more with Amp, some of it while they sleep. Data Workers makes the data changes those threads ship right the first time: fewer wrong dashboards, fewer experiments called on bad numbers, less senior time spent tracing who reads what.

The risk story is plain. At L1, agents only read and explain. At L2, they propose and a named owner approves. At L3, they act only on pre-approved, reversible classes of change, with rollback ready. Every action leaves a receipt with who asked, who approved, what changed and how to undo it. Our safety guide covers what an agent can and can't do at each level, and the security and deployment guide covers where your data goes.

Why now: Amp threads already run unattended in orbs and on schedules, so their data changes should land with context and an owner's approval. Our build-vs-buy guide lays out what building that layer in house takes.

The first win is impact reports on every dbt pull request in one repository, at L1, with no change to how anyone works. What stays the same: Amp, Databricks or Snowflake, dbt, Looker and your permission systems. Nothing is migrated. Our ROI guide shows how to size the return.

The sentence to repeat upstairs: "Our engineers keep Amp; Data Workers makes every data change its threads ship scoped, approved, verified and recorded."

Getting started

Start with a pilot: add the Data Workers skill to one repository, connect a workspace-scoped server for orbs, and run one domain such as freshness incidents or schema changes from L1 to L2 with your own engineers and owners. See pricing for the pilot terms; the pilot is credited in full against the first year.

FAQ

Does Data Workers replace Amp's agent, orbs or Automations? No. Amp keeps writing, testing and shipping the code, in the CLI, on the web or in an orb. Data Workers adds the data context and carries the approved data change into production, then verifies and records it.

Amp doesn't ask before running tools. Can a thread change production data on its own? Only within the autonomy level you set for that domain. An Amp plugin can ask a human before any locally configured Data Workers tool that writes, and Data Workers' guardrail gates every change itself, however it was called. At L2 a named owner approves every change; at L3 only pre-approved, reversible classes run, each with rollback ready and a receipt.

Can our Amp admins control whether engineers use Data Workers? Yes. Workspace admins add the shared server and publish the skill and plugin for everyone. Servers committed to a repository need explicit approval before they run, and on Enterprise an MCP registry allowlist limits members to approved servers.

How do we keep warehouse credentials out of threads and repositories? Data Workers uses its own scoped credentials to each platform, so engineers never paste warehouse keys into Amp. The Amp side needs only the token for the Data Workers server, passed through --bearer-token-file or Amp's secrets settings for orbs, never as a command argument or a committed file.

Which data platforms does this work with? Data Workers runs control-plane connectors for Snowflake, Databricks and BigQuery, and reads dbt manifests and runs. For Segment, Looker and the rest of your stack, Data Workers connects over each tool's API or MCP server today.

Sources

Amp capabilities and statuses are current as of October 2, 2026, from Amp's own pages: documentation home (surfaces: web, macOS and iOS apps, CLI; updated 2026-09-02; checked 2026-10-02), MCP (amp.mcpServers, amp mcp add, remote definitions by scope, workspace server approval, orbs; updated 2026-09-28; checked 2026-10-02), Skills (.agents/skills/, mcp.json, includeTools; updated 2026-09-27; checked 2026-10-02), Tools (default permissions, oracle, Librarian; updated 2026-09-28; checked 2026-10-02), Plugins (tool.call, ctx.ui.confirm, requireHuman, permissions example plugin; updated 2026-10-02; checked 2026-10-02), Plugin API (MCP tools named mcp__<server>__<tool>; checked 2026-10-02), Configuration (settings locations; checked 2026-10-02), Orbs (checked 2026-10-02), Automations (checked 2026-10-02), Puck (checked 2026-10-02), The Dial (modes; checked 2026-10-02), Workspaces (roles, shared configuration, Enterprise controls; checked 2026-10-02), MCP Registry Allowlist (checked 2026-10-02), Workspace Entitlements (checked 2026-10-02), Minimal Data Retention (checked 2026-10-02), Amp Frontier Corporation (spin-out from Sourcegraph, Dec 2, 2025; checked 2026-10-02), The Coding Agent Is Dead (VS Code and Cursor extensions retired March 5, 2026; Feb 19, 2026; checked 2026-10-02) and About (checked 2026-10-02). Data Workers setup follows the client setup docs and the open-source dataworkers-claw-community repository (start-agent.sh and the agent tool definitions for trace_cross_platform_lineage, blast_radius_analysis, diagnose_incident, get_root_cause, remediate, run_quality_check, scan_pii, apply_migration and rollback_migration; checked 2026-10-02). Product names and statuses change quickly; if we've got something wrong, tell us and we'll fix it.