Product
Product12 min readBy The Data Workers Team

You're on goose. Here's how Data Workers builds on it

Your engineers already run goose recipes on their machines and on a schedule. Add Data Workers as an extension and every data recipe answers from governed context, with approvals kept in Data Workers.

Your engineers already work in goose, in goose Desktop or the goose CLI. They turn on extensions for the tools they need, keep project hints in .goosehints, and pick the model that suits the task, from a frontier provider to a local one. The workflows that proved useful became recipes: YAML files with a prompt, parameters, the extensions to load and a structured response, started with a slash command or a link. The best ones run on their own with goose schedule, or headless with goose run in CI. goose started at Block and is now a project of the Agentic AI Foundation (AAIF) at the Linux Foundation, next to MCP itself. For automating engineering work, that setup is excellent. The open question for a data team is what a recipe should be allowed to do when the task touches the warehouse, the DAGs and the dashboards finance reads at 9:00.

That's where Data Workers comes in. goose is the open agent that runs your team's recipes. Data Workers is the agentic data platform those recipes call: it knows what a change touches, does the data work, and keeps the approval for every data change inside Data Workers. Added as goose extensions over MCP, Data Workers answers from one governed context graph, proposes fixes with their blast radius, and applies them after the domain owner approves, with rollback ready and a receipt. Your engineers keep their agent, their models, their recipes and their schedules.

Key takeaways

  • •goose runs the recipe; Data Workers carries the data change. A recipe calls Data Workers for load lag, lineage and diagnosis. Data Workers proposes the fix, routes it to the owner, queues the approved rerun and verifies it.
  • •Approvals stay in Data Workers. Scheduled and headless runs have nobody at the keyboard to click Allow, so a named owner approves data changes in Spellbook, set per domain.
  • •Three settings to start. Data Workers extensions in config.yaml, Smart Approve with write tools on Ask Before, and available_tools on each recipe.
  • •Admins vet the servers. goose's GOOSE_ALLOWLIST admits the Data Workers commands you approve.
  • •Climb one domain at a time. Start with a morning check recipe at L1 observe, move to L2 propose, then L3 act reversibly where the record earns it.

goose is the recipe runner. Data Workers is the data crew behind it.

goose is a general agent by design: it runs on the engineer's machine, works with any model, and speaks to any MCP server as an extension. The live estate sits outside the repo and the prompt: which Airflow DAG loads a Snowflake table, which dbt models read it, which Looker Explore the leadership review uses, and who owns each one. Data Workers keeps that in one governed context graph and does the data side of the work through 20+ specialist agents.

One morning, end to end. This is an illustration, not a customer case.

TimeSystemWhat happens
06:00goosegoose schedule runs the team's morning-data-check recipe, which loads the Data Workers extensions with read tools only.
06:01SnowflakeThe recipe calls get_incident_history. Data Workers reports analytics.fct_orders is 9 hours behind the load-lag baseline the team records with monitor_metrics.
06:02Airflowdiagnose_incident traces it to the load_orders DAG run that failed at 01:40 after the source added a discount_code column.
06:03dbttrace_cross_platform_lineage shows fct_orders feeds fct_revenue and two marts.
06:03LookerThe Revenue Explore reads fct_revenue and is on the agenda for the 9:00 review.
06:04gooseThe recipe returns its structured summary. Data Workers has queued a fix (add the column in staging, rerun the DAG, rebuild the models) for the owner.
07:42SpellbookThe platform owner reads the blast radius and approves.
07:50AirflowThe owner applies the drafted migration, with its rollback SQL on file; Data Workers queues the rerun of load_orders.
08:18dbtThe affected models rebuild and their tests pass.
08:31LookerRevenue totals match the source counts, and the receipt is recorded.
Incident timeline across the stack: what goose, your team and Data Workers each do, step by step

The owner made one decision at 07:42, and the numbers were right by 08:31.

JobWhat goose doesWhat Data Workers does
Start the workRuns a recipe on a schedule, headless in CI, or when an engineer asks in Desktop or the CLIAnswers each tool call from live data context: freshness, lineage, owners and usage
Understand the problemFollows the recipe's instructions and .goosehints, calls the tools it is givenDiagnoses the cause across Snowflake, Airflow, dbt and Looker and computes the blast radius
Write the changeEdits DAG code, SQL and dbt models when an engineer asksProposes the data fix (migration, rerun, backfill) with its rollback plan
Gate the actionPermission modes and per-tool Always Allow, Ask Before or Never Allow; the allowlist controls which extensions installRoutes every data change to the domain owner by policy, at the autonomy level set for that domain
Apply and verifyRuns shell checks in a recipe's retry block and returns a structured responseApplies the approved change in the warehouse, reruns what depends on it and checks the numbers downstream
RecordKeeps the session history and exports agent telemetry over OpenTelemetryWrites a receipt to the audit trail: who asked, who approved, what changed, how to undo it

Why doesn't goose just do this itself?

Focus and risk. goose built an excellent product for one job: an open, local agent that works with any model and any MCP server, runs on a laptop or in CI, and packages good workflows as recipes. Its design follows from that. It lets each person choose how much to approve, classifies tools as read or write on a best-effort basis, and leaves what an extension does on the far side to that extension. As a foundation project, it stays neutral about which systems you connect. That is exactly right for a general agent.

Changing production data across systems is a different product. It needs to know that a column feeds a finance model in another repo, compute a blast radius before anything runs, send the approval to the domain owner instead of whoever started the recipe, keep a rollback path, check the numbers afterwards and leave a receipt an auditor can read. It also means answering for changes in Snowflake, Airflow and Looker, which goose doesn't run. A sensible open agent leaves that to the server on the other end of the extension. That server is Data Workers.

Every tool owns a slice. Data Workers covers the whole lifecycle

goose goes deepest on writing and automating pipeline code. Data Workers covers every stage of the lifecycle around that work. Each point tool adds another console, contract and handoff; Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, goose goes deep on its own area
StageData WorkersgooseWhy we scored it this way
Catalog & Context94goose reads .goosehints, files and Memory on demand, and MCP Roots share the working directory. Data Workers keeps one governed graph of tables, owners, lineage and usage across platforms.
Analytics & Insights85goose can query through an extension and chart results with Auto Visualiser. Data Workers answers business questions from governed metric definitions.
Data Quality84goose writes dbt tests and checks when asked, and recipe retry checks can gate a run. Data Workers runs quality checks, writes the missing tests and repairs failing ones.
Observability & Incidents8.53goose's OpenTelemetry export watches the agent itself. Data Workers detects data incidents, traces the cause and closes them with a receipt.
Pipelines & Ingestion8.59goose's home stage: it writes and runs pipeline code, and recipes, goose run and goose schedule automate it headless. Data Workers builds and reruns pipelines behind approval.
Schema & Migration86goose writes migration scripts and DDL on request. Data Workers assesses the impact, drafts each migration with its rollback SQL for the owner to apply in approved waves.
Governance & Access8.53Permission modes, tool permissions and the extension allowlist govern the agent. Data Workers proposes least-privilege grants on your data platforms behind approvals.
Security & Privacy85goose checks extensions for known malware, keeps secrets in the keyring and can detect prompt injection. Data Workers' pull request review flags new columns whose names or annotations look sensitive, and leaves a receipt on every data change.
Cost / FinOps82goose can show estimated model cost in the CLI. Data Workers traces Snowflake credits to the query and dbt model behind them and drafts the fix for the model's owner.
MLOps & Models7.54goose writes training and feature code. Data Workers keeps the data under models healthy and connects to MLflow and W&B.

How goose and Data Workers work together

goose stays on top, where engineers chat and schedule recipes. Data Workers sits underneath as goose extensions: the Data Context Wizard answers from the governed graph, the Data-Agents Swarm does the work, the Autonomous Data-Conductor runs each fix end to end, and Spellbook Data Catalog (in preview) is where people review, approve, roll back and audit.

How Data Workers fits with goose: your coding agent on top, Data Workers in the middle, your estate underneath

Add the extensions. Every Data Workers agent is a standard MCP stdio server, so goose connects with the same pattern as every other client. Clone the open-source core (dataworkers-claw-community, Apache 2.0), run npm install, and point each extension at start-agent.sh, as our client setup docs describe. Use Add custom extension (type Standard IO) in goose Desktop, Add Extension in goose configure, or edit ~/.config/goose/config.yaml directly. Example:

GOOSE_MODE: smart_approve
extensions:
  dw-context-catalog:
    type: stdio
    name: dw-context-catalog
    enabled: true
    cmd: /path/to/dataworkers-claw-community/start-agent.sh
    args: ["dw-context-catalog"]
    timeout: 300
  dw-incidents:
    type: stdio
    name: dw-incidents
    enabled: true
    cmd: /path/to/dataworkers-claw-community/start-agent.sh
    args: ["dw-incidents"]
    timeout: 300

Set the gates. goose applies Autonomous mode by default, which suits an engineer working on their own files. For data work, set Smart Approve (as above, or /mode smart_approve in a session), then in goose configure, goose settings, Tool Permission, mark Data Workers' read tools (get_incident_history, trace_cross_platform_lineage, get_root_cause, diagnose_incident) as Always Allow and any tool that changes data, such as remediate, as Ask Before. That is the first lock. The second is Data Workers' own guardrail for the domain, which holds whatever mode a session runs in. goose performs best with fewer than 25 tools enabled, and available_tools keeps the list short.

Write the recipe. A recipe names the extensions it needs, narrows each with available_tools, takes parameters and returns a structured response. Example, saved as recipes/morning-data-check.yaml:

version: "1.0.0"
title: "Morning data check"
description: "Check key marts with Data Workers and diagnose anything stale"
parameters:
  - key: domain
    input_type: string
    requirement: required
    description: "Data domain to check, for example orders"
extensions:
  - type: stdio
    name: dw-context-catalog
    cmd: /path/to/dataworkers-claw-community/start-agent.sh
    args: ["dw-context-catalog"]
    timeout: 300
    available_tools: [trace_cross_platform_lineage]
  - type: stdio
    name: dw-incidents
    cmd: /path/to/dataworkers-claw-community/start-agent.sh
    args: ["dw-incidents"]
    timeout: 300
    available_tools: [get_incident_history, diagnose_incident, get_root_cause]
instructions: |
  Use Data Workers tools only. Call get_incident_history for the {{ domain }} marts.
  For anything stale, call diagnose_incident and get_root_cause, then
  trace_cross_platform_lineage to list downstream readers. Never change data.
prompt: "Check the {{ domain }} domain: what is stale, why, and who is affected."
response:
  json_schema:
    type: object
    properties:
      stale_tables: { type: array, items: { type: string } }
      root_cause: { type: string }
      downstream: { type: array, items: { type: string } }
    required: [stale_tables, root_cause, downstream]

Run it with goose run --recipe recipes/morning-data-check.yaml --params domain=orders, then schedule it: goose schedule add --schedule-id orders-morning-check --cron "0 0 6 *" --recipe-source ./recipes/morning-data-check.yaml. Scheduled and headless runs finish without anyone at the keyboard, and goose's headless guide runs them in auto mode. That is why the recipe carries read tools only and the approval for the fix lives in Data Workers. Share the recipe through GOOSE_RECIPE_GITHUB_REPO, or map it to a slash command in config.yaml so an engineer can type /orders-check in any session.

Vet the servers. For a team, publish an allowlist YAML file over HTTPS and set the GOOSE_ALLOWLIST environment variable to its URL. Each entry under extensions is an id and the command goose should accept; goose checks an extension's install command against the list and blocks anything not on it. goose's guide recommends exact command strings with full paths, so list the start-agent.sh path of your managed install with each agent name.

For a shared deployment, run Data Workers over Streamable HTTP and add it as a streamable_http extension; goose handles OAuth for remote extensions.

One request, from L0 to L4. An engineer types into goose Desktop: "the orders dashboard looks low today, find out why and fix it". Here is what happens at each level, set per domain.

LevelWhat happens when the engineer asks
L0 manualgoose helps the engineer read the Airflow log and write the fix by hand. Data Workers isn't in the loop.
L1 observegoose calls get_incident_history and diagnose_incident. Data Workers answers in the session: fct_orders is stale because load_orders failed on a new source column, the Revenue Explore reads it, and the platform team owns it.
L2 proposeData Workers drafts the fix (add the column in staging, rerun the DAG, rebuild the models) with its blast radius. The owner approves in Spellbook.
L3 act reversiblyFor this pre-approved class of fix, Data Workers drafts the migration with rollback SQL, queues the rerun itself once the owner applies it, and records the receipt. The engineer sees the result in goose.
L4 autonomousNew-column failures in this domain run end to end without waiting: detect, fix, verify, record. The scheduled recipe reports what was done. The owner reviews receipts in Spellbook and can dial the domain back at any time.

What changes for your team

Engineers keep their agent, models and recipes. What changes is the work around each data problem: the early hunt through DAG logs, the thread to find an owner, the manual backfill, the evidence an auditor asks for later.

Six jobs that run on autopilot with Data Workers next to goose, with a concrete example of each

The data platform team stops being the human lookup service for lineage and ownership. A morning check ends in a diagnosis and a queued fix instead of a list of red tables. Owners approve changes in one place with the blast radius in front of them. Admins govern Data Workers with the allowlist and permission settings they already use for every goose extension, and the audit trail builds itself from receipts.

Keep goose, or consolidate?

Keep goose if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

For most data teams, goose stays: it is the open agent engineers chose, and its recipes are a good home for repeatable work. What teams consolidate are the extra tools around data problems: the cron script that checks freshness, the separate lineage lookup, the spreadsheet of who owns what, the runbook for backfills. Data Workers runs that work with one context, one approval flow and one audit trail, and serves every other MCP client from the same agents. For other tools, read You're on OpenCode, You're on Cline and You're on Claude, or start from the hub, Your company just rolled out AI assistants. Now what?.

The case for your CFO

Your engineers already use goose to automate their work, and some of it now runs every morning without them. Data Workers makes the data side of that automation trustworthy: stale tables found and fixed before the leadership review, fewer wrong numbers in front of finance, and less senior engineering time spent tracing who reads what.

The risk story is plain. At L1, agents read and explain. At L2, they propose and a named owner approves. At L3, they act on pre-approved, reversible classes of change, with rollback ready. Every action leaves a receipt: who asked, who approved, what changed and how to undo it. Because the approval lives in Data Workers, a scheduled recipe never approves its own data change. Our safety guide covers each level, and the security and deployment guide covers where your data goes.

Why now: recipes and schedules move agent work from a person at a keyboard to jobs that run on their own, and those jobs should touch data with context and an approval trail. Our build-vs-buy guide lays out what building that layer in house takes.

The first win is one scheduled morning check for one domain, at L1. What stays the same: goose, your model providers, Snowflake, Airflow, dbt, Looker and your permission systems. Nothing is migrated. Our ROI guide shows how to size the return.

The sentence to repeat upstairs: "Our engineers keep goose; Data Workers makes every data change a recipe starts scoped, approved, verified and recorded."

Getting started

Start with a pilot: add the Data Workers extensions to goose, schedule one morning check recipe, and run one domain such as freshness failures from L1 to L2 with your own engineers and owners. See pricing for the pilot terms; the pilot is credited in full against the first year.

FAQ

Is goose still a Block project? goose started at Block. Block contributed it to the Agentic AI Foundation (AAIF) at the Linux Foundation as a founding project, alongside MCP and AGENTS.md, when the foundation formed on December 9, 2025. In April 2026 the docs moved to goose-docs.ai and the code to github.com/aaif-goose/goose, under the Apache 2.0 license. Nothing changes for the Data Workers setup: extensions are still MCP servers in config.yaml.

Where does the Data Workers config go in goose? In ~/.config/goose/config.yaml under extensions (on Windows, %APPDATA%\Block\goose\config\config.yaml). Recipes can also declare the extensions they need, so a shared recipe brings its own Data Workers setup.

Which goose permission mode should we use for data work? Smart Approve or Manual Approval, with Data Workers' read tools on Always Allow and data-changing tools on Ask Before. Give scheduled and headless recipes read tools only through available_tools, and let Data Workers hold the approval for every data change.

Can a scheduled goose recipe change production data on its own? Only within the autonomy level you set for that domain in Data Workers. At L1 and L2 the recipe reads and proposes, and a named owner approves in Spellbook. At L3 only pre-approved, reversible classes run, each with rollback ready and a receipt.

Can our admins control which servers engineers install? Yes. Publish an allowlist YAML of approved extension commands and set GOOSE_ALLOWLIST to its URL; goose then blocks installation of any extension whose command isn't listed. goose also checks external extensions for known malware before activation.

Does it matter which model goose uses? No. goose works with many providers, including local models, and a recipe can set its own provider and model in settings. The approval and receipt live in Data Workers either way.

Which data platforms does this work with? Data Workers runs control-plane connectors for Snowflake, Databricks and BigQuery, and reads dbt manifests and runs. For Airflow, Looker and the rest of your stack, Data Workers connects over each tool's API or MCP server today.

Sources

goose facts checked October 2, 2026, from its own docs and repository: llms.txt overview (checked 2026-10-02), Linux Foundation announces the formation of the Agentic AI Foundation (December 9, 2025; checked 2026-10-02), goose has a new home, the Agentic AI Foundation (April 7, 2026; checked 2026-10-02), aaif-goose/goose on GitHub (Apache 2.0, release v1.53.0 on 2026-10-02; checked 2026-10-02), Install goose (checked 2026-10-02), Using Extensions (checked 2026-10-02), Configuration Files (checked 2026-10-02), Permission Modes (checked 2026-10-02), Managing Tool Permissions (checked 2026-10-02), Extension Allowlist (checked 2026-10-02), Recipes and the Recipe Reference Guide (checked 2026-10-02), CLI Commands (checked 2026-10-02), Headless mode (checked 2026-10-02) and Using .goosehints (checked 2026-10-02). Data Workers setup follows the client setup docs and the open-source dataworkers-claw-community repository (start-agent.sh and the registered tools named here; checked 2026-10-02).