You're on Gemini CLI: give it governed data context and approved data changes
Your engineers already work in Gemini CLI. Add Data Workers to mcpServers in settings.json and Gemini CLI answers from governed data context, while every data change goes through approvals and receipts.
Your engineers already live in Gemini CLI. They type gemini in the terminal, sign in with Google or a Code Assist license, and work against a 1M-token context window with built-in file, shell, web fetch and Google Search tools. They keep project notes in GEMINI.md, install extensions from the gallery that package MCP servers, commands, sub-agents and agent skills (Google ships ones for BigQuery, Looker, Cloud SQL, Cloud Composer and Knowledge Catalog), and add MCP servers to mcpServers in settings.json. Your platform team decides which folders are trusted, writes policy rules that allow, ask or deny each tool, and may enforce an MCP allowlist from system settings or the admin console. Gemini CLI is open source under Apache 2.0, and it is very good at the job it was built for: reading a repository, writing the change and running it from the terminal, asking before it acts. Data Workers is the agentic data platform behind it. Connected over MCP, Gemini CLI answers from your governed data context, and every change to production data goes through Data Workers' approvals and leaves a receipt.
Key takeaways
- •Gemini CLI is the coding agent. Data Workers is the data team it calls. Gemini CLI writes and runs the code; Data Workers brings the context, owns the change in production and proves it worked.
- •A few `mcpServers` entries connect them. Add the Data Workers agents to the project's
.gemini/settings.jsonor to~/.gemini/settings.json, and/mcplists their tools. - •Two locks on every write. Gemini CLI asks before any tool its policy marks
ask_user; Data Workers applies its per-domain guardrail, keeps a rollback point and writes the receipt. - •Admins stay in charge. System settings, policy rules and the enterprise MCP allowlist all apply to the Data Workers server the same way they apply to any other.
- •Climb the ladder one domain at a time. Start at L1 observe, move to L2 propose, then L3 act reversibly, and let trusted domains run at L4 autonomous.
Gemini CLI is the coding agent. Data Workers is the data team it calls.
Gemini CLI reads your dbt project, your Composer DAGs and your SQL, and writes the change you ask for. With Google's data extensions it can also query BigQuery and Looker directly. What it doesn't hold is the state of the estate: which table is canonical, who owns it, what reads it downstream, what changed overnight in a source system, and who must approve a fix to finance data. That is what Data Workers holds. The Data Context Wizard keeps one governed context graph across warehouses, dbt, orchestration and BI. The Data-Agents Swarm (20+ specialist agents) does the data work. The Autonomous Data-Conductor runs each fix end to end: detect, diagnose, fix, review, verify, remember. Spellbook Data Catalog (in preview) is where people review, approve, roll back and audit.
Here is one request, end to end. It is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| 01:55 | Cloud SQL | An app release starts storing orders.discount in cents instead of dollars, as INT64 |
| 02:10 | BigQuery | Datastream lands the new column type in raw.orders |
| 06:00 | dbt on Managed Airflow (formerly Cloud Composer) | The run succeeds; fct_revenue subtracts cents as if they were dollars |
| 08:40 | Looker | Net revenue in the Revenue Explore drops overnight |
| 09:02 | Gemini CLI | An analytics engineer types: "Why did net revenue drop in Looker last night?" |
| 09:03 | Gemini CLI | Gemini CLI calls Data Workers over MCP (trace_cross_platform_lineage, run_quality_check, diagnose_incident) |
| 09:04 | Data Workers | run_quality_check passes on raw.orders, so diagnose_incident looks upstream and ties the drop to the 01:55 release that stores discount in cents; blast radius is 4 dbt models and 3 Looker Explores |
| 09:07 | Gemini CLI | Gemini CLI asks to run remediate; the engineer confirms the tool call |
| 09:20 | Spellbook | The finance data owner approves the change, as the finance domain requires |
| 09:22 | dbt / BigQuery | Data Workers pins a cents-to-dollars cast and rebuilds the affected models, keeping a rollback point |
| 09:41 | BigQuery | Totals match Cloud SQL; dbt tests pass; receipt written |
| 10:00 | Looker | Revenue is right before the morning review |

The engineer never left the terminal. Gemini CLI did what it does best: took the question, picked the right tools, showed what it was about to run and asked first. Data Workers did the data work: it found the cause in the source system, measured the blast radius through dbt into Looker, routed the change to the person who owns finance data, made the fix reversibly and proved it against the source. Meanwhile Gemini CLI writes the lasting fix where it belongs, in the repository: a typed staging model and a dbt test on the discount unit, opened as a pull request your team reviews as usual.
| What Gemini CLI does | What Data Workers does |
|---|---|
Takes the question in the terminal, with your repo and GEMINI.md in context | Answers it from governed context: lineage, ownership, definitions, recent changes |
| Reads and edits files and runs shell commands, sandboxed when you turn it on | Works on BigQuery and dbt directly, and on Composer and Looker over their APIs, through each system's own permissions |
Asks before tool calls its policy rules mark ask_user | Applies per-domain guardrails and routes the change to the data owner |
| Writes the code fix and the dbt test | Owns the production change, keeps a rollback point and verifies downstream |
| Shows the result in this session | Writes a receipt that outlives the session: who, why, what changed, how to undo |
Why doesn't Gemini CLI just do this itself?
Focus and risk. Google built Gemini CLI as an open-source agent for the terminal, the most direct path from a prompt to the model, and its design is right for that job. Its controls are built around the session. Every MCP tool call asks for confirmation unless the server is trusted. Policy rules allow, ask or deny each tool, with admin policies outranking user ones. Trusted folders stop a project's MCP servers, hooks and settings from loading until someone approves the folder. Enterprise admin controls can switch MCP off, enforce an allowlist of servers and keep users out of yolo mode. Every one of those controls answers one question: what may this agent, on this machine, call right now?
Changing production data is a different product. A fix to fct_revenue touches systems Gemini CLI doesn't run, owned by people who aren't at the keyboard, with downstream Explores nobody in the repository can see. Doing it safely takes context about every other system, a blast radius, an approver per domain, a rollback plan, verification after the fact and a receipt an auditor can read next quarter. It also means carrying the liability for changes in tools Google's CLI doesn't own. A general coding agent is right to leave that to a server built for it. That is the product Data Workers is, and Gemini CLI is a great place for your engineers to ask for it.
Every tool owns a slice. Data Workers covers the whole lifecycle
Each point tool adds another console, another contract and another handoff. Data Workers covers the whole data lifecycle with one context, one approval flow and one audit trail. Gemini CLI owns a valuable slice of it: writing and running code in the terminal. It leads where that is the work, in pipeline code.

| Stage | Data Workers | Gemini CLI | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 4 | Gemini CLI reads the repository, GEMINI.md context files and the files you point it at. Data Workers keeps one governed context graph across warehouses, dbt, orchestration and BI. |
| Analytics & Insights | 8 | 5 | Gemini CLI explains code and query results in the session, with Google data extensions for BigQuery and Looker. Data Workers answers business questions from governed definitions, lineage and ownership. |
| Data Quality | 8 | 5 | Gemini CLI writes dbt tests and SQL checks when you ask. Data Workers writes, runs and repairs checks across platforms and keeps them running. |
| Observability & Incidents | 8.5 | 3 | Gemini CLI debugs the error in front of it; it doesn't watch production. Data Workers detects, traces and closes incidents across systems with a receipt. |
| Pipelines & Ingestion | 8.5 | 9 | Gemini CLI leads. It writes pipeline code, dbt models and DAGs, runs them from the terminal and hands back the change. Data Workers owns the run in production behind approval. |
| Schema & Migration | 8 | 6 | Gemini CLI writes migrations and refactors well. Data Workers checks the blast radius across every downstream system before a schema change lands. |
| Governance & Access | 8.5 | 3.5 | Gemini CLI governs Gemini CLI: confirmations, policy rules, folder trust and admin MCP allowlists. Data Workers proposes least-privilege grants on your data platforms behind approvals. |
| Security & Privacy | 8 | 5 | Gemini CLI sandboxes shell and file operations and redacts sensitive environment variables. Data Workers flags sensitive column names in pull request review across the estate. |
| Cost / FinOps | 8 | 2 | Gemini CLI doesn't work on your warehouse bill. Data Workers traces Snowflake credits to the dbt model behind them and drafts the fix for its owner. |
| MLOps & Models | 7.5 | 6 | Gemini CLI writes training, feature and evaluation code on request. Data Workers keeps the data under the models healthy and traced. |
How Gemini CLI and Data Workers work together
Gemini CLI stays on top, where engineers ask, delegate and confirm. Data Workers sits underneath as MCP servers: the Context Wizard, the Swarm, the Conductor and the guardrails. Spellbook is where people look: data owners review and approve changes, roll them back and read the audit trail.

Setup over MCP today. Every Data Workers agent is a standard MCP stdio server, so Gemini CLI starts it like any other server in mcpServers. Clone the open-source core (dataworkers-claw-community, Apache 2.0), run npm install, and add a start-agent.sh entry for each agent you want, as our client setup docs describe. Put the block in the project's .gemini/settings.json, or in ~/.gemini/settings.json to use it in every project. Then start gemini and run /mcp to see the servers and their tools. Here is a version that exposes only the tools one team needs and keeps confirmations on. Example:
{
"mcpServers": {
"dw-context-catalog": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-context-catalog"],
"includeTools": ["trace_cross_platform_lineage", "blast_radius_analysis"]
},
"dw-schema": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-schema"],
"includeTools": ["assess_impact"]
},
"dw-incidents": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-incidents"],
"includeTools": ["diagnose_incident", "get_incident_history", "remediate"]
},
"dw-quality": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-quality"],
"includeTools": ["run_quality_check", "get_quality_score"]
}
},
"mcp": {
"allowed": ["dw-context-catalog", "dw-schema", "dw-incidents", "dw-quality"]
}
}Leave trust off (it defaults to false): with it on, Gemini CLI skips every confirmation for that server. Keep server names hyphenated, as here, because Gemini CLI's policy parser splits tool names on the first underscore. To let the read tools run without a prompt while remediate always asks, add a policy file. Example:
# Example: ~/.gemini/policies/data-workers.toml
[[rule]]
mcpName = "dw-context-catalog"
toolName = ["trace_cross_platform_lineage", "blast_radius_analysis"]
decision = "allow"
priority = 100
[[rule]]
mcpName = "dw-schema"
toolName = "assess_impact"
decision = "allow"
priority = 100
[[rule]]
mcpName = "dw-incidents"
toolName = ["diagnose_incident", "get_incident_history"]
decision = "allow"
priority = 100
[[rule]]
mcpName = "dw-incidents"
toolName = "remediate"
decision = "ask_user"
priority = 100
[[rule]]
mcpName = "dw-quality"
toolName = ["run_quality_check", "get_quality_score"]
decision = "allow"
priority = 100The open-source agents run against a built-in sample estate, so the whole flow works on day one. Your Data Workers deployment points the same tools at BigQuery, dbt, Composer and Looker and adds the per-domain guardrails, approvals and receipts. For a shared deployment, Gemini CLI also connects to Streamable HTTP servers through httpUrl, with OAuth or Google credentials. Admins who manage Gemini CLI centrally can define the servers once in system settings, where they win over workspace and user definitions, and list them in mcp.allowed. With Enterprise Admin Controls, the admin console can switch MCP on, enforce an MCP server allowlist (preview) and inject required servers for everyone (preview); both take remote servers by URL, which makes the shared HTTP deployment the natural fit there. Turn on trusted folders so a cloned repository can't bring its own MCP servers without a prompt; in an untrusted folder Gemini CLI connects to no MCP servers at all, so trust the project folder that holds the config. Teams building their own agents on Google's stack can use the same tools through ADK, covered in Data Workers in Gemini Enterprise and ADK.
The same request on the autonomy ladder. Autonomy is set per domain, on the ladder L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous.

- •L0 manual. The engineer queries BigQuery by hand and greps the dbt project. Gemini CLI helps write the SQL; nobody has the lineage or the source change.
- •L1 observe. Gemini CLI calls Data Workers and gets the cause, the blast radius and the owner. Nothing changes.
- •L2 propose. Data Workers drafts the fix with its blast radius and rollback plan. The finance owner approves in Spellbook; Gemini CLI asks the engineer before the tool call.
- •L3 act reversibly. For a domain you trust, Data Workers handles unit and type changes like this one on its own, with a rollback point and a receipt. The engineer reads it in the terminal.
- •L4 autonomous. The Conductor catches the type change at 02:10, before dbt runs at 06:00, fixes it and verifies it. Gemini CLI is where an engineer reads the receipt the next morning.
Gemini CLI confirmations and Data Workers guardrails stack. Turning one up never turns the other down.
What changes for your team
Engineers keep the agent they already use. What changes is the work behind it: the jobs that used to wait for a person with the right access and the right context now run through Data Workers, at the autonomy level each domain sets.

Data owners stop being pinged in chat for "is this table right?" and start approving scoped changes in Spellbook. On-call stops starting from zero, because the first question typed into Gemini CLI already returns the cause and the blast radius. The platform team stops writing a different MCP setup for every team: one server definition in system settings, one allowlist entry, one policy file and one audit trail.
Keep Gemini CLI, or consolidate?
Keep Gemini CLI if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too. With a coding agent, keeping it is the usual answer: Gemini CLI is where your engineers write code, and Data Workers makes every data question and data change they send from it governed. The same Data Workers server also answers in Claude, Codex and, for business users on Google's stack, Gemini Enterprise, so teams that use more than one assistant share one context and one audit trail. For the whole category, see Your company just rolled out AI assistants. Now what?.
The case for your CFO
Your engineers already use Gemini CLI on data work, on Code Assist licenses or on their own sign-in. The outcome to buy now is that the questions they ask get governed answers, and the changes they ask for get made safely: fewer wrong numbers reaching leadership, incidents closed in the same morning, and an audit trail for every change to production data.
The risk story is plain. At L1, Data Workers only reads and explains. At L2, it proposes and a named owner approves. At L3, it acts only where a change is reversible, and keeps the rollback point. L4 is reserved for domains you've trusted on evidence. Every action leaves a receipt: who asked, who approved, what changed, what it touched, and how to undo it. Gemini CLI keeps its own confirmations, policy rules and admin controls, so there are two locks on every write. Nothing migrates: Gemini CLI, BigQuery, dbt, Composer and Looker all stay. The cost case and how to model it are on the ROI page, and deployment and data-handling answers are on the security page.
Why now: coding agents are already the front door for data work, and an open-source agent spreads fast. Every week without governed context is another week of engineers guessing which table is canonical, and of production changes made with personal credentials and no receipt. The safety page walks through what agents can and can't do at each level, and the build-vs-buy page covers what it takes to wire this yourself.
The first win is small and visible: read-only Data Workers tools in Gemini CLI for one domain, so "why is this number wrong?" gets a traced answer. Start with a pilot; pricing has the details, and the pilot is credited in full against the first year.
The sentence to repeat upstairs: "Our engineers keep Gemini CLI; Data Workers gives it our governed data context and routes every data change through one approval flow with a receipt."
Getting started
Start with a pilot. Pick one domain where engineers already ask Gemini CLI about data, often finance or product analytics, add the Data Workers server to settings.json with read tools allowed by policy and change tools set to ask_user, and run it at L1 for the first weeks. Move that domain to L2 when the answers hold up, and to L3 for the fixes your team approves every time. See pricing; the pilot is credited in full against the first year.
FAQ
Does Data Workers bypass Gemini CLI's confirmations or policies? No. Gemini CLI still decides when to ask, from the server's trust setting and your policy rules, and admin policies still outrank user ones. Data Workers adds its own per-domain guardrail on top, so a change needs both.
Should we set `trust: true` on the Data Workers server? Not for change tools. trust skips every confirmation for that server. Keep it off and use a policy file to allow the read tools, so only remediate and other change tools prompt.
Can our admins control which MCP servers engineers use? Yes. Define the servers in system settings, where they take precedence, and list them in mcp.allowed. With Enterprise Admin Controls, MCP can be switched on or off from the admin console, and two preview controls go further: an allowlist that ignores any locally configured server not on it, and required servers injected for every user. Both take remote servers by URL, so point them at your shared Data Workers deployment. Admin policy rules also outrank user ones.
Can we use Data Workers alongside Google's BigQuery and Looker extensions? Yes. They answer different questions. The extensions query BigQuery and Looker directly; Data Workers adds lineage across systems, ownership, the blast radius and the approved change with a receipt.
Won't more tools crowd the context? Use includeTools to expose only what a team needs, or add only the agents a team uses. A good start is the lineage, schema and incident tools, adding more as domains move up the ladder.
Whose credentials touch BigQuery? Data Workers' own credentials, scoped per domain, so no personal warehouse keys sit in a terminal. The engineer's identity is recorded on the receipt as the person who asked and confirmed.
We could wire a BigQuery MCP server into Gemini CLI ourselves. Why buy? You can connect a warehouse. The hard parts are the context across systems, the blast radius, the approvals by domain, the rollback and the receipts. The build-vs-buy page walks through it.
Sources
- •Google, Gemini CLI repository and README (open source, Apache 2.0, auth options, release channels): https://github.com/google-gemini/gemini-cli (checked Oct 2, 2026)
- •Google, Gemini CLI docs, "MCP servers with Gemini CLI" (mcpServers, trust, includeTools, httpUrl, gemini mcp add, OAuth): https://geminicli.com/docs/tools/mcp-server/ (checked Oct 2, 2026)
- •Google, Gemini CLI docs, "Policy engine" (allow, ask_user, deny; mcpName; tiers): https://geminicli.com/docs/reference/policy-engine/ (checked Oct 2, 2026)
- •Google, Gemini CLI docs, "Gemini CLI for the enterprise" (system settings, precedence): https://geminicli.com/docs/cli/enterprise/ (checked Oct 2, 2026)
- •Google, Gemini CLI docs, "Enterprise Admin Controls" (MCP toggle, MCP allowlist in preview, strict mode): https://geminicli.com/docs/admin/enterprise-controls/ (checked Oct 2, 2026)
- •Google, Gemini CLI docs, "Trusted Folders": https://geminicli.com/docs/cli/trusted-folders/ (checked Oct 2, 2026)
- •Google, Gemini CLI docs, "Extensions": https://geminicli.com/docs/extensions/ (checked Oct 2, 2026)
- •Google, Gemini CLI extensions for BigQuery, Looker and other Google data services: https://github.com/gemini-cli-extensions (checked Oct 2, 2026)
- •npm, @google/gemini-cli (0.62.0 stable, Sep 29, 2026): https://www.npmjs.com/package/@google/gemini-cli (checked Oct 2, 2026)
- •Data Workers, open-source client setup (clone,
start-agent.shentries inmcpServers): https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026) - •Data Workers, open-source configuration (sample estate, transports): https://dataworkers.io/opensource-docs/configuration/ (checked Oct 2, 2026)
- •Data Workers, open-source core and agent tool definitions: https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)