You're on JetBrains AI and Junie. Here's how Data Workers builds on it
Your data engineers already work in PyCharm and DataGrip. Connect Data Workers over MCP and Junie answers from governed data context, while every data change goes through approvals and receipts.
Many data engineers never left JetBrains. They write SQL in DataGrip against the warehouse connection they set up years ago, build dbt models and Airflow DAGs in PyCharm, and keep the platform services in IntelliJ IDEA. With AI Assistant installed, the AI Chat in each IDE now hosts coding agents: Junie, JetBrains' own agent (out of Beta since June 17, 2026), plus Claude Agent, Codex and GitHub Copilot. Junie plans in Advanced Plan Mode before it touches code, follows the team's AGENTS.md, asks before sensitive actions unless the Action Allowlist says otherwise, and also runs from the terminal (Junie CLI), on GitHub and in GitLab CI/CD. In DataGrip you can add database objects to the chat as context. For writing and refactoring data code, that setup is excellent. The open question for a data team is what happens when the change Junie writes alters a column that a finance export, a Tableau workbook and two other models depend on.
That's where Data Workers comes in. JetBrains AI and Junie are where your engineers write the change. Data Workers, the agentic data platform, knows what that change touches and carries it into production: scoped, approved, verified and recorded. Connected over MCP, Junie and AI Assistant answer data questions from Data Workers' governed context, and every change to data goes through Data Workers' approvals and receipts. Your engineers keep the IDE they chose.
Key takeaways
- •Junie writes the change; Data Workers carries the data change. PyCharm, DataGrip and IntelliJ IDEA stay where engineers ask, plan and delegate. Data Workers supplies lineage, owners and blast radius, then applies the approved change and verifies it downstream.
- •Two files to start. One entry per Data Workers agent in Junie's
mcp.json(or under AI Assistant's Model Context Protocol settings) connects the agents, and a few lines in AGENTS.md tell Junie when to check lineage, freshness, quality, PII and schema impact. - •Two locks on every write. Junie's Action Allowlist (or the operation mode of the agent you picked in the AI Chat) gates the tool call; Data Workers' per-domain guardrail gates the change to the warehouse. No agent approves its own work.
- •One configuration, IDE and terminal. Junie CLI uses the same
mcp.jsonas Junie in the IDE, so a request typed in PyCharm or in the terminal gets the same impact checks, and Data Workers posts the blast radius on every pull request Junie opens. - •Climb one domain at a time. Start at L1 observe (answers, lineage, impact reports on pull requests), move to L2 propose, then L3 act reversibly where the record earns it.
JetBrains AI is the IDE's agent. Data Workers is the data crew behind it.
JetBrains IDEs know your code deeply: inspections, refactorings, the debugger, test runners, and in DataGrip the live schema of every connection. Junie uses that IDE intelligence when it plans and edits. What the project can't show on its own is the rest of the estate: which Snowflake objects a dbt model really feeds, which Tableau workbook extracts a column, which Airflow task exports it to finance, and who owns each one. Data Workers keeps that in one governed context graph and does the data side of the work through 20+ specialist agents.
Here is one change, end to end. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| 09:10 | PyCharm | An analytics engineer asks Junie in the AI Chat to change amount in stg_orders from FLOAT to NUMBER(18,2) to stop rounding drift. |
| 09:11 | PyCharm | Junie drafts the plan. The team's AGENTS.md says to check impact before editing a model or DDL, so Junie calls assess_impact on the Data Workers MCP server. |
| 09:12 | dbt | Data Workers reports that fct_revenue and fct_refunds read the column. |
| 09:13 | Airflow | It also finds the finance_close DAG's export task casting amount in Python. |
| 09:13 | Tableau | And the Revenue workbook's extract, which rounds the field in a calculated field. |
| 09:31 | GitHub | Junie opens a pull request that changes the staging model and fixes all three readers together. |
| 09:33 | GitHub | Data Workers posts the blast radius on the pull request: two models, one DAG task, one workbook, with the rollback plan. |
| 09:52 | GitHub | dbt CI passes and the analytics lead approves the pull request. |
| 09:58 | Spellbook | The platform owner approves the Snowflake change and the rebuild of the affected models. |
| 10:05 | Snowflake | The owner applies the drafted change, with its rollback SQL on file; Data Workers queues the rebuild of the affected models. |
| 10:40 | Tableau | Revenue totals match the pre-change baseline to the cent, the next finance_close run succeeds, and the receipt is recorded. |

Without that context, the same request is a tidy pull request that passes CI and quietly changes a finance export at month end. With it, one pull request and one owner approval cover the whole change.
| Job | What JetBrains AI and Junie do | What Data Workers does |
|---|---|---|
| Understand the request | Read the project, AGENTS.md and any database objects added to the chat; plan in Advanced Plan Mode | Adds live data context: lineage across dbt, Snowflake, Airflow and Tableau, owners and usage |
| Write the change | Edit dbt models, SQL, DAGs and scripts using the IDE's inspections, tests and refactorings | Supplies the impact report and the list of every reader the change must fix first |
| Gate the action | Junie asks before MCP tools and terminal commands by default; the Action Allowlist sets what runs without asking | Routes data changes to the domain owner by policy, with blast radius attached |
| Apply to production | Commit to a branch and open the pull request | Applies the approved migration or rebuild in the warehouse, with rollback ready |
| Verify | Run tests, builds and run configurations in the IDE | Checks the changed tables against baselines, then workbooks and downstream runs |
| Record | Keep the plan in .junie/plans and the chat history | Writes a receipt to the audit trail: who asked, who approved, what changed, how to undo it |
Why doesn't JetBrains just do this itself?
Focus and risk. JetBrains has spent more than two decades building tools for working with code, and Junie carries that focus forward: it works on the project in front of it, uses the IDE's own engine to understand that code, and treats terminal commands, MCP tools and reads of secret files as sensitive actions that need approval. JetBrains even ships an MCP server in the IDE (since IntelliJ IDEA 2025.2), so other agents can use its inspections, refactorings and, with the Database Tools plugin, its database connections. For data analysis, JetBrains also offers Databao, a desktop data-analyst app that shows where every number comes from. That is exactly the right focus for an IDE vendor: help people write code and answer questions well.
Changing production data across systems is a different product. It needs to know that a Tableau workbook nobody keeps in the repo extracts a column, and that a DAG task reads it at month end. It needs a blast radius before anything runs, an approval routed to the domain owner, a rollback path inside Snowflake, a check that the numbers are right afterwards, and a receipt an auditor can read. It also means taking responsibility for changes in systems JetBrains doesn't run. A sensible IDE vendor leaves that to the server on the other end of the MCP connection. That server is Data Workers.
Every tool owns a slice. Data Workers covers the whole lifecycle
JetBrains AI and Junie go deepest on writing and refactoring pipeline code with real IDE intelligence behind them. Data Workers covers every stage of the data lifecycle around that code. Each point tool adds another console, contract and handoff; Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail.

| Stage | Data Workers | JetBrains AI | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 5 | JetBrains AI reads the project, AGENTS.md and the database objects you add to the chat. Data Workers keeps one governed graph of tables, owners, lineage and usage across platforms. |
| Analytics & Insights | 8 | 5 | AI Chat in DataGrip writes and explains SQL against your connections. Data Workers answers business questions from governed metric definitions. |
| Data Quality | 8 | 5 | Junie writes tests and checks when asked. Data Workers runs quality checks, writes the missing tests and repairs failing ones. |
| Observability & Incidents | 8.5 | 3 | Junie fixes CI failures and reviews pull requests. Data Workers detects data incidents, traces the cause and closes them with a receipt. |
| Pipelines & Ingestion | 8.5 | 9 | JetBrains' home stage: Junie plans, writes and refactors pipeline code with the IDE's own inspections, tests and debugger. Data Workers builds and reruns pipelines behind approval. |
| Schema & Migration | 8 | 6 | DataGrip and Junie write DDL and migration code. Data Workers assesses the impact, drafts each migration with its rollback SQL for the owner to apply in approved waves. |
| Governance & Access | 8.5 | 3 | IDE Services and JetBrains Central govern AI in the IDE. Data Workers proposes least-privilege grants on your data platforms behind approvals. |
| Security & Privacy | 8 | 5 | Junie asks before reading secret files and respects .aiignore. Data Workers' pull request review flags new columns whose names or annotations look sensitive, and leaves a receipt on every data change. |
| Cost / FinOps | 8 | 2 | JetBrains AI plans meter agent usage in AI credits. Data Workers traces Snowflake credits to the query and dbt model behind them and drafts the fix for the model's owner. |
| MLOps & Models | 7.5 | 4 | PyCharm and Junie write training and feature code. Data Workers keeps the data under models healthy and connects to MLflow and W&B. |
How JetBrains AI and Data Workers work together
The IDE stays on top, where engineers ask, plan, delegate and approve. Data Workers sits underneath over MCP: the Data Context Wizard answers from the governed graph, the Data-Agents Swarm does the work, the Autonomous Data-Conductor runs each fix end to end, and Spellbook Data Catalog (in preview) is where people review, roll back and audit.

Connect the MCP server. Every Data Workers agent is a standard MCP stdio server and connects with the pattern in our client setup docs. Clone the open-source core (dataworkers-claw-community, Apache 2.0), run npm install, and add the agents you want to Junie's MCP configuration. In the IDE, open Junie's MCP Settings and edit the mcp.json that opens; from the terminal, run /mcp in Junie CLI. Project scope lives in .junie/mcp/mcp.json at the repo root and can be committed so the whole team gets the same servers; user scope lives in ~/.junie/mcp/mcp.json. AI Assistant takes the same mcpServers JSON under Settings | Tools | AI Assistant | Model Context Protocol (MCP), at global or project level. Example:
{
"mcpServers": {
"dw-context-catalog": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-context-catalog"] },
"dw-schema": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-schema"] },
"dw-incidents": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-incidents"] },
"dw-quality": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-quality"] },
"dw-governance": { "command": "/path/to/dataworkers-claw-community/start-agent.sh", "args": ["dw-governance"] }
}
}For a team, run Data Workers as a remote server over Streamable HTTP, which both AI Assistant and Junie accept by URL. If your organization manages AI Assistant through JetBrains IDE Services or JetBrains Central, an administrator can preconfigure the Data Workers server for everyone and decide whether engineers may add their own. Keep tokens out of committed project files; Junie's docs warn against sharing secrets in a committed .junie/mcp/mcp.json. Example (placeholder host):
{
"mcpServers": {
"data-workers": {
"url": "https://<your-data-workers-host>/mcp",
"headers": { "Authorization": "Bearer <token from your secret store>" }
}
}
}Add Data Workers lines to AGENTS.md. Junie reads .junie/AGENTS.md, or the root AGENTS.md together with any .junie/rules/*.md files; Codex and GitHub Copilot in the AI Chat read AGENTS.md too (Claude Agent reads CLAUDE.md). A short section tells the agent when to call Data Workers. Example:
## Data changes
- Before editing a dbt model or DDL, call assess_impact and list every downstream reader.
- Before trusting a query result, call get_incident_history for open incidents on the tables it reads.
- To find a table or its owner, call search_across_platforms, then trace_cross_platform_lineage.
- Before merge, call run_quality_check on changed models.
- When something is described as broken, call diagnose_incident before patching.
- On code that adds a column holding personal data, annotate the column and ask the data owner before merge.Project guidelines win over a personal ~/.junie/AGENTS.md when the two conflict, so the team's data rules hold.
Set the gates. Junie asks before running MCP tools by default; for the other agents in the AI Chat, the operation mode you pick decides what runs without approval. We recommend adding read tools such as lineage, freshness and impact to Junie's Action Allowlist with an MCP rule so they run without a prompt, and keeping anything that changes data behind Junie's confirmation and Data Workers' guardrail. Leave brave mode off on machines with production connections.
A note on DataGrip. With Data Workers connected, DataGrip's AI Chat answers the questions a connection can't: who owns this table, what reads it, is it fresh, and what breaks if this column changes. The engineer's DataGrip connection stays theirs; Data Workers uses its own scoped credentials to each platform.
One request, from L0 to L4. Take one request an engineer types into PyCharm: "the orders_daily freshness check failed again, fix it". Here is what happens at each level, set per domain.
| Level | What happens when the engineer asks |
|---|---|
| L0 manual | Junie helps the engineer read logs and write a patch by hand. Data Workers isn't in the loop. |
| L1 observe | Junie calls diagnose_incident. Data Workers answers in the chat: the Airflow load task timed out after a source API change, two workbooks are stale, and the owner is the finance data team. |
| L2 propose | Data Workers drafts the fix (a retry and timeout change on the DAG plus a backfill plan) with its blast radius. Junie opens the pull request; the owner approves. |
| L3 act reversibly | For this pre-approved class of fix, Data Workers reruns the load and the backfill itself, with rollback ready, and records the receipt. The engineer sees the result in the AI Chat. |
| L4 autonomous | Freshness incidents in this domain run end to end without waiting: detect, fix, verify, record. The owner reviews receipts in Spellbook and can dial the domain back at any time. |
If part of your team works in other tools, the wiring is the same pattern: see You're on Cursor, You're on GitHub Copilot (Copilot also runs inside JetBrains IDEs) and You're on Claude.
What changes for your team
Engineers keep their IDE and their habits. What changes is the work around each data change: finding who reads a column, finding the owner, the manual backfill, the evidence an auditor asks for later.

The data platform team stops being the human lookup service for lineage and ownership. Analytics engineers ship dbt changes from PyCharm with the impact attached, owners approve in one place with the blast radius in front of them, and the audit trail builds itself from receipts.
Keep JetBrains AI and Junie, or consolidate?
Keep JetBrains AI and Junie if you love them; Data Workers works with them from day one. Many teams consolidate once Data Workers runs that slice too.
For most data teams, the JetBrains IDEs stay, and Data Workers is built to sit underneath them. What teams tend to consolidate are the extra tools around data changes: the separate lineage lookup, the impact-analysis script someone maintains, the spreadsheet of who owns what, the manual runbook for backfills. Data Workers runs that work with one context, one approval flow and one audit trail, and it serves Junie, Claude Agent, Codex, GitHub Copilot and every other MCP client from the same agents. For the wider picture across every assistant your company bought, start from the hub, Your company just rolled out AI assistants. Now what?.
The case for your CFO
You already pay for JetBrains licences and AI, and your engineers write data code faster because of them. The outcome Data Workers adds is that the data changes they ship are right the first time: fewer broken reports, fewer month-end surprises in finance exports, and less senior engineering time spent tracing who reads what.
The risk story is plain. At L1, agents only read and explain. At L2, they propose and a named owner approves. At L3, they act only on pre-approved, reversible classes of change, with rollback ready. Every action leaves a receipt with who asked, who approved, what changed and how to undo it. Our safety guide covers exactly what an agent can and can't do at each level, and the security and deployment guide covers where your data goes.
Why now: since Junie left Beta, your engineers can hand an agent data code every day, and each of those changes should land with context and an approval trail. Building that layer in house is possible, and our build-vs-buy guide lays out what it takes.
The first win is impact reports on every dbt pull request in one repo, at L1, with no change to how anyone works. What stays the same: PyCharm, DataGrip, Junie, Snowflake, dbt, Airflow, Tableau and your permission systems. Nothing is migrated. Our ROI guide shows how to size the return.
The sentence to repeat upstairs: "Our engineers keep PyCharm, DataGrip and Junie; Data Workers makes every data change they write scoped, approved, verified and recorded."
Getting started
Start with a pilot: connect Data Workers to Junie in one repo, add the AGENTS.md lines, and run one domain such as freshness incidents or schema changes from L1 to L2 with your own engineers and owners. See pricing for the pilot terms; the pilot is credited in full against the first year.
FAQ
Does Data Workers replace Junie or AI Assistant? No. Junie and the other agents in the AI Chat keep writing the code. Data Workers adds the data context and carries the approved data change into the warehouse, then verifies and records it.
Where does the configuration go: AI Assistant settings or Junie's mcp.json? Both take the same mcpServers JSON. For Junie (IDE and CLI), use .junie/mcp/mcp.json in the repo for the team or ~/.junie/mcp/mcp.json for one engineer. For the other agents in the AI Chat, use the AI Assistant MCP settings page.
Can our admins control whether engineers use Data Workers? Yes. Organizations that manage AI Assistant through JetBrains IDE Services or JetBrains Central can preconfigure the MCP servers available to engineers and control whether they can add their own. Inside Data Workers, each domain's autonomy level is set centrally as well.
How is this different from JetBrains' own MCP server? They point in opposite directions. JetBrains' MCP server lets external agents use the IDE: its inspections, run configurations and, with the Database Tools plugin, its database connections. Data Workers gives the agents in the IDE governed context about the whole estate (lineage, owners, freshness, impact) and carries approved changes into production with a receipt. Many teams run both.
Can Junie change production data on its own? Only within the autonomy level you set for that domain. Junie's confirmation prompt or Action Allowlist gates the tool call, and Data Workers' guardrail gates the change itself. At L2 a named owner approves every change; at L3 only pre-approved, reversible classes run, each with rollback ready and a receipt.
Does this work when Junie runs on GitHub? Yes. The Junie GitHub Action runs Junie CLI on your own GitHub runners and follows the repository's AGENTS.md, and Data Workers posts the blast radius on the pull request itself, so a pull request Junie opens from an issue gets the same impact review and owner approval as one opened from PyCharm.
Which data platforms does this work with? Data Workers runs control-plane connectors for Snowflake, Databricks and BigQuery, and reads dbt manifests and runs. For Airflow, Tableau and the rest of your stack, Data Workers connects over each tool's API or MCP server today.
Sources
JetBrains capabilities and statuses are current as of October 2, 2026, from JetBrains' own documentation: AI Assistant: Model Context Protocol (AI Assistant 2026.2; settings path, transports, global or project level, IDE Services and JetBrains Central admin controls; checked 2026-10-02), AI Assistant: Agents (Junie, Claude Agent, Codex, GitHub Copilot and ACP agents in the AI Chat; operation modes, AGENTS.md, .aiignore; page dated 2026-08-03, checked 2026-10-02), AI Assistant in DataGrip (database objects as chat context; checked 2026-10-02), IntelliJ IDEA MCP server (built in since 2025.2, database tools; checked 2026-10-02), Junie leaves Beta (JetBrains blog, June 17, 2026; checked 2026-10-02), Junie (surfaces incl. GitLab CI/CD, Advanced Plan Mode, .junie/plans, AI credits; checked 2026-10-02), Databao (JetBrains' desktop data-analyst app; checked 2026-10-02), and the Junie documentation, which moved to junie.jetbrains.com/docs: MCP Settings, Junie CLI MCP configuration, Action Allowlist, Guidelines and memory, Junie GitHub Action and Junie IDE plugin (all checked 2026-10-02). Data Workers setup follows the client setup docs and the open-source dataworkers-claw-community repository (start-agent.sh and the agent tool registrations for assess_impact, search_across_platforms, trace_cross_platform_lineage, run_quality_check, diagnose_incident, scan_pii and rollback_migration; checked 2026-10-02). Product names and statuses change quickly; if we've got something wrong, tell us and we'll fix it.