You're on Neo4j: Keep Your Knowledge Graph True, From the Warehouse Feed to the GraphRAG Answer
Already on Neo4j? Data Workers brings your graph's model in as governed context and keeps the warehouse feeds behind it true, with approvals and receipts.
Your team models the business as a graph. Customers, accounts, products, plants, suppliers and contracts are nodes; who buys what, which part comes from which supplier and which document mentions which product are relationships. The graph lives in AuraDB or a self-managed Neo4j cluster, it is loaded from the warehouse through Aura Import or a nightly load DAG that runs MERGE on business keys, and it answers questions tables struggle with: which suppliers are a single source for a critical part, which accounts sit inside one fraud ring, which documents ground a GraphRAG answer. Graph Data Science runs the algorithms, Aura Agent and neo4j-graphrag serve GraphRAG, and the official Neo4j MCP server (v1.6.0 since September 2026) lets Claude Code, Cursor or VS Code read the schema and run Cypher. Neo4j holds the relationships. Data Workers keeps them true: Data Context Wizard brings the graph's model in as context with provenance, joins it to lineage, quality and usage, and when a feed breaks or a definition drifts, Data Workers proposes and carries the fix with approvals and a receipt.
The graph is only as good as what feeds it. An ERP key format change, a renamed dbt column or a moved definition of "active customer" all reach the graph through the nightly load, and a GraphRAG agent answers from whatever landed. That seam between the warehouse and the graph is the job this guide covers.
Key takeaways
- •Neo4j keeps its job. AuraDB, Cypher, Graph Data Science, Aura Agent and your load jobs stay as they are. Data Workers works next to them, from day one.
- •The graph's model becomes governed context. Data Context Wizard reads labels, relationship types and the keys each load merges on, and puts them next to warehouse lineage, quality, usage and a named owner.
- •Feeds are watched before the load. Upstream schema changes, key-match drops and stale tables are caught in the warehouse, traced to the graph and fixed before the nightly
MERGEruns. - •Writes stay gated. Your team's assistant reads the graph through the Neo4j MCP server in read-only mode, side by side with Data Workers. Approved fixes land in dbt and in your own load pipeline, with a blast radius, an approver and a rollback path.
- •Start with a pilot. One graph, read-only, with its feeds watched; then one fix class in one domain, on the ladder from L0 manual to L4 autonomous.
Neo4j holds the relationships. Data Workers keeps them true.
Neo4j is a superb place to model and traverse relationships. Its MCP server exposes get-schema to "introspect labels, relationship types, property keys", read-cypher for read queries, write-cypher for writes and list-gds-procedures for Graph Data Science. Aura hosts an MCP server per instance. Aura Agent builds GraphRAG agents "contextualized by your own knowledge graph in AuraDB" and can publish them over an API or an MCP endpoint. Aura Import loads from Snowflake, Databricks, BigQuery, Redshift, Postgres and more without code.
Outside the graph sits everything that decides whether it is right: the ERP and SaaS sources, ingestion, dbt models, the orchestrator that runs the load and the semantic layer that defines its terms. Data Workers covers that side with one context, one approval flow and one audit trail.
Here is an evening with Data Workers next to a supplier graph. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| 18:40 | SAP S/4HANA | A new supplier ID format with a plant prefix goes live for newly onboarded suppliers and existing records updated in the change |
| 19:10 | Fivetran | The sync lands the updated supplier records in Snowflake |
| 19:30 | dbt | dim_supplier builds; its unique and not_null tests pass, because the new IDs are unique and present |
| 19:35 | Data Workers | It detects the format change on supplier_id and a drop in key matches against yesterday's snapshot: 412 suppliers would no longer match their existing graph nodes |
| 19:40 | Neo4j AuraDB | The on-call's assistant reads the graph model over the Neo4j MCP server and hands it to Data Workers: Supplier nodes merge on supplier_id, and SUPPLIES relationships hang off them |
| 19:45 | Airflow | Data Workers opens an incident ahead of the 23:00 graph load, with the cause, the 412 affected suppliers and the downstream Aura Agent named |
| 20:05 | dbt | Data Workers proposes a diff that normalizes the prefixed IDs back to the business key the graph uses, with its blast radius: two dbt models, one load DAG, one graph |
| 21:15 | Spellbook | The graph owner reviews the diff and the impact and approves; dbt CI passes |
| 23:00 | Airflow | The load runs; MERGE matches every supplier to its existing node |
| 23:20 | Neo4j AuraDB | Data Workers verifies node counts, finds no duplicate keys and writes the receipt |
| 08:30 | Aura Agent | A category buyer asks which parts depend on a single supplier; the GraphRAG answer is right |

Without that check, the load would have created 412 new Supplier nodes with no relationships, and the morning answer about single-source risk would have been wrong in a way nobody could see from the graph alone. Neo4j did exactly what it was asked to do. The fix was upstream, in systems the graph doesn't run, and it went through an owner's approval first.
| Job | What Neo4j does | What Data Workers does |
|---|---|---|
| The model | Stores entities and relationships, with constraints and indexes where you want them | Reads the model as context and joins it to warehouse lineage, quality, usage and owners |
| The questions | Answers traversals, runs graph algorithms and serves GraphRAG through Aura Agent | Makes sure the data and definitions behind those answers are correct and current |
| The feed | Imports from the warehouse with Aura Import or your load job | Watches the tables and keys that feed the graph and catches breaks before the load |
| The break | Loads what it is given | Detects the upstream change, traces it across SAP, Fivetran, dbt and Airflow to the graph |
| The fix | Runs the Cypher your team approves | Proposes the change with its blast radius, routes it to a named owner, lands it in dbt and the load pipeline |
| The proof | Holds the graph state after the load | Verifies node counts and keys after the reload and writes a receipt with the cause, diff, approver and rollback |
| The meaning | Holds labels and relationship types | Keeps business definitions consistent across the graph, the semantic layer and the catalog, with conflicts routed to an owner |
Why doesn't Neo4j just do this itself?
Because Neo4j built a database for the hardest part of its job: storing and traversing connected data at scale, for any domain. The systems that feed the graph belong to other vendors and your data team, and changing SAP, Fivetran, dbt and Airflow is a different product with a different liability.
Neo4j's own guidance draws that line sensibly. Its MCP documentation says of write-cypher: "LLM-generated queries can cause harm. Use only in development environments." Its security page advises using "a restricted Neo4j user for exploring the database" and reviewing generated Cypher "before executing them against a database, especially when it is a production database." NEO4J_READ_ONLY=true removes the write tool entirely. That is the right design for a database that serves many teams and many agents.
Data Workers is the product on the other side of that line. It knows what a change will touch across systems, routes it to the owner, applies it reversibly through the pipelines you already run, verifies the result and keeps the record.
Every tool owns a slice. Data Workers covers the whole lifecycle
Neo4j owns one slice of the data lifecycle, and owns it outright: modeling and traversing relationships, and the graph analytics and GraphRAG built on them. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the graph you already run.

| Stage | Data Workers | Neo4j | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 9.5 | Neo4j's home stage: customers, products, suppliers and documents as nodes and relationships, served to agents through the Neo4j MCP server and GraphRAG. Data Workers brings that model in as context with provenance, next to lineage, quality and usage. |
| Analytics & Insights | 8 | 9 | Neo4j's other home stage: Cypher traversals, Graph Data Science algorithms and Aura Graph Analytics sessions find patterns tables hide. Data Workers' Insights agent answers through governed metric definitions. |
| Data Quality | 8 | 4 | Constraints and indexes keep node keys and properties in shape inside the graph. Data Workers checks the warehouse tables that feed it and repairs the breaks before a load. |
| Observability & Incidents | 8.5 | 3 | Aura reports on the health of the database itself. Data Workers detects a bad feed, traces it across systems, fixes it and verifies the graph after the reload. |
| Pipelines & Ingestion | 8.5 | 4 | Aura Import loads from Snowflake, Databricks, BigQuery, Redshift, Postgres and more with no code. Data Workers builds, reruns and backfills the pipelines behind those loads with approvals. |
| Schema & Migration | 8 | 3 | The graph is schema-optional by design, with constraints where you want them. Data Workers detects upstream schema changes and assesses their impact on every model and load before they land. |
| Governance & Access | 8.5 | 6 | Strong over its own graph: database users and roles, and project admin, member and viewer roles in Aura. Data Workers proposes and applies grants across your data platforms by policy. |
| Security & Privacy | 8 | 5 | Strong for its own estate: tool authentication with Aura users, and Neo4j's guidance to give agents a restricted user. Data Workers flags sensitive column names in pull request review and proposes masking for the owner. |
| Cost / FinOps | 8 | 2 | Aura bills its own instances and sessions. Data Workers traces credits in the warehouse that feeds the graph to the dbt model behind them and drafts the fix for its owner. |
| MLOps & Models | 7.5 | 5 | Graph Data Science adds node embeddings and graph machine learning. Data Workers keeps the data under your models healthy and connects to MLflow and W&B. |
How Neo4j and Data Workers work together
Your engineers and agents stay where they are: Claude Code, Cursor or VS Code for Cypher and dbt, Aura Agent for GraphRAG. Spellbook Data Catalog (in preview) is where the data team looks: each proposed change, its blast radius, its approver and its rollback. Between them run four layers: Data Context Wizard keeps one governed context graph, the Data-Agents Swarm does the work with more than 20 specialist agents, the Autonomous Data-Conductor runs each fix end to end (detect, diagnose, fix, review, verify, remember), and per-domain guardrails hold approvals, receipts and rollback.

What Context Wizard does with your graph. It reads the graph's schema (labels, relationship types, property keys) and the business keys each load merges on, and records them as context with provenance: where each fact came from, when it was observed and who owns it. It joins that to warehouse lineage, so Supplier.supplier_id in the graph traces back through dim_supplier and Fivetran to SAP. It scores trust with quality, freshness, usage and ownership, and serves one governed view to every agent: explain_table returns definition, lineage and trust score, trace_cross_platform_lineage walks the hops, and get_quality_score on dw-quality returns the quality score. When the graph's idea of "active customer" disagrees with your semantic layer's, the conflict goes to a named owner; promoting a fact to authoritative takes a named person.
Setup over MCP today. Data Workers connects to Neo4j over the Neo4j MCP server today, alongside its own agents, in the same client. Data Workers' agents are MCP servers from the open-source repository: clone it and add start-agent.sh entries to your client config, as the client setup docs show. Run the Neo4j MCP server with a restricted Neo4j user and NEO4J_READ_ONLY=true, so agents can read the graph and nothing can write to it from the chat. For Aura, enable Tool authentication in the console and use the instance's hosted MCP URL instead of the local binary.
// Example: .mcp.json for Claude Code (Cursor uses the same mcpServers shape)
{
"mcpServers": {
"neo4j": {
"command": "neo4j-mcp",
"env": {
"NEO4J_URI": "neo4j+s://<your-instance>.databases.neo4j.io",
"NEO4J_USERNAME": "graph_reader",
"NEO4J_PASSWORD": "<from your secret store>",
"NEO4J_DATABASE": "neo4j",
"NEO4J_READ_ONLY": "true"
}
},
"dw-context-catalog": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-context-catalog"]
},
"dw-schema": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-schema"]
},
"dw-quality": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-quality"]
}
}
}List the tools in your client with its own command (for example /mcp in Claude Code). With this in place, an engineer can ask "what feeds the Supplier label, and is it healthy tonight?" and get the graph schema from Neo4j, the lineage from trace_cross_platform_lineage, and the latest run_quality_check results on dim_supplier, in one answer.
Where writes go. Fixes land where your team already reviews change: a dbt diff for the owner to merge, a change to the load DAG or the Aura Import mapping. Your load pipeline runs the Cypher, as it does today, after the owner approves. That keeps graph writes on the path Neo4j recommends: reviewed queries, run by a known job with a known user.
One request, L0 to L4. The autonomy ladder is set per domain.

- •L0 manual. Data Workers is connected but not acting. Your engineer traces the feed by hand.
- •L1 observe. Data Workers watches the tables and keys that feed the graph, flags breaks before the load and explains the cause with lineage and owners. Nothing changes.
- •L2 propose. Data Workers drafts the dbt diff or the load-job change with its blast radius. The graph owner approves in Spellbook before anything runs.
- •L3 act reversibly. For change classes with a proven record, such as rerunning a failed load for affected partitions, Data Workers applies the change, verifies node counts and keys, and can roll it back.
- •L4 autonomous. For a scoped, trusted class like late feeds into one graph, Data Workers fixes and verifies on its own and posts the receipt for review.
For the safety model behind each step, read is it safe to let AI agents change production data, and for where data and credentials live, read where does our data go.
The same pattern holds for every source of meaning. See the hub, bring your own context, and the guides for teams on Palantir Ontology, Fabric IQ and Glean. If you are weighing a graph against a semantic layer for grounding, our explainer on semantic layer vs knowledge graph for LLM grounding covers the tradeoffs, and what is a context graph covers the idea behind Context Wizard.
What changes for your team

Graph teams spend much of the week on work that isn't graph work: chasing a jumped node count, reconciling an upstream key, rebuilding after a bad load. With Data Workers next to Neo4j, those jobs run on autopilot at the level you set.
- •Incidents. An upstream key or schema change that would duplicate or orphan nodes is caught, traced and fixed before the nightly load, not discovered from a wrong GraphRAG answer.
- •Data quality. Every business key the graph merges on gets uniqueness, null and match-rate checks in the warehouse, so the graph's constraints never have to catch a break that started upstream.
- •Cloud spend. Cleanups of staging tables and extracts built for retired graph loads are proposed to the owner after a dependency check.
- •Access. A request to read a sensitive part of the graph, such as customer or HR relationships, arrives as a time-boxed grant proposal for its owner, with the policy that justifies it.
- •Audits. Every change to a feed, a load or a definition carries an approver, a diff, the verification and a rollback path, so "where did this relationship come from?" has an answer.
- •Migrations. When the warehouse under the graph moves, the feeds move in parity-checked waves while the graph keeps loading.
The graph team gets its week back for modeling, new domains and the GraphRAG work only it can do.
Keep Neo4j, or consolidate?
Keep Neo4j if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
For most teams on Neo4j, the answer is to keep it: the graph model, the Cypher, the Graph Data Science work and the GraphRAG agents belong there. What teams consolidate is the tooling around the graph: a separate data-quality tool for the feeds, a homegrown script that compares node counts, a metadata spreadsheet that maps graph labels to warehouse tables. Data Workers runs those jobs with one context, one approval flow and one audit trail. Data Workers can also run its own context graph on Neo4j, so one graph technology serves the whole stack. If you are weighing building this layer yourself on top of the Neo4j MCP server, read build it ourselves with Claude Code and MCP servers: the connection is the easy part; the cross-system context, approvals and rollback are where the work is.
The case for your CFO
The outcome: the knowledge graph the company paid to build gives answers people can act on. Supplier risk, fraud rings and GraphRAG answers all rest on the data that feeds the graph, and Data Workers keeps it right every night.
The risk story is plain. Your team's assistant reads the graph through the Neo4j MCP server in read-only mode, with a restricted user; Data Workers' agents never call it. Every change Data Workers proposes shows its blast radius, goes to a named approver, lands through the pipelines your team already runs, is verified after the reload and leaves a receipt: who approved it, what it touched and how to undo it. Autonomy is set per domain from L0 manual to L4 autonomous and can be dialled back at any time. There is zero migration: Neo4j, the warehouse, dbt and your load jobs stay where they are.
Why now: agents read the graph directly, so a bad load reaches every agent and every person who asks. The first win is one graph with its feeds watched, read-only, so the next upstream change is caught before the load instead of after the answer. What stays the same: your graph model, your Cypher, your Aura plan, your warehouse permissions and your review process. For the numbers, see the ROI of agentic data operations.
The sentence to repeat upstairs: "Our graph is only as good as what feeds it; Data Workers keeps the feed and the meaning right, and fixes it with an approval and a receipt."
Getting started
Start with a pilot. Pick one graph that agents or analysts already rely on, such as the supplier or customer graph, connect the Neo4j MCP server read-only next to Data Workers, and let Data Workers watch its feeds for a few weeks before enabling the first fix class. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.
FAQ
Does Data Workers write to our Neo4j database? Your team's assistant reads the graph through the Neo4j MCP server with NEO4J_READ_ONLY=true and a restricted user. Approved fixes land in dbt, the load DAG or the import job, and your pipeline runs the Cypher as it does today, after the graph owner approves in Spellbook.
Can Data Workers run its own context graph on Neo4j? Yes. Data Workers' context graph runs on Neo4j when NEO4J_URI is set, with every query filtered by tenant, so one graph technology serves your graph and Data Workers' context.
We load with Aura Import. Does that change anything? No. Data Workers watches the Snowflake, Databricks or BigQuery tables the import reads from, catches schema and key changes before the import runs, and proposes the change to the source model or the import mapping for the owner to approve.
How does lineage reach from the graph back to the source? Context Wizard records which warehouse tables and columns feed each label and relationship type, and joins that to lineage across dbt, ingestion and the source systems. trace_cross_platform_lineage follows a graph property back to the ERP field it came from.
Does this replace Graph Data Science or Aura Agent? No. They stay where your graph analytics and GraphRAG run; Data Workers keeps the data and definitions under them correct and current.
What happens when the graph and the semantic layer disagree on a definition? Context Wizard routes the conflict to a named owner. Agents see both candidates and their sources until the owner decides, and the decision is recorded.
Sources
- •Neo4j, Neo4j MCP overview, https://neo4j.com/docs/mcp/current/ (checked Oct 2, 2026)
- •Neo4j, Neo4j MCP tools and read-only mode, https://neo4j.com/docs/mcp/current/tools/ (checked Oct 2, 2026)
- •Neo4j, Neo4j MCP configuration reference, https://neo4j.com/docs/mcp/current/configuration/ (checked Oct 2, 2026)
- •Neo4j, Neo4j MCP client configuration, https://neo4j.com/docs/mcp/current/client-configuration/ (checked Oct 2, 2026)
- •Neo4j, Neo4j MCP security, https://neo4j.com/docs/mcp/current/security/ (checked Oct 2, 2026)
- •Neo4j, MCP for Aura, https://neo4j.com/docs/mcp/current/mcp-for-aura/ (checked Oct 2, 2026)
- •Neo4j, Neo4j MCP releases (v1.0.0 Nov 24, 2025; v1.6.0 Sept 10, 2026), https://github.com/neo4j/mcp/releases (checked Oct 2, 2026)
- •PyPI, neo4j-graphrag 1.22.0 (Oct 1, 2026), https://pypi.org/project/neo4j-graphrag/ (checked Oct 2, 2026)
- •Neo4j, Aura user management (project roles), https://neo4j.com/docs/aura/user-management/ (checked Oct 2, 2026)
- •Neo4j, Aura Agent, https://neo4j.com/docs/aura/aura-agent/ (checked Oct 2, 2026)
- •Neo4j, GraphRAG for Python (neo4j-graphrag 1.22.0, Oct 1, 2026), https://neo4j.com/docs/neo4j-graphrag-python/current/ (checked Oct 2, 2026)
- •Neo4j, Graph Data Science manual v2026.09, https://neo4j.com/docs/graph-data-science/current/ (checked Oct 2, 2026)
- •Neo4j, Graph analytics in Aura, https://neo4j.com/docs/aura/graph-analytics/ (checked Oct 2, 2026)
- •Neo4j, Aura Import, https://neo4j.com/docs/aura/import/introduction/ (checked Oct 2, 2026)
- •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)
- •Data Workers open-source repository, agent tools and start-agent.sh, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)