You're on Memgraph: Keep the Live Graph True, From the Kafka Event to the GraphRAG Answer
Already on Memgraph? Data Workers keeps the streams and loads behind your Memgraph graph true, next to the Memgraph MCP server in your team's client, with approvals and receipts.
Your graph has to be right within seconds of each event. Accounts, cards, devices, payments and merchants are nodes in Memgraph; who paid whom, which device touched which account and which product followed which view are relationships. Kafka, Redpanda or Pulsar topics arrive through Memgraph streams and their transformation modules, reference data comes in from Postgres or the warehouse through a batch job or the cross_database module, and Cypher answers in milliseconds: is this payment part of a ring, what should this shopper see next, which documents ground this GraphRAG answer. MAGE runs the algorithms, GraphChat in Memgraph Lab lets people explore, and the Memgraph MCP server (mcp-memgraph 0.4.1, September 14, 2026) gives Claude Code, Cursor or any MCP client schema search and Cypher. Memgraph is the live graph. Data Workers is the control plane that keeps it true: Data Context Wizard takes the graph's model in as context with provenance, ties it to the producers, topics and tables upstream, and when an event contract or a feed breaks, Data Workers carries the fix through an owner's approval and leaves a receipt.
A producer renames a field or a reference table loses a key, and the stream keeps committing whatever the transformation makes of it. That seam between producers, warehouse and graph is the job this guide covers.
Key takeaways
- •Memgraph keeps its job. The in-memory engine, Cypher, streams, MAGE, vector search, GraphRAG and Memgraph Lab stay as they are. Data Workers works next to them from day one.
- •The graph's model becomes governed context. Context Wizard reads labels, relationship types, enums and the keys each stream matches on, and places them beside event contracts, warehouse lineage, quality, usage and a named owner.
- •Feeds are watched at stream speed. A producer field change, a drop in key matches or a stall in relationship creation is caught while the stream runs and traced to its source.
- •Cypher writes stay gated. Your team's assistant reads through the Memgraph MCP server in its default read-only mode, side by side with Data Workers. Approved fixes land in your repositories and pipelines with a blast radius, an approver and a rollback path.
- •Start with a pilot. One stream-fed graph, watched read-only; then one fix class in one domain, on the ladder from L0 manual to L4 autonomous.
Memgraph is the live graph. Data Workers is the control plane that keeps it true.
Memgraph calls itself a "high-performance, in-memory graph database that powers real-time AI context," and the engine backs that up. Kafka streams commit "at least once": the queries a transformation returns are "committed to the database for every batch of messages, and only then is the message offset committed." Memgraph 3.13 (September 2026) added global vertex-property indexes and made per-database OpenMetrics the default metrics format; 3.12 (July) added suspend and resume for Enterprise tenant databases; 3.11 (June) added the cross_database module, which pulls rows from Postgres, MySQL, Oracle, S3, DuckDB and more, and is an Enterprise feature since 3.13. The MCP server's default build exposes seven tools: run_cypher_query, search_schema, get_node_schema, get_relationship_schema, get_enum_schema, list_databases and use_database. Memgraph also ships Memgraph Zero, whose first component, MemGQL (0.12.0, September 14, 2026), is a federated GQL engine that queries Postgres, ClickHouse, Iceberg, Snowflake and other stores in place.
Around the graph sit the producers, their schemas, the warehouse tables, the dbt models and DAGs, and the repository of transformation modules. Data Workers runs that side. Here is one afternoon beside a fraud graph, as an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| 14:02 | Payments service | A deploy renames merchant_id to merchant_ref in the payment event |
| 14:02 | Kafka | The payments topic starts carrying the new shape |
| 14:03 | Memgraph | The stream keeps committing. The transformation module reads merchant_id with a null default, matches no merchant, and payments land without a PAID_TO relationship |
| 14:11 | Data Workers | Flags the field change against the registered contract for the payments topic |
| 14:13 | Memgraph | Reads PAID_TO over the Memgraph MCP server (get_relationship_schema) and runs a read-only count: new PAID_TO relationships have fallen to zero since 14:02 |
| 14:20 | GitHub | Opens an incident naming the stream, the unlinked payments and the fraud agent downstream, plus a diff to the module that accepts both field names. Blast radius: one module, one stream, one graph, one backfill |
| 14:50 | Spellbook | The fraud platform owner reviews diff and impact and approves; CI passes and the team's deploy job reloads the module |
| 15:05 | Airflow | A backfill DAG links 9,400 payments since 14:02 to their merchants from the Snowflake copy of the topic |
| 15:20 | Memgraph | A read-only check finds zero payments since 14:02 without PAID_TO and no duplicate merchants. Data Workers writes the receipt |
| 15:40 | Fraud agent | An analyst asks whether a cluster of merchants forms a ring; the GraphRAG answer is right |

Nothing in the graph looked broken that afternoon. It simply had fewer edges, and every ring detection that runs through merchants went quiet. Memgraph did exactly what it was asked. The cause sat upstream, in a producer contract and a module the team owns, and the fix went through that team's approval.
| Job | What Memgraph does | What Data Workers does |
|---|---|---|
| The model | Holds entities and relationships in memory, with constraints, enums and indexes where you want them | Reads the model as context and joins it to event contracts, warehouse lineage, quality, usage and owners |
| The questions | Runs real-time traversals, MAGE algorithms, vector search and GraphRAG | Keeps the events and definitions behind those answers correct and current |
| The feed | Consumes Kafka, Redpanda and Pulsar through streams; pulls rows with cross_database | Watches the topics, schemas and tables that feed the graph and catches breaks as they happen |
| The break | Commits what the transformation returns | Traces the change from the producer through the topic to the relationships it stopped creating |
| The fix | Runs the modules and Cypher your team deploys | Proposes the change with its blast radius, routes it to an owner, lands it in the repository and the backfill DAG |
| The proof | Holds the graph state after the backfill | Verifies counts and keys with read-only queries and records cause, diff, approver and rollback |
| The meaning | Holds labels, relationship types and enums | Keeps definitions consistent across graph, event contracts, semantic layer and catalog |
Why doesn't Memgraph just do this itself?
Memgraph put its engineering into its own hardest problem: a large, fast-changing graph held in memory, answering deep traversals as events land. The payments service, topic schemas, warehouse and DAGs belong to other teams and vendors, and changing them safely is a separate product with separate liability.
Memgraph's defaults draw that line well. Its MCP server "runs in read-only mode to prevent accidental data modifications," blocking CREATE, MERGE, DELETE, SET, DROP and REMOVE unless an operator sets MCP_READ_ONLY=false. With OIDC turned on, each session reaches only the tenant databases its token lists, and use_database "cannot expand authorization beyond what the JWT grants." Streams favor never losing a message, so the graph keeps pace with whatever the topic carries. Those are the right choices for a database serving many teams, tenants and agents at once. MemGQL follows the same idea for reads: query data where it lives.
Cross-system change is the job Data Workers was built for: it knows what a fix touches, routes it to the owner, applies it reversibly through your own repositories and pipelines, verifies the graph and keeps the record.
Every tool owns a slice. Data Workers covers the whole lifecycle
Memgraph owns real-time graph traversal and analytics, and the AI context built on them, outright. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the graph you already run.

| Stage | Data Workers | Memgraph | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 9.5 | Memgraph's home stage: accounts, devices, payments and merchants as a live in-memory graph, served to agents through the Memgraph MCP server with schema search and GraphRAG. Data Workers brings that model in as context with provenance, next to lineage, quality and usage. |
| Analytics & Insights | 8 | 9.5 | Memgraph's other home stage: real-time multi-hop Cypher traversals, MAGE algorithms and vector search find fraud rings and recommendations as events arrive. Data Workers' Insights agent answers through governed metric definitions. |
| Data Quality | 8 | 3 | Existence, uniqueness and data type constraints keep nodes in shape inside the graph. Data Workers checks the events and warehouse tables that feed it and repairs breaks at the source. |
| Observability & Incidents | 8.5 | 3.5 | Memgraph exports OpenMetrics and Prometheus metrics per database for the engine's own health. Data Workers detects a bad feed, traces it across systems, fixes it and verifies the graph afterwards. |
| Pipelines & Ingestion | 8.5 | 5 | Streams consume Kafka, Redpanda and Pulsar through transformation modules, and the Enterprise cross_database module pulls rows from Postgres, MySQL, Oracle, S3 and DuckDB. Data Workers builds, reruns and backfills the pipelines behind those feeds with approvals. |
| Schema & Migration | 8 | 3 | The graph is schema-flexible, with constraints and enums where you want them. Data Workers catches producer schema changes in Schema Registry and warehouse schema changes in the dbt manifest and in review, and assesses their impact on every stream and load. |
| Governance & Access | 8.5 | 6.5 | Strong over its own graph: role-based and label-based access control and multi-tenancy in Enterprise, plus JWT tenant routing in the MCP server. Data Workers proposes and applies grants across your data platforms by policy. |
| Security & Privacy | 8 | 5 | Strong for its own estate: user authentication, OIDC on the MCP server and read-only mode by default. Data Workers flags sensitive column names in pull request review and proposes masking for the owner. |
| Cost / FinOps | 8 | 2 | Memgraph lets Enterprise tenants suspend databases to free memory. Data Workers traces Snowflake credits to the dbt model behind them and drafts the fix for its owner. |
| MLOps & Models | 7.5 | 4.5 | MAGE adds graph algorithms and embeddings for machine learning on the graph. Data Workers keeps the data under your models healthy and connects to MLflow and W&B. |
How Memgraph and Data Workers work together
Engineers and agents keep their tools: Claude Code or Cursor for Cypher, modules and dbt, GraphChat for exploring, your fraud or recommendation agents for GraphRAG. Spellbook Data Catalog (in preview) is where the data team looks: each proposed change, its blast radius, approver and rollback. Underneath, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm's 20+ specialist agents do the work, the Autonomous Data-Conductor runs each fix end to end (detect, diagnose, fix, review, verify, remember), and per-domain guardrails hold approvals, receipts and rollback.

From stream to context. Context Wizard calls search_schema, get_node_schema and get_relationship_schema to learn labels, relationship types, properties, indexes, constraints and enums, then adds what Memgraph cannot see from inside: which topic and module create each relationship, which warehouse table supplies each reference label, and who owns the producer. Every fact carries its source, when it was observed and its owner. So Merchant.merchant_id traces back through the transformation module and the payments topic to the payments service, and through dim_merchant to Snowflake. Agents get that view through explain_table and trace_cross_platform_lineage, and a fact becomes authoritative only when a named person promotes it.
Wiring it up. Point your MCP client at both servers. Data Workers' agents run from the open-source repository: clone it and add start-agent.sh entries, as the client setup docs describe. Give the Memgraph server a restricted database user and keep MCP_READ_ONLY at its default of true. For a shared deployment, switch to the streamable HTTP transport with MCP_AUTH_ENABLED=true and your identity provider as issuer, so each person reaches only the tenant databases their token lists.
// Example: .mcp.json for Claude Code (Cursor uses the same mcpServers shape)
{
"mcpServers": {
"memgraph": {
"command": "uv",
"args": ["run", "--with", "mcp-memgraph", "--python", "3.13", "mcp-memgraph"],
"env": {
"MCP_TRANSPORT": "stdio",
"MEMGRAPH_URL": "bolt://<your-memgraph-host>:7687",
"MEMGRAPH_USER": "graph_reader",
"MEMGRAPH_PASSWORD": "<from your secret store>",
"MEMGRAPH_DATABASE": "memgraph",
"MCP_READ_ONLY": "true"
}
},
"dw-context-catalog": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-context-catalog"]
},
"dw-schema": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-schema"]
},
"dw-quality": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-quality"]
}
}
}Check the tool list with your client's own command (/mcp in Claude Code). An engineer can then ask "what feeds PAID_TO, and is it healthy right now?" and get one answer that combines the relationship schema from Memgraph, lineage from trace_cross_platform_lineage, and the latest run_quality_check results on the tables behind it. Topic contracts go into Confluent Schema Registry through register_kafka_schema on the connectors agent, so a contract change is a reviewed change.
The write path. Approved fixes go wherever your team already reviews change: a diff to the transformation module, a dbt model, a backfill DAG or a compatibility change on the topic's schema. Your deploy job reloads the module and your DAG runs the Cypher with its own service user. The MCP server stays a reader, which is the split Memgraph's defaults suggest.
Autonomy for one stream, set per domain.

- •L0 manual. Data Workers is connected and idle; an engineer traces the stream by hand.
- •L1 observe. It watches topics, schemas, tables and match keys, flags breaks as they happen and explains cause, lineage and owner. Nothing changes.
- •L2 propose. It drafts the module diff or the backfill with its blast radius; the graph owner approves in Spellbook before anything runs.
- •L3 act reversibly. For change classes with a proven record, such as rerunning a backfill for an affected window, it applies the change, verifies counts and keys, with the undo recorded before it runs.
- •L4 autonomous. For a scoped, trusted class like late reference data into one graph, it fixes, verifies and posts the receipt for review.
The safety model behind each rung is in is it safe to let AI agents change production data; where data and credentials live is in where does our data go.
The hub, bring your own context, covers the pattern for every graph and source of meaning, with guides for teams on Neo4j, PuppyGraph and TigerGraph. For background, see semantic layer vs knowledge graph for LLM grounding, what is a context graph and the streaming agent for Kafka and Flink.
What changes for your team

Chasing a drop in edges, decoding a producer's change and replaying a bad batch eat a graph team's week. With Data Workers beside Memgraph, those jobs run on autopilot at the level you set.
- •Incidents. A renamed event field that stops linking payments to merchants is caught the same afternoon and fixed at the source, before a quiet fraud dashboard gives it away.
- •Data quality. Every key a stream or load matches on gets null, uniqueness and match-rate checks upstream of the graph's own constraints.
- •Cloud spend. Cleanups of topics, extracts and staging tables left from retired graph feeds are proposed to the owner after a dependency check.
- •Access. A request to query a sensitive tenant database, such as one region's customers, arrives as a time-boxed grant proposal for its owner, with the policy behind it.
- •Audits. Every change to a stream, module, load or definition carries an approver, a diff, the verification and a rollback path.
- •Migrations. When the warehouse or broker under the graph moves, feeds move in parity-checked waves while the graph keeps streaming.
Keep Memgraph, or consolidate?
Keep Memgraph if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
On Memgraph, consolidation usually means the scaffolding around the engine: a lag-and-edge-count monitor someone wrote, a script that diffs event schemas, a wiki page mapping labels to topics. The in-memory graph, streams, MAGE, vector search and GraphRAG agents stay put. If you are tempted to build the surrounding layer yourself on the Memgraph MCP server, read build it ourselves with Claude Code and MCP servers: connecting is quick; cross-system context, approvals and rollback are the real build.
The case for your CFO
The outcome: the fraud, recommendation and AI decisions that run on the real-time graph rest on events that are correct when they arrive. Data Workers keeps those feeds right and fixes breaks before they cost money.
The risk story. Your team's assistant reads Memgraph through its MCP server in default read-only mode under a restricted user, and in shared deployments each session sees only the tenants its token allows. Each proposed change shows its blast radius, goes to a named approver, ships through your own repositories and pipelines, is verified in the graph and leaves a receipt: who approved, what changed, how to undo it. Autonomy is set per domain from L0 manual to L4 autonomous and can be turned down at any time. Zero migration: Memgraph, Kafka, the warehouse, dbt and your DAGs stay where they are.
Why now: scoring services and agents act on the graph seconds after an event, so a bad feed reaches every decision before anyone looks. The first win is one stream-fed graph watched read-only, so the next producer change is caught the same day. What stays the same: your graph model, Cypher, modules, Memgraph plan, Kafka setup and review process. The numbers are in the ROI of agentic data operations.
The line for upstairs: "Our real-time graph is only as good as what streams into it; Data Workers keeps those feeds and their meaning right, with an approval and a receipt for every fix."
Getting started
Start with a pilot. Choose the graph your agents or analysts lean on most, often fraud or recommendations, connect the Memgraph MCP server read-only beside Data Workers, and let it watch that graph's streams and loads for a few weeks before you switch on the first fix class. Plans and the pilot path are on the pricing page, and the pilot is credited in full against the first year.
FAQ
Can Data Workers' agents change data in Memgraph? They read through the Memgraph MCP server, read-only by default, under a restricted user. Approved fixes ship through the module repository, dbt or a backfill DAG, and your own jobs run the Cypher after the graph owner approves in Spellbook.
Streams commit at least once. Does that complicate backfills? Memgraph commits each batch before the offset, so replays can repeat messages. Data Workers designs backfills to match on stable keys, so a replay leaves no duplicate nodes, and it watches both the topic against its contract and the graph's relationship creation rate.
We run multi-tenant Memgraph with OIDC. Does Data Workers respect that? Yes. With MCP_AUTH_ENABLED=true and your identity provider as issuer, each session reaches only the databases its JWT allows, and list_databases and use_database stay inside that grant. Data Workers routes its own access proposals to each tenant's owner.
Can a graph property be traced back to the service that produced it? Yes. Context Wizard records which topic, module and warehouse table feed each label and relationship type, and trace_cross_platform_lineage follows a property through the module and topic to the producing field and service.
We are trying MemGQL. Where does Data Workers fit? MemGQL reads across stores in place with one GQL query. Data Workers keeps those stores and the feeds between them correct, and carries fixes with approvals, so federated answers rest on sound data.
Does this replace MAGE, GraphChat or our GraphRAG agents? No. They keep running your graph analytics and AI answers; Data Workers keeps the data and definitions under them correct.
Sources
- •Memgraph, Memgraph documentation overview, https://memgraph.com/docs (checked Oct 2, 2026)
- •Memgraph, Model Context Protocol (MCP), https://memgraph.com/docs/ai-ecosystem/mcp (checked Oct 2, 2026)
- •Memgraph, Memgraph MCP Server README (tools, read-only mode, multi-tenant authentication), https://github.com/memgraph/ai-toolkit/tree/main/integrations/mcp-memgraph (checked Oct 2, 2026)
- •PyPI, mcp-memgraph 0.4.1 (Sept 14, 2026), https://pypi.org/project/mcp-memgraph/ (checked Oct 2, 2026)
- •Memgraph, Release notes (Memgraph 3.13.1 Sept 14, 2026; 3.13.0 Sept 9, 2026; 3.12.0 July 15, 2026; 3.11.0 June 17, 2026; Lab 3.13.2 Sept 18, 2026), https://memgraph.com/docs/release-notes (checked Oct 2, 2026)
- •Memgraph, Memgraph Zero and MemGQL (changelog: MemGQL 0.12.0 Sept 14, 2026), https://memgraph.com/docs/memgraph-zero and https://memgraph.com/docs/memgraph-zero/memgql/changelog (checked Oct 2, 2026)
- •Memgraph, Data streams, https://memgraph.com/docs/data-streams (checked Oct 2, 2026)
- •Memgraph, AI ecosystem (GraphRAG, agents, integrations, GraphChat), https://memgraph.com/docs/ai-ecosystem (checked Oct 2, 2026)
- •Memgraph, Constraints, https://memgraph.com/docs/fundamentals/constraints (checked Oct 2, 2026)
- •Memgraph, Monitoring, https://memgraph.com/docs/database-management/monitoring (checked Oct 2, 2026)
- •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)
- •Data Workers open-source repository, agent tools and start-agent.sh, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)