You're on Cognee: Keep What Your Agents Recall About Data True Against the Warehouse
Already on Cognee? Data Workers checks what your agents recall about tables, schemas and metrics against the live data, and feeds approved facts back to Cognee.
Your agents remember now. Your team runs Cognee, the open-source AI memory platform, self-hosted from the Apache 2.0 repository or in Cognee Cloud. You call remember on documents, Slack threads, code and database rows, and Cognee builds a knowledge graph with embeddings on top. Agents call recall through the Cognee MCP server in Claude Code, Cursor or Codex, the Claude Code plugin, the REST API or the Python SDK. improve folds session lessons into the permanent graph, forget removes an item or a dataset, and the Company Brain in Cognee Cloud reads Slack, GitHub and Google Drive into shared memory. Cognee holds what your agents remember. Data Workers, the agentic data platform, keeps what they remember about your data true: Data Context Wizard holds the governed facts about the data estate (schemas, metric definitions, lineage, owners, freshness), checks recalled facts against the live data, and feeds the approved ones back to Cognee through your team's own client.
Memory is faithful to what it was given. It cannot know that the app team split a column this morning or that finance changed a metric in dbt. An agent that recalls the old fact gives an answer that sounds right and is wrong. That seam is the job this guide covers.
Key takeaways
- •Cognee keeps its job.
remember,recall,improve,forget, your datasets, ontologies and Company Brain stay as they are. Data Workers works next to them from day one. - •Facts about data get one governed source. Schemas, metric definitions, lineage, owners and freshness live in Data Context Wizard with provenance and a named owner.
- •Recalled facts are checked against the live data. Your team's assistant recalls what your memory says about a table or a metric over the Cognee MCP server, and Data Workers flags anything the source systems contradict.
- •Approved facts flow back into Cognee. After a named owner approves, your team's load job or Cognee client writes the update. Your team's assistant runs the next recall, and Data Workers checks it and writes a receipt.
- •Start with a pilot. One memory dataset, read-only, checked against one domain; then one update class, on the ladder from L0 manual to L4 autonomous.
Cognee is your agents' memory. Data Workers is the agentic data platform that keeps that memory true to the data.
Cognee is built for one hard problem: turning messy knowledge into memory an agent can reason over. Its docs call it "a memory layer for AI agents and tools", "not an LLM" and "more than plain RAG", because recall can follow connections in the graph instead of only matching text. With an OWL ontology, Cognee grounds entities to your canonical terms and marks them ontology_valid=True. Through the dlt integration it loads database tables and CSVs deterministically, with no LLM in the loop: tables become schema nodes, rows become documents and foreign keys become edges. Release v1.6.2 (Sept 29, 2026) added Slack conversation import and sync to the open-source project.
That makes Cognee a natural home for what agents know about your own data: the app database schema, the Slack thread on how to count Pro customers. Those facts hold only while the source systems agree. Data Workers covers that side.
Here is a morning with Data Workers next to a product-data dataset in Cognee. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| 08:10 | GitHub | The app team merges a Postgres migration that splits accounts.plan into plan_tier and billing_cycle; new signups no longer fill plan |
| 08:30 | Postgres | The migration runs in production |
| 08:36 | Data Workers | run_quality_check on Postgres profiles accounts and shows the split; impact analysis shows the dbt model dim_accounts reads plan |
| 08:44 | Snowflake | Data Workers checks the replicated table: plan is empty on every account created since 08:30 |
| 08:52 | Cognee | The agent owner's assistant recalls what product-data and the synced Slack history say about plans, over the Cognee MCP server, and hands it to Data Workers |
| 08:55 | Cognee | Recall returns the accounts.plan schema node from last night's dlt load and a Slack thread: "count Pro customers with plan = 'pro'" |
| 09:05 | Spellbook | Data Workers routes one item to the dim_accounts owner: the schema diff, a proposed dbt change and proposed memory text |
| 09:40 | Spellbook | The owner approves the dbt change and the memory update |
| 09:50 | Cognee | The team's load job re-ingests the accounts source through dlt; the team's Cognee client remembers the approved note with its source and owner |
| 10:05 | Cognee | The assistant recalls again, both answers name plan_tier, and Data Workers writes the receipt |
| 14:20 | Cursor | A product manager's agent counts this week's Pro signups with plan_tier; the number matches the dashboard |

Without that check, the product manager's agent would have filtered on plan and reported that Pro signups stopped at 08:30. Cognee did what it was asked: it remembered the schema it loaded and what people said. The fix needed the Postgres diff, the warehouse, dbt and an owner's approval, and it landed in Cognee through the team's own pipeline.
| Job | What Cognee does | What Data Workers does |
|---|---|---|
| The memory | Turns documents, conversations, code and rows into a knowledge graph with embeddings | Holds the governed facts about the data estate with provenance and a named owner |
| The recall | Routes each question to the right retrieval strategy across graph, vectors and sessions | Checks recalled facts about tables, schemas and metrics against the source systems |
| The vocabulary | Grounds entities to your OWL ontology, or infers types from the text | Keeps business definitions consistent across dbt, the catalog, BI and memory, with conflicts routed to an owner |
| The drift | Keeps what it was told until it is told otherwise | Detects the schema change, renamed table or changed metric, and finds the memories it makes wrong |
| The fix | Remembers, re-ingests and forgets what your client sends | Proposes the update with sources and a diff, routes it to a named owner, hands the approved fact to your pipeline |
| The proof | Serves the updated memory on the next recall | Checks the recall after the update and writes a receipt with the cause, diff, approver and rollback |
Why doesn't Cognee just do this itself?
Because Cognee built a memory engine for any domain, and a memory engine that second-guessed what it was told would be a worse one. Deciding whether a remembered column still means what it meant yesterday needs the source database, the warehouse, the dbt project and an owner's sign-off, which belong to other vendors and to your data team.
Cognee's design draws sensible lines. Ontology grounding standardizes types and relationships, and its docs note that it does not copy data properties or run an OWL reasoner. A dlt load reflects the source at load time, and picking up upstream changes is an explicit re-ingest your team controls. In API mode, each MCP instance authenticates with one token, so everything it writes belongs to one Cognee user. That is the right shape for a memory layer many agents share.
Data Workers is the product on the other side of that line: it flags the recalled memories a change makes wrong, routes the update to the owner, checks the next recall and keeps the record.
Every tool owns a slice. Data Workers covers the whole lifecycle
Cognee owns one slice, and owns it outright: turning documents, conversations, code and rows into memory agents can recall. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the memory you already run.

| Stage | Data Workers | Cognee | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 9.5 | Cognee's home stage: remember turns documents, conversations, code and database rows into a knowledge graph with embeddings, and recall serves it to agents over MCP, the HTTP API and SDKs. Data Workers adds governed facts about the data estate next to it, with lineage, quality, usage and a named owner. |
| Analytics & Insights | 8 | 5 | Recall answers questions with graph completion and auto-routed search over memory; it does not compute metrics. Data Workers' Insights agent answers through governed metric definitions. |
| Data Quality | 8 | 3 | Ontology grounding standardizes entity types, and improve applies feedback weights to graph elements. Data Workers checks the tables and definitions those memories describe and repairs the breaks. |
| Observability & Incidents | 8.5 | 2 | Cognee tracks the status of its own ingestion pipelines. Data Workers detects a broken table or a changed definition, traces it across systems and fixes it. |
| Pipelines & Ingestion | 8.5 | 5 | Remember ingests text, files, URLs, repositories and database rows through dlt, and Cognee Cloud reads Slack, GitHub and Google Drive. Data Workers builds, reruns and backfills the data pipelines themselves with approvals. |
| Schema & Migration | 8 | 3 | dlt ingestion turns tables into schema nodes and foreign keys into edges at load time. Data Workers detects upstream schema changes and assesses their impact before they land. |
| Governance & Access | 8.5 | 5 | Multi-user mode adds users, tenants, roles and dataset-level permissions. Data Workers proposes and applies grants across your data platforms by policy. |
| Security & Privacy | 8 | 5 | Apache 2.0, self-hosted, and able to run on local models with no API key. Data Workers flags sensitive column names in pull request review and proposes masking for the owner. |
| Cost / FinOps | 8 | 2 | Cognee Cloud manages its own compute and storage, and quiet workspaces scale to zero. Data Workers traces Snowflake credits to the dbt model behind them and drafts the fix for its owner. |
| MLOps & Models | 7.5 | 6 | Memory makes agents consistent across sessions, and session lessons distill into durable knowledge. Data Workers keeps the data under your models healthy and connects to MLflow and W&B. |
How Cognee and Data Workers work together
Your engineers and agents stay in Claude Code, Cursor, Codex or your own agents, with Cognee as their memory. Spellbook Data Catalog (in preview) is where the data team looks: each contradiction, its sources, approver and receipt. Underneath, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm does the work with more than 20 specialist agents, the Autonomous Data-Conductor runs each fix end to end (detect, diagnose, fix, review, verify, remember), and per-domain guardrails hold approvals, receipts and rollback.

What Context Wizard does with your memory. Your Cognee datasets stay yours. Your team's assistant brings facts in with recall, scoped to the datasets you name, and Data Workers compares each fact about the data estate with the governed source. explain_table returns a table's meaning, owner and trust score; resolve_metric returns the canonical definition, or every candidate when a name is ambiguous; trace_cross_platform_lineage follows a table back to its source; monitor_metrics baselines say whether the data is current. When a recalled fact disagrees, Data Workers raises a contradiction with both sources, and flag_stale_context marks the governed side for review when the source moved first. Correct decisions that only lived in Slack go the other way: import_tribal_knowledge records them as business rules with their author. Promoting any fact to authoritative takes a named person; an agent cannot promote its own proposal.
Setup over MCP today. Cognee connects over the Cognee MCP server or its REST API today, in your team's client beside Data Workers' agents, which never call it. Run the Cognee MCP server in API mode against your shared Cognee backend with the Streamable HTTP transport, under a Cognee user with read access to the datasets in scope. Data Workers' agents are MCP servers from the open-source repository: clone it and add start-agent.sh entries to your client config, as the client setup docs show.
// Example: .mcp.json for Claude Code (Cursor uses the same mcpServers shape)
{
"mcpServers": {
"cognee": {
"type": "http",
"url": "http://localhost:8000/mcp"
},
"dw-context-catalog": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-context-catalog"]
},
"dw-quality": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-quality"]
},
"dw-schema": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-schema"]
}
}
}The Cognee entry matches Cognee's own Claude Code instruction, claude mcp add --transport http cognee http://localhost:8000/mcp -s project. List the tools in your client with its own command (for example /mcp in Claude Code). An engineer can then ask "what does our memory say about accounts, and is it still true?" and get the recalled facts from Cognee, the latest assess_impact results from dw-schema, the lineage from trace_cross_platform_lineage and a run_quality_check on the replicated table, in one answer.
Where writes go. Data Workers does not write into your Cognee graph. Approved facts reach Cognee through the path your team already controls: the dlt load that re-ingests a source, or the client that calls remember and forget under your Cognee user, with the approved text, its source and its owner. Fixes to the data itself land as a dbt diff for the owner to merge or a proposal to the table's owner.
One request, L0 to L4. The autonomy ladder is set per domain.

- •L0 manual. Data Workers is connected; your engineer compares memory with the sources by hand.
- •L1 observe. Data Workers takes the facts about data your team's assistant recalls from your Cognee datasets, checks them against the governed definitions and the live data, and reports contradictions with sources and owners.
- •L2 propose. Data Workers drafts the memory update with sources and a diff. The owner approves in Spellbook before anything reaches Cognee.
- •L3 act reversibly. For update classes with a proven record, such as a re-ingest after an approved schema change, the update runs through your pipeline, Data Workers checks the next recall your team's assistant runs, and the change can be rolled back.
- •L4 autonomous. For a scoped, trusted class in one domain, Data Workers keeps memory in step on its own and posts the receipt for review.
For the safety model, read is it safe to let AI agents change production data; for where data and credentials live, read where does our data go. The same pattern holds for every source of meaning: see the hub, bring your own context, and the guides for teams on Mem0, Zep and Letta. Our explainers on agent memory for data pipelines and what a context graph is cover the ideas behind Context Wizard.
What changes for your team

Teams that run agent memory spend real time tracing why an agent gave an old answer and re-loading sources after a change. With Data Workers next to Cognee, those jobs run on autopilot at the level you set.
- •Incidents. A schema change, table rename or metric change that makes a remembered fact wrong is caught and routed before an agent repeats it.
- •Data quality. The tables your memory datasets load from through dlt get freshness, volume and schema checks of their own, so a bad load never becomes remembered truth.
- •Cloud spend. Snowflake credits are traced to the query and dbt model behind them, and each fix goes to its owner drafted.
- •Access. An agent that needs a sensitive table gets a time-boxed grant proposal for the table's owner, with the policy that justifies it.
- •Audits. Every fact an agent recalled about data carries its source, owner, approver and history.
- •Migrations. When the warehouse moves, the memories that name old tables are found and updated in the same parity-checked waves.
Keep Cognee, or consolidate?
Keep Cognee if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
For most teams the answer is to keep it: conversation memory, document graphs, code graphs and session lessons belong in a memory layer built for them. What teams consolidate is the tooling around it: a script that diffs remembered schemas against the database, a separate quality tool for the tables memory loads from. If you are weighing building that checking layer yourself on the Cognee MCP server, read build it ourselves with Claude Code and MCP servers: the connection is the easy part; governed definitions, cross-system lineage, approvals and receipts are where the work is.
The case for your CFO
The outcome: the agents the company relies on give answers people can act on, because what they remember about customers, revenue and tables matches the live data and what the data team approved.
The risk story is plain. Your team's assistant reads Cognee under a Cognee user scoped to named datasets; Data Workers' agents never call Cognee and do not write into the memory graph. Every update it proposes shows the sources on both sides, goes to a named approver, reaches Cognee through your own pipeline, is verified on the next recall and leaves a receipt: who approved it, what changed and how to undo it. Autonomy is set per domain from L0 manual to L4 autonomous and can be dialled back at any time. There is zero migration: Cognee, the warehouse, dbt and your agents stay where they are.
Why now: agents answer without a person in between, so one stale fact reaches every agent that shares the memory. The first win is one memory dataset checked against the live data, read-only, so the next schema or definition change is caught before an agent repeats the old one. What stays the same: your Cognee deployment, datasets and ontologies, your warehouse permissions and your review process. For the numbers, see the ROI of agentic data operations.
The sentence to repeat upstairs: "Our agents remember what the company told them; Data Workers makes sure what they remember about our data is still true, with an owner and a receipt."
Getting started
Start with a pilot. Pick one Cognee dataset that agents already recall for data questions, such as a product-data dataset loaded through dlt or a Company Brain workspace, connect the Cognee MCP server next to Data Workers, and let Data Workers check its facts against the live data and your governed definitions before enabling the first update class. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.
FAQ
Does Data Workers write to our Cognee graph? No. Your team's assistant recalls from the datasets you name and hands the facts to Data Workers. Approved updates reach Cognee through your own load job or client, which calls remember and forget under your Cognee user after the owner approves in Spellbook.
Which facts does Data Workers check? Facts about the data estate: table and column meanings, schemas, metric definitions, owners, refresh times and lineage. Conversation memory, user preferences and code memory stay Cognee's alone.
We ground our graph with an OWL ontology. Does that change anything? It helps. Grounded entity types such as "Table" or "Metric" map cleanly to governed entries in Context Wizard, and Data Workers checks whether each specific fact is still true.
We load database tables into Cognee with dlt. What does Data Workers add? Data Workers watches those source tables for quality, schema changes in the dbt manifest and lateness against a baseline your team records, tells you when memory needs a re-ingest, and traces each table to where it is used.
Can decisions from Slack become governed definitions? Yes. A correct decision found in memory can be recorded in Context Wizard as a business rule with its author and source, then promoted to authoritative by a named owner. Every agent then reads the same rule.
Sources
- •Cognee, Documentation index (llms.txt), https://docs.cognee.ai/llms.txt (checked Oct 2, 2026)
- •Cognee, Introduction: what is Cognee, remember, recall, improve, forget, https://docs.cognee.ai/getting-started/introduction (checked Oct 2, 2026)
- •Cognee, Cognee MCP overview (standalone and API modes), https://docs.cognee.ai/cognee-mcp/mcp-overview (checked Oct 2, 2026)
- •Cognee, MCP tools reference, https://docs.cognee.ai/cognee-mcp/mcp-tools (checked Oct 2, 2026)
- •Cognee, MCP quickstart and transports, https://docs.cognee.ai/cognee-mcp/mcp-quickstart (checked Oct 2, 2026)
- •Cognee, Claude Code with Cognee MCP, https://docs.cognee.ai/cognee-mcp/integrations/claude-code (checked Oct 2, 2026)
- •Cognee, Remember, https://docs.cognee.ai/core-concepts/main-operations/remember (checked Oct 2, 2026)
- •Cognee, Improve, https://docs.cognee.ai/core-concepts/main-operations/improve (checked Oct 2, 2026)
- •Cognee, Forget, https://docs.cognee.ai/core-concepts/main-operations/forget (checked Oct 2, 2026)
- •Cognee, Ontology quickstart, https://docs.cognee.ai/guides/ontology-support (checked Oct 2, 2026)
- •Cognee, Ontologies concepts, https://docs.cognee.ai/core-concepts/further-concepts/ontologies (checked Oct 2, 2026)
- •Cognee, dlt integration (write dispositions, re-ingesting a source), https://docs.cognee.ai/integrations/dlt-integration (checked Oct 2, 2026)
- •Cognee, Cognee Cloud overview, https://docs.cognee.ai/cognee-cloud/overview (checked Oct 2, 2026)
- •Cognee, Permissions setup, https://docs.cognee.ai/setup-configuration/permissions (checked Oct 2, 2026)
- •Cognee, GitHub repository and README (Apache 2.0, Claude Code and Codex plugins), https://github.com/topoteretes/cognee (checked Oct 2, 2026)
- •Cognee, GitHub releases (v1.6.2, Sept 29, 2026: Slack conversation import and sync), https://github.com/topoteretes/cognee/releases (checked Oct 2, 2026)
- •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)
- •Data Workers open-source repository, agent tools and start-agent.sh, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)