Product
Product11 min readBy The Data Workers Team

You're on TrustGraph: Keep Your Context Cores True When the Warehouse Data Changes

Self-hosting TrustGraph for GraphRAG? Data Workers traces each context core to its warehouse source, catches stale facts and proposes the rebuild with approvals.

Your team self-hosts TrustGraph, the open-source (Apache 2.0) platform that calls itself "the Semantic Intelligence Layer for Ontologies". You generated a deployment with npx @trustgraph/config, run it on Docker Compose or Kubernetes, and feed it through flows: documents through the library, warehouse rows through tg-load-structured-data, with an OWL ontology guiding extraction so entities and relationships come out typed. Each workspace holds collections, and each processed document becomes a Context Core (also called a knowledge core): a portable package of graph edges, schema and graph embeddings you can download, upload and load. Agents answer with Graph RAG, Ontology RAG or Document RAG, query with SPARQL and show an explainability trace back to source. TrustGraph builds the knowledge. Data Workers, the agentic data platform, keeps it true: Data Context Wizard traces each core back to the tables it was built from, and when the warehouse changes under it, Data Workers catches the stale facts and proposes the rebuild with an approval and a receipt.

A core is a snapshot of what the data said on the day it was extracted. When a billing table is renamed or a plan is split, the warehouse changes in the afternoon and the core keeps answering from last week. That seam is the job this guide covers.

Key takeaways

  • •TrustGraph keeps its job. Flows, ontologies, collections, Context Cores and Graph RAG stay as they are. Data Workers works next to them from day one.
  • •Every core gets a source. Data Context Wizard records which tables, dbt models and export jobs built each collection and core, with lineage, quality, freshness and an owner.
  • •Stale facts are caught before agents read them. An upstream schema change, null jump or changed definition is traced to the cores it affects, and graph facts are checked over SPARQL.
  • •Writes stay with your team. Your team's assistant reads TrustGraph with a reader-role API key, side by side with Data Workers. Fixes land in dbt and the export job after a named owner approves.
  • •Start with a pilot. One collection, read-only, sources watched; then one fix class in one domain, on the ladder from L0 manual to L4 autonomous.

TrustGraph is the knowledge layer. Data Workers is the agentic data platform that keeps it true.

TrustGraph is a strong place to turn documents and data into knowledge agents can use. It stores RDF 1.2 named graphs with OWL ontologies and PROV-O provenance on Cassandra, Neo4j, Memgraph or FalkorDB, and ships Python and TypeScript libraries, nearly 100 tg- CLI commands and an MCP server (the trustgraph-mcp container, with bearer-token authentication since v2.5 in June 2026) whose 31 tools include graph_rag, sparql_query, get_knowledge_cores and get_kg_core. Version 2.10.2 shipped on October 2, 2026.

Outside the graph sits everything that decides whether it is right: app databases, the CDC stream, the lakehouse, dbt models, the export job and the governed definitions. Data Workers covers that side.

Here is a Tuesday next to a support knowledge graph. This is an illustration, not a customer case.

TimeSystemWhat happens
10:12PostgresThe billing team ships a migration: plans.max_seats becomes seat_limit, and Pro splits into pro_monthly and pro_annual, with annual seats raised from 10 to 25
10:13Debezium and KafkaDebezium streams the change events with the new schema to the billing.public.plans topic
10:40Databricks and dbtdim_plan builds; its unique and not_null tests on plan_code pass, and max_seats is null on every new row
10:46Data WorkersIts monitor_metrics baseline flags the null jump on dim_plan.max_seats, and the producer team confirms the column rename in the topic's new schema version
10:52DagsterData Workers traces dim_plan to the support_kb_export asset that runs at 13:00 and rebuilds the support-kb collection in TrustGraph
10:58TrustGraphOver the TrustGraph MCP server, a SPARQL query shows the loaded core still says Pro allows 10 seats and has no annual plan; the 13:00 rebuild would load null seat limits
11:05Data WorkersThe core's fact contradicts the governed definition the billing owner approved last month; the conflict goes to that owner with both sources
11:20dbtData Workers proposes a diff that maps seat_limit into dim_plan, with its blast radius: one dbt model, one Dagster asset, one collection, one support agent
11:45SpellbookThe billing data owner reviews the diff and the impact and approves; dbt CI passes
13:00DagsterThe export runs with the team's own TrustGraph CLI steps and rebuilds the collection and its core
13:40TrustGraphThe agent owner's assistant checks all six plans over SPARQL, Data Workers compares them with dim_plan, finds them matching and writes the receipt
15:10Support agentA customer asks how many seats Pro annual includes; the Graph RAG answer says 25 and cites the plan record
Incident timeline across the stack: what TrustGraph, your team and Data Workers each do, step by step

Without that check, the rebuild would have loaded plans with no seat limit, and the support agent would have answered from a graph that looked complete. TrustGraph did exactly what it was asked: it served what the export gave it. The fix was upstream, in systems TrustGraph doesn't run, and it went through an owner's approval first.

JobWhat TrustGraph doesWhat Data Workers does
The knowledgeExtracts entities and relationships into collections, guided by your ontologyReads the collections as context and joins them to warehouse lineage, quality, usage and owners
The answersServes Graph RAG, Ontology RAG, Document RAG and SPARQL, with an explainability traceMakes sure the data and definitions behind those answers are correct and current
The coresPackages each document's knowledge as a portable Context CoreRecords which tables and jobs built each core and flags the cores a source change makes stale
The breakLoads what the export gives itDetects the upstream change and traces it across Postgres, Debezium, Databricks, dbt and Dagster to the collection
The meaningHolds the ontology and the facts extracted under itChecks facts against governed definitions and routes contradictions to a named owner
The fixRebuilds when your job runs the CLIProposes the change with its blast radius, lands it in dbt and the export job after approval
The proofHolds the rebuilt graph and its provenanceVerifies the rebuilt facts over SPARQL and writes a receipt with the cause, diff, approver and rollback

Why doesn't TrustGraph just do this itself?

Because TrustGraph built a platform for the hardest part of its job: turning any source, in any domain, into typed, traceable knowledge on infrastructure the customer controls. It is designed to serve faithfully what it is given. The tables that feed it belong to your data team, and the definitions it should agree with live in dbt, the semantic layer or a catalog. Changing Postgres migrations, dbt models and Dagster jobs is a different product with a different liability.

TrustGraph's own design draws that line sensibly. Workspaces isolate data, and the open-source access model has three roles, reader, writer and admin, with the gateway checking every request's capability with the IAM service. The MCP server forwards each caller's bearer token to that gateway. A reader key can call 22 of the 31 MCP tools: the query tools (graph_rag, sparql_query, triples_query, graphql_query, agent) and the reads of cores, flows, documents and config. Loading, deleting or adding documents and cores needs a writer; changing config and starting or stopping flows needs an admin. That is the right design for a platform serving many teams from one deployment. Data Workers is the product on the other side of that line: it knows what a change will touch across systems, routes it to the owner, applies it reversibly through your pipelines, verifies the result and keeps the record.

Every tool owns a slice. Data Workers covers the whole lifecycle

TrustGraph owns one slice of the lifecycle outright: turning documents and data into ontology-grounded knowledge and answering from it. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the knowledge graph you already run.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, TrustGraph goes deep on its own area
StageData WorkersTrustGraphWhy we scored it this way
Catalog & Context99.5TrustGraph's home stage: Graph RAG, Ontology RAG and Document RAG over RDF named graphs and OWL ontologies, packaged as Context Cores and served through SPARQL, GraphQL, its APIs and an MCP server. Data Workers reads that knowledge as context next to lineage, quality, usage and a named owner.
Analytics & Insights88.5TrustGraph's other home stage: natural-language questions over the graph and over structured rows, with an explainability trace from answer back to source. Data Workers' Insights agent answers through governed metric definitions.
Data Quality84Ontologies guide extraction so entities and relationships come out typed and consistent. Data Workers checks the warehouse tables those facts are built from and repairs the breaks.
Observability & Incidents8.54Prometheus metrics, Grafana dashboards and audit events report on TrustGraph's own pipelines. Data Workers detects a changed source, traces it across systems and fixes it.
Pipelines & Ingestion8.56Flows ingest PDF, DOCX, XLSX, HTML, Markdown and CSV, and tg-load-structured-data loads CSV, JSON and XML rows. Data Workers builds, reruns and backfills the warehouse pipelines behind those exports with approvals.
Schema & Migration84Schemas describe structured rows and the Ontology Workbench edits classes and properties. Data Workers detects upstream schema changes and assesses their impact on every model and export before they land.
Governance & Access8.56Workspaces isolate data, and reader, writer and admin roles are gated per capability at the API gateway. Data Workers proposes and applies grants across your data platforms by policy.
Security & Privacy86Self-hosted with no telemetry, hashed API keys and JWT login. Data Workers flags sensitive column names in pull request review and proposes masking for the owner.
Cost / FinOps83TrustGraph reports token costs for its own LLM calls and runs on local models. Data Workers traces Snowflake credits to the dbt model behind them and drafts the fix for its owner.
MLOps & Models7.55Over 40 LLM providers and self-hosted inference with vLLM, TGI and Ollama. Data Workers keeps the data under your models healthy and connects to MLflow and W&B.

How TrustGraph and Data Workers work together

Engineers stay in Claude Code or Cursor for dbt and export code, and in TrustGraph's Workbench and agent runtime for questions over the graph. Spellbook Data Catalog (in preview) is where the data team looks: each proposed change, its blast radius, approver and rollback. Underneath, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm does the work with more than 20 specialist agents, the Autonomous Data-Conductor runs each fix end to end (detect, diagnose, fix, review, verify, remember), and per-domain guardrails hold approvals, receipts and rollback.

How Data Workers fits with TrustGraph: your coding agent on top, Data Workers in the middle, your estate underneath

What Context Wizard does with your collections. It reads which collections and cores exist, their ontology classes and the facts agents rely on, and records them with provenance: source, time observed and owner. It joins that to warehouse lineage, so a plan's seat limit in the support collection traces back through the Dagster export and dim_plan to the Postgres table and the CDC topic. Agents get definitions, lineage and a trust score from explain_table, resolve_metric and trace_cross_platform_lineage, and quality from get_quality_score. When a fact in a core disagrees with a governed definition, the conflict goes to a named owner; promoting a fact to authoritative takes a named person.

Setup over MCP today. Data Workers connects to TrustGraph over TrustGraph's MCP server today, in the same client as its own agents. Run the trustgraph-mcp container next to your API gateway (it serves streamable HTTP on port 8000 by default), create an API key for a user with the reader role in the workspace you want covered (tg-create-user, then tg-create-api-key), and pass it as a bearer token. Data Workers' agents come from the open-source repository: clone it and add start-agent.sh entries, as the client setup docs show.

// Example: .mcp.json for Claude Code (Cursor uses the same mcpServers shape)
{
  "mcpServers": {
    "trustgraph": {
      "type": "http",
      "url": "http://<trustgraph-mcp-host>:8000/mcp",
      "headers": { "Authorization": "Bearer <reader-role API key from your secret store>" }
    },
    "dw-context-catalog": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-context-catalog"]
    },
    "dw-schema": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-schema"]
    },
    "dw-quality": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-quality"]
    }
  }
}

List the tools with your client's own command (for example /mcp in Claude Code). An engineer can then ask "which tables built the support-kb core, and are they still the same shape?" and get the cores from get_knowledge_cores, lineage from trace_cross_platform_lineage, column history from the dbt manifest and Schema Registry, null rates on dim_plan against their monitor_metrics baseline and the graph facts from sparql_query, in one answer. TrustGraph's agent runtime is also an MCP client (tg-set-mcp-tool), so a TrustGraph agent can ask Data Workers which definition is governed.

Where writes go. Fixes land where your team already reviews change: a dbt diff for the owner to merge, or the export job that runs tg-load-structured-data, tg-start-library-processing and tg-load-kg-core with its own writer key after the owner approves. Because a core is a file, the previous core is the rollback: reload it and the graph is back where it was.

One request, L0 to L4. The autonomy ladder is set per domain.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. Connected, not acting. Your engineer traces a wrong answer by hand.
  • •L1 observe. Data Workers watches each collection's source tables, flags changes that would make a core stale and explains the cause. Nothing changes.
  • •L2 propose. It drafts the dbt diff or export-job change with its blast radius; the owner approves in Spellbook first.
  • •L3 act reversibly. For proven classes, such as rerunning a failed export, it triggers the team's job, verifies over SPARQL and can restore the previous core.
  • •L4 autonomous. For a scoped, trusted class like a late upstream table, it reruns, verifies and posts the receipt for review.

The safety model behind each step is in is it safe to let AI agents change production data. The same pattern holds for every source of meaning: see the hub, bring your own context, and the guides on Cognee, Contextual AI and Neo4j. For graph versus semantic layer, read semantic layer vs knowledge graph for LLM grounding and building a context graph with MCP.

What changes for your team

Six jobs that run on autopilot with Data Workers next to TrustGraph, with a concrete example of each

Tracing wrong answers to tables, finding which cores a change touched and rebuilding by hand now run on autopilot at the level you set.

  • •Incidents. A source change that would leave a core stale is caught, traced and fixed before the next rebuild.
  • •Data quality. Every table a collection is built from gets null and key checks in the warehouse and a load-lag baseline.
  • •Cloud spend. Cleanups of exports and staging tables for retired collections are proposed to the owner after a dependency check.
  • •Access. A request to read a sensitive workspace arrives as a time-boxed grant proposal for its owner.
  • •Audits. Every rebuild carries the cause, approver, diff, verification and rollback path, next to TrustGraph's own explainability trace.
  • •Migrations. A warehouse move runs in parity-checked waves while the collections keep loading.

Keep TrustGraph, or consolidate?

Keep TrustGraph if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

Most keep it: the ontologies, flows, cores and agents belong there, and running it yourself is often why you chose it. What teams consolidate is the tooling around it: a script that diffs exports before a rebuild, a spreadsheet that maps collections to source tables, a separate quality tool for the warehouse feeds. Data Workers runs those jobs with one context, one approval flow and one audit trail, on an Apache 2.0 core. If you are weighing building this layer yourself on TrustGraph's MCP server, read build it ourselves with Claude Code and MCP servers: the connection is the easy part; cross-system context, approvals and rollback are the work.

The case for your CFO

The outcome: the knowledge graph the company built for its agents gives answers people can act on, because Data Workers keeps each core in step with the data it was built from.

The risk story is plain. Data Workers reads TrustGraph with a reader-role key, so it cannot load, delete or change a core. Every proposed change shows its blast radius, goes to a named approver, lands through your dbt project and export job, is verified over SPARQL after the rebuild and leaves a receipt: who approved it, what it touched and how to undo it. The previous core stays available as the rollback. Autonomy is set per domain from L0 manual to L4 autonomous and can be dialled back at any time. There is zero migration: TrustGraph, the lakehouse, dbt and your orchestrator stay where they are.

Why now: agents answer customers from the graph, so a stale core reaches everyone who asks. The first win is one collection with its sources watched, read-only, so the next upstream change is caught before the rebuild. What stays the same: your ontologies, deployment, LLM choices, workspaces, roles and review process. For the numbers, see the ROI of agentic data operations.

The sentence to repeat upstairs: "Our knowledge graph is only as current as its data; Data Workers keeps the two in step, with an approval and a receipt on every rebuild."

Getting started

Start with a pilot. Pick one collection agents rely on that is built from warehouse data, connect TrustGraph's MCP server with a reader-role key, and let Data Workers watch its sources for a few weeks before enabling the first fix class. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.

FAQ

Does Data Workers write to our TrustGraph deployment? No. Your team's assistant reads over TrustGraph's MCP server with a reader-role API key, which the gateway checks on every call. Approved fixes land in dbt or the export job, which rebuilds the core with its own writer key, as it does today.

How does Data Workers know a context core is stale? Context Wizard records which tables, models and export runs built each collection. When a source changes shape, goes stale or changes a governed definition, Data Workers lists the affected cores and checks the facts your team's assistant reads over SPARQL against the warehouse.

Can we roll back a rebuild? Yes. Download the Context Core with tg-get-kg-core before the rebuild. If verification fails, your job reloads it, and the receipt records both.

Does this replace our ontology or TrustGraph's explainability? No. The ontology keeps guiding extraction, and TrustGraph's explainability trace still shows which nodes and documents an answer used. Data Workers adds which warehouse table and which approved definition those facts came from.

We self-host everything. Where does Data Workers run? Data Workers' agents run from the open-source repository on your own infrastructure, next to TrustGraph, over MCP inside your network. See where does our data go.

Sources

  • •TrustGraph, documentation index and self-description, https://docs.trustgraph.ai/llms.txt (checked Oct 2, 2026)
  • •TrustGraph, How does TrustGraph work?, https://docs.trustgraph.ai/learn/how-it-works.html (checked Oct 2, 2026)
  • •TrustGraph, Open source, https://docs.trustgraph.ai/learn/open-source.html (checked Oct 2, 2026)
  • •TrustGraph, Working with Context Cores, https://docs.trustgraph.ai/guides/context-cores/ (checked Oct 2, 2026)
  • •TrustGraph, Working with Context Cores using CLI, https://docs.trustgraph.ai/guides/context-cores-cli/ (checked Oct 2, 2026)
  • •TrustGraph, Structured data processing, https://docs.trustgraph.ai/guides/structured-processing/ (checked Oct 2, 2026)
  • •TrustGraph, MCP Integration, https://docs.trustgraph.ai/guides/mcp-integration/ (checked Oct 2, 2026)
  • •TrustGraph, Architecture and storage options, https://docs.trustgraph.ai/architecture/architecture.html (checked Oct 2, 2026)
  • •TrustGraph, Security (API keys, roles, MCP service), https://docs.trustgraph.ai/architecture/security.html (checked Oct 2, 2026)
  • •TrustGraph, Containers (trustgraph-mcp), https://docs.trustgraph.ai/reference/containers.html (checked Oct 2, 2026)
  • •TrustGraph, CLI reference, https://docs.trustgraph.ai/reference/cli/ (checked Oct 2, 2026)
  • •TrustGraph, Changelog (v2.9 Sept 16, 2026; v2.8 Aug 24, 2026; v2.5 MCP server authentication June 10, 2026), https://docs.trustgraph.ai/changelog/trustgraph.html (checked Oct 2, 2026)
  • •TrustGraph on GitHub (README, Apache 2.0, v2.10.2), https://github.com/trustgraph-ai/trustgraph (checked Oct 2, 2026)
  • •TrustGraph MCP server source (31 tools, bearer token forwarded to the gateway), https://github.com/trustgraph-ai/trustgraph/blob/master/trustgraph-mcp/trustgraph/mcp_server/mcp.py (checked Oct 2, 2026)
  • •TrustGraph open-source IAM role table (reader, writer, admin capabilities), https://github.com/trustgraph-ai/trustgraph/blob/master/trustgraph-flow/trustgraph/iam/service/iam.py, and gateway operation registry, https://github.com/trustgraph-ai/trustgraph/blob/master/trustgraph-flow/trustgraph/gateway/registry.py (checked Oct 2, 2026)
  • •TrustGraph, Monitoring (Prometheus and Grafana), https://docs.trustgraph.ai/guides/monitoring/ (checked Oct 2, 2026)
  • •PyPI, trustgraph 2.10.2 (Oct 2, 2026), https://pypi.org/project/trustgraph/ (checked Oct 2, 2026)
  • •TrustGraph, homepage, https://trustgraph.ai/ (checked Oct 2, 2026)
  • •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)
  • •Data Workers open-source repository, agent tools and start-agent.sh, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)