Product
Product11 min readBy The Data Workers Team

You're on TigerGraph: Keep Your Deep-Link Analytics Right, From the Warehouse Table to the Fraud-Ring Query

Already on TigerGraph? Data Workers brings the graph's schema in as governed context, keeps the loading jobs and their source tables true, and fixes breaks with approvals.

Your team runs TigerGraph for the questions that take many hops. Accounts, cards, devices and merchants sit in one distributed graph so a GSQL query can find a fraud ring in seconds; suppliers, parts and shipments sit in another; customer records from five systems meet in an entity resolution graph. The graph runs on TigerGraph DB 4.2 or on Savanna, the managed service, and fills through loading jobs that read Snowflake, BigQuery or PostgreSQL with a SQL query, or Kafka and object storage. The TigerGraph MCP server (1.0 since April 2026) lets Claude Code, Cursor and VS Code work with the graph in plain language. TigerGraph runs the deep-link analytics. Data Workers keeps the graph's inputs true: Data Context Wizard brings the graph schema in as context with provenance, joins it to the lineage, quality and owners of the tables each loading job reads, and when a feed breaks, Data Workers proposes and carries the fix with approvals and a receipt.

A fraud ring found by a query is only as real as the edges under it. A placeholder value, a renamed column or a changed key format reaches the graph through the next loading job, and that seam between the warehouse and the graph is what this guide covers.

Key takeaways

  • •TigerGraph keeps its job. GSQL, installed queries, the Graph Data Science Library, Savanna and your loading jobs stay as they are.
  • •The graph schema becomes governed context. Data Context Wizard reads vertex and edge types and the loading jobs that fill them, next to warehouse lineage, quality and a named owner.
  • •Feeds are checked before the load. Placeholder keys, schema changes and fan-out spikes are caught in the warehouse and fixed before the loading job runs.
  • •GSQL writes stay gated. Your team's assistant reads the graph through the TigerGraph MCP server's read-only tool set, side by side with Data Workers; approved fixes land in dbt and the loading job your team owns, with an approver and a rollback path.
  • •Start with a pilot. One graph with its feeds watched, then one fix class, on the ladder from L0 manual to L4 autonomous.

TigerGraph is built to traverse large graphs fast and in parallel, which is why teams pick it for fraud, supply chain and entity resolution at scale. Loading jobs can point at a warehouse: "We currently support BigQuery, Snowflake, and PostgreSql data warehouses," with the SQL inline in the job's DEFINE FILENAME statement. The tigergraph-mcp package (1.0.3, September 16, 2026) serves 69 tools by default, from get_graph_schema and get_loading_jobs to get_vertex_count and get_node_degree, and Savanna's docs walk through it in Claude Code, Cursor, VS Code and Claude Desktop.

TigerGraph is where your relationships are analyzed at scale. Data Context Wizard is where every agent reads that graph's schema, next to the lineage, quality and owners of the tables that load it.

Here is an afternoon with Data Workers next to a fraud graph. This is an illustration, not a customer case.

TimeSystemWhat happens
14:00Mobile appRelease 7.4 ships; for users who decline tracking, the app now sends an all-zero device ID instead of leaving the field empty
15:10FivetranThe sync lands the new app sessions in Snowflake
16:00dbtfct_login_device builds; its not_null test passes, because the placeholder is a value, not a null
16:05Data WorkersIt flags a distribution break: one device ID now links to 38,412 accounts, where yesterday no device linked to more than six
16:12TigerGraphThe on-call's assistant reads the graph over the TigerGraph MCP server, read-only, and hands it to Data Workers: Device vertices key on device_id, LOGGED_IN_FROM edges come from Account, and the loading job load_login_device reads fct_login_device from Snowflake
16:20AirflowData Workers opens an incident ahead of the 18:00 load, naming the cause, the 38,412 accounts and the installed fraud-ring query and case queue downstream
16:40dbtData Workers proposes a diff that maps the placeholder to null and adds a fan-out test, plus a one-line change to the loading job's SQL that skips null device IDs, with its blast radius: one dbt model, one loading job, one graph, one installed query
17:30SpellbookThe fraud graph owner reviews the diff and the impact and approves; dbt CI passes
18:00AirflowThe DAG runs the loading job; no placeholder device reaches the graph
18:25TigerGraphData Workers checks the maximum Device degree and the LOGGED_IN_FROM edge count against the last seven loads, finds both in range and writes the receipt
21:00Case queueThe nightly fraud-ring query flags the usual handful of real rings for analysts
Incident timeline across the stack: what TigerGraph, your team and Data Workers each do, step by step

Without that check, the load would have created one Device vertex with 38,412 edges, the connected-components step of the fraud-ring query would have merged tens of thousands of honest customers into one "ring", and the real rings would have been buried in a flooded case queue. TigerGraph did exactly what it was asked to do. The fix was upstream, in systems the graph doesn't run, and it went through an owner's approval first.

JobWhat TigerGraph doesWhat Data Workers does
The modelStores vertex and edge types with primary keys, typed attributes and vector attributesReads the schema as context and joins it to warehouse lineage, quality, usage and owners
The questionsRuns multi-hop GSQL and openCypher queries, graph algorithms and hybrid graph and vector searchMakes sure the data and definitions behind those answers are correct and current
The feedLoads from Snowflake, BigQuery, PostgreSQL, Kafka and object storage through loading jobsWatches the tables, keys and value distributions each loading job reads and catches breaks before the run
The breakLoads what the job's SQL returnsDetects the upstream change and traces it across the app, Fivetran, dbt and Airflow to the graph
The fixRuns the loading job and the GSQL your team approvesProposes the change with its blast radius, routes it to a named owner, lands it in dbt and the loading job
The proofHolds the graph state after the load, with vertex and edge countsVerifies degrees and counts after the reload and writes a receipt with the cause, diff, approver and rollback
The meaningHolds vertex and edge types and installed query descriptionsKeeps business definitions consistent across the graph, the semantic layer and the catalog, with conflicts routed to an owner

Why doesn't TigerGraph just do this itself?

Because TigerGraph built a distributed graph engine for the hardest part of its job: storing and traversing huge, connected data in parallel, for any domain. The app that emits the device ID, the Fivetran connector, the dbt project and the Airflow DAG belong to other vendors and to your data team. Changing them safely is a different product with a different liability.

TigerGraph's own MCP guidance draws that line sensibly. Savanna's setup page says: "Actions performed through TigerGraph MCP run against the connected database and can modify or permanently delete data and graph resources. Review destructive tool calls carefully before approving them, particularly in production." The server ships a read-only selector for "the tools that change nothing", a destructive selector you can block, and MCP hints on every tool so the client can ask before a write. It even marks gsql, run_query and run_installed_query as destructive, "because what they do depends on the text they are given." That is the right design for a database serving many teams and agents.

Data Workers is the product on the other side of that line: it knows what a change touches across systems, routes it to the owner, applies it reversibly, verifies the graph and keeps the record.

Every tool owns a slice. Data Workers covers the whole lifecycle

TigerGraph owns one slice of the data lifecycle outright: deep-link graph analytics at scale. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on the graph you already run.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, TigerGraph goes deep on its own area
StageData WorkersTigerGraphWhy we scored it this way
Catalog & Context99.5TigerGraph's home stage: accounts, devices, suppliers and parts as vertices and edges in one distributed graph, served to agents through the TigerGraph MCP server. Data Workers brings that schema in as context with provenance, next to lineage, quality and usage.
Analytics & Insights89.5TigerGraph's other home stage: multi-hop GSQL and openCypher queries, the Graph Data Science Library and hybrid graph+vector search find fraud rings, supply risks and duplicate entities at scale. Data Workers' Insights agent answers through governed metric definitions.
Data Quality84Primary keys and typed attributes keep vertices in shape inside the graph. Data Workers checks the warehouse tables that loading jobs read and repairs the breaks before a load.
Observability & Incidents8.53TigerGraph reports on loading job status and the health of the database itself. Data Workers detects a bad feed, traces it across systems, fixes it and verifies the graph after the reload.
Pipelines & Ingestion8.55Loading jobs read Snowflake, BigQuery and PostgreSQL with a SQL query, plus object storage and Kafka. Data Workers builds, reruns and backfills the pipelines behind those loads with approvals.
Schema & Migration83GSQL schema changes are explicit and owned by designers. Data Workers detects upstream schema changes and assesses their impact on every model and loading job before they land.
Governance & Access8.56.5Strong over its own graph: role-based access with built-in roles, type- and attribute-level scopes and row policies. Data Workers proposes and applies grants across your data platforms by policy.
Security & Privacy85Strong for its own estate: SSO and identity provider support in Savanna, and a read-only tool set for agents. Data Workers flags sensitive column names in pull request review and proposes masking for the owner.
Cost / FinOps82Savanna bills its own workspaces. Data Workers traces credits in the warehouse that feeds the graph to the dbt model behind them and drafts the fix for its owner.
MLOps & Models7.55Graph algorithms and vector attributes feed features and embeddings to models. Data Workers keeps the data under your models healthy and connects to MLflow and W&B.

How TigerGraph and Data Workers work together

Your engineers and agents stay in Claude Code, Cursor or VS Code for GSQL and dbt, and in GraphRAG or your own apps for answers. Spellbook Data Catalog (in preview) is where the data team looks: each proposed change, its blast radius, its approver and its rollback. Between them, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm does the work with more than 20 specialist agents, the Autonomous Data-Conductor runs each fix end to end (detect, diagnose, fix, review, verify, remember), and per-domain guardrails hold approvals, receipts and rollback.

How Data Workers fits with TigerGraph: your coding agent on top, Data Workers in the middle, your estate underneath

What Context Wizard does with your graph. It reads the graph schema (vertex types, primary keys, edge types, attributes) and the loading jobs that fill them, including the SQL each job runs against the warehouse, and records them as context with provenance: where each fact came from, when it was observed and who owns it. It joins that to warehouse lineage, so Device.device_id traces back through the load_login_device job, fct_login_device and Fivetran to the mobile app, and serves one governed view to every agent through tools like explain_table and trace_cross_platform_lineage. When the graph's idea of an "active account" disagrees with your semantic layer's, the conflict goes to a named owner.

Setup over MCP today. Data Workers connects to TigerGraph over the TigerGraph MCP server today, alongside its own agents, in the same client. Clone the open-source repository and add start-agent.sh entries to your client config, as the client setup docs show. Run the TigerGraph MCP server with TG_ALLOWED_TOOLS=read-only, which serves only the 37 tools that change nothing, so agents can read the schema, the loading jobs and the counts and nothing can write from the chat.

// Example: .mcp.json for Claude Code (Cursor uses the same mcpServers shape)
{
  "mcpServers": {
    "tigergraph": {
      "command": "uvx",
      "args": ["tigergraph-mcp"],
      "env": {
        "TG_HOST": "<your Savanna workspace URL>",
        "TG_GRAPHNAME": "FraudGraph",
        "TG_SECRET": "<database secret, from your secret store>",
        "TG_ALLOWED_TOOLS": "read-only"
      }
    },
    "dw-context-catalog": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-context-catalog"]
    },
    "dw-schema": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-schema"]
    },
    "dw-quality": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-quality"]
    }
  }
}

List the tools in your client with its own command (for example /mcp in Claude Code). An engineer can then ask "what feeds the Device vertex, and is it safe to load tonight?" and get the schema and loading job from TigerGraph, lineage from trace_cross_platform_lineage, the latest run_quality_check and get_anomalies results on fct_login_device, and its load lag against the monitor_metrics baseline, in one answer. On a shared HTTP server, the X-TG-Tools header applies the same selector per session.

Where writes go. Fixes land where your team already reviews change: a dbt diff for the owner to merge, or a change to the loading job's GSQL kept in your repository. Your DAG runs the loading job, as it does today, after the owner approves, and blast_radius_analysis shows every model, loading job and installed query a change touches before anyone approves it.

One request, L0 to L4. The autonomy ladder is set per domain.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. Data Workers is connected but not acting; your engineer traces the feed by hand.
  • •L1 observe. Data Workers watches the tables each loading job reads, flags breaks before the run and explains the cause. Nothing changes.
  • •L2 propose. Data Workers drafts the dbt diff or loading-job change with its blast radius; the graph owner approves in Spellbook.
  • •L3 act reversibly. For proven classes, such as holding a loading job until a late table lands and rerunning it, Data Workers applies the change, verifies counts and degrees, and can roll it back.
  • •L4 autonomous. For a scoped, trusted class like late feeds into one graph, Data Workers fixes and verifies on its own and posts the receipt.

For the safety model behind each step, read is it safe to let AI agents change production data, and for where data and credentials live, read where does our data go.

The same pattern holds for every graph: see the hub, bring your own context, and the guides for teams on Neo4j, Memgraph and Amazon Neptune. For grounding agents, see semantic layer vs knowledge graph for LLM grounding and what is a context graph.

What changes for your team

Six jobs that run on autopilot with Data Workers next to TigerGraph, with a concrete example of each

Graph teams lose much of the week to work that isn't graph work: a vertex that suddenly has forty thousand edges, a key that changed format upstream, a rerun after a bad source. With Data Workers next to TigerGraph, those jobs run on autopilot at the level you set.

  • •Incidents. A placeholder ID, key format change or dropped column that would merge or orphan vertices is fixed before the loading job runs.
  • •Data quality. Every key a loading job maps to a vertex gets null, placeholder and fan-out checks, so a supernode never forms from a value that was never real.
  • •Cloud spend. Cleanups of extracts and staging tables built for retired loading jobs are proposed to the owner after a dependency check.
  • •Access. A request to query a sensitive graph arrives as a time-boxed grant proposal for its owner, with the policy behind it.
  • •Audits. Every change to a feed, loading job or definition carries an approver, a diff, the verification and a rollback path, so "why is this account in this ring?" traces to the data that put it there.
  • •Migrations. When the warehouse under the graph moves, the feeds move in parity-checked waves while the graph keeps loading.

The graph team gets its week back for new queries, new domains and the deep-link work only it can do.

Keep TigerGraph, or consolidate?

Keep TigerGraph if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

For most teams, the graph itself stays: the schema, the GSQL, the installed queries and the scale belong in TigerGraph. What teams consolidate is the tooling around it: a separate data-quality tool for the feeds, a homegrown script that compares vertex counts after each load, a spreadsheet mapping vertex types to tables and owners. If you are weighing building this layer yourself on the TigerGraph MCP server, read build it ourselves with Claude Code and MCP servers: the connection is the easy part; the cross-system context, approvals and rollback are where the work is.

The case for your CFO

The outcome: the graph analytics the company paid for give answers people can act on. Fraud decisions, supplier risk and merged customer records rest on the tables that feed the graph, and Data Workers keeps them right every load.

The risk story is plain. Your team's assistant reads TigerGraph through its MCP server with only the read-only tool set enabled; Data Workers' agents never call it. Every change Data Workers proposes shows its blast radius, goes to a named approver, lands through the pipelines your team already runs, is verified after the reload and leaves a receipt: who approved it, what it touched and how to undo it. Autonomy is set per domain from L0 manual to L4 autonomous and can be dialled back at any time. There is zero migration: TigerGraph, the warehouse, dbt and your loading jobs stay where they are.

Why now: agents now query the graph directly through TigerGraph's MCP server, so a bad load reaches every agent and every analyst who asks. One false ring can cost a fraud team a night of reviews. The first win is one graph with its feeds watched, so the next upstream change is caught before the load. What stays the same: your schema, your GSQL, your Savanna workspaces, your warehouse permissions and your review process. For the numbers, see the ROI of agentic data operations.

The sentence to repeat upstairs: "Our graph is only as good as the tables that load it; Data Workers keeps those tables and their meaning right, and fixes them with an approval and a receipt."

Getting started

Start with a pilot. Pick one graph that analysts or agents already rely on, such as the fraud or entity resolution graph, connect the TigerGraph MCP server read-only next to Data Workers, and let Data Workers watch the tables behind its loading jobs before you enable the first fix class. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.

FAQ

Does Data Workers write to our TigerGraph database? Your team's assistant reads the graph through the TigerGraph MCP server with TG_ALLOWED_TOOLS=read-only. Approved fixes land in dbt or in the loading job's definition, and your DAG runs the loading job as it does today, after the graph owner approves in Spellbook. On self-managed TigerGraph you can also give that connection a user whose role grants read privileges only.

Our loading jobs read straight from Snowflake with SQL. Does that change anything? No, it makes the lineage clearer. Context Wizard records the SQL each loading job runs, so the tables behind every vertex and edge type are known, and Data Workers checks them before the job runs.

We load from Kafka, not the warehouse. Is that covered? Yes. Data Workers watches the topic's registered schema and the upstream tables or services that produce it, flags breaking changes and placeholder values, and traces them to the vertex and edge types the job fills.

How does Data Workers verify a load? After the loading job runs, Data Workers checks the source tables the job reads, your team's assistant compares vertex and edge counts and the degree of high-risk vertex types against recent loads over the read-only MCP tools, and the receipt records both. Out-of-range numbers open an incident.

Does it work with Savanna and self-managed TigerGraph? Yes. The TigerGraph MCP server connects to both with the same tools: a Savanna workspace URL and secret, or your own host with a user or API token.

Sources

  • •TigerGraph, Documentation home (TigerGraph DB 4.2.5, Sep 2, 2026; 4.3.0-rc1, Jul 8, 2026), https://www.tigergraph.com/docs/home/ (checked Oct 2, 2026)
  • •TigerGraph, Savanna overview, https://www.tigergraph.com/docs/savanna/main/overview/ (checked Oct 2, 2026)
  • •TigerGraph, Connect AI tools with MCP (Savanna), https://www.tigergraph.com/docs/savanna/main/get-started/connect-agent-mcp (checked Oct 2, 2026)
  • •PyPI, tigergraph-mcp 1.0.3 (Sep 16, 2026), https://pypi.org/project/tigergraph-mcp/ (checked Oct 2, 2026)
  • •TigerGraph, tigergraph-mcp README and warehouse loading guide, https://github.com/tigergraph/tigergraph-mcp (checked Oct 2, 2026)
  • •TigerGraph, Load from a data warehouse (TigerGraph DB 4.2), https://www.tigergraph.com/docs/tigergraph-server/4.2/data-loading/load-from-warehouse (checked Oct 2, 2026)
  • •TigerGraph, Access control model and row policies (TigerGraph DB 4.2), https://www.tigergraph.com/docs/tigergraph-server/4.2/user-access/access-control-model (checked Oct 2, 2026)
  • •TigerGraph, Create a database secret (Savanna), https://www.tigergraph.com/docs/savanna/main/administration/settings/how2-create-database-secret (checked Oct 2, 2026)
  • •TigerGraph, GraphRAG repository and release notes (v2.0.2, Aug 28, 2026), https://github.com/tigergraph/graphrag (checked Oct 2, 2026)
  • •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)
  • •Data Workers open-source repository, agent tools and start-agent.sh, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)