Product
Product11 min readBy The Data Workers Team

You're on Contextual AI: Ground the Documents There, Govern the Numbers With Data Workers

Already on Contextual AI? Its agents ground answers in your documents. Data Workers supplies the metrics and tables over MCP, keeps them true and fixes what breaks.

Your experts run on Contextual AI, the context layer for expert AI. Datastores sync specifications, contracts, test reports and runbooks from SharePoint, Confluence, Box, Google Drive and OneDrive, with each source's entitlements enforced at query time. Agents built in Agent Composer (public preview since January 27, 2026) start from the Basic Search or Agentic Search templates, or, on enterprise plans, from a custom YAML workflow. They retrieve with the instruction-following reranker, generate with the Grounded Language Model (GLM), and return answers with attributions; LMUnit tests those answers. Contextual AI is where your documents become grounded answers. Data Context Wizard is where every agent reads the numbers, next to lineage, quality and usage, with a named owner on every fact.

That split matters the first time an expert question touches a table. "Is this customer within its SLA?" needs the contract clause and the on-time rate. Contextual AI's own blog (March 12, 2026) puts it plainly: a semantic layer makes structured data understandable to BI tools, and a context layer makes all enterprise data usable by AI agents, so enterprise AI needs both. Data Workers is the agentic data platform, and here it owns the structured side: governed definitions, verified tables, and a crew that repairs the data when it breaks. Over MCP, it hands the agent a number it can stand behind.

Key takeaways

  • •Contextual AI keeps its job. Datastores, Agent Composer, the GLM, the reranker and LMUnit stay as they are. Data Workers joins as an MCP server your workflows call.
  • •Documents and numbers, one answer. Contextual AI grounds the clause, the spec and the runbook. Data Workers supplies the metric with its definition, owner, lineage, quality score and freshness.
  • •A grounded answer needs a true input. The GLM is built to prioritize the context it is given. Data Workers checks that the numbers in that context are right before an agent quotes them.
  • •Fixes land upstream, behind approvals. Your datastores stay with your Contextual AI admins. Approved fixes land in dbt and the pipelines, with a blast radius, a named approver, a rollback path and a receipt.
  • •Start with a pilot. One Agent Composer workflow, read tools first, one metric domain, then one fix class, on the ladder from L0 manual to L4 autonomous.

Contextual AI is the grounded-answer layer. Data Workers is the context and crew for the data estate.

Contextual AI does the hardest part of document AI well: parsing complex files, retrieving the few passages that matter, reranking them by your instructions on recency, source or document type, and generating an answer faithful to what it retrieved, with citations a reviewer can follow. Agent Composer adds multi-step reasoning and tools, including QueryStructuredDatastoreStep for "questions about metrics, statistics, or tabular data" and MCPClientStep for calling any MCP server.

The data leader's next question is which definition the agent used, whether the table was fresh, and who owns it. Those answers live in dbt, the warehouse, the orchestrator and the semantic layer, and Data Workers covers that side.

Here is a morning with both connected. This is an illustration, not a customer case.

TimeSystemWhat happens
22:30NetSuiteOperations adds a new fulfillment status, Partially Shipped, for split orders
23:15FivetranThe sync lands fulfillment records with the new status in Snowflake
01:00AirflowThe nightly DAG run succeeds; dbt's fct_shipments counts any unknown status as late, so on_time_rate drops from 98.4% to 95.1% for accounts with split orders
01:20Data WorkersThe accepted-values check on status fails and the rate shifts outside its range; Data Workers traces lineage to the NetSuite change and opens an incident with the metric owner
01:25Data Workerson_time_rate is marked under incident, so any agent that reads it gets the flag with the number
08:40Contextual AIAn account manager asks the SLA agent whether a key customer is within its 97% on-time commitment this quarter
08:41Contextual AI + Data WorkersThe agent retrieves and cites the SLA clause from the contract in SharePoint; its MCPClientStep tools call resolve_metric for the governed definition and owner, explain_table for definition and lineage, get_quality_score for quality and get_incident_history for the open incident, and the answer says the rate is under repair
08:50Data WorkersData Workers proposes a dbt diff mapping Partially Shipped to its on-time rule, with its blast radius: three models, one Looker Explore and two Contextual AI workflows that read the metric
09:30SpellbookThe analytics engineer who owns on_time_rate reviews the diff and approves; dbt CI passes
09:45AirflowData Workers reruns the affected partitions
10:05SnowflakeThe rate is back on its monitor_metrics baseline and the owner confirms it against NetSuite shipment dates; Data Workers clears the flag and writes the receipt
10:10Contextual AIThe account manager asks again; the answer cites the clause and the verified 98.3% rate
Incident timeline across the stack: what Contextual AI, your team and Data Workers each do, step by step

Contextual AI did its job: it found the right clause, kept the account manager inside the documents they are entitled to see, and grounded the answer. The number broke upstream, in systems a document platform doesn't run. Data Workers caught it at 01:20, the agent held the number, and the fix went through an owner's approval.

JobWhat Contextual AI doesWhat Data Workers does
The documentsSyncs, parses, chunks and indexes them, with source entitlements enforcedRecords the definitions those documents set as governed rules, with author and source
The retrievalHybrid search, instruction-following rerank and attributionsServes the metric's definition, owner, lineage, quality score and freshness
The answerGenerates with the GLM, faithful to the retrieved contextMakes sure the structured context it is given is correct and current
The evaluationTests answers with LMUnit and collects annotated feedbackTests the data with quality checks, freshness and volume monitors
The breakAnswers from what its datastores and tools returnDetects the upstream change and traces it across NetSuite, Fivetran, dbt and Airflow
The fixUpdates its own datastores and agent configs through its APIsProposes the fix with its blast radius, routes it to a named owner, applies it reversibly and writes a receipt

Why doesn't Contextual AI just do this itself?

Because Contextual AI built its platform for one demanding job, expert answers over complex documents, and its design is right for that job. The GLM is built, in Contextual AI's words, "to prioritize faithfulness to in-context retrievals over parametric knowledge." That is exactly what you want from a model that cites a contract, and it means the model trusts the number a tool hands it. Checking that number is a separate job.

That separate job is a different product. It needs metric definitions reconciled across dbt, semantic layers and BI tools; quality and freshness checks on warehouse tables; lineage from the ERP through ingestion and transformation; blast-radius scoping, owner approvals and rollback for every change; and liability for changes inside Snowflake, dbt and Airflow, which Contextual AI doesn't run. Its write paths sensibly cover its own estate: datastores, documents, chunks and agent configurations. Agent Composer gives you the hook: MCPClientStep calls an external MCP server with encrypted credentials. Data Workers is the server on the other side, and it knows what a fix will touch and how to undo it.

Every tool owns a slice. Data Workers covers the whole lifecycle

Contextual AI owns one slice of the data lifecycle, and owns it well: grounded answers over enterprise documents. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on Contextual AI where your experts already ask.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, Contextual AI goes deep on its own area
StageData WorkersContextual AIWhy we scored it this way
Catalog & Context99.5Contextual AI's home stage for documents: datastores fed by Box, Confluence, Google Drive, OneDrive and SharePoint connectors, parsed, chunked and retrieved with citations. Data Workers' context covers the data estate: definitions, lineage, owners, quality and freshness.
Analytics & Insights87.5Agents answer from retrieved documents with attributions, and Agent Composer can add structured query steps and MCP tools. Data Workers answers metric questions through governed definitions and verified tables.
Data Quality83LMUnit tests the answer, not the data under it. Data Workers writes, runs and repairs the quality checks behind the numbers an agent quotes.
Observability & Incidents8.54The Root Cause Analysis template and example agents investigate device logs for engineering teams. Data Workers detects data pipeline breaks, traces them across systems, fixes them and verifies the result.
Pipelines & Ingestion8.55Connectors sync documents into datastores with auto and manual syncing, and Parse turns complex files into Markdown or structured blocks. Data Workers builds, reruns and backfills the warehouse pipelines behind approvals.
Schema & Migration82Warehouse schema work is outside a document platform's job. Data Workers catches upstream schema changes in the dbt manifest and in review, assesses impact and drafts each migration with rollback SQL for the owner to apply.
Governance & Access8.56Source entitlements are enforced at query time, and RBAC (Provisioned Throughput plans) scopes roles to specific agents and datastores. Data Workers proposes and applies grants on your data platforms by policy.
Security & Privacy86Strong for its own estate: entitlement-aware retrieval, encrypted secrets in Agent Composer and VPC deployment on Enterprise. Data Workers flags sensitive column names in pull request review and proposes masking for the owner.
Cost / FinOps82A usage page and billing APIs cover Contextual AI's own spend. Data Workers traces Snowflake credits to the dbt model behind them and drafts the fix for its owner.
MLOps & Models7.58.5Contextual AI's second home stage: the Grounded Language Model (GLM, v1 and v2), the instruction-following reranker, LMUnit natural-language unit tests and feedback annotation. Data Workers keeps the data under your models healthy.

How Contextual AI and Data Workers work together

Experts stay in Contextual AI agents or your own app, engineers in Claude or Cursor. Spellbook Data Catalog (in preview) is where the data team looks: each proposed change, its blast radius, who approved it and how to roll it back. Between them, Data Context Wizard keeps one governed context graph across every platform, the Data-Agents Swarm does the work with 20+ specialist agents, the Autonomous Data-Conductor runs each fix end to end (detect, diagnose, fix, review, verify, remember), and per-domain guardrails hold approvals, receipts and rollback.

How Data Workers fits with Contextual AI: your coding agent on top, Data Workers in the middle, your estate underneath

Bring your own context, both ways. Contextual AI holds the documents that set many of your definitions: the contract that defines "on time", the quality manual that sets a yield threshold. The data team records each as a governed rule with define_business_rule, with its author and source document, and an owner marks the canonical table with mark_authoritative. Every agent then reads the same approved rule next to its lineage, quality score and freshness. For the full pattern, see the hub, bring your own context.

Setup in Agent Composer. Data Workers connects to Contextual AI over its API or MCP server today, and Agent Composer calls Data Workers tools directly. The product's remote endpoint serves every Data Workers agent's tools over HTTP with API-key (Bearer) authentication, which is what MCPClientStep sends in auth_headers; Contextual AI encrypts the key on save. Custom YAML workflows are in public preview for enterprise users.

# Example: an Agentic Search workflow that can call Data Workers for metrics
research:
  type: AgenticResearchStep
  config:
    tools_config:
      - name: search_contracts
        description: Search customer contracts and SLA schedules.
        step_config:
          type: SearchUnstructuredDataStep
          config:
            top_k: 50
      - name: governed_metric
        description: |
          Look up a business metric in Data Workers. Returns the governed
          definition, formula and owner. Use for any rate, total or count.
        step_config:
          type: MCPClientStep
          config:
            server_url: "https://<your-data-workers-host>/mcp"
            tool_name: "resolve_metric"
            tool_args: '{"metricName": "$query", "customerId": "<your-tenant>"}'
            auth_headers:
              Authorization: "Bearer ${DW_API_KEY}"
      - name: table_status
        description: |
          Check the table behind a metric in Data Workers. Returns freshness,
          lineage, documentation and trust score.
        step_config:
          type: MCPClientStep
          config:
            server_url: "https://<your-data-workers-host>/mcp"
            tool_name: "explain_table"
            tool_args: '{"assetId": "$query", "customerId": "<your-tenant>"}'
            auth_headers:
              Authorization: "Bearer ${DW_API_KEY}"
      - name: table_quality
        description: Get the quality score (0-100) for a table in Data Workers.
        step_config:
          type: MCPClientStep
          config:
            server_url: "https://<your-data-workers-host>/mcp"
            tool_name: "get_quality_score"
            tool_args: '{"datasetId": "$query", "customerId": "<your-tenant>"}'
            auth_headers:
              Authorization: "Bearer ${DW_API_KEY}"
      - name: table_incidents
        description: List recent Data Workers incidents related to a table.
        step_config:
          type: MCPClientStep
          config:
            server_url: "https://<your-data-workers-host>/mcp"
            tool_name: "get_incident_history"
            tool_args: '{"similarTo": "$query", "customerId": "<your-tenant>"}'
            auth_headers:
              Authorization: "Bearer ${DW_API_KEY}"
    agent_config:
      agent_loop:
        num_turns: 10
        research_guidelines_prompt: |
          Cite the contract for terms. Use governed_metric for every number,
          and table_status, table_quality and table_incidents for its table.
          If the table has an open incident, say so and do not quote the number.
  input_mapping:
    message_history: create_message_history#message_history

In Cursor or Claude, engineers can list Contextual AI's hosted MCP server (https://mcp.app.contextual.ai/mcp/, whose query tool routes a question to a Contextual AI agent) next to Data Workers' agents, started with the open-source repo's start-agent.sh entries in the client's MCP config, as the client setup docs show.

List tools with the client's own command (for example /mcp). This guide uses resolve_metric, explain_table, trace_cross_platform_lineage and blast_radius_analysis on dw-context-catalog, get_quality_score on dw-quality, and get_incident_history, diagnose_incident and remediate on dw-incidents.

Where writes go. Your datastores and agent configurations stay under your Contextual AI admins, by design. Fixes land where your data team already reviews change: a dbt diff for the owner to merge, a rerun of the Airflow DAG, a mapping in the load job.

One request, L0 to L4. The autonomy ladder is set per domain.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. The Contextual AI agent cites the contract; an analyst looks up the rate and checks it by hand.
  • •L1 observe. Read tools only. The agent calls resolve_metric, explain_table and get_incident_history, and the answer carries the governed rate, its owner, its lineage and any open incident.
  • •L2 propose. When a check fails, Data Workers drafts the dbt diff with its blast radius. Nothing reaches production until the owner approves in Spellbook and CI passes.
  • •L3 act reversibly. For proven change classes, such as reruns and backfills of failed partitions, Data Workers applies the fix, re-runs the checks on the changed tables, with the undo recorded before it runs.
  • •L4 autonomous. For a scoped domain like freshness failures in the shipment marts, Data Workers fixes overnight, so morning answers read verified numbers.

Each step up is a per-domain decision backed by receipts. For the safety model, read is it safe to let AI agents change production data, and for where data and credentials live, read where does our data go.

The same Data Workers server serves every other context tool you run. See you're on Pinecone for vector search, you're on TrustGraph for open-source GraphRAG, and you're on Glean for enterprise search.

What changes for your team

Contextual AI gave your experts grounded answers from thousands of documents. Data Workers gives the data team a crew, so expert questions that reach a table don't become a queue of "can I trust this number?" tickets.

Six jobs that run on autopilot with Data Workers next to Contextual AI, with a concrete example of each
  • •Incidents. A status change that would skew a metric is caught, traced and fixed before an agent quotes it.
  • •Data quality. Every metric an agent cites gets accepted-value, volume and freshness checks, and every break that reached an answer becomes a test.
  • •Cloud spend. Cleanups of staging tables built for retired agent experiments are proposed to the owner after a dependency check.
  • •Access. A request for a governed table becomes a time-boxed grant proposal to its owner.
  • •Audits. Contextual AI's attributions show which documents grounded an answer; Data Workers' receipts show who changed the numbers in it and how to undo it.
  • •Migrations. A warehouse move runs in approved, parity-checked waves while agents keep reading the same definitions.

Keep Contextual AI, or consolidate?

Keep Contextual AI if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

Most teams keep Contextual AI: your datastores, workflows and evaluations are tuned, and Data Workers adds the slice those workflows call out to. Where teams consolidate, it is usually a separate data catalog, a metric glossary kept in a spreadsheet, or a standalone observability tool, now that Data Workers runs those jobs. Weighing building this layer yourself? Read build it ourselves with Claude Code and MCP servers: the MCP endpoint is the easy part; the context graph, approvals and rollback are where the work is.

The case for your CFO

The outcome: Contextual AI gets experts answers from complex documents fast. Data Workers makes the numbers inside those answers correct, current and auditable, and repairs the data when it isn't. An SLA credit or a warranty reserve decided from an agent's answer rests on the table behind it.

The risk story: Contextual AI's admins decide which workflows call Data Workers. Data Workers sets autonomy per domain from L0 manual to L4 autonomous, routes each change to a named approver, applies it reversibly, verifies it downstream and writes a receipt: who approved it, what it touched, how to undo it. There is zero migration: your warehouse, dbt project, orchestration, BI and Contextual AI stay where they are.

Why now: expert agents are moving from pilots to everyday use, and more questions mix a document with a number. A grounded model faithfully repeats the number it is handed, so a broken table reaches every expert who asks. The first win is one Agent Composer workflow calling read-only Data Workers tools for one metric domain, so every number arrives with an owner, a load-lag check and an incident flag. What stays the same: your Contextual AI workspace, datastores, entitlements, warehouse grants and dbt review process. For the numbers, see the ROI of agentic data operations. Start with a pilot; the path is on the pricing page, and the pilot is credited in full against the first year.

The sentence to repeat upstairs: "Contextual AI grounds the answer in our documents; Data Workers makes sure the numbers in that answer are right, and fixes them with an approval and a receipt when they aren't."

Getting started

Start with a pilot. Pick one Contextual AI workflow whose answers already quote numbers, such as SLA or warranty questions, add Data Workers as an MCP tool with read tools only, record the definitions your contracts set, and run it for a few weeks before enabling the first fix class. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.

FAQ

Doesn't Contextual AI already query structured data? Yes. Agent Composer includes structured query steps for metrics and tabular data, Contextual AI is available as a Snowflake Native App, and Agent Composer can call any MCP server. It works from the data and definitions it is given. Data Workers keeps those correct: it holds the governed definition, checks quality and load lag against baselines, and fixes and verifies breaks.

Will Data Workers change our datastores or agents? No. Datastores, documents, chunks and agent configurations stay under your Contextual AI admins. Data Workers reads definitions and context and lands fixes in dbt, the orchestrator or the load job, each with an approver and a rollback path.

How does this help with hallucinations? Contextual AI's GLM and attributions keep answers faithful to what was retrieved. Data Workers addresses a different cause of wrong answers: a correct citation of a wrong number. The metric arrives with its definition, owner and freshness, and a flag when it is under repair.

Can LMUnit and Data Workers checks work together? Yes. LMUnit tests whether an answer meets your criteria, such as "cites the SLA clause". Data Workers tests the tables behind the answer, so a failing LMUnit test that traces to a data break becomes an incident with an owner.

Whose credentials does Data Workers use? Agent Composer authenticates with the key stored (encrypted) in the workflow. Data Workers acts on the warehouse, dbt and orchestration with the credentials you set per connection, scoped to each domain; warehouse permissions stay the system of record, by design.

Sources

All links checked Oct 2, 2026.

  • •Contextual AI, homepage, https://contextual.ai/
  • •Contextual AI blog, Introducing Agent Composer (Jan 26, 2026), https://contextual.ai/blog/introducing-agent-composer
  • •Contextual AI blog, Semantic Layer vs. Context Layer: Why Enterprise AI Needs Both (Mar 12, 2026), https://contextual.ai/blog/semantic-layer-vs-context-layer
  • •Contextual AI docs, 2026 release updates, https://docs.contextual.ai/release-notes/2026
  • •Contextual AI docs, Why Contextual AI, https://docs.contextual.ai/quickstarts/why-contextual-ai
  • •Contextual AI docs, Agent Composer overview, https://docs.contextual.ai/reference/ac-overview
  • •Contextual AI docs, Agent Composer MCP integration, https://docs.contextual.ai/reference/ac-mcp-integration
  • •Contextual AI docs, Agent Composer tools configuration, https://docs.contextual.ai/reference/ac-tools-config
  • •Contextual AI docs, Agent Composer YAML step reference, https://docs.contextual.ai/how-to-guides/ac-yaml-reference
  • •Contextual AI docs, Rerank and Parse, https://docs.contextual.ai/api-reference/rerank/rerank and https://docs.contextual.ai/quickstarts/parse
  • •Contextual AI docs, Databases integration (Snowflake Native App), https://docs.contextual.ai/integration/databases
  • •Contextual AI docs, Contextual AI MCP Server quickstart, https://docs.contextual.ai/quickstarts/mcp-server
  • •Contextual AI docs, Generate (Grounded Language Model), https://docs.contextual.ai/how-to-guides/generate
  • •Contextual AI API reference, Generate, Rerank, Parse, LMUnit, https://docs.contextual.ai/api-reference/openapi.json
  • •Contextual AI docs, LMUnit, https://docs.contextual.ai/how-to-guides/lmunit
  • •Contextual AI docs, Entitlements enforcement, https://docs.contextual.ai/connectors/entitlements-enforcement
  • •Contextual AI docs, Role-based access control, https://docs.contextual.ai/admin-setup/rbac
  • •Contextual AI docs, Pricing and billing, https://docs.contextual.ai/admin-setup/pricing-billing
  • •Contextual AI docs, Collecting and annotating feedback, https://docs.contextual.ai/how-to-guides/feedback
  • •Data Workers, client setup, https://dataworkers.io/opensource-docs/client-setup/
  • •Data Workers open-source repository, tool registrations, https://github.com/DataWorkersProject/dataworkers-claw-community