Which model does Data Workers use, and what does it cost to run?
Data Workers is model-agnostic: bring Claude, GPT, Gemini or a local model on Ollama or vLLM and pay your provider directly with no markup. Here is what drives model spend, with a worked example at October 2026 list prices.
Data Workers uses the model you choose: Claude, GPT or Gemini through their APIs or your cloud tenant (Amazon Bedrock, Azure OpenAI, Google Cloud's Gemini Enterprise Agent Platform, formerly Vertex AI), or an open model you run yourself on Ollama or vLLM. You pay your model provider directly with no markup, and the Data Workers platform fee is flat: a $7,500 pilot credited in full against the first year, then Scale from $1,000/month or Enterprise from $3,000/month billed annually, with unlimited seats and no usage meter.
Data Workers is the agentic data platform, built so the model is a choice you make per task. Below: the providers it runs on, what drives the model bill, and a worked example at this month's list prices that finance can check.
Key takeaways
- •Bring your own model. Anthropic, OpenAI, Azure OpenAI, Amazon Bedrock, Google Cloud (Vertex AI, now Gemini Enterprise Agent Platform), Ollama or vLLM. Switch with configuration, not a migration.
- •Two bills, both predictable. Your model provider bills tokens at its own list or negotiated price; Data Workers bills a flat platform fee with no meter and no markup. Automating more next quarter changes only the first.
- •Most steps run as code, not as model calls. Lineage walks, blast-radius checks, quality scores, pattern-based PII checks and search embeddings run deterministically; the model is called for reasoning, and several tools escalate to it only when rule-based confidence is low.
- •You control the spend levers. Route routine steps to a small or local model, keep hard diagnosis on a frontier model, keep the built-in budget guards on, and see spend by provider, model and agent.
- •At October 2026 list prices, a mid-size month of agent work costs from about $3 to about $160 in model tokens in our illustration below, depending on the model; ten times that volume stays in the hundreds of dollars on a mid-tier model. Your own numbers come from your pilot.
Which models Data Workers runs on
We describe this from the product itself, the data-workers-agent-swarm repository, and the public Data Workers repository.
Providers. The LLM layer ships adapters for Anthropic (Claude), OpenAI (GPT), Azure OpenAI (now part of Microsoft Foundry), Amazon Bedrock, Google Cloud's Gemini Enterprise Agent Platform (formerly Vertex AI, for Gemini), Ollama and vLLM, plus an OpenAI-compatible adapter for local endpoints that speak that API. You pick the provider and model with environment settings: DW_LLM_PROVIDER and DW_LLM_MODEL for the agents' chat model, and the provider's own settings (an AWS region, a Google Cloud project, an Azure OpenAI endpoint) for the cloud adapters. Bedrock, Azure OpenAI and Google Cloud keep calls inside your own cloud account or tenant.
Sovereign mode. Turn on sovereign mode and Data Workers blocks every outbound call to an external model API and routes requests to your local Ollama or vLLM endpoint instead. That is the setting teams use for Data Workers in your VPC or air-gapped, and it makes the token bill zero: you pay for the hardware the model runs on.
Routing and fallback. A model router lets you send a type of request to a preferred model and name a fallback model for it. A provider fallback chain with a circuit breaker moves traffic to the next provider when one is down, so an outage at one vendor does not stop the agents.
When you work through an assistant. You can drive Data Workers from Claude, ChatGPT, Gemini or a coding agent over MCP. In that setup the assistant's model does the conversation and the reasoning on the plan you already pay for, and Data Workers tools such as trace_cross_platform_lineage or blast_radius_analysis run as code against your context graph. The few tools that escalate to a model use the key you configured for Data Workers. AI assistants rolled out, now what? covers that pattern.
What you pay, and to whom

There are two lines on the budget, and they move for different reasons.
The platform fee is the rate card on /pricing/: a $7,500 pilot, credited in full against the first year; Scale from $1,000/month and Enterprise from $3,000/month, billed annually; unlimited seats; Apache 2.0 core. There are no credits, runs, tasks or units, so the fee does not rise when you turn on another domain. Our post on pricing without a usage meter explains the reasoning.
The model bill goes from your provider to you at your provider's price, including any discount, committed-use deal or cloud credit you already have. Data Workers adds no markup. Every model call comes from the agents in your infrastructure, on your key; the Autonomous Data-Conductor, which we host as part of the platform fee, orchestrates from workflow metadata and makes no model calls of its own.
What drives model spend
Model spend is tokens times price, so the bill comes down to how many agent runs you do, how many of a run's steps call a model, how much context each call carries, and which model it calls.

Deterministic code first. Most of what Data Workers does on an incident is graph and SQL work. trace_cross_platform_lineage walks lineage, blast_radius_analysis computes the downstream assets and owners a change touches, get_quality_score reads check results and run_quality_check runs plain SQL. None of those needs a model call. The pull request privacy check reads column names and annotations as code, with no model call. Search embeddings run on-device with an ONNX model, with no cloud call. Other tools escalate to a model the same way, only when local logic is unsure: pipeline generation calls a model only when its parser's confidence falls below 0.8, and SQL dialect translation only when rule-based confidence falls below 0.7.
The relevant slice of context. The context graph assembles what a step needs (the failing model, its upstream sources, recent changes, owners) rather than pasting the estate into a prompt. Smaller prompts are cheaper prompts. Most of what a call carries is structure: schema, SQL, error text and the request. When a step needs values, such as up to five samples to classify an unclear column, they go to the model account you control; the public LLM data disclosure lists what is sent, and where does our data go? covers each boundary.
Caching. Providers price repeated prompt prefixes far below fresh input: Anthropic charges 10% of the input price for a cache hit on most Claude models (5% on Claude Opus 5.5), OpenAI lists cached input at $0.10 per million tokens for gpt-6.1-sol, and Google lists Gemini 3.8 Flash caching at $0.075. Data Workers generates its orchestration system prompts deterministically, so the same response spec always produces byte-identical prompt text, which is what a prompt cache needs.
Budget guards. The pipeline agent checks cumulative spend against a per-request budget (default $0.10) before calling a model; catalog enrichment runs stop at a per-run dollar ceiling (default $0.10) and a table cap; and a cost dashboard breaks spend down by provider, model and agent.
Autonomy changes volume, not price. As a domain moves up the ladder, agents close more of the work themselves, so runs per month rise.

That is the point of the product, and it is why the platform fee is flat. Autonomy levels L0 to L4 explained covers when a domain moves up.
A worked cost example
This is an illustration, not a measured Data Workers result. Every token count below is an assumption we chose to make the arithmetic visible; your pilot measures the real numbers from your own runs. Prices are the providers' public list prices on their own pricing pages, checked October 2, 2026, standard (non-batch) tier.
The estate: Snowflake, dbt, Airflow, Fivetran and Looker, with a data team that wants agents on three kinds of work.
| Workload (assumed) | Runs per month | Model calls per run | Input tokens per call | Output tokens per call | Input per month | Output per month |
|---|---|---|---|---|---|---|
| Incidents and failed loads (a Fivetran sync gap, a failed Airflow DAG run, a dbt test failure, a Looker tile that drifted) | 300 | 6 | 8,000 | 800 | 14.4M | 1.44M |
| Access and data requests | 400 | 3 | 4,000 | 400 | 4.8M | 0.48M |
| Pipeline and SQL drafting | 150 | 4 | 10,000 | 2,000 | 6.0M | 1.2M |
| Total | 850 | 25.2M | 3.12M |
Lineage walks, blast-radius checks, quality reads and privacy checks in those runs are code and add no tokens. The cached column assumes 60% of input tokens are served from the provider's prompt cache; cache-write and cache-storage charges are left out because they are small at this volume. Results are rounded to the dollar.
| Model (list price per million tokens, input / cached input / output) | No caching, per month | 60% of input cached, per month |
|---|---|---|
| Claude Opus 5.5 ($4 / $0.20 / $20) | $163 | $106 |
| Claude Sonnet 5.5 ($2 / $0.20 / $10) | $82 | $54 |
| OpenAI gpt-6.1-sol ($2 / $0.10 / $10) | $82 | $53 |
| Claude Haiku 4.5 ($1 / $0.10 / $5) | $41 | $27 |
| Gemini 3.8 Flash, global endpoint ($0.75 / $0.075 / $3.75, through Dec 31, 2026) | $31 | $20 |
| OpenAI gpt-6-luna ($0.10 / $0.01 / $0.50) | $4 | $3 |
| Open model on your Ollama or vLLM server | $0 in tokens | $0 in tokens |
How to read it. Each figure is input tokens times the input price plus output tokens times the output price; for Claude Sonnet 5.5 uncached that is 25.2 × $2 + 3.12 × $10 = $81.60. If your runs carry ten times the context or make ten times the calls, multiply by ten: Claude Sonnet 5.5 with caching becomes about $544 a month. Two notes from the providers' own pages: Anthropic says Claude 4.7 and later models use a tokenizer that produces about 30% more tokens for the same text, and Google's Gemini 3.8 Flash price doubles to $1.50 / $7.50 on January 1, 2027, on both the Gemini API and Agent Platform (non-global endpoints on Agent Platform list about 10% higher). Routing raises the savings further: send access requests and drafting to a small model and keep incident diagnosis on a frontier model.
The ROI calculator puts these numbers next to the hours your team gets back, and the ROI of agentic data operations shows what to baseline.
The alternatives buyers weigh
Build it yourself. A coding agent plus vendor MCP servers gets you model choice too, and you pay the same token prices. The cost that grows is engineering time: lineage across engines, approvals, rollback, receipts and evaluation. Build it ourselves with Claude Code and MCP servers? walks through that build.
Platform-native agents. Snowflake bills AI Functions, Cortex Agents and Snowflake CoWork in AI Credits per million tokens, and Databricks bills its Foundation Model APIs in DBUs per million tokens or per hour of provisioned throughput, according to each vendor's pricing pages this month. That is a sensible design for agents that live inside one platform, and it puts model use on that platform's meter. Data Workers keeps model spend on your provider's bill and works across Snowflake, Databricks and the rest of the estate at once.
Coding agents and assistants. Claude Code, Cursor, Codex and enterprise assistants bill on their own plans. They stay the best place to write code and to chat; connected to Data Workers over MCP, they add governed context and approved changes without a second model contract.
The case for your CFO
The outcome is fewer engineer hours on tickets and incidents at a cost finance can forecast, split into two lines with clear owners. The platform fee is flat and annual, with unlimited seats, so automating another domain next quarter does not change it. The model bill goes straight from the provider you already chose to you, at the price procurement already negotiated, with no markup in between.
The risk story is the same one your security team asks about. Agents work inside approvals set per domain on the ladder from L0 manual to L4 autonomous, every change carries a rollback path and a tamper-evident receipt, and nothing migrates: your warehouse, dbt project and orchestrator stay where they are. If data residency matters, sovereign mode keeps every model call on a local endpoint. Is it safe to let AI agents change production data? and where does our data go? cover both.
Why now: model prices keep falling, and a platform that switches models with configuration captures each cut the day it lands.
A strong first win is failed loads and reruns in one domain, observe-only first, then proposals. Start with a pilot on /pricing/; the pilot is credited in full against the first year, and it gives you measured token use per run instead of our assumptions.
The sentence for upstairs: "We pay a flat fee for the platform and our existing model provider for tokens, with no markup, and we can move to a cheaper model whenever one is good enough."
FAQ
Which model do you recommend? Start with a mid-tier frontier model (Claude Sonnet, GPT or Gemini Flash class) for incident diagnosis and a smaller model for routine requests, then compare cost per closed item in the pilot. The choice is yours, and you can change it per request type.
Can we use our existing Azure OpenAI, Bedrock or Google Cloud agreement? Yes. Data Workers has adapters for Azure OpenAI, Amazon Bedrock and Google Cloud's Gemini Enterprise Agent Platform (formerly Vertex AI), so calls run in your tenant and bill against the agreement and any committed spend you already have.
Can Data Workers run with no external model calls at all? Yes. Sovereign mode blocks external model APIs in code and routes every request to your Ollama or vLLM endpoint; if no local model is configured, inference stops rather than falling back to an external provider. It suits VPC deployments and Enterprise on-premise air-gapped installs.
Does Data Workers send our table data to the model? Mostly it sends structure: schema, SQL text, error messages and natural-language requests, when a step needs them. When a task needs values, such as a query result or up to five samples to classify an unclear column, they go only to the model account you hold the key for. In sovereign mode that sample pass is skipped, so samples never leave the box.
How do we keep model spend from running away? Keep the default budget guards on (a per-request budget on pipeline generation and a per-run dollar ceiling on catalog enrichment), route routine steps to a small or local model, and watch the cost dashboard by agent. Runs rise as you automate more; cost per closed item should fall.
Does using Data Workers from Claude or ChatGPT cost extra? The assistant's model runs on the plan you already pay for. Most Data Workers tools called over MCP run as code, and the few that escalate to a model bill to your own provider key, so the platform fee does not change and there is no second model contract.
Sources
- •Data Workers product repository,
data-workers-agent-swarm@ 871ae3df (Sep 24, 2026): core/llm-provider/src (provider adapters, llm-provider-factory, agent-migration, sovereign-mode, model-router, provider-fallback, vercel-ai-adapter), agents/dw-conductor/src (dtp.ts, backends.ts: no model calls, no warehouse connector), agents/dw-governance/src/pii-scanner.ts (pass 3 model classification, skipped in sovereign mode), core/metering/src (llm-cost-dashboard), agents/dw-pipelines/src/tools/generate-pipeline.ts, agents/dw-migration/src/tools/translate-sql.ts, agents/dw-context-catalog/src (onnx-embedding-backend, codex-enrichment/budget-guard), agents/dw-orchestration/src/structured-response.ts. Checked Oct 2, 2026. - •Data Workers public repository and LLM data disclosure: https://github.com/DataWorkersProject/dataworkers-claw-community/blob/main/docs/LLM-DATA-DISCLOSURE.md (checked Oct 2, 2026)
- •Data Workers pricing, rate card as published Sep 10, 2026: https://dataworkers.io/pricing/ (checked Oct 2, 2026)
- •Anthropic, Claude pricing: https://platform.claude.com/docs/en/about-claude/pricing (checked Oct 2, 2026)
- •OpenAI, API pricing: https://developers.openai.com/api/docs/pricing (checked Oct 2, 2026)
- •Google, Gemini Developer API pricing: https://ai.google.dev/gemini-api/docs/pricing (updated Oct 1, 2026; checked Oct 2, 2026)
- •Google Cloud, Gemini Enterprise Agent Platform (formerly Vertex AI) pricing: https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing (checked Oct 2, 2026)
- •Google Cloud, Gemini Enterprise Agent Platform (formerly Vertex AI) product page: https://cloud.google.com/products/gemini-enterprise-agent-platform (checked Oct 2, 2026)
- •Microsoft, Azure OpenAI in Microsoft Foundry Models: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure (updated Sep 23, 2026; checked Oct 2, 2026)
- •Snowflake, AI pricing and billing: https://docs.snowflake.com/en/user-guide/snowflake-cortex/pricing (checked Oct 2, 2026)
- •Databricks, Foundation Model Serving pricing: https://www.databricks.com/product/pricing/foundation-model-serving (checked Oct 2, 2026)