You're on Starburst: It Queries Data Where It Lives. Data Workers Owns Whether the Data Behind Those Queries Is Right
Starburst federates queries across lakes and databases, and AIDA answers from data products. Data Workers traces a wrong number in a data product to its source and gets the fix approved.
Your platform team runs Starburst. Catalogs point Trino at Iceberg and Hive tables on object storage, Postgres, Oracle, Snowflake and a dozen other sources, and one SQL query joins them without copying a row. On Starburst Galaxy, the managed service, clusters autoscale, Icehouse ingests from Kafka and files into Iceberg, and Icehouse LakeOps (public preview since August 10, 2026) compacts files and expires snapshots. Self-managed teams run Starburst Enterprise, now on the 482-e LTS line announced September 24, 2026. Curated tables become data products with owners, business context and domains, and AIDA, Starburst's AI Data Assistant (generally available in Galaxy since May 28, 2026), answers plain-language questions from those data products, next to the Starburst AI Agent chatbot (public preview since February 17, 2026).
Starburst is excellent at its job: fast, governed queries across every source you have. Whether the data behind those queries is right is a different job. A Postgres table changes shape, a dbt model on Galaxy joins it the old way, and a data product, an AIDA answer and a capacity dashboard all report a wrong number while every query succeeds. Data Workers watches the data behind your data products, traces a wrong number to its cause, and gets the fix approved before a planning decision rests on it.
Key takeaways
- •Starburst keeps its job. Catalogs, clusters, data products, AIDA and access policy stay with your platform team. Data Workers works on the data those queries read and on the changes that move it.
- •Connected from day one. Starburst connects over its API or MCP server today, and Data Workers reads the systems around it natively: PostgreSQL, dbt, Tableau and Datadog.
- •Wrong numbers, traced to the source. When a source change leaves a data product wrong, Data Workers raises one incident with the cause, the blast radius and the fix proposed.
- •Every fix goes through a named person. The owner approves the change and merges it; data product refreshes and access policy stay with your team.
- •Autonomy is set per domain. Start at L1 observe, move to L2 propose, and open L3 act reversibly for narrow classes once the record supports it.
Starburst is the engine that queries data where it lives. Data Workers owns whether the data behind those queries is right.
One week at a regional mobile operator (an illustration, not a customer case).
The operator runs Starburst Galaxy. The catalog lake reads Iceberg tables on S3 registered in AWS Glue, including usage.data_sessions, which network mediation fills every night. The catalog billing_pg reads the billing team's Postgres database. A dbt Cloud job, usage_nightly, builds analytics.fct_daily_usage_by_plan on Galaxy by joining sessions to billing.rate_plans on rate_plan_code. The Network Usage data product publishes a view and a materialized view over that table, AIDA answers questions from it, and the Tableau workbook "5G Capacity Plan" reads it for Thursday's capacity review, where the network team sets next quarter's site spend. The dbt project declares the workbook and the data product as exposures, and the team has recorded both in the context graph.
| Time | System | What happens |
|---|---|---|
| Mon 17:40 | Postgres | The pricing team ships effective-dated plans for an October price change. billing.rate_plans gains valid_from and valid_to, and three 5G plans now have two rows each: the old price ending September 30 and the new one from October 1. rate_plan_code is no longer unique |
| Tue 01:00 | Iceberg + AWS Glue | Mediation lands Monday's sessions in usage.data_sessions, on time and complete |
| Tue 02:15 | dbt Cloud + Starburst Galaxy | usage_nightly rebuilds fct_daily_usage_by_plan. The join on rate_plan_code now matches two plan rows for every subscriber on the three plans, so 386,000 subscribers get two rows each. The not-null tests pass; nobody had a uniqueness test on the table's grain |
| Tue 02:40 | Starburst Galaxy + Datadog | The data product's materialized view refreshes. Galaxy's quality check on it (rows present, no null subscriber IDs) passes, and the green result reaches Datadog |
| Tue 06:00 | Data Workers | The table's daily row count, a metric the team records with monitor_metrics, reads 14% above its baseline. Data Workers opens an incident |
| Tue 06:08 | Data Workers + Postgres | run_quality_check on billing.rate_plans, read from Postgres natively, finds rate_plan_code no longer unique, and the pricing team's merged pull request from Monday 17:40 added valid_from and valid_to. diagnose_incident ranks that change first. trace_cross_platform_lineage follows rate_plan_code from Postgres through the dbt model, and blast_radius_analysis returns the two recorded readers: the Network Usage data product and the Tableau workbook due Thursday |
| Tue 06:15 | Claude Code + Starburst MCP | The on-call data engineer asks in Claude Code. Her client calls Starburst Galaxy's own read-only MCP server, and a read-only count confirms 386,000 subscribers with two rows each |
| Tue 06:20 | Data Workers + Slack | Data Workers proposes a diff for the owner to merge: join plans on rate_plan_code with the session date between valid_from and valid_to, plus a dbt uniqueness test on subscriber and date. The approval request goes to the usage analytics lead who owns the model |
| Tue 08:10 | Starburst AIDA | A network VP asks AIDA, with the executive persona, how 5G usage moved this week. AIDA answers from the data product: up 96% on the three plans. It answered correctly from the data it was given |
| Tue 08:30 | Spellbook + GitHub | The owner reviews the diff, the blast radius and the rollback (revert the commit and rerun), approves, and merges |
| Tue 08:36 | dbt Cloud | The team's scheduler queues the approved usage_nightly run, and Data Workers records the approval. dbt Cloud finishes it at 08:58 and the new uniqueness test passes |
| Tue 09:05 | Starburst Galaxy | The owner refreshes the materialized view, then reads AIDA's agent logs to find every answer that used the data product since 02:40, and sends the VP the corrected figure |
| Tue 09:20 | Data Workers | The row-count metric is back on baseline. The receipt records the cause, the approval, the run, the checks and the undo |
| Thu 10:00 | Tableau | The capacity review runs on corrected numbers |

Every part of the stack did its job: Starburst ran each query correctly, the quality check tested what it was written to test, and AIDA answered faithfully from a governed data product. Catching the problem took knowledge outside any one engine: a billing table changed shape, a model joins it on a key that stopped being unique, and a Thursday decision reads the result.
| Job | What Starburst does | What Data Workers does |
|---|---|---|
| The query | Federates SQL across lakes, warehouses and databases through catalogs, on Galaxy or Enterprise | Connects to Starburst over its API and reads the sources, models and dashboards around it |
| The meaning | Data products carry owners, business context and domains; AIDA answers from them | Keeps one governed context graph from each source column through dbt to every data product and dashboard |
| The checks | Runs the SQL quality checks the team writes and exports results to Datadog | Holds key metrics to a baseline and ties a jump to the schema change or release behind it |
| The fix | Runs whatever SQL, refresh or policy the owner chooses | Proposes the change as a diff to a named owner, then queues the approved run through the orchestrator, or onto the team's dbt Cloud schedule |
| The proof | Logs queries and AIDA interactions | Re-checks the data and writes a receipt: what changed, who approved it, how it was checked, how to undo it |
Why doesn't Starburst just do this itself?
Because Starburst is a query engine by design, and that focus is why platform teams trust it in front of every source. Its job is to answer the query fast and enforce who may run it. Whether a Postgres table's grain changed under a dbt model is a question about the data's meaning, which a federation layer rightly leaves to each source's owners.
Starburst's AI follows the same scope. AIDA answers from the data product you select, using LLMs on a Starburst-managed Amazon Bedrock deployment. Galaxy's hosted MCP server (public preview since January 26, 2026) and the integrated MCP server in Starburst Enterprise accept read-only SQL only: the Enterprise server rejects INSERT, UPDATE, DELETE, MERGE, GRANT, REVOKE and DDL. Galaxy's lineage covers "only workloads that have been processed by Galaxy". Sound choices for an engine in front of systems it does not own.
Owning whether the data is right from a Postgres source through dbt, a data product, AIDA and Tableau is a different product: a context graph across every system, blast-radius scoping, named approvers, a recorded undo and receipts an auditor can read. That is Data Workers. See is it safe to let AI agents change production data.
Every tool owns a slice. Data Workers covers the whole lifecycle
Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, on the Starburst estate already there.

| Stage | Data Workers | Starburst | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 7 | Data products add owners, business context and domains, and Starburst Metastore holds table metadata. Data Workers keeps one governed context graph across Postgres, the lake, dbt, BI and the data products on top. |
| Analytics & Insights | 8 | 9 | Starburst's home stage: one SQL query across lakes, warehouses and databases, with AIDA answering questions from data products. Data Workers answers from governed definitions with lineage behind every number. |
| Data Quality | 8 | 5 | Galaxy runs SQL quality checks the team writes and exports results to Datadog. Data Workers ties a check or metric to the change that broke it, across the systems the query touches. |
| Observability & Incidents | 8.5 | 4 | Query details, failed-query views and cluster metrics show the engine's health. Data Workers diagnoses the data incident from the source change to the report, proposes the fix and verifies it. |
| Pipelines & Ingestion | 8.5 | 7 | Icehouse ingest from Kafka and files, LakeOps table upkeep and materialized views. Data Workers plans the repair and queues the reruns through the orchestrator after approval; dbt Cloud jobs run on the team's scheduler. |
| Schema & Migration | 8 | 6 | Table format migration moves Hive tables to Iceberg. Data Workers catches a source schema change first, sizes it downstream and plans larger moves in approved waves. |
| Governance & Access | 8.5 | 7 | Fine-grained RBAC and ABAC, row filters and column masks decide who can query what. Data Workers routes every data change to a named approver and records the decision. |
| Security & Privacy | 8 | 6 | Encryption, private connectivity and AI-powered classification protect the data. Data Workers leaves a receipt on every data change. |
| Cost / FinOps | 8 | 5 | Autoscaling, cluster schedules, billing views and AI token consumption pages. Data Workers attributes Snowflake credits to dbt models, totals BigQuery spend and reads AWS Cost Explorer. |
| MLOps & Models | 7.5 | 4 | AI functions and AIDA sit on top of the data. Data Workers keeps the data under models and agents healthy. |
How Starburst and Data Workers work together

Data Workers connects to Starburst over its API today. Starburst stays the place where queries run and access policy lives.
Most of what sits under and around Starburst connects natively. Sources: PostgreSQL, Snowflake and BigQuery. Around the queries: dbt and dbt Cloud, Airflow, Dagster, Tableau and Looker (read), Datadog (read), GitHub (pull-request review) and Slack, among 50+ connectors. The lake catalogs (Iceberg REST, Polaris, Nessie, Glue + Lake Formation and Hive Metastore) connect over their APIs today. Enterprise 482-e removes the packaged Hive Metastore in favor of Starburst Metastore, which adds an Iceberg REST catalog endpoint in public preview; for teams that adopt it, that endpoint connects over its API like the other lake catalogs.
Starburst's own MCP server belongs in the same client as Data Workers'. The team's client calls it for read-only queries; Data Workers' agents work from their own connectors. By design, Data Workers does not change catalogs, data products, materialized views, cluster settings or access policy: the owner runs those steps, and Data Workers proposes them, scopes what they touch and verifies afterwards.
Setup follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry to your client. Add Galaxy's MCP server per Starburst's docs, with an OAuth client holding the galaxy.mcp scope.
# Example: Data Workers agents in Claude Code, from a clone of the open-source repo
claude mcp add --scope user dw-incidents -- "$(pwd)/start-agent.sh" dw-incidents
claude mcp add --scope user dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
claude mcp add --scope user dw-schema -- "$(pwd)/start-agent.sh" dw-schema
claude mcp add --scope user dw-connectors -- "$(pwd)/start-agent.sh" dw-connectors
# Starburst Galaxy's read-only MCP server, per Starburst's docs (OAuth client with the galaxy.mcp scope)
claude mcp add --transport http --client-id "$GALAXY_OAUTH_CLIENT_ID" \
--callback-port 8765 starburst https://<account-name>.mcp.galaxy.starburst.ioList the tools with your client's own command (/mcp in Claude Code). In this incident: monitor_metrics and diagnose_incident (dw-incidents) flag the jump and rank the cause; run_quality_check (dw-quality) finds the broken grain in Postgres; trace_cross_platform_lineage and blast_radius_analysis (dw-context-catalog) map what reads the table; the approved run goes to the team's dbt Cloud scheduler with the approval on record.
The agents run in your infrastructure with your credentials and model key; your data stays in your systems and the hosted Conductor sees workflow metadata only. See where does our data go.
One incident, L0 to L4, set per domain:

- •L0 manual. Thursday's review sets site spend on 5G usage that looks twice its real size.
- •L1 observe. Data Workers flags the jump at 06:00 with Monday's plan change as the cause and both readers named. Nothing changes.
- •L2 propose. Data Workers proposes the join fix and the test as a diff; nothing changes until the named owner approves.
- •L3 act reversibly. For a class with a clean record, Data Workers moves the fix from approval to the team's dbt Cloud schedule without a second ticket, in the domain you open.
- •L4 autonomous. Once the owner's fix is merged and the scheduled rebuild runs, Data Workers re-checks the metric and posts the receipt. Data product refreshes and access policy stay with the owner at every level.
Who sets those levels: who owns the agents, how approvals work for AI data agents and autonomy levels L0 to L4 explained.
What changes for your team

- •One record for every team. The platform team, data product owners and analysts see the same incident and receipt in Spellbook Data Catalog (in preview), linked from Slack.
- •Data products get a guard on their sources. A schema change in any source a data product reads is weighed against everything on top before the next refresh.
- •AIDA answers rest on checked data. The key metrics behind the data products AIDA uses hold to baselines the team records, so a jump is explained before an executive repeats it.
- •Migrations have a plan. A Hive to Iceberg move or a legacy warehouse retirement runs in approved waves; Data Workers plans each wave with its parity checks and holds the completion gate for the owner's sign-off.
See also Data Workers vs a data catalog and Data Workers vs data observability. For agents on Trino directly, see the Claude Code Trino integration and MCP servers for Trino and Presto.
Keep Starburst, or consolidate?
Keep Starburst if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
For almost every Starburst team the answer is keep it: federation, data products, AIDA and Icehouse are why it was bought. What teams consolidate is the tooling around the data: a separate observability tool, spreadsheets of who reads which table, and runbooks that say "rerun the model and eyeball the dashboard". Related stacks: you're on Teradata, you're on Postgres, you're on dbt and Data Workers integrations.
Weighing a build on Starburst's MCP server and a coding agent? Read build it ourselves with Claude Code and MCP servers.
The case for your CFO
The outcome. When a source change leaves a data product wrong, the error is caught within hours and corrected before a planning meeting, a board pack or an AI answer repeats it, with a record of how it was checked.
The risk story. At L1 agents only read. At L2 a named person approves every change; an unanswered request expires and escalates, never auto-grants. At L3 they act on reversible changes inside the domains you open. Starburst catalogs, data products and access policy stay with your team. No agent can promote its own work, an org-wide stop halts all autonomous dispatch, and every change carries a receipt.
Why now. AIDA and MCP clients put data products in front of executives and assistants directly, so a wrong number travels further than a dashboard ever did.
The first win. L1 on the data products that feed planning: each one's key metrics checked against baseline after every source change.
What stays the same. Starburst, your catalogs, data products, dbt, Tableau and sources. Zero migration. For the numbers, see the ROI of agentic data operations.
The sentence for upstairs: "Starburst lets us query every source in one place; Data Workers makes sure the data behind those answers is right, and gets it fixed with our approval when it isn't."
Getting started
Start with a pilot. Pick the data products that feed planning, connect their sources, the dbt project and the dashboards on top, record the metrics that define "right", and run at L1 through a few weeks of source changes. Then turn on L2 for one domain. The pilot path is on the pricing page, and the pilot is credited in full against the first year.
FAQ
How does Data Workers connect to Starburst? Over Starburst's API today. The sources, dbt project, Tableau and Datadog around it connect natively; the lake catalogs connect over their APIs today.
Does Starburst have an MCP server? Yes. Galaxy has a hosted MCP server (public preview since January 26, 2026), and Starburst Enterprise has an integrated, licensed MCP server with read-only queries, data product search and administrator-defined parameterized queries. Both are read-only. Add it to the same client as Data Workers; the client calls it, and Data Workers' agents work from their own connectors.
Our Galaxy quality checks passed. How was the number wrong? A quality check tests what it is written to test, such as rows present and no null IDs; a duplicated join passes both. Data Workers holds the table's row count and key metrics to a baseline and ties a jump to the source change behind it. A uniqueness test on the grain closes the gap.
Will Data Workers change our data products, materialized views or access policies? No. Data products, refreshes, catalogs, clusters and policies stay with the owner. Data Workers proposes fixes where they live, such as a dbt model diff, and queues the approved run.
How do AIDA and Data Workers fit together? AIDA answers from the data products you select. Data Workers keeps the data under them right, and the owner can use AIDA's agent logs to find answers given during an incident window.
Can Data Workers help with a Hive to Iceberg move, or with retiring the warehouse Starburst sits beside? Yes. Data Workers maps the readers of each table, plans each migration wave with its parity checks, tracks them and holds the completion gate for the owner's sign-off. Generated schema migrations come with rollback SQL for the owner to apply.
Sources
- •Starburst, "Starburst Enterprise Delivers Incremental Iceberg Materialized Views, Query Speed improvements and Enhanced AI Intelligence" (482-e LTS announcement, Sep 24, 2026; AIDA Skills, external MCP OAuth, Starburst Metastore Iceberg REST and incremental materialized views in public preview; packaged Hive Metastore removed), https://www.starburst.io/blog/starburst-enterprise-delivers-incremental-iceberg-materialized-views-query-speed-improvements-and-enhanced-ai-intelligence/ (checked Oct 3, 2026)
- •Starburst Enterprise release notes (482-e LTS, 481-e STS), https://docs.starburst.io/latest/release.html (checked Oct 3, 2026)
- •Starburst Galaxy release notes (MCP server public preview Jan 26, 2026; Starburst AI Agent public preview Feb 17, 2026; cluster metrics and data quality checks export to Datadog Feb 18, 2026; AIDA GA May 28, 2026; guardrails Jul 13; Icehouse LakeOps public preview Aug 10; data product domains Aug 11; Agent logs Aug 17, 2026), https://docs.starburst.io/starburst-galaxy/get-started/release-notes.html (checked Oct 3, 2026)
- •Starburst Galaxy product page (five layers, 50+ data sources, RBAC/ABAC, AI-powered classification), https://www.starburst.io/platform/starburst-galaxy/ (checked Oct 3, 2026)
- •Starburst Galaxy docs, AI Data Assistant (AIDA) (answers from the selected data product; personas; Starburst-managed Bedrock deployment), https://docs.starburst.io/starburst-galaxy/starburst-ai/ai-aida.html (checked Oct 3, 2026)
- •Starburst Galaxy docs, Model Context Protocol (MCP) on Starburst Galaxy (read-only SQL; OAuth with the galaxy.mcp scope; Claude Code setup), https://docs.starburst.io/starburst-galaxy/starburst-ai/mcp-server.html (checked Oct 3, 2026)
- •Starburst Enterprise docs, Starburst MCP server (read-only query tool; searchDataProducts, getDataProductDetails; parameterized queries), https://docs.starburst.io/latest/starburst-ai/mcp-server.html (checked Oct 3, 2026)
- •Starburst Galaxy docs, Starburst AI Agent (public preview; natural language to SQL), https://docs.starburst.io/starburst-galaxy/starburst-ai/ai-agent.html (checked Oct 3, 2026)
- •Starburst Galaxy docs, Data quality (SQL boolean checks, severity, schedules), https://docs.starburst.io/starburst-galaxy/data-engineering/optimization-performance-and-quality/observability/profile-and-quality.html (checked Oct 3, 2026)
- •Starburst Galaxy docs, Data lineage (public preview; "Only workloads that have been processed by Galaxy are captured"), https://docs.starburst.io/starburst-galaxy/data-engineering/optimization-performance-and-quality/observability/data-lineage.html (checked Oct 3, 2026)
- •Starburst Galaxy docs, Icehouse LakeOps (public preview; compaction, snapshot expiration, orphan file removal), https://docs.starburst.io/starburst-galaxy/data-engineering/optimization-performance-and-quality/observability/icehouse-lakeops.html (checked Oct 3, 2026)
- •Trino releases (483, Jul 18, 2026), https://github.com/trinodb/trino/releases (checked Oct 3, 2026)
- •Data Workers open-source repository (tool registrations in dw-incidents, dw-schema, dw-context-catalog, dw-connectors;
run_quality_checkreads PostgreSQL through the native connector), https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026) - •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)