You're on ClickHouse: It Answers Real-Time Analytics Questions in Milliseconds. Data Workers Owns Whether the Data Behind Those Answers Is Right
ClickHouse answers real-time analytics in milliseconds. Data Workers catches a materialized view counting the wrong rows, traces the cause and gets the fix approved.
Your team runs ClickHouse, open source or as a ClickHouse Cloud service on AWS, Google Cloud, Azure or your own cloud through BYOC. ClickPipes lands events from Kafka, object storage or Postgres CDC in MergeTree tables, and materialized views roll them up on insert, so product dashboards, Grafana panels and internal apps aggregate billions of rows in milliseconds. You model with dbt-clickhouse and watch the service in Managed ClickStack, and since January 2026 the same company also runs ClickHouse Managed Postgres and Langfuse. Engineers query it from Claude Code or Cursor through mcp-clickhouse, read-only by default.
ClickHouse computes exactly what your SQL asks for, at a speed few engines match, which is why one class of incident is so quiet. When an upstream release changes which rows arrive, a materialized view keeps doing what it was told, and every customer-facing number built on it is wrong within seconds, with no error anywhere. Data Workers watches both sides of those views, the database they replicate and the metrics your team records from them, traces a wrong number to its cause and gets the fix approved before invoices, payouts or reports run on it.
Key takeaways
- •ClickHouse keeps its job. Services, MergeTree tables, ClickPipes, materialized views and the MCP servers stay with your platform team. Data Workers works on what they mean and what reads them.
- •Fast but wrong becomes an incident. When an upstream change makes a rollup count rows twice, Data Workers ties the anomaly to the change and lists every dashboard, job and invoice it reaches.
- •It reads the systems around ClickHouse natively. The upstream PostgreSQL, the dbt-clickhouse manifest, AWS Step Functions and Slack connect natively. ClickHouse's own MCP servers sit next to Data Workers in your team's MCP client.
- •Every fix goes through a named person. The owner approves each proposal and ships the ClickHouse change through your dbt deploy.
- •It stacks with ClickHouse's own AI. ClickHouse Agents and
mcp-clickhouseanswer questions; Data Workers keeps the data under the answers right.
ClickHouse answers real-time analytics questions in milliseconds. Data Workers owns whether the data behind them is right.
The job on top of ClickHouse is to notice that a rollup now counts something else, find the cause, size the damage, hold anything that bills on it, get the fix made where it lives and prove it. Here is one Tuesday at a food delivery marketplace, an illustration, not a customer case.
The setup: the order service writes to Amazon RDS for PostgreSQL. A ClickPipes Postgres CDC pipe replicates orders into ClickHouse Cloud, where, as ClickPipes documents, it lands as a ReplacingMergeTree table: every update arrives as a new version of the row (_peerdb_version), and background merges later keep the latest one. A dbt-clickhouse model partner_daily is an incremental materialized view: on each insert into orders with status = 'delivered', it adds the order's subtotal to a SummingMergeTree table per restaurant per day. That table feeds the "today's sales" panel in the partner portal that 3,870 restaurants use, and at 02:00 an AWS Step Functions state machine, partner_commissions, reads yesterday's totals and posts 15% commission invoices to NetSuite. Both readers are dbt exposures in the project, and the team has recorded them in the context graph.
| Time | System | What happens |
|---|---|---|
| Tue 11:00 | RDS PostgreSQL | The consumer app ships post-delivery tipping. The migration adds tip_amount numeric DEFAULT 0 to orders, and customers can add a tip for two hours after delivery, which updates the order row |
| 11:02 | ClickPipes | As documented, the added column with a supported default "propagated automatically once the table gets an insert/update/delete". Each tip update arrives in ClickHouse as a new version of a delivered order |
| 11:02 | ClickHouse Cloud | The incremental view is "an insert trigger", so it adds the subtotal again for every new delivered version. orders will merge down to one version per order; partner_daily keeps both sums. No query fails and nothing errors |
| 11:15 | Data Workers + PostgreSQL | The run_quality_check profile of orders on the replica shows a new, non-breaking column. assess_impact finds no dbt model selecting it. Data Workers logs it to the timeline |
| 12:00 to 14:00 | Partner portal | Through the lunch rush, restaurants see sales up 40 to 60%. Nobody complains about good news. Managed ClickStack shows insert latency, query p99 and errors all normal |
| 14:30 | Data Workers | GMV per delivered order, which the team records hourly from partner_daily with monitor_metrics, is 1.57 times its baseline while order counts are normal. Data Workers opens an incident |
| 14:40 | Data Workers | diagnose_incident ties the jump to the 11:00 column. trace_cross_platform_lineage follows the dbt-clickhouse manifest from orders to partner_daily; blast_radius_analysis adds the portal panel and partner_commissions, which bills at 02:00 |
| 14:45 | Slack | send_slack_alert reaches the analytics platform owner, billing operations copied, with the evidence and a version-count query to run |
| 14:55 | mcp-clickhouse | The owner runs that query read-only from her own MCP client: 21,400 of today's orders have two delivered versions. At 02:00 that is $91,800 of commission on tips counted as sales, across 3,870 restaurants |
| 15:05 | Spellbook | The owner reviews two proposals: hold tonight's commission run; and replace the incremental view with a refreshable materialized view that rebuilds the last 35 days from orders FINAL WHERE _peerdb_is_deleted = 0, shipped as a dbt-clickhouse diff with its rollback. She approves both |
| 15:10 | AWS | She disables the EventBridge schedule that starts partner_commissions |
| 15:40 | GitHub + CI | She merges the diff. CI runs dbt build against a staging service with a replayed sample of the day's CDC |
| 16:20 | ClickHouse Cloud | She deploys to production. The first refresh rebuilds partner_daily, and the portal panel drops back to real sales |
| 16:40 | Data Workers | GMV per delivered order is back on its baseline, and the comparison query it proposed, which the owner runs, shows partner_daily matching the deduplicated orders for every day in the window. The receipt records cause, approvals, deploy, checks and the undo (revert the diff) |
| 17:00 | AWS | The owner re-enables the schedule. The receipt names the window when the portal showed inflated sales, for the partner success team |
| Wed 02:00 | Step Functions + NetSuite | partner_commissions runs on schedule and invoices commission on food, not tips |

Every part of ClickHouse did its job. ReplacingMergeTree is the documented way to apply CDC updates, an insert-triggered view is what makes rollups fast, and the CDC docs recommend FINAL or a refreshable view for deduplicated results. Knowing this rollup now needed one took facts outside ClickHouse: an upstream feature made delivered orders update again, and a billing job reads the rollup at 02:00.
| Job | What ClickHouse does | What Data Workers does |
|---|---|---|
| Ingestion | ClickPipes streams Postgres CDC, Kafka and object storage into MergeTree tables, and propagates added columns | Watches the upstream schema natively, such as new columns on a Postgres replica, and asks what they mean for the views that read the table |
| Rollups | Incremental views transform each insert; refreshable views rebuild on a schedule | Checks that what a rollup produces still matches what it should count, against baselines the team records |
| Agent access | mcp-clickhouse and the Cloud remote MCP server answer questions read-only; ClickHouse Agents chat over your service | Keeps the tables and definitions behind those answers right, and traces which numbers a wrong rollup reached |
| Service health | Managed ClickStack and system tables show inserts, queries, merges and errors | Diagnoses the data incident across the source, the pipe, the view and every reader |
| The fix | Runs the view change and deploy the owner makes | Proposes the hold and the change set to a named owner, in order, with the blast radius |
| The proof | Answers the comparison query in milliseconds | Re-checks the baseline and writes a receipt: what changed, who approved it, how it was checked, how to undo it |
Why doesn't ClickHouse just do this itself?
Because ClickHouse is built to store data efficiently and answer queries fast, and it scopes its decisions to the SQL you wrote. A materialized view is a trigger on insert, by design: that is how rollups stay cheap at any scale. Whether that trigger still matches the business after an upstream release depends on the app roadmap, the commission policy and the finance calendar, which a database has no reason to hold.
ClickHouse's AI work follows the same clear line, and it is good work. mcp-clickhouse (v0.7.0, Sep 21 2026) gives agents run_query, list_databases and list_tables, and "queries run in read-only mode by default". The ClickHouse Cloud remote MCP server at mcp.clickhouse.cloud signs in with OAuth or an API key, and its run_select_query tool permits SELECT only. ClickHouse Agents, in beta in ClickHouse Cloud, lets people query a service in plain conversation; the open-source Agentic Data Stack pairs ClickHouse, the MCP server, LibreChat and Langfuse; and since 25.7, clickhouse-client and clickhouse-local generate SQL from plain language. All of it answers questions about the data ClickHouse holds, read-only by default, which is right for a database.
Owning whether the numbers are right from the order database to the invoice is a different product with a different liability: a context graph across systems, blast-radius scoping, named approvals, a recorded undo and receipts. That is Data Workers. More in is it safe to let AI agents change production data.
Every tool owns a slice. Data Workers covers the whole lifecycle
Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, on top of the ClickHouse service already there.

| Stage | Data Workers | ClickHouse | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 4 | Table and column metadata live in system tables, and agents read them over MCP. Data Workers keeps one governed context graph from the source database through ClickPipes, the MergeTree tables and views, to every reader. |
| Analytics & Insights | 8 | 9 | ClickHouse's home stage: a column store that answers aggregations over billions of rows in milliseconds, for dashboards, apps and agents. Data Workers answers data questions from governed definitions with lineage behind every number. |
| Data Quality | 8 | 4 | Engines such as ReplacingMergeTree and FINAL give the tools to deduplicate, and the docs explain when to use them. Data Workers checks that what a rollup serves still matches what it should count, against baselines the team records. |
| Observability & Incidents | 8.5 | 6 | Managed ClickStack and system tables show how the service runs: inserts, queries, merges, errors. Data Workers diagnoses the data incident across systems and verifies the fix. |
| Pipelines & Ingestion | 8.5 | 7 | ClickPipes streams Kafka, object storage and Postgres CDC into MergeTree tables, and materialized views transform on insert. Data Workers sees the upstream change first and plans the view change for the owner. |
| Schema & Migration | 8 | 3 | ClickPipes propagates added columns, and table changes run as SQL the owner writes. Data Workers notices when an added column changes what an existing view counts, and generates migrations with rollback SQL for the owner. |
| Governance & Access | 8.5 | 4 | Users, roles, row policies and grants govern who reads what. Data Workers routes every data change to a named approver. |
| Security & Privacy | 8 | 5 | Cloud adds SSO, private networking and encryption, and the MCP servers default to read-only. Data Workers leaves a receipt on every data change. |
| Cost / FinOps | 8 | 5 | Cloud usage and query logs help size services. Data Workers traces Snowflake spend to the dbt model and reads BigQuery and AWS spend. |
| MLOps & Models | 7.5 | 4 | Vector search, chDB and Langfuse serve AI apps and LLM tracing. Data Workers keeps the data under models and agents healthy. |
How ClickHouse and Data Workers work together

Data Workers reads the systems around ClickHouse natively: the upstream PostgreSQL databases ClickPipes replicates, the dbt-clickhouse manifest (dbt is a native connector, so sources, models and materialized views arrive as lineage), the orchestrators that read ClickHouse, such as Airflow, Dagster, Prefect and Step Functions, Kafka topics and Schema Registry on the way in, and Slack, among 50+ connectors. Rollup health comes from metrics your team records with Data Workers, such as GMV per order each hour. ClickHouse itself serves your apps and your team's MCP client through its own MCP servers and HTTP interface. On AWS, see Data Workers on AWS.
Services, tables, ClickPipes, views, roles and every deploy stay with the owner, by design, in SQL, dbt-clickhouse or Terraform. Data Workers proposes the change as a diff for the owner to merge and verifies afterwards; data cleanups, such as removing duplicate rows, are proposed for the owner to apply.
Both sit in the same client: ClickHouse answers "what does this table return now", Data Workers "is it still right, and what's the fix". The Data Workers side follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry. Keep mcp-clickhouse on its read-only default, with a ClickHouse user that holds a read-only role.
# Example: ClickHouse's MCP server (read-only by default) plus Data Workers agents in Claude Code
claude mcp add mcp-clickhouse \
-e CLICKHOUSE_HOST=<your-service-host> -e CLICKHOUSE_USER=<read-only-user> \
-e CLICKHOUSE_PASSWORD=<password> -e CLICKHOUSE_SECURE=true \
-- uv run --with mcp-clickhouse --python 3.12 mcp-clickhouse
# Data Workers agents, from a clone of the open-source repo
claude mcp add --scope user dw-connectors -- "$(pwd)/start-agent.sh" dw-connectors
claude mcp add --scope user dw-schema -- "$(pwd)/start-agent.sh" dw-schema
claude mcp add --scope user dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
claude mcp add --scope user dw-incidents -- "$(pwd)/start-agent.sh" dw-incidentsOn ClickHouse Cloud the team can use the remote MCP server instead: enable it on the service under Connect, add https://mcp.clickhouse.cloud/mcp to the client and sign in with OAuth. List the tools with your client's own command (/mcp in Claude Code). In this incident: run_quality_check (dw-quality) and assess_impact (dw-schema) record the new column and its readers; monitor_metrics and diagnose_incident (dw-incidents) score GMV per order and name the cause; trace_cross_platform_lineage and blast_radius_analysis (dw-context-catalog) map what the rollup reached; send_slack_alert (dw-connectors) reaches the owner; and trigger_step_function (dw-connectors) queues a state machine rerun after approval.
In production the agents run in your infrastructure and hold the database credentials and model key; your data stays in your systems, and the hosted Conductor sees workflow metadata only. More in where does our data go. For ClickHouse setup patterns, see connecting Claude Code to ClickHouse and an MCP server for ClickHouse analytics.
One incident, L0 to L4, set per domain:

- •L0 manual. A restaurant disputes its invoice on Thursday. Finance reconciles for a day, then credits 3,870 invoices.
- •L1 observe. Data Workers flags the anomaly at 14:30 with the cause and blast radius. Nothing changes.
- •L2 propose. Data Workers proposes the hold and the change set. Nothing reaches production until the named owner approves; an unanswered request expires and escalates, never auto-grants.
- •L3 act reversibly. For a class with a clean record, Data Workers takes reversible steps in the domain you open, such as queuing the held run once checks pass. Views, tables and deploys stay with the owner.
- •L4 autonomous. For a scoped domain, migrations on the databases ClickPipes replicates reach Data Workers for review before they ship, checked against the views that read the table, so the tipping release is flagged before the first tip lands.
Who sets those levels: who owns the agents, how approvals work for AI data agents and autonomy levels L0 to L4 explained.
What changes for your team

- •App releases get a second reader. A new column on a replicated table is checked against every view that reads it, by the Schema Evolution agent.
- •Fast stops meaning right by default. Rollups that pass every service check are still scored against the metrics the team records, so a quiet double count becomes an incident with an owner, by the Incident Debugging agent.
- •Producers and consumers share one record. The app team, the analytics platform team and billing operations see the same incident and receipt in Spellbook Data Catalog (in preview).
- •Customer-facing numbers have a paper trail. The receipt names the view, the window and the panel affected, so partner success knows whom to contact.
Keep ClickHouse, or consolidate?
Keep ClickHouse if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
Most ClickHouse teams keep it: sub-second aggregations served into products are hard to give back. What teams consolidate is the tooling around the service: a separate data observability tool, hand-written "compare ClickHouse to Postgres" queries and "check the rollups after every release" runbooks. Neighbours: you're on Postgres, you're on Firebolt and you're on Materialize; Data Workers integrations lists what connects natively, and Data Workers vs data observability explains why service health alone misses this class.
Building it yourself with a coding agent and mcp-clickhouse? See build it ourselves with Claude Code and MCP servers: the query is easy; the context graph, approvals, undo and receipts are the work.
The case for your CFO
The outcome. When an upstream release quietly changes what a customer-facing rollup counts, it is caught within hours and corrected before invoices, payouts or board numbers run on it.
The risk story. At L1 agents only read. At L2 a named person approves each change; an unanswered request expires and escalates. L3 covers reversible steps in domains you open. Views, tables and deploys stay with the owner. No agent can promote its own work, an org-wide stop halts all autonomous dispatch, and every change carries a receipt with its blast radius and undo. Nothing migrates.
Why now. ClickHouse now sits under products, agents and billing jobs at once. A wrong rollup reaches customers in seconds and invoices overnight.
The first win. The rollups that bill or pay: commissions, usage invoices, payouts. A pilot at L1 shows within weeks what the agents would have caught.
What stays the same. ClickHouse, ClickPipes, your views, dbt-clickhouse, ClickStack and your on-call rota. For the numbers, see the ROI of agentic data operations.
The sentence for upstairs: "ClickHouse makes our numbers fast; Data Workers makes sure they are right, and gets them fixed with our approval when they aren't, before we bill or pay on them."
Getting started
Start with a pilot. Pick the ClickHouse rollups behind invoices, payouts or customer-facing dashboards, give Data Workers read access to the upstream databases and the dbt-clickhouse project, record the rollup metrics that matter, and run at L1 for a few weeks. Then turn on L2 for one domain, and L3 for a narrow class once the receipts earn it. The pilot path is on the pricing page, and the pilot is credited in full against the first year.
FAQ
Does Data Workers query ClickHouse directly? No, by design. Data Workers reads the systems around it natively: upstream PostgreSQL, the dbt-clickhouse manifest, Kafka and Schema Registry, and orchestrators such as Airflow, Dagster, Prefect and Step Functions. Rollup health comes from metrics your team records with monitor_metrics. ClickHouse's own MCP servers serve your team's MCP client, next to Data Workers.
Why did the materialized view count orders twice? An incremental materialized view runs on each block inserted into its source table, and ClickPipes writes every Postgres update as a new row version, so a view that sums delivered rows sums each delivered version. ClickHouse's docs recommend FINAL, a view with FINAL or a refreshable materialized view for deduplicated results.
Will Data Workers change our tables, views or ClickPipes? No. Tables, views, pipes, roles and deploys stay with the owner. Data Workers proposes the change as a diff, and your team ships it through dbt-clickhouse and CI. Duplicate-row cleanups are proposed for the owner to apply.
We already run ClickStack. Isn't that enough? ClickStack shows how the service and your apps behave: logs, traces, metrics and sessions. In this incident all of them were healthy. Data Workers watches what the data means across the source, the pipe, the view and the readers, and owns the fix.
How does this relate to ClickHouse Agents and the Agentic Data Stack? They let people ask ClickHouse questions in plain language. Data Workers checks that the data under those answers is right and gets fixes approved by a named owner.
We also use ClickHouse Managed Postgres. Does that change anything? It is Postgres: Data Workers reads it the way it reads any PostgreSQL database, so new columns are caught at the source the same way. CDC into ClickHouse runs on ClickPipes either way.
Sources
- •ClickHouse blog, "ClickHouse raises $400 million Series D, acquires Langfuse, launches Postgres" (Jan 16, 2026), https://clickhouse.com/blog/clickhouse-raises-400-million-series-d-acquires-langfuse-launches-postgres (checked Oct 3, 2026)
- •ClickHouse, Managed Postgres product page (available on AWS; Google Cloud in private preview), https://clickhouse.com/cloud/postgres (checked Oct 3, 2026)
- •ClickHouse docs, ClickHouse Managed Postgres overview (Beta badge, AWS Public Beta; the product page says available on AWS, so this guide gives no GA or Beta label), https://clickhouse.com/docs/products/managed-postgres/overview (checked Oct 3, 2026)
- •ClickHouse docs, ClickStack overview (Managed ClickStack), https://clickhouse.com/docs/use-cases/observability/clickstack/overview (checked Oct 3, 2026)
- •ClickHouse docs, ClickHouse Agents (Beta), https://clickhouse.com/docs/products/cloud/features/ai-ml/agents (checked Oct 3, 2026)
- •ClickHouse docs, Agentic Data Stack, https://clickhouse.com/docs/products/agentic-data-stack/overview (checked Oct 3, 2026)
- •ClickHouse docs, AI-powered SQL generation, https://clickhouse.com/docs/use-cases/AI/ai-powered-sql-generation (checked Oct 3, 2026)
- •ClickHouse docs, Enable and connect ClickHouse Cloud remote MCP server (OAuth or API key;
run_select_querySELECT only), https://clickhouse.com/docs/products/cloud/features/ai-ml/mcp/remote-mcp (checked Oct 3, 2026) - •ClickHouse, mcp-clickhouse README and releases (v0.7.0, Sep 21, 2026), https://github.com/ClickHouse/mcp-clickhouse (checked Oct 3, 2026)
- •ClickHouse docs, ClickPipes for Postgres: deduplication strategies, https://clickhouse.com/docs/integrations/clickpipes/postgres/deduplication (checked Oct 3, 2026)
- •ClickHouse docs, ClickPipes for Postgres: schema changes propagation support, https://clickhouse.com/docs/integrations/clickpipes/postgres/schema-changes (checked Oct 3, 2026)
- •ClickHouse docs, Incremental materialized view, https://clickhouse.com/docs/materialized-view/incremental-materialized-view (checked Oct 3, 2026)
- •ClickHouse docs, dbt-clickhouse materialized_view materialization, https://clickhouse.com/docs/integrations/dbt/materialization-materialized-view (checked Oct 3, 2026)
- •ClickHouse docs, Langfuse in ClickHouse Cloud, https://clickhouse.com/docs/products/cloud/features/ai-ml/langfuse (checked Oct 3, 2026)
- •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)