You're on Postgres: It Runs Your Applications and Often Your First Analytics. Data Workers Owns Whether the Data In and Out of It Is Right
Postgres runs your product and often your first analytics. Data Workers catches a CDC feed that goes quiet after a migration, traces it and gets the fix approved.
Your product writes to Postgres: PostgreSQL 18 on Cloud SQL (where 18.6 is the default), Amazon RDS or Aurora, Azure Database for PostgreSQL, AlloyDB, Supabase, Neon (built around Databricks' Lakebase Postgres) or Snowflake Postgres. The app team ships migrations with Flyway, Alembic or Rails. A read replica serves the first reports, a few dbt-postgres marts live in a reporting schema, and pgvector holds embeddings next to the rows they describe. For everything bigger, a publication and a replication slot feed a CDC tool that copies tables into BigQuery, Snowflake or Databricks. Engineers query it through a Postgres MCP server, read-only if the DBA has a say.
Postgres is so reliable that one incident class is very quiet: the app team changes a table in a clean transaction, replication publishes exactly what the publication names, the CDC stream reports healthy, and the warehouse stops getting rows. Data Workers watches what lands, reads the source natively, traces the gap to its cause and gets the fix approved before the business reads it.
Key takeaways
- •Postgres keeps its job. Instances, migrations, publications, slots and roles stay with the app team and the DBAs.
- •A green CDC stream with no rows becomes an incident. Data Workers scores what lands against the volume baselines your team records, ties the gap to what changed in the source and lists every model and report it reaches.
- •It reads Postgres natively. The Context & Catalog agent reads schemas and their owners, tables and views, and column types, nullability and primary keys, and searches tables by name across every schema. The Data Quality agent runs null, uniqueness and row-count checks on Postgres tables directly.
- •Every fix goes through a named person. The DBA runs the Postgres change, the data engineer runs the stream and the dbt rerun, and Data Workers verifies the result and writes the receipt.
- •It stacks with your Postgres MCP server. That server answers questions about one database; Data Workers keeps the numbers built from it right.
Postgres is the database your product runs on. Data Workers owns whether the data in and out of it is right.
The job is to notice when what leaves Postgres stops matching what is in it, find the cause, size the damage, get the fix made by each owner and prove it. Here is one Tuesday at a meal-kit subscription company, an illustration, not a customer case.
The setup: the order service writes to Cloud SQL for PostgreSQL 18. Google Datastream streams public.orders and twenty other tables into the BigQuery dataset raw_app through the publication ds_pub, created FOR TABLE on those tables as Datastream's docs suggest, and a pgoutput replication slot. A dbt Cloud job, hourly_orders, builds stg_orders, fct_orders and fct_daily_bookings. Looker's "Daily bookings" dashboard, a dbt exposure the team has also recorded in the context graph, goes to leadership at 18:00 and sets the next day's ingredient orders. A scheduled query posts orders per quarter hour from raw_app to Data Workers' monitor_metrics, about 2,200 on an afternoon.
| Time | System | What happens |
|---|---|---|
| Tue 13:52 | GitHub | The app team merges Flyway migration V187: partition orders by month on created_at, so old months can be archived cheaply |
| 14:05 | Cloud SQL for PostgreSQL 18 | Flyway runs V187 in one transaction: it creates a partitioned table with monthly partitions, copies the rows, renames orders to orders_legacy and the new table to orders. The app keeps writing, now into the partitioned orders |
| 14:05 | Logical replication | A publication tracks tables, not names. ds_pub now publishes orders_legacy, which gets no more writes. The new orders is in no publication |
| 14:06 | Datastream | The stream stays Running with no errors. It streams what the publication emits, so raw_app.public_orders stops growing |
| 14:40 | Data Workers + BigQuery | monitor_metrics flags orders per quarter hour at 0 against the baseline of about 2,200, and Data Workers opens an incident |
| 14:46 | Data Workers + PostgreSQL | search_across_platforms reads the source natively: in the public schema it finds orders, orders_legacy and partitions orders_2026_01 to orders_2026_12. diagnose_incident classifies a schema change at the source, and Data Workers names the likely cause: the publication followed the renamed table |
| 14:50 | Data Workers + dbt | trace_cross_platform_lineage follows the dbt manifest from the orders source to fct_daily_bookings; blast_radius_analysis adds the Looker "Daily bookings" dashboard and its 18:00 delivery |
| 14:52 | Slack | send_slack_alert reaches the data platform owner and the app team's DBA with the evidence and a check to run: SELECT * FROM pg_publication_tables WHERE pubname = 'ds_pub'; |
| 15:05 | Cloud SQL | The DBA runs it: ds_pub lists public.orders_legacy, not public.orders. Cause confirmed |
| 15:20 | Spellbook | The owners review the plan. For the DBA: a new publication ds_orders_root for public.orders with publish_via_partition_root = true, so the partitions stream as one table, plus its own slot (Datastream documents that changing that setting on a publication a stream already uses fails the stream permanently, so ds_pub is left alone). For the data engineer: a new Datastream stream on that publication, with backfill, into raw_app_v2; a dbt diff that points the orders source at raw_app_v2.public_orders, with its rollback; then a rerun of hourly_orders. Both approve |
| 15:30 | Cloud SQL | The DBA creates the publication and slot |
| 15:40 | Datastream | The data engineer creates the stream and starts backfill for orders; it completes at 16:55 |
| 17:00 | GitHub + dbt Cloud CI | The data engineer merges the dbt diff after CI passes |
| 17:05 | dbt Cloud | The data engineer reruns the approved hourly_orders job; it finishes at 17:25 |
| 17:30 | Data Workers + BigQuery | run_quality_check on raw_app_v2.public_orders finds no null or duplicate order IDs, and orders per quarter hour are back on baseline, including the roughly 31,000 orders placed since 14:04. The receipt records the cause, both approvals, the SQL the DBA ran, the stream, the diff, the run, the checks and the undo (revert the diff; the old stream is untouched) |
| 18:00 | Looker | The "Daily bookings" delivery goes out complete, and tomorrow's ingredient orders are set on real demand |

Every part of the stack did its job. Seeing that a renamed table had cut the warehouse off took facts outside any one of them: what the warehouse should receive, where the rows went in the source, and which report reads them at 18:00.
| Job | What Postgres does | What Data Workers does |
|---|---|---|
| Storing the data | Commits every write safely, with constraints, transactional DDL and replicas | Reads the catalog natively: schemas, owners, tables, views, columns, types and primary keys |
| Feeding the warehouse | Publishes changes through publications and replication slots to the CDC tool | Scores what lands against the team's volume baselines and checks keys and nulls |
| Migrations | Runs whatever migration the app team ships, atomically | Ties a downstream gap or wrong number to what changed in the source, with the blast radius |
| Agent access | Postgres MCP servers answer questions over one database, read-only when configured | Keeps the data behind those answers right across the database, the CDC tool, the warehouse and BI |
| The fix | Runs the SQL the DBA writes | Proposes the change set to named owners, in order, with the undo for each step |
| The proof | Answers the check query | Re-checks the landed data and writes a receipt: what changed, who approved it, how it was checked, how to undo it |
Why doesn't Postgres just do this itself?
Because Postgres is built to store data safely and answer SQL correctly, and it does that better than almost anything, scoped to one database. PostgreSQL 18 added asynchronous I/O, virtual generated columns, uuidv7(), OAuth 2.0 sign-in, faster major upgrades, logical replication write-conflict reporting and the option to drop idle replication slots automatically so a publisher doesn't fill up with write-ahead log. PostgreSQL 19 is in beta 4. None of that is meant to know that a table on a publication feeds a dashboard that sets tomorrow's purchasing.
Managed vendors build the same way, and well. Cloud SQL, RDS, Azure and AlloyDB run the instance, and AlloyDB adds conversational analytics over operational data (in preview). Supabase, Neon, Lakebase and Snowflake Postgres give apps and agents a Postgres backend and, for some, a mirror into analytics. The MCP servers lean read-only: Supabase's server takes read_only=true, Neon's read-only mode disables writes such as running migrations, Postgres MCP Pro has a restricted mode for production, and Google's MCP Toolbox for Databases ships a prebuilt Postgres toolset. Each operates or answers questions about one database.
Owning whether the data is right from the app database to the report is a different product: a context graph across systems, blast-radius scoping, named approvals, a recorded undo and receipts, and the judgment to route a Postgres change to the DBA and a stream change to the data engineer. That is Data Workers. More in is it safe to let AI agents change production data.
Every tool owns a slice. Data Workers covers the whole lifecycle
Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, on top of the Postgres databases already there.

| Stage | Data Workers | Postgres | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 4 | The system catalogs and information_schema describe every schema, table and column in one database. Data Workers keeps one governed context graph from the app database through CDC, the warehouse and dbt to every report. |
| Analytics & Insights | 8 | 8.5 | Postgres's home stage: a mature SQL engine that runs the app and, for many teams, the first marts and reports, with extensions such as pgvector. Data Workers answers data questions from governed definitions with lineage behind every number. |
| Data Quality | 8 | 4 | Constraints, foreign keys and check constraints protect each row on write. Data Workers runs null, uniqueness and row-count checks on Postgres, Snowflake and BigQuery tables and flags drift against baselines the team records. |
| Observability & Incidents | 8.5 | 4 | Statistics views and server logs show how the database runs: connections, locks, replication lag. Data Workers diagnoses the data incident across systems and verifies the fix. |
| Pipelines & Ingestion | 8.5 | 6 | Logical replication publishes changes through publications and slots, and foreign data wrappers read other systems. Data Workers sees when a feed goes quiet and plans the fix with the owners. |
| Schema & Migration | 8 | 5 | Transactional DDL makes each migration safe inside the database. Data Workers reviews what a migration means for the feeds and models that read the table, and generates migrations with rollback SQL for the owner. |
| Governance & Access | 8.5 | 6 | Roles, grants and default privileges govern who reads what. Data Workers routes every data change to a named approver. |
| Security & Privacy | 8 | 7 | Row-level security, SCRAM and OAuth 2.0 sign-in in PostgreSQL 18 protect the data. Data Workers leaves a receipt on every data change. |
| Cost / FinOps | 8 | 4 | Managed services bill the instance; Postgres itself has no spend view. Data Workers traces Snowflake spend to the dbt model and reads BigQuery and AWS spend. |
| MLOps & Models | 7.5 | 4 | pgvector stores embeddings and serves similarity search next to the app's rows. Data Workers keeps the data under models and agents healthy. |
How Postgres and Data Workers work together

Data Workers reads Postgres natively. With a read-only role, the Context & Catalog agent reads each database's schemas and their owners, its tables and views, and every column's type (with length or precision), nullability and primary key, and it searches tables by name across all schemas, which is how the renamed table and its partitions surfaced right after the alert. The Data Quality agent runs null, uniqueness and row-count checks on Postgres tables over the same connection. Data Workers reads the rest of the path natively too: BigQuery and Snowflake where the Postgres data lands, the dbt project (sources, models and tests arrive as lineage), Airflow, Dagster and Prefect, Looker and Tableau, Kafka and Schema Registry, and Slack, among 50+ connectors. CDC tools such as Datastream, Fivetran, Airbyte and Debezium connect over their API or MCP server today. See Data Workers on Google Cloud.
Publications, slots, roles, streams and dbt jobs stay with their owners, by design. Data Workers does not write to your Postgres: it proposes SQL for the DBA, dbt diffs for the owner to merge, migrations with rollback SQL, and a generated GRANT for each access request, with its expiry on record.
Your Postgres MCP server answers "what is in this table now", Data Workers "is what we built from it still right". The Data Workers side follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry. Give Data Workers a read-only Postgres role, read access to the warehouse and the dbt project, and, for checks on BigQuery tables, a service-account key. Point the Postgres MCP server at a replica where you can.
# Example: Postgres MCP Pro in restricted (read-only) mode plus Data Workers agents in Claude Code
claude mcp add postgres -e DATABASE_URI="postgresql://<read-only-role>@<replica-host>:5432/<db>" \
-- uvx postgres-mcp --access-mode=restricted
# Data Workers agents, from a clone of the open-source repo
claude mcp add --scope user dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
claude mcp add --scope user dw-quality -- "$(pwd)/start-agent.sh" dw-quality
claude mcp add --scope user dw-incidents -- "$(pwd)/start-agent.sh" dw-incidents
claude mcp add --scope user dw-connectors -- "$(pwd)/start-agent.sh" dw-connectorsList the tools with your client's own command (/mcp in Claude Code). The incident used monitor_metrics and diagnose_incident (dw-incidents), search_across_platforms, trace_cross_platform_lineage and blast_radius_analysis (dw-context-catalog), send_slack_alert (dw-connectors) and run_quality_check (dw-quality).
The agents run in your infrastructure and hold the credentials and model key; your data stays in your systems, and the hosted Conductor sees workflow metadata only (where does our data go). For access patterns, see how to connect an AI agent to Postgres safely, an MCP server for Postgres, running a Postgres MCP server in production and Claude Code for Postgres data engineering.
One incident, L0 to L4, set per domain:

- •L0 manual. Leadership orders ingredients on a bookings number missing three hours of orders; someone notices on Thursday.
- •L1 observe. Data Workers flags the gap at 14:40 with the cause and the blast radius. Nothing changes.
- •L2 propose. Data Workers proposes the publication, the stream, the diff and the rerun to named owners; an unanswered request expires and escalates, never auto-grants.
- •L3 act reversibly. For a class with a clean record, Data Workers takes reversible steps in the domain you open, such as queuing an Airflow or Prefect rerun once checks pass. Postgres changes and streams stay with their owners.
- •L4 autonomous. For a scoped domain, migrations that touch replicated tables reach Data Workers in pull request review before they ship, checked against the feeds and models that read them, so V187 arrives with its publication change already planned with the DBA.
More: who owns the agents, how approvals work for AI data agents and autonomy levels L0 to L4 explained.
What changes for your team

- •App migrations get a second reader. A change to a replicated table is reviewed against the feeds and models that read it, by the Schema Evolution agent.
- •Healthy streams stop meaning complete. A quiet gap becomes an incident with an owner, by the Incident Debugging agent.
- •DBAs and data engineers share one record. Everyone sees the same incident, plan and receipt in Spellbook Data Catalog (in preview), and each owner acts on their piece.
- •The fix is remembered. The Context Wizard records that
ordersnow replicates throughds_orders_root.
Keep Postgres, or consolidate?
Keep Postgres if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
Almost every team keeps Postgres. What teams consolidate is the tooling around it: a separate observability tool, hand-written row-count comparisons, and the "tell the data team before you rename a table" runbook nobody reads. If you are moving analytics off a Postgres replica, Data Workers translates PostgreSQL SQL into Snowflake SQL and plans each wave with its parity checks; see Data Workers on Snowflake. Neighbours: you're on ClickHouse, you're on MongoDB Atlas, you're on Airbyte and you're on dbt. See also Data Workers integrations, what is an agentic data platform, Data Workers vs data observability and, if you plan to build it with a coding agent, build it ourselves with Claude Code and MCP servers.
The case for your CFO
The outcome. When an app change quietly cuts the warehouse off from the product database, it is caught within the hour and fixed before purchasing, forecasts or board numbers run on the gap.
The risk story. At L1 agents only read, and Data Workers does not write to Postgres. At L2 a named person approves each change; an unanswered request expires and escalates. L3 covers reversible steps in domains you open. No agent can promote its own work, an org-wide stop halts all autonomous dispatch, and every change carries a receipt with its blast radius and undo. Nothing migrates.
Why now. Every cloud and warehouse vendor now sells a Postgres and mirrors it into analytics, so more of the numbers you run the company on start in a database your data team does not own.
The first win. The replicated tables that feed revenue, bookings or customer reporting; a pilot at L1 shows within weeks what the agents would have caught.
What stays the same. Postgres, your managed service, migrations tool, CDC tool, warehouse, dbt and on-call rota. For the numbers, see the ROI of agentic data operations.
The sentence for upstairs: "Postgres runs our product; Data Workers makes sure every number we build from it is right, and gets it fixed with our approval when it isn't, before we act on it."
Getting started
Start with a pilot. Pick the Postgres tables that feed revenue or customer reporting, connect them read-only with the warehouse and the dbt project, record the volume metrics that matter, and run at L1 for a few weeks. Then turn on L2 for one domain, and L3 once the receipts earn it. The pilot path is on the pricing page, and the pilot is credited in full against the first year.
FAQ
Does Data Workers write to our Postgres database? No. It reads with a read-only role. Changes to the database, such as publications, slots, grants and migrations, are proposed for the DBA to run, with the SQL and the rollback in the proposal.
What does Data Workers read from Postgres? The catalog: each schema and its owner, tables and views, every column's type, nullability and primary key, and table search by name across schemas. run_quality_check adds null, uniqueness and row-count checks on Postgres tables. It does not read query history.
Why did the stream stay healthy with no new rows? A publication created FOR TABLE tracks the table itself, so after the rename it published orders_legacy, and the new partitioned orders was in no publication. Datastream's docs replicate partitioned tables as one table through a publication with publish_via_partition_root = true, and warn that changing that setting on a publication a stream already uses fails the stream permanently.
We run Supabase, Neon, Lakebase or Snowflake Postgres. Does that change anything? No: they are Postgres, and Data Workers reads Postgres natively. Where the vendor mirrors data into Snowflake or BigQuery, Data Workers checks what lands there.
We already monitor replication slots and lag. Isn't that enough? Slot and lag monitoring shows that replication is flowing. Here it was: lag stayed near zero, because nothing was published for the new table. Data Workers checks the data where it is used and owns the fix across Postgres, the CDC tool, the warehouse and BI. PostgreSQL 18 can also drop idle slots automatically, one more way a feed can stop; a volume baseline catches that too.
Sources
- •PostgreSQL Global Development Group, postgresql.org home page (PostgreSQL 19 Beta 4, Sep 24, 2026; 18.6, Aug 13, 2026), https://www.postgresql.org/ (checked Oct 3, 2026)
- •PostgreSQL Global Development Group, "PostgreSQL 18 Released!" (Sep 25, 2025), https://www.postgresql.org/about/news/postgresql-18-released-3142/ (checked Oct 3, 2026)
- •Google Cloud, Cloud SQL for PostgreSQL database versions (PostgreSQL 18 default, 18.6), https://cloud.google.com/sql/docs/postgres/db-versions (checked Oct 3, 2026)
- •Google Cloud, AlloyDB for PostgreSQL release notes (conversational analytics, Preview, Mar 30, 2026), https://cloud.google.com/alloydb/docs/release-notes (checked Oct 3, 2026)
- •AWS, Amazon RDS for PostgreSQL updates (18.6; PostgreSQL 19 Beta 4 in the RDS Preview environment), https://docs.aws.amazon.com/AmazonRDS/latest/PostgreSQLReleaseNotes/postgresql-versions.html (checked Oct 3, 2026)
- •Microsoft Learn, Supported versions of PostgreSQL in Azure Database for PostgreSQL flexible server (PostgreSQL 18, 18.6), https://learn.microsoft.com/en-us/azure/postgresql/flexible-server/concepts-supported-versions (checked Oct 3, 2026)
- •Neon home page ("Built around Lakebase Postgres, by Databricks"), https://neon.com/ (checked Oct 3, 2026)
- •Databricks, Lakebase product page, https://www.databricks.com/product/lakebase (checked Oct 3, 2026)
- •Snowflake documentation, About Snowflake Postgres (data mirroring), https://docs.snowflake.com/en/user-guide/snowflake-postgres/about (checked Oct 3, 2026)
- •Supabase home page, https://supabase.com/ (checked Oct 3, 2026)
- •pgvector, GitHub repository and changelog (0.8.7, Oct 1, 2026), https://github.com/pgvector/pgvector (checked Oct 3, 2026)
- •Supabase MCP server (
read_only=true), https://github.com/supabase/mcp (checked Oct 3, 2026) - •Neon MCP server (read-only mode), https://github.com/neondatabase/mcp-server-neon (checked Oct 3, 2026)
- •Postgres MCP Pro (restricted mode), https://github.com/crystaldba/postgres-mcp (checked Oct 3, 2026)
- •Google, MCP Toolbox for Databases (prebuilt Postgres tools), https://github.com/googleapis/mcp-toolbox (checked Oct 3, 2026)
- •Google Cloud, Datastream: configure a Cloud SQL for PostgreSQL source (publications, replication slots), https://cloud.google.com/datastream/docs/configure-cloudsql-psql (checked Oct 3, 2026)
- •Google Cloud, Datastream: work with PostgreSQL partitioned tables (
publish_via_partition_root; changing it on a used publication fails the stream), https://cloud.google.com/datastream/docs/work-with-postgresql-partitioned-tables (checked Oct 3, 2026) - •Google Cloud, Datastream: manage backfill for the objects of a stream, https://cloud.google.com/datastream/docs/manage-backfill-for-the-objects-of-a-stream (checked Oct 3, 2026)
- •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)