You're on MongoDB Atlas: It Runs Your Applications' Document Data. Data Workers Owns Whether the Data That Leaves It for Analytics and AI Is Right
MongoDB Atlas runs your app's documents. Data Workers catches a document change that quietly empties a warehouse column, traces it and gets the fix approved.
Your product teams build on MongoDB Atlas. Each service writes documents to its own collections on an Atlas cluster in AWS, Azure or Google Cloud, and the document model lets a release add a field, nest an object or turn a value into an array without a migration. Since September 29, 2026 the engine is MongoDB 9.0, generally available, with Atlas Infinite (public preview) as a new deployment option for spiky load. The AI work sits next to the data: MongoDB Vector Search with Automated Embedding on Voyage AI by MongoDB models, and Atlas Agent Engine (public preview) to run agents with memory and governance. Engineers query clusters through the MongoDB MCP Server; the data team lands collections in a warehouse with a CDC tool and flattens them in dbt.
That last step is where the flexibility that makes Atlas good for apps turns into risk. A release changes a document's shape, the CDC tool lands it faithfully, and a JSON path in a dbt model starts returning NULL. Nothing errors, and every report, tax extract and model built on that column goes wrong. Data Workers watches what leaves Atlas for analytics and AI, catches the change by what it does to the data, traces its reach and gets the fix approved before anyone files, bills or retrains on it.
Key takeaways
- •Atlas keeps its job. Clusters, collections, validators, Vector Search and Agent Engine stay with your app and platform teams. Data Workers works on what leaves the cluster and everything that reads it.
- •A new document shape becomes an incident. When a release moves a field and a warehouse column quietly empties, Data Workers opens one incident with the onset, the evidence and every report, extract and model it reaches.
- •It reads the systems around Atlas natively. BigQuery, Snowflake, Databricks and PostgreSQL, the dbt project and Slack connect natively. MongoDB and Fivetran connect over their API or MCP server today.
- •Every fix goes through a named person. Data Workers proposes the change as a diff for the owner to merge; the owner approves the hold and the rerun.
- •It stacks with MongoDB's own AI. Agent Engine and the MCP Server serve agents on Atlas data; Data Workers keeps the warehouse tables and features built from that data right.
MongoDB Atlas runs your applications' document data. Data Workers owns whether the data that leaves it for analytics and AI is right.
The job on top of Atlas is to notice that what a collection means downstream has changed, size the damage, hold what files or trains on it, get the fix made and prove it. Here is one Tuesday at an online home goods retailer, an illustration, not a customer case.
The setup: the checkout service writes each order to the orders collection. Fivetran syncs it to BigQuery on a 15-minute schedule in packed mode, the mode Fivetran recommends, so each document lands whole in a data JSON column of mongo_shop.orders. A dbt view, stg_shop__orders, pulls JSON_VALUE(data, '$.shipping_address.country') into ship_country. Downstream, fct_orders feeds a US sales-tax extract finance files on the 20th, the Looker "Orders" Explore and a BigQuery ML model, demand_by_region, that retrains at 03:00. All three are declared as dbt exposures and recorded in the context graph. The team records one metric hourly with Data Workers: the share of new orders that carry a ship country. Data Workers reads BigQuery with a service account.
| Time | System | What happens |
|---|---|---|
| Tue 13:30 | MongoDB Atlas | Checkout release 5.8 adds split shipments. New orders carry fulfillment.shipments, an array where each shipment has its own address; the top-level shipping_address is no longer written. No validator covers that field |
| 13:45 | Fivetran | The sync succeeds. The new documents land in the same data column, so the BigQuery table schema is unchanged |
| 14:00 | dbt on BigQuery | stg_shop__orders returns NULL ship_country for every order since 13:30. No query fails |
| 15:00 | Data Workers | The ship-country share, recorded with monitor_metrics, falls from a 99.6% baseline to 3.1% while order counts and revenue hold. Data Workers opens an incident |
| 15:02 | Data Workers + BigQuery | run_quality_check on stg_shop__orders counts 6,240 NULL ship_country values; the morning check counted 12 |
| 15:05 | Data Workers | diagnose_incident dates the onset to the 13:00 hour. trace_cross_platform_lineage follows the dbt manifest from mongo_shop.orders to fct_orders; blast_radius_analysis adds the tax extract, the Explore and demand_by_region through the team's context-graph notes |
| 15:10 | Slack | send_slack_alert reaches the analytics engineering owner, with checkout and finance copied, plus a document query to run |
| 15:25 | MongoDB MCP Server | The owner runs the query from her own client, read-only: the new orders carry fulfillment.shipments and no shipping_address; 9% ship to two addresses |
| 15:45 | Spellbook | She reviews two proposals: a dbt diff that adds stg_shop__order_shipments, one row per shipment, and falls back to the first shipment's country in stg_shop__orders, with a not-null test and its rollback; and a hold on tonight's retrain. She approves both and pauses the retrain's scheduled query |
| 16:30 | GitHub + dbt Cloud | She merges the diff; the dbt Cloud CI job builds the changed models and runs the new test. She reports the test green in the incident |
| 16:40 | dbt Cloud | She reruns the production job, as approved |
| 17:20 | Data Workers | run_quality_check finds no new NULL countries and the share is back on baseline. Her BigQuery count by country matches her read-only count in Atlas. The receipt records cause, approvals, merge, checks and the undo (revert the diff) |
| 17:30 | BigQuery | She resumes the retrain and asks checkout to flag future orders shape changes in the release checklist |
| Wed 03:00 | BigQuery ML | demand_by_region retrains on orders with their real regions |

Every part of the stack did its job: Atlas stored valid documents, Fivetran delivered them as written, and BigQuery correctly returned NULL for a missing JSON path. Knowing that a NULL country would mis-file sales tax and teach a forecast that a region stopped buying took facts outside the database: which models read the field, what reads those, and when they run.
| Job | What MongoDB Atlas does | What Data Workers does |
|---|---|---|
| Documents | Stores flexible documents; JSON Schema validation is opt-in per collection | Checks that what the warehouse builds from them still means what it should, against baselines the team records |
| Change capture | Change streams and Atlas Stream Processing emit and reshape events | Reads where the data lands, natively in BigQuery, Snowflake, Databricks or PostgreSQL, and through the dbt project |
| Agent access | The MCP Server answers questions, with read-only controls; Agent Engine runs agents on Atlas data | Keeps the tables and features behind analytics and AI right, and traces which numbers a bad field reached |
| The fix | Runs the release and validator changes the app team makes | Proposes the hold and the dbt change to a named owner, with the blast radius |
| The proof | Answers the owner's read-only count | Re-runs the checks and writes a receipt: what changed, who approved it, how it was checked, how to undo it |
Why doesn't MongoDB Atlas just do this itself?
Because Atlas is built to keep applications fast and flexible, and letting a release reshape data without a migration is the product. Whether a new shape breaks a dbt model in BigQuery depends on the warehouse, the finance calendar and a forecast's retrain schedule, none of which a database has a reason to hold.
MongoDB's AI work is strong and draws the same line. MongoDB Vector Search generates and syncs embeddings with Voyage AI models through Automated Embedding. Atlas Agent Engine, in public preview since September 29, 2026, is "a unified execution, memory, and governance layer for production AI agents". The MongoDB MCP Server comes as Local MCP, where MongoDB's docs say the "default is to allow cluster write operations. Typically, always enable read-only mode", and as the Atlas Managed MCP Server, with "org-level opt-in and read-only enforcement for AI clients" and a per-configuration read-only flag enabled by default. All of it serves apps and agents on Atlas data.
Owning whether the data is right from the collection to the tax filing and the model is a different product with a different liability: a context graph across systems, blast-radius scoping, named approvals, a recorded undo and receipts. That is Data Workers. More in is it safe to let AI agents change production data.
Every tool owns a slice. Data Workers covers the whole lifecycle
Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, on top of the Atlas clusters already there.

| Stage | Data Workers | MongoDB Atlas | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 3 | Collections, indexes and a schema Compass infers from samples describe one cluster. Data Workers keeps one governed context graph from the collection through the warehouse and dbt to every report and model. |
| Analytics & Insights | 8 | 7 | The aggregation pipeline, the read-only SQL Interface and Atlas Charts answer questions on live documents. Data Workers answers from governed definitions with lineage behind every number. |
| Data Quality | 8 | 3 | JSON Schema validation is opt-in per collection, and flexible documents are the point. Data Workers checks what the warehouse builds from them against baselines the team records. |
| Observability & Incidents | 8.5 | 4 | Cluster metrics, the Performance Advisor and OpenTelemetry show how the database runs. Data Workers diagnoses the data incident across systems and verifies the fix. |
| Pipelines & Ingestion | 8.5 | 6 | Change streams and Atlas Stream Processing move and reshape events continuously. Data Workers sees what a document change does downstream and plans the fix for the owner. |
| Schema & Migration | 8 | 3 | Relational Migrator moves relational workloads into MongoDB. Data Workers catches a document-shape change by its effect on the models and generates migrations with rollback SQL for the owner. |
| Governance & Access | 8.5 | 5 | Database roles, IP access lists and Managed MCP read-only controls govern the cluster. Data Workers routes every data change to a named approver. |
| Security & Privacy | 8 | 7 | Queryable Encryption and private networking protect the documents themselves. Data Workers leaves a receipt on every data change. |
| Cost / FinOps | 8 | 4 | Atlas billing and usage views size clusters. Data Workers traces Snowflake spend to the dbt model and reads BigQuery spend as an account total. |
| MLOps & Models | 7.5 | 8 | MongoDB's home stage: Vector Search, Automated Embedding and Agent Engine serve AI apps on operational data. Data Workers keeps the data under models and agents healthy. |
How MongoDB Atlas and Data Workers work together

Data Workers reads the systems your Atlas data flows into natively: BigQuery, Snowflake, Databricks and PostgreSQL, the dbt project (sources, models and tests arrive as lineage from the manifest), orchestrators such as Airflow, Dagster and Prefect, Kafka and Schema Registry when change streams feed a topic, and Slack, among 50+ connectors. MongoDB and Fivetran connect over their API or MCP server today; field health comes from warehouse checks and the metrics your team records. On BigQuery, see Data Workers on Google Cloud.
Clusters, validators, Vector Search indexes and releases stay with the app team. Models, tests, schedules and syncs stay with the data owner. Data Workers proposes the change as a diff for the owner to merge, queues approved reruns through orchestrators such as Airflow or Prefect, and leaves dbt Cloud jobs to the owner or the team's scheduler.
The MongoDB MCP Server answers "what do the documents look like now"; Data Workers answers "is what we built from them still right, and what's the fix". Data Workers' agents never call MongoDB's MCP server; your team's client uses it side by side with Data Workers' servers. The Data Workers side follows the documented client setup: clone the open-source repo and add each agent's start-agent.sh entry.
# Example: MongoDB MCP Server in read-only mode plus Data Workers agents in Claude Code
claude mcp add mongodb \
-e MDB_MCP_CONNECTION_STRING="<atlas-connection-string-for-a-read-only-user>" \
-e MDB_MCP_READ_ONLY=true \
-- npx -y mongodb-mcp-server
# Data Workers agents, from a clone of the open-source repo
claude mcp add --scope user dw-connectors -- "$(pwd)/start-agent.sh" dw-connectors
claude mcp add --scope user dw-quality -- "$(pwd)/start-agent.sh" dw-quality
claude mcp add --scope user dw-catalog -- "$(pwd)/start-agent.sh" dw-context-catalog
claude mcp add --scope user dw-incidents -- "$(pwd)/start-agent.sh" dw-incidentsTeams on the Atlas Managed MCP Server point the client at https://mcp.mongodb.com with the service-account credentials an Atlas owner creates and keep its read-only flag on. List the tools with your client's own command (/mcp in Claude Code). In this incident: monitor_metrics and diagnose_incident (dw-incidents) score the ship-country share and date the onset; run_quality_check (dw-quality) counts the NULLs; trace_cross_platform_lineage and blast_radius_analysis (dw-context-catalog) map the reach; send_slack_alert (dw-connectors) reaches the owner.
In production the agents run in your infrastructure and hold the warehouse credentials and model key; your data stays in your systems, and the hosted Conductor sees workflow metadata only. More in where does our data go and an MCP server for MongoDB data.
One incident, L0 to L4, set per domain:

- •L0 manual. Finance spots a short tax filing on the 20th; two analysts trace it for a day.
- •L1 observe. Data Workers flags the drop at 15:00 with the evidence and blast radius. Nothing changes.
- •L2 propose. Data Workers proposes the hold and the dbt change. Nothing reaches production until the named owner approves; an unanswered request expires and escalates, never auto-grants.
- •L3 act reversibly. For a class with a clean record, Data Workers takes reversible steps in the domain you open, such as queuing an approved rerun through Airflow or Prefect. Models, schedules and releases stay with their owners.
- •L4 autonomous. For a scoped domain, Data Workers runs the loop on this class end to end, each step logged with its receipt.
Who sets those levels: who owns the agents, how approvals work for AI data agents and autonomy levels L0 to L4 explained.
What changes for your team

- •Silent NULLs become incidents with owners. Fields your reports and models depend on are scored against the baselines the team records, so a quiet gap gets a cause and a fix, by the Incident Debugging agent.
- •The fix gets a reviewer. The Schema Evolution agent reviews the dbt change against everything that reads the model before it merges.
- •App and data teams share one record. Checkout, analytics engineering and finance see the same incident and receipt in Spellbook Data Catalog (in preview).
- •Models train on data someone vouched for. A retrain that reads a broken field is held, and the receipt shows when it resumed and why.
Keep MongoDB Atlas, or consolidate?
Keep MongoDB Atlas if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
Most Atlas teams keep it. What teams consolidate is the tooling around the data that leaves it: a separate observability tool, hand-written shape checks and "check the models after every release" runbooks. Teams moving reports off the BI Connector, which reached end of life in September 2026, onto the read-only SQL Interface or a warehouse can bring them under the same checks. Neighbours: you're on Fivetran, you're on Postgres and you're on Pinecone; Data Workers integrations lists what connects natively, and Data Workers vs data observability explains why healthy clusters and green syncs miss this class.
Building it yourself with a coding agent? See build it ourselves with Claude Code and MCP servers: the query is easy; the context graph, approvals, undo and receipts are the work.
The case for your CFO
The outcome. When an app release quietly changes what the warehouse gets from MongoDB, it is caught within hours and corrected before tax filings, revenue reports or model retrains run on it.
The risk story. At L1 agents only read. At L2 a named person approves each change; an unanswered request expires and escalates. L3 covers reversible steps in domains you open. Clusters, releases, models and schedules stay with their owners. No agent can promote its own work, an org-wide stop halts all autonomous dispatch, and every change carries a receipt with its blast radius and undo. Nothing migrates.
Why now. MongoDB 9.0 and Atlas Agent Engine put more of the business, and more AI, on Atlas data. Each release can now break more downstream overnight.
The first win. The tables built from the collections that feed tax, finance or a model. A pilot at L1 shows within weeks what the agents would have caught.
What stays the same. Atlas, your releases, your CDC tool, BigQuery or Snowflake, dbt and your on-call rota. For the numbers, see the ROI of agentic data operations.
The sentence for upstairs: "Atlas keeps our apps fast and flexible; Data Workers makes sure the reports and models built from it stay right, and gets them fixed with our approval when they aren't."
Getting started
Start with a pilot. Pick the collections behind tax, finance or model features, give Data Workers read access to the warehouse they land in and the dbt project, record the field metrics that matter, and run at L1 for a few weeks. Then turn on L2 for one domain, and L3 for a narrow class once the receipts earn it. The pilot path is on the pricing page, and the pilot is credited in full against the first year.
FAQ
Does Data Workers connect to MongoDB directly? Data Workers connects to MongoDB over its MCP server or the Atlas Admin API today, and reads the warehouses your collections land in natively, plus the dbt project. Your team keeps the MongoDB MCP Server in its own client, in read-only mode, for checks on the documents.
Why didn't Fivetran or BigQuery raise an error? Neither had anything to report. In packed mode Fivetran delivers each document whole as JSON, so a moved field leaves the table schema unchanged, and BigQuery returns NULL for a JSON path that isn't there. The break shows up only in what the data means, which is what Data Workers checks.
Will Data Workers change our collections, validators or releases? No. They stay with the app team. Data Workers proposes the warehouse-side fix as a diff for the data owner to merge, and can suggest a release-checklist step or validator for the app team to decide on.
How does this relate to Atlas Agent Engine? Agent Engine runs your agents on Atlas data. Data Workers runs data operations across the warehouse, dbt and every reader, keeping the analytics and features those agents depend on right.
The BI Connector reached end of life. Does that change anything? MongoDB points BI Connector users to the SQL Interface, read-only SQL for Tableau and Power BI. Reports that move onto a warehouse instead come under Data Workers' lineage and checks the same way as this incident.
Is MongoDB 9.0 required? No. Data Workers works from the warehouses and tools downstream of Atlas, so the MongoDB version and deployment option make no difference to it. MongoDB notes that its MCP Server is not currently supported on Atlas Infinite clusters during the public preview.
Sources
- •MongoDB press release, "MongoDB Launches MongoDB 9.0, the Best Version Ever Built, and Atlas Infinite for AI-Scale Demand" (Sep 29, 2026), https://www.mongodb.com/company/newsroom/press-releases/mongodb-launches-mongodb-9-0-the-best-version-ever-built-and-atlas-infinite-for-ai-scale-demand (checked Oct 3, 2026)
- •MongoDB press release, "MongoDB Launches Atlas Agent Engine to Put AI Agents in Production Without a New Stack" (Sep 29, 2026; public preview), https://www.mongodb.com/company/newsroom/press-releases/mongodb-launches-atlas-agent-engine-to-put-ai-agents-in-production-without-a-new-stack (checked Oct 3, 2026)
- •MongoDB docs, Release Notes (9.0 current stable; 8.3 and 8.0 previous), https://www.mongodb.com/docs/manual/release-notes/ (checked Oct 3, 2026)
- •MongoDB docs, MongoDB MCP Server Overview (Local MCP and Atlas Managed MCP Server; read-only controls; not supported on Atlas Infinite during preview), https://www.mongodb.com/docs/mcp-server/overview/ (checked Oct 3, 2026)
- •MongoDB docs, Enable or Disable MongoDB MCP Server Features (read-only mode), https://www.mongodb.com/docs/mcp-server/local-mcp/configuration/enable-or-disable-features/ (checked Oct 3, 2026)
- •MongoDB docs, How Automated Embedding Works and Models for Automated Embedding (Voyage AI models), https://www.mongodb.com/docs/vector-search/crud-embeddings/automated-embedding/overview/ (checked Oct 3, 2026)
- •MongoDB docs, Voyage AI by MongoDB, https://www.mongodb.com/docs/voyageai/ (checked Oct 3, 2026)
- •MongoDB docs, Atlas Stream Processing, https://www.mongodb.com/docs/atlas/atlas-stream-processing/ (checked Oct 3, 2026)
- •MongoDB docs, Transition from BI Connector to SQL Interface (BI Connector end of life September 2026; SQL Interface read-only), https://www.mongodb.com/docs/sql-interface/transition-bic-to-atlas-sql/ (checked Oct 3, 2026)
- •MongoDB docs, Relational Migrator Overview, https://www.mongodb.com/docs/relational-migrator/getting-started/ (checked Oct 3, 2026)
- •Fivetran docs, MongoDB connector (pack modes; packed mode recommended), https://fivetran.com/docs/connectors/databases/mongodb (checked Oct 3, 2026)
- •MongoDB, mongodb-mcp-server releases on GitHub (v3.0.5, Oct 1, 2026), https://github.com/mongodb-js/mongodb-mcp-server/releases (checked Oct 3, 2026)
- •Data Workers client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)