Product
Product11 min readBy The Data Workers Team

Data Workers + Cube, AtScale, MetricFlow and Apache Ossie (OSI): One Governed View Across Every Semantic Layer

How Data Context Wizard plugs Cube, AtScale, dbt MetricFlow and Apache Ossie (formerly OSI) in as first-class sources, joins them to lineage, quality and usage, and flags conflicts.

Semantic layers hold the meaning of your metrics. Data Context Wizard plugs each of them in as a first-class source with provenance, joins every definition to the lineage, quality and usage underneath it, flags conflicts across layers, and gives every agent one governed view. Cube, AtScale and dbt keep defining the metrics. Data Workers, the agentic data platform, makes sure every agent and every person gets the same approved answer.

Most estates we see don't run one semantic layer. Product built customer-facing analytics on Cube. Finance runs AtScale so Excel and Power BI users can query Snowflake through DAX and MDX. The analytics engineers moved metric logic into dbt MetricFlow. Each layer is right inside its own runtime, each now ships an MCP server, and Apache Ossie, formerly Open Semantic Interchange (OSI), promises a common format. The question nobody's tool answers is the one the CFO asks: which arr is the real one, and is the data under it healthy today?

Key takeaways

  • •Every layer keeps its job. Cube, AtScale and dbt MetricFlow stay where metrics are defined, deployed and queried. Data Workers reads them and never edits a model without its owner.
  • •Each layer becomes a first-class source. Context Wizard imports dbt MetricFlow directly, reads Cube's model files and AtScale's SML from Git, checks both live over their MCP servers, and takes Ossie YAML through the Ossie project's own dbt converter.
  • •Definitions get joined to the data. Every metric carries its lineage, freshness and quality checks, and who queries it, so a conflict arrives with its blast radius attached.
  • •Conflicts go to a named owner. When two layers disagree, the owner approves one version in Spellbook; agents see every variant with its source until then.
  • •Every change goes back through the layer's own repo, as a diff the owner merges, with a receipt.

What each layer does, and why teams keep it

Each of these tools earned its place, and the reason teams keep it is the reason Data Workers builds on it.

LayerWhat it does bestAgent and MCP surfaceStatus, October 2026
CubeAgentic analytics platform built on a semantic layer: data model of cubes and views, pre-aggregations, Analytics Chat, Workbooks, Dashboards, embedded analyticsCube MCP server with 30 tools for query, discovery, dashboards, data model editing and Git commits; REST /v1/meta and /v1/load; SQL APIMCP server available on all plans; model editing and branch workflows over MCP added in 2026
AtScaleUniversal semantic layer defined in SML, queried through SQL, MDX, DAX, Python, REST and MCP; its AI Computation Engine computes each metric the same way for every toolRead-only MCP tools (list_models, explore_columns, focus_columns, run_query, get_outbound_queries) over governed models, SELECT onlyMCP server documented for AtScale deployments; SML is Apache 2.0
dbt Semantic LayerMetricFlow metrics next to the dbt models they read, under dbt's review and CIdbt MCP Semantic Layer tools (list_metrics, get_dimensions, query_metrics and more)Starter, Enterprise and Enterprise+; MetricFlow Apache 2.0
OssieVendor-neutral YAML spec for semantic models, with converters for many layersConverters for Cube, dbt, Databricks, Snowflake, Microsoft and others in the Ossie repoIncubating at the ASF since July 2026; spec 0.1.1 tagged as a release candidate, 0.2.0 in draft

A few details matter for the wiring. Cube's MCP server marks nine tools as destructive, including writeDataModelFile and mergeToDefaultBranch, and clients that honor those annotations, Claude among them, ask for confirmation first. Its model editing and commit tools only register for roles that can edit the model, so a Viewer session reads and queries without them. AtScale goes further by design: every tool on its MCP server is read-only and the server rejects anything but a SELECT. AtScale's SML project ships converters to Snowflake Cortex semantic models, Databricks UC metrics and Power BI. And the Ossie project now counts more than 50 organizations, Cube, AtScale, dbt Labs and Microsoft among them, and Microsoft has contributed a Power BI and Fabric converter. For a wider survey of the category, see semantic layer tools compared.

What Data Workers reads, and what it sends back

Data Workers readsData Workers sends back
Cube: cubes, views, measures and dimensions from the model files in Git, checked against the deployed model with the Cube MCP server (searchDataModel, runQuery) or REST /v1/metaA conflict report: each variant side by side, with its layer, file and the time it was read
AtScale: metrics, dimensions and hierarchies from the SML files in your Git repository, checked with AtScale's MCP server (list_models, run_query)An owner review in Spellbook, routed to the metric's named owner
dbt MetricFlow: metrics, entities and dimensions from the semantic manifest or the dbt MCP Semantic Layer toolsA proposed change as a diff for the layer's own repo
Ossie YAML, converted with the Ossie project's convertersThe approved definition, served to every agent that asks Data Workers over MCP
Warehouse signals: lineage, freshness, quality checks and query historyA receipt for every decision: variants, approver, checks and time
What Data Workers reads from Cube, AtScale, MetricFlow and Ossie and what it writes back through Cube, AtScale, MetricFlow and Ossie

Here is how each source lands. dbt MetricFlow goes straight in: import_semantic_definitions with format: metricflow loads metrics, dimensions and entities into the Context Wizard metric store, the same importer that handles Databricks metric views and Wren MDL. Context Wizard can also read a compiled dbt manifest with Cube's dimension typing (format: cube_dbt), so dbt models arrive as typed, key-aware dimensions. Cube data models and AtScale SML come in from their Git repositories and over their MCP servers today: your coding agent or a scheduled run reads each definition and records it with define_metric, with its source, owner, expression and grain, then uses each layer's MCP server to confirm what is deployed. Ossie YAML takes one conversion step: the Ossie project's apache-ossie-dbt converter (ossie-dbt ossie-to-msi) turns Ossie into a MetricFlow semantic manifest, and Context Wizard imports that.

The join is what makes this more than a list of definitions. The Data Context & Catalog agent ties each metric to the tables it reads, the quality checks and lateness baselines on those tables, and the query history that shows who uses it. When arr disagrees across layers, you see which dashboards, Excel models and agents read each version, and whether the table under it passed its checks this morning.

How it fits together

How Data Workers fits with Cube, AtScale, MetricFlow and Ossie: your coding agent on top, Data Workers in the middle, your estate underneath

Your team works in its coding agent (Claude Code, Codex or Cursor) and reviews in Spellbook Data Catalog, which is in preview. The Data-Agents Swarm does the reading and checking, and the Autonomous Data-Conductor sequences each run and holds every change at the autonomy level you set. Semantic definitions stay in Cube, AtScale and dbt. Nothing migrates, and Data Workers keeps metadata and scrubbed facts, not copies of your tables.

One incident, every handoff

Here is a run most teams with more than one semantic layer will recognize. It's an illustration, not a customer case. Seven systems are involved: GitHub, the dbt platform, Cube Cloud, AtScale, Snowflake, Spellbook and a coding agent.

TimeSystemWhat happensWho decides
08:40GitHubA pull request changes the MetricFlow arr metric to exclude paused subscriptions. It merges.An engineer
09:00dbt platformThe job builds; the semantic manifest now carries the new arr.dbt
09:05Cube Cloud, AtScaleData Workers imports the manifest and reads arr from the Cube model file and the SML file in Git, and your team's assistant confirms the deployed Cube member over Cube's MCP server. Both still include paused subscriptions. scan_for_contradictions flags the conflict.Data Workers, read-only
09:08SnowflakeLineage and query history show AtScale's arr feeds the finance Excel model behind the board pack, and Cube's feeds the customer-facing usage dashboard. The fct_subscriptions table passed its freshness and quality checks, so the gap is definitional.Data Workers, read-only
09:10SpellbookThe conflict goes to the Promotion Inbox, tagged with the metric's named owner, the finance analytics lead, with all three expressions, their readers and the commit that changed dbt.Data Workers routes
10:30Cube CloudA product manager asks Analytics Chat for ARR. Cube answers from its own model, as designed.Cube
10:35Coding agentAn analyst asks through Data Workers. resolve_metric returns all three candidates with their sources instead of picking one silently.Data Workers, read-only
13:00SpellbookThe owner approves the dbt definition: paused subscriptions are not recurring revenue.The named owner
13:10GitHubData Workers proposes the matching change as a diff for the Cube model repo and for the SML repo, each citing the approval and the blast radius.Data Workers proposes
15:20Cube Cloud, AtScaleThe Cube and AtScale owners merge and deploy. Cube rebuilds the affected pre-aggregations.Each layer's owner
16:00SnowflakeData Workers re-reads all three definitions, runs the same arr query through each layer, confirms the values match, and closes the conflict with a receipt.Data Workers, read-only
Incident timeline across the stack: what Cube, AtScale, MetricFlow and Ossie, your team and Data Workers each do, step by step

Every change went through the layer that owns it: dbt's CI, Cube's branch and deploy, AtScale's SML repo. Data Workers supplied what no single layer owns: the cross-layer comparison, the readers and data health behind each version, one approval, the matching proposals and the check that the numbers agree.

The receipt. resolve_contradiction only accepts a named human approver, and the write goes through the governed path: PII scrub, tenant isolation, an authority guard that stops any agent approving its own work, and a hash-chained log. The receipt, fetched with get_change_receipt, holds:

FieldValue in this run
Metricarr, finance domain
Variantsdbt MetricFlow (excludes paused, commit at 08:40); Cube model (includes paused, read 09:05); AtScale SML (includes paused, read 09:05)
Data healthfct_subscriptions fresh and passing checks at 09:08
ReadersBoard-pack Excel model (AtScale), customer usage dashboard (Cube), two agents
ApprovedThe dbt definition, as authoritative
ApproverThe finance analytics lead, 13:00
ChangesCube and SML diffs, merged by their owners 15:20
ChecksAll three re-read and queried at 16:00; values matched

Why doesn't any one semantic layer just do this itself?

Each layer built a great product for one job: defining metrics once and serving them well to the tools that query it. That focus is the right design. Cube's model is authoritative for Cube's APIs, its chat and its embedded analytics. AtScale's SML is authoritative for every DAX, MDX and SQL query that runs through AtScale. MetricFlow is authoritative inside the dbt project, under dbt's review. Each MCP server serves its own model, and that's exactly what a good semantic layer should do.

Settling a disagreement between layers is a different job. Cube's MCP server has no input for an AtScale definition to compare against, and AtScale's has none for Cube's. Ossie standardizes the format, which makes comparison easier; it doesn't decide which version wins, and it doesn't say whether the table underneath passed its checks this morning. Deciding needs an approval a business owner signs, proposals into repositories the layer doesn't own, rollback for each one, and someone accountable for a call that crosses vendors. A semantic layer vendor taking that on would be reaching into its competitors' repos, which is a sensible thing for each of them to avoid. That neutral, cross-layer authority, joined to lineage, quality and usage, is the product Data Workers is.

Setup

Data Workers agents are MCP servers. Register Context Wizard next to the semantic layer MCP servers your team already uses.

Prerequisites

  • •Cube: a Cube Cloud user with the Viewer role for the MCP server (OAuth with PKCE), which sees searchDataModel and runQuery but none of the model editing tools, plus read access to the Git repository that holds the Cube model; developer access to that repo only when you turn on proposals.
  • •AtScale: the AtScale MCP server enabled on your instance (atscale-mcp: enabled: true in the Helm values), an AtScale API token or the atscale-mcp OAuth client, and read access to the Git repository that holds your SML.
  • •dbt: the semantic manifest from your CI artifacts, or for the dbt remote MCP server a personal access token (dbt's recommendation for Semantic Layer tools) or a service token with Semantic Layer Only, Metadata Only and Developer permissions, plus your production environment ID.
  • •Ossie, if you use it: the apache-ossie-dbt converter from the Ossie repository (converters/dbt) to turn Ossie YAML into a MetricFlow manifest.
  • •Warehouse: a read-only role for lineage, freshness checks and query history.
  • •An owner per metric, named in Spellbook, so every conflict has somewhere to go.

Example (.mcp.json for Claude Code; replace every <your-...> value with your own):

{
  "mcpServers": {
    "cube": { "type": "http", "url": "https://cubecloud.dev/mcp" },
    "atscale": {
      "command": "npx",
      "args": ["-y", "mcp-remote", "https://<your-atscale-host>/mcp",
               "--header", "Authorization: Bearer ${AUTH_TOKEN}"],
      "env": { "AUTH_TOKEN": "<your-atscale-api-token>" }
    },
    "dbt": {
      "type": "http",
      "url": "https://<your-dbt-host>/api/ai/v1/mcp/",
      "headers": {
        "Authorization": "Token <your-dbt-access-token>",
        "x-dbt-prod-environment-id": "<your-prod-environment-id>"
      }
    },
    "dw-context-catalog": {
      "command": "./start-agent.sh",
      "args": ["dw-context-catalog"],
      "env": {
        "SNOWFLAKE_ACCOUNT": "<your-org>-<your-account>",
        "SNOWFLAKE_USER": "<your-dw-reader-user>",
        "SNOWFLAKE_PASSWORD": "<your-dw-reader-password>",
        "SNOWFLAKE_ROLE": "DW_SEMANTIC_READER",
        "SNOWFLAKE_WAREHOUSE": "ANALYTICS_XS"
      }
    }
  }
}

Then ask in plain words: "Read arr, net_revenue and active_customers from Cube, AtScale and our dbt manifest, and show me every metric defined differently." The coding agent reads the Cube model file and the SML file from Git, checks the deployed Cube members with searchDataModel, imports dbt with import_semantic_definitions, records the Cube and AtScale definitions with define_metric, and runs scan_for_contradictions. Everything starts at L1, observe: Data Workers reads and reports, and changes nothing. Connect Cube with a Viewer role for this work: the model editing and commit tools never register, and Data Workers doesn't need them to read.

The per-platform how-tos go deeper on the native layers: Snowflake semantic views over MCP and Unity Catalog metric views and lineage. For the dbt side of the wiring, read Data Workers + dbt.

Guardrails: approvals, and what each layer owns

What the layers stay responsible for. Metric definitions, their repositories and review. Cube's deploys, pre-aggregations and access policies. AtScale's models, query engines and security policies. dbt's CI and jobs. Ossie's spec and converters. Data Workers never bypasses any of them.

What Data Workers enforces on top.

  • •Read-only start. New deployments observe and report, one domain at a time.
  • •Autonomy per domain, L0 to L4. Description refreshes can move faster while definitions stay at "propose".
  • •Promotion is human. Only a named owner can make a definition authoritative; an authority guard in the write path stops any agent approving its own work.
  • •Changes through each layer's own path. A diff for that layer's repo; its owner merges and deploys.
  • •Receipts and rollback. Every decision records variants, readers, data health, approver and checks in a hash-chained log, and every merged change reverts like any commit.
The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous

The ladder runs L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous, set per domain. Most teams run metric conflicts at L1 for a few weeks, then move to L2 so Data Workers drafts the matching change in each layer. For how the levels work across an estate, read the autonomous data platform playbook.

What changes for your team

Analytics engineers stop finding out about metric drift from a finance question. Each conflict arrives with every variant, its readers and the data health underneath, ready for one decision. Semantic layer owners keep their models and review proposals like a colleague's pull request. The finance analytics lead becomes the named owner of the board metrics, with one inbox instead of three Slack threads. Data platform leads get one record of every definition decision across Cube, AtScale and dbt, which makes audits a matter of reading receipts. And the "which semantic layer do we standardize on" project can stop being a blocker: keep the layers each team loves, and govern the meaning across them.

The case for your CFO

The outcome. Every board metric has one approved definition across every semantic layer, and every agent, dashboard and Excel model that reads it gives the same number, with a record of who approved it.

The risk story. The exposure is two numbers for ARR in front of the board, found after the deck ships. At L1 Data Workers only reads definitions, lineage, quality checks and query history, and reports. At L2 it drafts changes as diffs; each layer's owner applies them. No agent can promote its own definition, every change has a rollback path, and every receipt shows the variants, the readers, the approver, the checks and the time. Nothing migrates.

Why now. Cube, AtScale and dbt all ship MCP servers, so agents are already reading each layer directly. Ossie entered the Apache Incubator in July, which makes definitions portable and easier to compare. More agents reading more copies of each metric means more ways to disagree.

The first win. The board-pack metrics. Read their definitions from every layer, and give each owner a list of conflicts with readers and data health attached inside the first week of a pilot.

What stays the same. Cube, AtScale, dbt, your BI tools and Excel, your repositories and your deploy process.

The pilot path. Start with a pilot on one domain, read-only first. The pilot is credited in full against the first year.

One sentence for upstairs: "Cube, AtScale and dbt keep defining our metrics; Data Workers checks every definition against the others and the data underneath, and a named owner approves the one version every agent uses."

FAQ

Do we have to standardize on one semantic layer? No. Keep the layers your teams chose. Data Context Wizard reads each one as a source and governs the meaning across them, so the standardization decision stops blocking agents.

Does Ossie solve metric conflicts on its own? Ossie gives every layer a common format, which makes definitions portable and easy to compare. Deciding which version wins, and checking the data under it, still needs an owner and a record. Data Workers adds both.

Will Data Workers change our Cube models or AtScale SML? Only through their owners. It proposes a diff for the layer's repo; the owner merges and deploys.

How does Data Workers read AtScale today? From the SML files in your Git repository, recorded in Context Wizard with source, owner and grain, with your team's assistant using AtScale's read-only MCP server to confirm the deployed model and run the same query.

What if Cube's Analytics Chat and a coding agent answer differently during a conflict? Cube answers from its own model, as designed. Agents that ask Data Workers get every variant with its source until the owner decides.

Does this work with dbt MetricFlow open source? Yes. MetricFlow is Apache 2.0, and Context Wizard imports its definitions directly, whether they come from the dbt platform or your own builds. For the dbt-specific view, see the dbt Semantic Layer as an agent layer.

Sources

Capabilities and statuses, checked October 2, 2026: Cube and its changelog (positioning, products, MCP server January 2026, MCP branch and Git workflow, entries to Sep 25, 2026), the Cube MCP server (available on all plans, OAuth, endpoint, 30 tools, role-gated model editing, nine destructive tools that prompt), the Cube REST API reference (/v1/meta, /v1/load), the AtScale platform (SML, MCP, Design Center), AtScale semantic context for AI over MCP, AtScale's MCP Server Tools Reference (read-only tools, SELECT only), enabling the AtScale MCP Server and its development environment setup (https://<atscale_url>/mcp, OAuth client or API token), SML on GitHub (Apache 2.0, converters), AtScale's Microsoft joins open semantic standard (Sep 29, 2026), the dbt Semantic Layer (plans; updated Sep 29, 2026), the dbt MCP server (Semantic Layer tools), dbt's remote MCP setup (https://YOUR_DBT_HOST_URL/api/ai/v1/mcp/, token and environment headers), Apache Ossie enters the Apache Incubator (July 10, 2026), the Ossie ecosystem and the Ossie repository (tag osi-0.1.1-rc1, core-spec/spec.md at 0.2.0.dev0 draft, Cube, dbt and Microsoft converters, apache-ossie-dbt README). Product names and statuses change quickly; if we've got something wrong, tell us and we'll fix it.