You're on OSI Semantics: Keep Your Ossie YAML True in Every Engine, With an Owner on Every Definition
Your Ossie (OSI) YAML defines metrics, dimensions and joins once. Data Context Wizard imports it, checks every engine's copy and the data under it, and routes drift to an owner.
Your team made the call: metrics, dimensions and joins are defined once, in YAML, in the repo. A semantic model lists its datasets with their sources and primary keys, the relationships between them, the fields people group and filter by, and the metrics, each expression written in ANSI SQL or a named dialect, with ai_context telling language models how to use it. CI runs the converters: Ossie to a Databricks metric view, Ossie into a Snowflake semantic view through SYSTEM$CREATE_SEMANTIC_VIEW_FROM_OSSIE_YAML (in preview), Ossie to a dbt semantic manifest with ossie-dbt ossie-to-msi. The project you adopted as Open Semantic Interchange (OSI) is now Apache Ossie (incubating), and your files didn't change. Ossie YAML is your source of meaning. Data Workers is the layer that keeps it true in every engine.
That second job starts the day after adoption. One YAML file becomes three or four engine objects, each owned by a different platform, each read by different agents, each sitting on tables that change under it. Data Context Wizard imports your Ossie definitions, compares them with what every engine actually holds and with the data underneath, and puts a named owner on every definition, so drift reaches a person with a proposed fix instead of reaching an agent's answer.
Key takeaways
- •Your Ossie YAML stays the source. The files, the converters and your pull request review stay exactly as they are. Data Workers reads the YAML and every engine's copy of it.
- •The import path uses Ossie's own tooling. Convert with the Ossie project's dbt converter (
ossie-dbt ossie-to-msi), and Data Context Wizard imports the resulting MetricFlowsemantic_manifest.jsonas governed definitions. - •Every engine is checked against the YAML. Data Workers imports Databricks metric views and dbt's semantic manifest, and takes Snowflake semantic views from the Ossie YAML your team exports, so a skipped field, a simplified metric or a native edit shows up as a conflict with its source.
- •The data is checked too. A renamed source column breaks the same expression in every engine at once. Data Workers catches the schema change, traces it to each engine and agent, and proposes the fix.
- •Owners decide, agents follow. Conflicts go to the definition's named owner in Spellbook. Agents never promote a definition to authoritative on their own, and every change leaves a receipt.
Ossie YAML is your source of meaning. Data Workers is the layer that keeps it true in every engine.
A semantic model in Ossie is portable by design, and each converter is faithful to its target. The work that remains is noticing when the copies stop agreeing with the source, or with the data. Here is a Tuesday on a stack that runs Ossie YAML into Databricks, Snowflake and dbt. This is an illustration, not a customer case.
| Time | System | What happens |
|---|---|---|
| 09:10 | GitHub | A pull request adds fiscal_quarter to semantics/revenue.ossie.yaml. The author writes the expression in the Databricks dialect only. Review passes and it merges |
| 09:25 | Databricks | CI runs the Ossie to metric view converter; the metric view revenue_mv gains fiscal_quarter |
| 09:26 | Snowflake | CI recreates the semantic view REVENUE_SV from the same model, stamped with version 0.1.1, the version Snowflake's procedure accepts. Snowflake prefers a SNOWFLAKE expression and falls back to ANSI_SQL; a field with neither is skipped, as Snowflake documents. The view has no fiscal_quarter |
| 09:40 | Data Context Wizard | The scheduled check re-imports the YAML and reads each engine. The Ossie model, the Databricks metric view and the dbt manifest all define fiscal_quarter; the Snowflake view does not. Blast radius: CoWork, a Sigma workbook and the finance board pack all read REVENUE_SV |
| 09:44 | Data Context Wizard | Data Workers proposes a one-line change to the YAML: an ANSI_SQL expression for fiscal_quarter, with a preview showing it returns the same quarter for every order date in the last eight quarters |
| 10:05 | Spellbook | The finance analytics lead, the named owner of the revenue model, approves and merges the change |
| 10:20 | Snowflake | CI recreates the view. Data Workers verifies that DESCRIBE SEMANTIC VIEW lists fiscal_quarter and that net revenue by fiscal quarter matches Databricks for the last eight quarters |
| 11:00 | CoWork | The CFO's team asks for net revenue by fiscal quarter. The answer matches the Databricks dashboard to the cent, and the receipt shows why |

Every tool in that chain did what it was designed to do. The YAML was valid, the Databricks converter was correct, and Snowflake behaved exactly as documented. Without the check at 09:40, CoWork would have answered by calendar quarter, or not at all, and nobody would have known which copy was out of step. What changed is that one layer was comparing every copy with the source and the data, and the fix went through the owner.
| Job | What Ossie YAML does | What Data Workers does |
|---|---|---|
| The definition | Holds datasets, fields, relationships and metrics in one vendor-neutral file | Imports it with provenance and gives it a named owner |
| The copies | Converters create each engine's native object | Reads every engine's live object and compares it with the YAML |
| The data | Expressions reference tables and columns | Watches those tables for schema changes, freshness and quality, and traces breaks to every engine |
| The conflict | Pull request review on the file | Routes each conflict to the owner with both versions, sources and blast radius |
| The fix | A new commit, converted again by CI | Proposes the change, verifies every engine afterwards and records the receipt |
| The agent's answer | ai_context guides the model | resolve_metric returns the approved definition, or every candidate when they disagree |
Why doesn't Ossie just do this itself?
Because Ossie is a specification and a set of converters, and that is exactly the right shape for it. A file format has to be neutral, stateless and easy for every vendor to implement. The converters translate one document into one target, offline, and they are careful about it: the dbt converter records every lossy step as a warning (conversion metrics and private metrics dropped on the way into Ossie, cumulative windows reduced to their base aggregation), and on the way back it gives every time dimension a day grain because Ossie carries no granularity field. Each engine owns its import, and Snowflake chose to skip fields it can't express rather than fail the whole view. Those are sensible decisions for a standard.
Keeping the copies true is a different job. It needs a live view into every engine after the import, a view of the tables and columns each expression reads, a record of who owns each definition, and the authority to propose changes in tools owned by other teams. When a fix touches production, it needs an approval, a rollback path, a verification across engines and a receipt. A neutral specification shouldn't carry runtime access to your warehouses or liability for changes in Snowflake and Databricks. That is the product Data Workers is.
There is also the spec itself moving forward. The repository's converters work on the 0.2.0 draft shape (one flat model per document, version 0.2.0.dev0), and ossie-to-msi reads only that flat shape. Snowflake's import procedure currently accepts version 0.1.1, and its export, SYSTEM$READ_OSSIE_YAML_FROM_SEMANTIC_VIEW, returns the 0.1.1 document with its semantic_model list. Teams that adopt early hold more than one shape for a while, which is normal for a young standard. A layer that reads what each engine actually built, regardless of the version that produced it, keeps that transition quiet.
Every tool owns a slice. Data Workers covers the whole lifecycle
Ossie YAML owns one slice of the data lifecycle, and it owns it well: portable metric, dimension and join definitions that every engine can import. Each point tool adds another console, another contract and another handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and builds on your Ossie YAML instead of replacing it.

| Stage | Data Workers | Ossie YAML | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 7 | Ossie YAML describes datasets, fields, relationships and metrics in one portable file, with ai_context for agents. Data Workers joins those definitions to lineage, quality, usage and a named owner in one governed context graph. |
| Analytics & Insights | 8 | 8.5 | A home stage: one definition, converted into Snowflake semantic views, Databricks metric views and dbt MetricFlow, so BI tools and agents query the same metric. Data Workers answers from the approved definition with lineage behind the number. |
| Data Quality | 8 | 2 | The spec declares keys, and the repo's validation tooling checks documents against the schema. Data Workers checks the data under each expression and turns breaks into tests. |
| Observability & Incidents | 8.5 | 1 | Ossie is a file format, so runtime monitoring belongs to the engines. Data Workers detects breaks in the tables a definition reads and traces them to every engine and agent. |
| Pipelines & Ingestion | 8.5 | 2.5 | Converters move definitions between tools; moving data is outside the spec's scope by design. Data Workers builds, reruns and backfills the pipelines under those tables, with approvals. |
| Schema & Migration | 8 | 8.5 | A second home stage: hub-and-spoke converters move a semantic model between engines without a rewrite. Data Workers catches upstream schema changes that break an expression and plans warehouse migrations in approved waves. |
| Governance & Access | 8.5 | 3 | Definitions live in git and change through pull request review; access control stays with each engine. Data Workers puts a named owner on every definition and routes conflicts to that owner. |
| Security & Privacy | 8 | 2 | Plain Apache 2.0 files in your repo, with security enforced by the engines that run them. Data Workers leaves a receipt on every change it proposes or makes. |
| Cost / FinOps | 8 | 1.5 | Cost sits with the engines that execute each metric. Data Workers traces Snowflake credits to the dbt model behind them and drafts the fix for its owner. |
| MLOps & Models | 7.5 | 2.5 | ai_context gives language models instructions and synonyms for a model. Data Workers keeps the data under your own models healthy and connects to MLflow and W&B. |
How Ossie YAML and Data Workers work together
Your YAML stays in git and your CI keeps running the converters. Analysts and agents stay where they ask questions: Claude Code, Codex or Cursor for the engineers, CoWork and Genie inside the platforms. Spellbook Data Catalog (in preview) is where owners look: each conflict, each proposed change, who approved it and how to roll it back. Between them, Data Context Wizard holds your Ossie definitions next to lineage, quality and usage, the Data-Agents Swarm does the checking and fixing, the Autonomous Data-Conductor runs each fix end to end (detect, diagnose, fix, review, verify, remember), and per-domain guardrails hold approvals, receipts and rollback.

Bringing Ossie definitions in. Data Context Wizard imports semantic definitions as MetricFlow, and Ossie ships its own bidirectional dbt converter, so the documented path uses tooling the Ossie community maintains. Clone the Ossie repository and install apache-ossie-dbt from source in its converters/dbt folder (the converter README documents uv sync for that folder), convert a flat 0.2.0-shape document, and import the manifest. Ossie to MetricFlow keeps simple aggregations, ratios and raw expressions, rejects composite keys, and sets time dimensions to a day grain; record any field the conversion simplifies, such as a fiscal-quarter grain, as a governed definition in Context Wizard with its owner. Then connect the read tools to the client your team uses, following the documented client setup path.
# Example: Ossie YAML into Data Context Wizard (paths are placeholders)
# 1. Install the Ossie project's own dbt converter from source
git clone https://github.com/apache/ossie.git
cd ossie/converters/dbt && uv sync
# 2. Convert a flat (0.2.0.dev0) Ossie document to a MetricFlow manifest
uv run ossie-dbt ossie-to-msi -i /path/to/semantics/revenue.ossie.yaml -o /path/to/build/revenue_semantic_manifest.json
# 3. Import build/revenue_semantic_manifest.json into Data Context Wizard
# as MetricFlow definitions in the "finance" domain
# 4. Connect the open-source agents to Claude Code (clone the repo first)
claude mcp add dw-context-catalog -- /path/to/dataworkers-claw-community/start-agent.sh dw-context-catalog
claude mcp add dw-schema -- /path/to/dataworkers-claw-community/start-agent.sh dw-schema
claude mcp add dw-quality -- /path/to/dataworkers-claw-community/start-agent.sh dw-qualityRun the conversion and import in the same CI job that runs your other converters, so Context Wizard always holds the version you just shipped. For the engine side, each copy comes in from the platform's own export today: Snowflake returns any semantic view as Ossie YAML through SYSTEM$READ_OSSIE_YAML_FROM_SEMANTIC_VIEW (or DESCRIBE SEMANTIC VIEW), Databricks metric views are YAML definitions in Unity Catalog, and dbt serves its metrics through the Semantic Layer API. From the client, ask in plain words: "List the revenue metrics and show me any engine whose definition differs from the Ossie model." The agent uses list_semantic_definitions and resolve_metric, which returns every candidate with its source when definitions disagree rather than picking one silently. trace_cross_platform_lineage and blast_radius_analysis show which tables, dashboards and agents read a metric; pull-request review and the dbt manifest catch changes to the columns your expressions reference; run_quality_check and get_quality_score cover the data. When the owner settles a question, mark_authoritative records the decision, and get_authoritative_source is what every agent reads afterwards.
One request, L0 to L4. The request: "Make sure fiscal_quarter means the same thing everywhere." The autonomy ladder is set per domain.

- •L0 manual. An analytics engineer runs
DESCRIBE SEMANTIC VIEW, opens the metric view YAML and diffs both against the Ossie file by hand. - •L1 observe. Data Workers reports that Snowflake's view lacks
fiscal_quarter, names the dialect cause and lists the three assets that read the view. Nothing changes. - •L2 propose. Data Workers drafts the ANSI_SQL expression with a value preview and routes it to the owner in Spellbook. Nothing merges until the owner approves.
- •L3 act reversibly. For change classes with a proven record, such as re-running the conversion after an approved merge, Data Workers runs it, verifies every engine and can roll back to the prior view.
- •L4 autonomous. For a scoped domain, Data Workers keeps engines in step with the approved YAML overnight and posts the receipt for review. Definition changes themselves always go to the named owner.
For the full safety model, read is it safe to let AI agents change production data; for where data and credentials live, read where does our data go. If you are still deciding whether to standardize on the spec, you're on Open Semantic Interchange covers the project and what a standard settles. For the engines themselves, see Snowflake semantic views over MCP into Data Context Wizard, Data Workers with Cube, AtScale, MetricFlow and OSI, you're on Snowflake semantic views and you're on the dbt Semantic Layer. The section hub is bring your own context, and our semantic layer tools comparison maps the wider field.
What changes for your team
Adopting Ossie gave your team one place to write a definition. Data Workers gives it one place to know whether every copy still holds.

- •Incidents. A renamed source column that breaks an Ossie expression is caught at the schema change, traced to every engine and agent, fixed and verified.
- •Data quality. Every metric in the YAML gets checks on the tables it reads, so a bad load is caught before an agent answers from it.
- •Cloud spend. Snowflake credits are traced to the query and dbt model behind them, and each fix goes to its owner drafted.
- •Access. A request to query a newly shared metric arrives as a scoped grant proposal for the owner in each engine.
- •Audits. Each definition carries a named owner, an approval history and a receipt for every change, in every engine.
- •Migrations. Moving a metric set to a new engine runs in waves, with values compared against the old engine before cutover.
The analytics engineers keep writing YAML. What they stop doing is the weekly diff between three engines, and the Slack thread that starts with "which revenue is right?"
Keep Ossie YAML, or consolidate?
Keep Ossie YAML if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
For Ossie, keeping it is the natural choice: it is your portable record of meaning, owned by you and readable by every engine you run. What teams consolidate is the work around it: the hand-written scripts that diff engines, the separate data-quality tool watching the same tables, and the spreadsheet of who owns which metric. If you are weighing whether to build that layer yourself on top of the converters, read build it ourselves with Claude Code and MCP servers first; the conversion is the easy part, and the live comparison, ownership, approvals and rollback are where the work is.
The case for your CFO
The outcome: the company standardized its metric definitions so every dashboard and every agent gives the same number. Data Workers makes sure that stays true after the files leave the repo: every engine's copy matches the approved definition, the data under it is healthy, and each definition has a named owner.
The risk story is plain. A standard makes definitions portable; it does not watch the engines that import them. Data Workers does, with autonomy set per domain on the ladder from L0 manual to L4 autonomous. Every change is routed to a named approver, applied reversibly, verified across engines and recorded in a receipt with who approved it, what it touched and how to undo it. Agents never approve their own changes. Zero migration: the YAML, the converters, Snowflake, Databricks and dbt stay where they are.
Why now: agents now query every engine directly, around the clock, and two engines that disagree produce two confident answers. The first win is the board metrics: import the Ossie models behind them, compare every engine, and resolve the conflicts with their owners in the first weeks. What stays the same: your repo, your review process, your converters and every platform's permissions. For the numbers, see the ROI of agentic data operations. Start with a pilot (pricing); the pilot is credited in full against the first year.
The sentence to repeat upstairs: "We define every metric once in Ossie; Data Workers makes sure every engine still matches it, and an owner signs off on every change."
Getting started
Start with a pilot. Pick the Ossie models behind your board or revenue metrics, add the ossie-to-msi conversion and the import to the CI job that already runs your converters, and connect Context Wizard to the engines those models feed. Within the first weeks you get a conflict list with owners, and a watch on the columns every expression reads. Then turn on the first proposal class in one domain. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.
FAQ
Does Data Workers read Ossie YAML directly? Data Context Wizard imports semantic definitions as MetricFlow, and Ossie's own apache-ossie-dbt converter turns Ossie YAML into a MetricFlow semantic_manifest.json. So the documented path is ossie-dbt ossie-to-msi, then a MetricFlow import into Context Wizard. It uses the converter the Ossie community maintains, which keeps the mapping in their hands.
What happens to fields the conversion simplifies? The dbt converter documents its choices: composite keys are rejected, time dimensions get a day grain, and anything that isn't a simple aggregation or a ratio is kept as a raw expression. Record anything that needs more, such as a fiscal-quarter grain, as a governed definition in Context Wizard with its owner, so agents read the full meaning.
Will Data Workers edit our Ossie files? Data Workers proposes changes; your owner approves them and they merge through your normal review. The YAML stays the source, and the engines are rebuilt from it by your CI as before.
We are on spec 0.1.1 and the repo is moving to 0.2.0. Does that matter? It matters for the tooling. The 0.2.0 draft changes the document shape (one flat model per file), the repository's converters, ossie-to-msi included, read the flat shape, and Snowflake's import procedure currently accepts version 0.1.1. Keep your source in the flat shape and stamp the version each target accepts in CI. Data Workers compares what each engine actually built, so it works the same on either side of the transition and shows you where shapes or results differ.
Which engine wins when two copies disagree? Neither, automatically. The conflict goes to the definition's named owner with both versions, their sources and the blast radius. Once the owner decides, mark_authoritative records it, and agents read the approved definition from then on.
Does this replace our semantic layers? No. Snowflake semantic views, Databricks metric views and the dbt Semantic Layer keep serving queries in their platforms. Data Workers keeps them consistent with the Ossie source and with the data underneath.
Sources
- •Apache Ossie (incubating), homepage (previously Open Semantic Interchange; YAML for metrics, dimensions and joins; news), https://ossie.apache.org/ (checked Oct 2, 2026)
- •Apache Ossie, "Apache Ossie (Incubating): The New Name for Open Semantic Interchange" (July 10, 2026), https://ossie.apache.org/updates/ossie-enters-apache-incubator/ (checked Oct 2, 2026)
- •Apache Ossie, Ecosystem, https://ossie.apache.org/ecosystem/ (checked Oct 2, 2026)
- •apache/ossie on GitHub (core-spec 0.2.0.dev0 draft, tag osi-0.1.1-rc1, converters), https://github.com/apache/ossie (checked Oct 2, 2026)
- •apache/ossie, converters/dbt README (ossie-to-msi, msi-to-ossie, conversion notes), https://github.com/apache/ossie/blob/main/converters/dbt/README.md (checked Oct 2, 2026)
- •apache/ossie, converters README (flat 0.2.0.dev0 format, hub and spoke), https://github.com/apache/ossie/blob/main/converters/README.md (checked Oct 2, 2026)
- •apache/ossie, converters/databricks README (Ossie and metric view converters), https://github.com/apache/ossie/blob/main/converters/databricks/README.md (checked Oct 2, 2026)
- •Snowflake, SYSTEM$CREATE_SEMANTIC_VIEW_FROM_OSSIE_YAML (Preview; currently supported version 0.1.1; dialect fallback; fields with only unsupported dialects omitted), https://docs.snowflake.com/en/sql-reference/stored-procedures/system_create_semantic_view_from_ossie_yaml (checked Oct 2, 2026)
- •Snowflake, SYSTEM$READ_OSSIE_YAML_FROM_SEMANTIC_VIEW (returns a semantic view as Ossie YAML, version 0.1.1 wrapper format), https://docs.snowflake.com/en/sql-reference/functions/system_read_ossie_yaml_from_semantic_view (checked Oct 2, 2026)
- •apache/ossie, converters/dbt pyproject.toml (package apache-ossie-dbt 0.2.0.dev0,
ossie-dbtcommand, apache-ossie from the repository's python folder), https://github.com/apache/ossie/blob/main/converters/dbt/pyproject.toml (checked Oct 2, 2026) - •Data Workers open-source repository (dw-context-catalog, dw-schema and dw-quality tool registrations) and client setup guide, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)