You're on Kaelio's ktx: Keep the Data Under Every Approved Measure Right, Across the Whole Lifecycle
Already on Kaelio's ktx? Keep its git-native context for your data agents, and add Data Workers for lineage, quality, incident fixes, approvals and receipts across engines.
Your analytics engineers ran ktx setup, pointed it at Snowflake, the dbt project, the MetricFlow metrics and the Looker repo, and let ktx ingest do the first pass. Now the project lives in git: semantic sources, joins and measures as YAML under semantic-layer/, definitions and caveats as Markdown under wiki/, every ingest landing as a diff someone reviews and merges. Claude Code, Codex and Cursor reach it through the ktx MCP server, search the wiki, read a source, compile bookings.total_amount with the fan traps already resolved, and run read-only SQL. ktx is the analytics agent's context, kept as code. Data Workers is the agentic data platform around it: Data Context Wizard reads that context with its provenance, next to catalogs, semantic layers and lineage across every engine, and the Data Workers agents keep the data under each measure right, with one approval flow and a receipt for every change.
That split shows up the first morning a source system changes overnight. ktx compiles exactly the measure your team approved. Making sure the rows under it arrived intact, finding the pipeline that broke and fixing it safely is the job this guide covers.
Key takeaways
- •ktx keeps its job. The
semantic-layer/andwiki/files, the ingest-and-review loop and the MCP server stay as they are. Data Workers runs next to them in the same agent client from day one. - •Your ktx context becomes one governed source among several. Data Workers connects to it over the ktx MCP server, and Data Context Wizard keeps it next to dbt, MetricFlow, your catalogs and cross-engine lineage, with the source and owner of every fact recorded.
- •The data under each measure is checked. Data Workers watches the tables ktx's measures read, traces breaks across Salesforce, Fivetran, dbt and Snowflake, and fixes them through approvals.
- •Fixes land upstream; context goes back through ktx's review. Data Workers proposes dbt changes as diffs for the owner to merge. Caveats return through ktx's own
memory_ingestand git history. - •Start with a pilot. Read tools first in one domain, then one write class, on the ladder from L0 manual to L4 autonomous.
ktx is the analytics agent's context. Data Workers is the agentic data platform around it.
ktx does one job very well, in the open. Kaelio describes it as "a self-improving context layer that teaches agents how to query your warehouse accurately - from approved metric definitions, joinable columns, and business knowledge it builds and maintains for you." It is Apache 2.0, runs locally with your own model keys, reads ten databases from PostgreSQL to Snowflake, BigQuery and Athena, and ingests dbt, MetricFlow, LookML, Looker, Metabase, Sigma, Notion and Google Drive. It flags contradictions for human review, builds a join graph that resolves chasm and fan traps, and keeps every change reviewable in git. Adoption is three commands: npm install -g @kaelio/ktx, ktx setup, ktx status.
Data Workers runs the whole lifecycle around that context: lineage across engines, data quality, incidents, pipelines, schema changes, access and cost, with more than 20 specialist agents, one approval flow and one audit trail. Data Context Wizard is where every agent reads ktx's context next to lineage, quality and usage across engines, with a named owner on every fact.
Here is a Monday with both connected (an illustration, not a customer case).
| Time | System | What happens |
|---|---|---|
| 22:10 Sun | Salesforce | Sales ops splits the Closed Won opportunity stage into Closed Won - New and Closed Won - Renewal |
| 00:30 | Fivetran | The connector syncs opportunities into Snowflake with the new stage values |
| 01:10 | dbt + Snowflake | The nightly dbt job succeeds; fct_bookings filters on stage = 'Closed Won', so new-business deals drop out |
| 01:30 | Data Workers | The volume and value checks on fct_bookings fail; Data Workers traces lineage to the Salesforce stage change and opens an incident with the model owner |
| 08:05 | Claude Code + ktx | An analyst asks for last week's bookings by segment; the agent finds the measure with ktx and compiles bookings.total_amount into SQL exactly as approved |
| 08:06 | Claude Code + Data Workers | Before answering, the agent checks the table with Data Workers: an open incident, two days of new-business deals missing, cause and owner attached |
| 08:20 | Data Workers | Data Workers proposes a dbt diff that maps both new stage values, with its blast radius: three models and two Looker dashboards |
| 08:45 | Spellbook | The fct_bookings owner reviews the diff and approves; dbt CI passes |
| 08:50 | dbt + Snowflake | Data Workers queues the rebuild of the affected models, confirms the totals are back on their monitor_metrics baseline and writes the receipt |
| 09:15 | ktx | With the receipt in hand, the agent records the caveat through ktx's memory_ingest; it lands as a git commit in the ktx project, reviewed in the team's usual ktx flow |
| 09:30 | Looker | Bookings are right before the 10:00 pipeline review, and the analyst's answer matches |

ktx did its job: the measure, joins and filters were the approved ones. The data under them was short, and Data Workers had been on it since 01:30. The fix went through one owner approval, and the lesson went back into ktx through ktx's own loop.
| Job | What ktx does | What Data Workers does |
|---|---|---|
| The question | Gives the agent one searchable surface across wiki pages, semantic sources, tables and columns | Adds lineage, quality score, freshness and open incidents for every table the answer touches |
| The SQL | Compiles approved measures, joins and dimensions into SQL, with fan and chasm traps resolved | Makes sure the rows under that SQL are complete, fresh and correct |
| The context | Builds and reconciles context from the warehouse, BI tools, modeling code and docs, as reviewable files | Reads that context with provenance next to catalogs, semantic layers and lineage across engines |
| The contradiction | Flags contradictions across sources for a person to review | Routes conflicts that span systems to a named owner and records the decision |
| The break | Answers from the data as it is, read-only by design | Detects the break, traces it upstream, proposes the fix and applies it reversibly after approval |
| The proof | Keeps the history of every context change in git | Keeps a receipt for every data change: cause, diff, approver, checks and rollback path |
Why doesn't ktx just do this itself?
Kaelio built ktx for one job, giving analytics agents accurate context, and made the right design choices for it. Connections are read-only: in the README's words, "ktx never writes to your database." Its sql_execution tool runs "one parser-validated read-only SQL query." It runs locally, and every change to its own context lives in git. That focus is why teams adopt it in an afternoon and trust it with their warehouse.
Repairing production data is a different product with a different liability. It needs lineage from Salesforce through dbt to Looker, quality checks on the data itself, blast-radius scoping, approvals routed to each model's owner, rollback for every change class, and accountability for changes in systems ktx doesn't run. Read-only is the property that makes a context layer safe to install anywhere, and Kaelio is right to keep it. ktx stays the accurate context your agents query. Data Workers is the agentic data platform that acts, with approvals and receipts.
Every tool owns a slice. Data Workers covers the whole lifecycle
ktx owns one slice of the data lifecycle and owns it well: giving analytics agents the context to write the right SQL. It leads on Analytics & Insights and ties on Catalog & Context. Each point tool adds another console, contract and handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, building on the ktx project you already run.

| Stage | Data Workers | ktx | Why we scored it this way |
|---|---|---|---|
| Catalog & Context | 9 | 9 | ktx's home stage, a tie: ktx builds query context from the warehouse, dbt, MetricFlow, LookML, BI tools and docs, flags contradictions and keeps it as YAML and Markdown in git. Data Workers covers the catalog around it: lineage across engines, owners, quality and usage, with ktx as one source. |
| Analytics & Insights | 8 | 9 | ktx's second home stage: agents search one surface and compile approved measures, joins and dimensions into SQL, with fan and chasm traps resolved. Data Workers' Insights agent answers through governed definitions. |
| Data Quality | 8 | 3 | ktx reads dbt tests as facts and validates its semantic sources against the warehouse. Data Workers runs the quality checks on the data itself and repairs the breaks. |
| Observability & Incidents | 8.5 | 2 | ktx answers from the data as it is, by design. Data Workers detects the break, traces it across systems, fixes it and verifies the number. |
| Pipelines & Ingestion | 8.5 | 2 | ktx reads the pipelines' outputs and never writes to your database. Data Workers reruns and backfills pipelines behind approvals and verifies the output. |
| Schema & Migration | 8 | 3 | Each ingest rescans the warehouse and reconciles changes into reviewable diffs. Data Workers catches upstream schema changes in the dbt manifest and in review and assesses their impact. |
| Governance & Access | 8.5 | 5 | Strong for its own files: every change lands in git for review, connections are read-only, and ktx Cloud adds review and approval workflows. Data Workers proposes and applies grants across your platforms by policy. |
| Security & Privacy | 8 | 5 | Runs locally with your own model keys, a token for any non-loopback MCP endpoint and redacted telemetry. Data Workers flags sensitive column names in pull request review and proposes masking for the owner. |
| Cost / FinOps | 8 | 2 | Query history informs ktx's context; warehouse cost is outside its job. Data Workers traces Snowflake credits to the dbt model behind them and drafts the fix for its owner. |
| MLOps & Models | 7.5 | 2 | Model training and monitoring are outside a context layer's job. Data Workers keeps the data under your models healthy and connects to MLflow and W&B. |
How ktx and Data Workers work together
Analysts and agents stay in Claude Code, Codex, Cursor or OpenCode, with ktx connected. Spellbook Data Catalog (in preview) is where the data team looks: each incident, proposed change, approver and rollback. Underneath, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm does the work, and the Autonomous Data-Conductor runs each fix end to end (detect, diagnose, fix, review, verify, remember) inside per-domain guardrails.

Bring your own context: ktx as one source with provenance. ktx holds what an analytics agent needs to write a query: sources, joins, measures, caveats. Context Wizard holds the estate around it: lineage across engines, quality scores, freshness, usage, owners and the decisions behind each definition. Data Workers connects to ktx over its MCP server today, so an agent reads a ktx source and the Data Workers context for the same table in one step. A ktx wiki page that settles a definition can be recorded with ingest_unstructured_context, author attached, and an owner marks the authoritative table with mark_authoritative. Context Wizard also reads dbt and MetricFlow directly, so when a dbt metric, a ktx measure and a Looker field disagree, resolve_metric returns every candidate and the owner decides. The context stays yours. For the full pattern, see bring your own context; for background, context layer vs semantic layer and open-source context layer tools.
Data Workers never writes into your ktx project, by design. ktx keeps every context change in git, and Data Workers respects that loop: data fixes land upstream as dbt diffs or approved pipeline reruns, and anything ktx should remember goes back through ktx's own memory_ingest tool or a diff to the ktx repo, for your ktx owner to review.
Setup over MCP today. ktx writes its own client config with ktx setup --agents; for Claude Code that is an HTTP entry in .mcp.json pointing at the local ktx server (ktx mcp start, default http://127.0.0.1:7878/mcp). Add Data Workers next to it the documented way: clone the open-source repository and add start-agent.sh entries, as the client setup docs show.
// Example: .mcp.json for Claude Code with ktx and Data Workers
{
"mcpServers": {
"ktx": {
"type": "http",
"url": "http://127.0.0.1:7878/mcp"
},
"dw-context-catalog": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-context-catalog"]
},
"dw-quality": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-quality"]
},
"dw-incidents": {
"command": "/path/to/dataworkers-claw-community/start-agent.sh",
"args": ["dw-incidents"]
}
}
}If ktx binds beyond loopback, it requires a bearer token (KTX_MCP_TOKEN), which its setup writes as an Authorization header. List the tools with the client's own command, for example /mcp. On the ktx side this guide uses discover_data, sl_query, wiki_search and memory_ingest. On the Data Workers side it uses explain_table (definition, lineage and trust score), trace_cross_platform_lineage and blast_radius_analysis on dw-context-catalog, run_quality_check and get_quality_score on dw-quality, and get_incident_history, diagnose_incident and remediate on dw-incidents. One line in the agent's rules ties them together: "Before answering from a ktx measure, check the tables it reads with Data Workers."
One request end to end, L0 to L4. Autonomy is set per domain.

- •L0 manual. ktx gives the agent the right SQL; your team fixes data breaks by hand.
- •L1 observe. Data Workers read tools only. Before it answers, the agent checks every table under a ktx measure:
get_quality_scorefor quality andget_incident_historyfor open incidents, late loads flagged againstmonitor_metricsbaselines included. - •L2 propose. Data Workers drafts dbt diffs with blast radius. Nothing changes until the owner approves in Spellbook and CI passes.
- •L3 act reversibly. For change classes with a clean record, such as reruns and backfills, Data Workers applies the fix after approval, re-runs the checks on the changed tables, with the undo recorded before it runs.
- •L4 autonomous. For a scoped domain like freshness failures in the bookings marts, Data Workers fixes overnight, so the agent's first query of the morning runs on complete data.
Each step up is a per-domain decision backed by receipts. For the safety model, read is it safe to let AI agents change production data; for where data and credentials live, read where does our data go.
Teams that also keep metrics in dbt should read you're on the dbt Semantic Layer; teams weighing another context graph for data agents should read you're on Jedify.
What changes for your team
ktx gave your agents the context to write the right SQL. Data Workers gives your data team a crew, so those answers rest on data that is right.

- •Incidents. A source change that empties a model at 1 a.m. is traced, fixed and verified before an agent queries it.
- •Data quality. Every break that reached an agent's answer becomes a dbt test, so the same failure is caught upstream next time.
- •Cloud spend. Snowflake credits are traced to the query and dbt model behind them, and each fix goes to its owner drafted.
- •Access. A request to bring a new table into an agent's scope arrives as a time-boxed grant for the data owner.
- •Audits. ktx keeps the history of its context in git. Data Workers keeps the other half: who changed the data, why, and how to undo it.
- •Migrations. A warehouse move runs in parity-checked waves; ktx re-ingests the new connection.
Analytics engineers keep owning the ktx project and stop explaining why a correct query returned a wrong number.
Keep ktx, or consolidate?
Keep ktx if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.
For most teams the answer is to keep ktx. It is open source, it lives in your git repo, and your agents already query it. Data Workers runs the rest of the lifecycle around it: lineage across engines, quality, incident repair, change control and evidence. What teams consolidate is usually a separate observability tool, a second glossary or an extra catalog, once Data Workers runs those jobs and reads ktx as one of its sources. If you are weighing building this yourself on top of ktx, read build it ourselves with Claude Code and MCP servers: connecting MCP servers is the easy part; cross-system lineage, approvals and rollback are the work.
The case for your CFO
The outcome: the team adopted ktx so agents answer data questions correctly without an analyst in the loop. Data Workers makes sure the data under those answers is right, and repairs it when it isn't. Every decision made from an agent's bookings or churn number rests on the rows under the SQL as much as on the SQL.
The risk story: Data Workers sets autonomy per domain, from L0 manual to L4 autonomous. Each change goes to a named approver with its blast radius, is applied reversibly, is verified downstream, and leaves a receipt: who approved it, what it touched and how to undo it. Zero migration: the warehouse, dbt, ingestion, BI and the ktx project stay where they are.
Why now: agents answer data questions for everyone, so when the data under a correct query is wrong, the wrong number reaches every agent at once. The first win is Data Workers read tools next to ktx in one domain, so every answer carries the table's quality, freshness and open incidents, then one write class behind approvals. What stays the same: ktx, your agent clients, warehouse grants and dbt review. For the numbers, see the ROI of agentic data operations.
The sentence to repeat upstairs: "ktx makes our agents write the right SQL; Data Workers makes sure the data under that SQL is right, and fixes it with an approval and a receipt when it isn't."
Getting started
Start with a pilot. Pick the domain where agents query ktx most, such as bookings, add Data Workers to the same agent client with read tools only, and run it for a few weeks before enabling the first write class. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.
FAQ
ktx and Data Context Wizard both give agents context. Do we need both? They hold different context. ktx holds what an analytics agent needs to write a query: semantic sources, joins, measures and wiki caveats, reviewed in git. Context Wizard holds the estate around it: lineage across engines, quality, freshness, usage, owners and decisions, with ktx as one source with provenance. Keep ktx as the query layer your agents already use.
Does Data Workers write into our ktx project? No, by design. Data fixes land upstream as dbt diffs or approved reruns. Anything ktx should remember goes back through ktx's own memory_ingest tool or a diff to the ktx repo, so it stays in ktx's git history for your ktx owner to review.
ktx is read-only. Why add something that writes? Because a correct query on broken data still returns the wrong number. ktx's read-only design is right for a context layer; Data Workers is the governed path to fix the data: blast radius first, a named approver, reversible changes, downstream verification and a receipt.
What happens when ktx, dbt and Looker define a metric differently? ktx flags contradictions it finds during ingest for review. Data Workers reads dbt and MetricFlow directly, and your team's assistant brings in ktx definitions over MCP; resolve_metric returns every candidate, and the metric's owner decides which is authoritative. The decision is recorded, and the ktx side is updated through ktx's own loop.
Which warehouses does this work on? ktx connects to ten databases, including PostgreSQL, Snowflake, BigQuery, ClickHouse, SQL Server, DuckDB and Athena. Data Workers connects to Snowflake, Databricks and BigQuery natively and to the rest of your stack over MCP, so one agent client covers both.
Is ktx open source, and does that matter here? Yes, ktx is Apache 2.0. Kaelio also offers ktx Cloud (hosted, multi-user, with review and approval workflows, SSO and audit logs) and a managed Data Agent that answers in web, Slack and email. Data Workers' core is Apache 2.0 too. This guide uses the open-source ktx server and the open-source Data Workers agents side by side; the same MCP pattern applies to ktx Cloud.
Sources
- •Kaelio, homepage, https://www.kaelio.com/ (checked Oct 2, 2026)
- •Kaelio, pricing (ktx, ktx Cloud, Data Agent), https://www.kaelio.com/pricing (checked Oct 2, 2026)
- •Kaelio, ktx Cloud, https://www.kaelio.com/products/ktx-cloud (checked Oct 2, 2026)
- •Kaelio, Data Agent, https://www.kaelio.com/products/data-agents (checked Oct 2, 2026)
- •ktx docs, Introduction, https://docs.kaelio.com/ktx/docs (checked Oct 2, 2026)
- •ktx docs, Reviewing Context, https://docs.kaelio.com/ktx/docs/guides/reviewing-context (checked Oct 2, 2026)
- •ktx docs, Context Sources, https://docs.kaelio.com/ktx/docs/integrations/context-sources (checked Oct 2, 2026)
- •ktx docs, ktx mcp, https://docs.kaelio.com/ktx/docs/cli-reference/ktx-mcp (checked Oct 2, 2026)
- •ktx docs, Agent Clients, https://docs.kaelio.com/ktx/docs/integrations/agent-clients (checked Oct 2, 2026)
- •GitHub, Kaelio/ktx README, license and MCP tool registrations, https://github.com/Kaelio/ktx (checked Oct 2, 2026)
- •npm, @kaelio/ktx 0.16.0 (Jul 3, 2026), https://www.npmjs.com/package/@kaelio/ktx (checked Oct 2, 2026)
- •Data Workers, client setup, https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 2, 2026)
- •Data Workers open-source repository, tool registrations, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 2, 2026)