Product
Product11 min readBy The Data Workers Team

You're on LangSmith: LangSmith Traces and Evaluates Your Agents. Data Workers Owns Whether the Data They Answer From Is Right

On LangSmith for tracing, evals, Engine and Deployment? Keep it. Data Workers checks, traces and fixes the warehouse tables your agents answer from, behind approvals.

Your agents report to LangSmith. Every run lands in a tracing project: model calls, tool calls, retrieved documents, the SQL a tool sent and what came back. Online evaluators score production traces, users leave thumbs up or down, and alerts on run count, errors, feedback score, latency and cost route to Slack, PagerDuty or a webhook. Engine, LangSmith's agent for agent engineering, files recurring issues from your traces, diagnoses them against your code and proposes the fix as a pull request; Engine v2, announced September 25, 2026, adds red teaming, and its fix validation is in private beta. Your agents run on LangSmith Deployment, the runtime once called LangGraph Platform, and the business builds no-code agents in Fleet, formerly Agent Builder.

LangSmith traces and evaluates your agents, and it does that very well. What it cannot see is whether the table a tool queried is right: whether the dbt model behind it loaded last night, or whether rows stopped arriving while every job stayed green. Data Workers owns that question. It checks and traces the tables your agents answer from, proposes the fix to the data owner, runs it behind approvals and leaves a receipt.

Key takeaways

  • •LangSmith keeps its job. Tracing, evals, Engine, Deployment, Fleet and the LangSmith Remote MCP stay where they are. Data Workers works with LangSmith from day one.
  • •A faithful agent on a broken table still gives a wrong answer. Groundedness evals pass when the agent repeats its tool. Data Workers watches the tables, with quality checks and the baselines your jobs post.
  • •The agent is a reader in the graph. Record each production agent as a reader of the tables it queries, and lineage and blast radius reach it from any upstream dbt model.
  • •Engine fixes the prompt; Data Workers fixes the data. Data Workers proposes the dbt change as a diff, queues the rebuild through your orchestrator after a named person approves, and re-checks.
  • •Every fix leaves a receipt, ready to link from the Engine issue.
  • •Start with a pilot on two production agents, read-only first.

LangSmith traces and evaluates your agents. Data Workers owns whether the data they answer from is right.

LangSmith answers what the agent did. Data Workers answers why the data was wrong: which table, which upstream change, what else it reaches, who owns the fix and whether it held. Here is how the two meet when an eval regression is really a frozen table.

A B2B software company runs an in-app account assistant, a LangGraph agent on LangSmith Deployment, that tells customer admins how many API calls they have used against their plan. Its get_account_usage tool queries marts.account_usage_daily in Snowflake. Events arrive through Segment into segment_prod.tracks; an Airflow 2.10 DAG, usage_hourly, runs dbt build every hour. fct_usage_events is an incremental model that loads rows newer than its own max(event_ts). The DAG's last task records the run's rows_inserted with monitor_metrics: the team's job posts the number, and Data Workers keeps the baseline and flags breaks. The team recorded the assistant and the billing team's overage export as readers of the mart in Data Workers' context graph. In LangSmith, an online LLM-as-judge evaluator scores groundedness and an alert watches the feedback score. This is an illustration, not a customer case.

TimeSystemWhat happens
Sun 21:40SegmentA platform engineer's load-test script, pointed at the production write key by mistake, sends 1,200 api_call events with timestamp hard-coded to 2027-03-01 to replay a fixture. Segment accepts them as sent
22:00Airflow, dbt, Snowflakeusage_hourly runs. fct_usage_events appends the hour's 138,000 real events and the 1,200 test events. max(event_ts) is now March 1, 2027
23:00Airflow and dbtThe next run succeeds and inserts zero rows: no real event is newer than 2027. Every dbt test passes; event_id is unique and nothing is null
23:05Data WorkersThe DAG posts rows_inserted = 0; monitor_metrics flags it against a baseline near 120,000 rows per run. run_quality_check on fct_usage_events passes: nulls, the event_id distinct ratio and the row floor are fine, and a max-timestamp lag would read as fresh. Data Workers posts the finding to #data-oncall; blast_radius_analysis reaches the mart, the assistant and the overage export
Mon 07:15LangSmith DeploymentAdmins in Europe start their week. The assistant tells one "You have used 0 API calls today" and another that they are at 41% of their plan, from usage frozen at Sunday 21:59. The groundedness evaluator passes: the answers match the tool output
07:50LangSmithThe Feedback Score alert fires to #ai-agents: the average user score has dropped below its threshold as thumbs-down on usage questions pile up
08:20LangSmith EngineEngine files an issue: get_account_usage returns no rows for the current day. It proposes a pull request that has the agent say usage can lag up to 24 hours and fall back to the previous day
08:35Claude CodeThe AI engineer asks, with the LangSmith Remote MCP and Data Workers' MCP servers in the same client: "Why does the usage tool return nothing after Sunday 22:00?"
08:45Data Workerstrace_cross_platform_lineage over the dbt manifest shows the mart built from fct_usage_events and segment_prod.tracks. The model's SQL in the manifest filters on its own max(event_ts); green runs with zero inserts are the signature of a high-water mark ahead of the clock. It links its 23:05 card and asks the data owner to check max(event_ts)
09:00SnowflakeThe analytics engineer finds the 1,200 rows dated 2027, all from the load-test library
09:20Spellbook and GitHubData Workers proposes one change set: a dbt diff that loads on Segment's server-side received_at with a three-hour lookback, merges on event_id, and routes events stamped more than a day after receipt to fct_usage_events_quarantine, with a test; a delete of the 1,200 rows for the owner to approve and apply, with Snowflake Time Travel as the undo; a rerun of usage_hourly after both; and a note for the platform owner to move the script to the staging key. blast_radius_analysis shows the change reaches the mart, the assistant and the overage export
10:05GitHub, Snowflake, SpellbookThe engineer merges the diff and applies the delete. The data owner approves the rerun
10:07AirflowData Workers queues a usage_hourly DAG run and reads its status. It succeeds at 10:41 and loads Sunday 22:00 to Monday 10:00
11:05Data WorkersThe 11:00 run posts 151,000 rows, inside its baseline; run_quality_check passes. Data Workers writes the receipt: baseline break, lineage, diff, cleanup, approver, rerun, before and after, undo path. The approved fact "the usage high-water mark keys on received_at" lands in the graph
11:20LangSmithThe AI engineer replays the examples Engine generated for the issue against the deployment: current usage comes back. She closes Engine's pull request unmerged and closes the issue with a link to the receipt; Engine reopens it if it resurfaces
Incident timeline across the stack: what LangSmith, your team and Data Workers each do, step by step

Every tool did its job: Segment delivered what it was sent, dbt ran the SQL as written, and LangSmith caught the users' frustration and named the failing tool call. The problem lived between them: one row from the future stopped a table with every job green. A trace shows that a tool returned nothing; it cannot show that a high-water mark in another team's dbt model moved to 2027. Data Workers caught the table hours earlier, traced the cause, routed the fix to its owners and left the agent decisions with the AI engineer.

JobWhat LangSmith doesWhat Data Workers does
The signalOnline evals, user feedback, alerts and Insights on what the agent doesQuality checks and posted baselines on the tables the agent queries
The agentTraces every run, tool call and retrieved document; runs it on LangSmith DeploymentHolds the agent as a reader in the graph, joined to its tables and their owners
The causeEngine clusters failing traces and diagnoses them against your codeTraces the change through the dbt manifest to the upstream model and its owner
The fixEngine proposes the prompt or code change as a pull requestProposes the dbt diff and cleanup; queues the rebuild through Airflow after approval; re-checks
The proofIssues, experiments and the dataset examples Engine generatesA receipt per fix: evidence, approver, checks, before and after, undo path

Why doesn't LangSmith just do this itself?

Because LangSmith is built to improve agents, and its design follows from that. Engine works where LangSmith's data lives: production traces, feedback and the code repository you connect, and its fixes arrive as pull requests to the agent's prompts and code. That scope is why AI engineers trust it: Engine can propose a fix without needing the keys to the rest of the company.

Fixing the table is a different product with a different liability. It needs warehouse and orchestrator credentials, lineage across another team's dbt models, a named data owner's approval, a cleanup with an undo, and proof the fix held: access no AI team wants its agent-improvement tool to hold. Data Workers is the product on that side of the line: it changes data, carefully, and its view starts at the table.

Every tool owns a slice. Data Workers covers the whole lifecycle

Each point tool adds another console, another contract and another handoff. Data Workers covers the whole lifecycle with one context, one approval flow and one audit trail, and it builds on the tools already there.

Spider chart of ten jobs a data team does: Data Workers covers the whole list, LangSmith goes deep on its own area

LangSmith's home stage is MLOps & Models, here meaning agent and LLM operations, where it leads by design.

StageData WorkersLangSmithWhy we scored it this way
Catalog & Context94Projects, datasets, prompts and Engine's agent overview document organize what AI teams know about an agent. Data Workers joins tables, lineage, quality, owners and the agents that read them in one governed graph.
Analytics & Insights84Dashboards, Insights clustering and Trajectories show how people use an agent. Data Workers answers questions about the data itself from governed context.
Data Quality83Online evaluators score the agent's answers, not the tables behind them. Data Workers checks the warehouse tables and posted baselines and fixes what fails where it starts.
Observability & Incidents8.56Alerts on run count, errors, feedback, latency and cost route to Slack, PagerDuty or a webhook, and Engine files issues from traces. Data Workers detects, diagnoses, fixes and verifies data incidents, with a receipt for each.
Pipelines & Ingestion8.52LangSmith Deployment runs agents durably; it does not build the tables they query. Data Workers queues reruns through your orchestrator after approval.
Schema & Migration82A tool's schema is LangSmith's view of the data. Data Workers shows a dbt change's reach and writes migrations with rollback for the owner.
Governance & Access8.54Workspaces, roles, SSO and SCIM govern who sees which project. Data Workers routes each data change to a named approver and keeps the record.
Security & Privacy84Self-hosted, BYOC and SmithDB in your VPC keep traces in your environment. Data Workers proposes classifications and masking for data owners to apply.
Cost / FinOps84Cost tracking prices every model call, and alerts catch spikes. Data Workers attributes Snowflake spend to dbt models and reads AWS Cost Explorer.
MLOps & Models7.59The home stage: tracing, online and offline evals, datasets, experiments, Engine, Fleet and LangSmith Deployment. Data Workers keeps the data under those agents right.

Scores are directional measures of scope, not benchmarks.

How LangSmith and Data Workers work together

Your people stay where they are, with Claude Code, Cursor or Codex for questions and code. Spellbook Data Catalog (in preview) is where the data team looks: each finding, proposed change, blast radius, approver and rollback, with approval requests reaching people in Slack or email. Underneath, Data Context Wizard keeps one governed context graph, the Data-Agents Swarm does the work with 20+ specialist agents, and the Autonomous Data-Conductor runs each fix (detect, diagnose, fix, review, verify, remember) behind per-domain guardrails.

How Data Workers fits with LangSmith: your coding agent on top, Data Workers in the middle, your estate underneath

Where LangSmith fits. LangSmith stays the system of record for agent behavior: traces, evaluations, datasets and Engine issues. Data Workers connects to LangSmith over its API today, and your client runs the LangSmith Remote MCP side by side with Data Workers' MCP servers, so one assistant sees the failing traces and the tables behind them. Data Workers reads natively what sits under the agent: Snowflake, BigQuery and Postgres checks, the dbt manifest, GitHub pull requests and Airflow, Dagster, Prefect or ADF runs, among 50+ connectors. Your own agents can call explain_table before writing SQL to get a table's definition, lineage, documentation and trust score. See the ML engineers guide and Inside the MLOps & Models Agent.

What Data Workers writes, and where. To LangSmith, nothing: prompts, evaluators, datasets and deployments belong to the AI engineer. dbt fixes go to the model's owner as a diff to merge, and Data Workers opens the pull request when your team turns on the GitHub pull-request target. Cleanups such as removing bad rows are proposed for the owner to approve and apply. Rebuilds are queued through your orchestrator after approval. Approved facts land in the Context Wizard graph.

Setup today. Clone the open-source repository and add start-agent.sh entries to your client config, as the client setup docs show. Add the LangSmith Remote MCP next to them; interactive clients sign in to LangSmith over OAuth. Data Workers' agents never call it.

// Example: .mcp.json for Claude Code
{
  "mcpServers": {
    "dw-context-catalog": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-context-catalog"]
    },
    "dw-quality": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-quality"]
    },
    "dw-incidents": {
      "command": "/path/to/dataworkers-claw-community/start-agent.sh",
      "args": ["dw-incidents"]
    },
    "langsmith": {
      "type": "http",
      "url": "https://api.smith.langchain.com/mcp"
    }
  }
}

List the tools with /mcp, then ask: "Engine says get_account_usage returns nothing today. Which dbt models build marts.account_usage_daily, and what else reads them?"

One request, L0 to L4. The autonomy ladder is set per domain.

The autonomy ladder: L0 manual, L1 observe, L2 propose, L3 act reversibly, L4 autonomous
  • •L0 manual. Connected, not acting. The team finds the frozen table after the Engine issue.
  • •L1 observe. Data Workers flags each failed check or baseline break on a table an agent reads, with blast radius and owners.
  • •L2 propose. Data Workers drafts the dbt diff, the cleanup and the rerun plan; owners approve before anything runs.
  • •L3 act reversibly. For proven classes, such as queuing the rerun once a fix merges, Data Workers acts, verifies and records it.
  • •L4 autonomous. For a scoped, trusted class in one domain, Data Workers runs the loop end to end and posts the receipt. Prompts, evaluators and deployments stay with the AI engineer at every level.

What changes for your team

Six jobs that run on autopilot with Data Workers next to LangSmith, with a concrete example of each

Teams on LangSmith know the moment an agent starts failing. What costs them time is the hour after: the AI engineer reading another team's dbt repo, or a prompt patch that hides a data break. With Data Workers, these jobs run on autopilot at the level you set.

  • •Incidents. An eval drop that is really a data break arrives with its cause, owner and proposed fix, often before users notice.
  • •Data quality. Agent tables get nulls, keys and row floors checked, plus the baselines your jobs post.
  • •Cloud spend. Snowflake spend is tied to the dbt models your agents read, through query tags.
  • •Access. A new agent's read role arrives as a dry run, with the sensitive columns it would reach.
  • •Audits. Each data fix links the alert, the approver, the diff and the undo path.
  • •Migrations. Each wave is planned with its parity checks; the owner signs off before it closes.

Running other agent and model tools too? See You're on Arize, You're on Comet and You're on MLflow.

Keep LangSmith, or consolidate?

Keep LangSmith if you love it; Data Workers works with it from day one. Many teams consolidate once Data Workers runs that slice too.

Nearly every team keeps LangSmith: it is where agent decisions get made. What teams consolidate is the tooling around it: freshness asserts bolted into tool code, prompt lines that apologize for late data, and the separate data observability seat the AI team never opens. For background, see data pipelines for LLMs, data freshness monitoring and how Data Workers differs from data observability. You're on Segment, You're on dbt and You're on Airflow cover the rest of this stack, and the integrations page lists every connector. Building this layer yourself? Read build it ourselves with Claude Code and MCP servers: reading a table is the easy part; approvals, rollback and receipts are the work.

The case for your CFO

The outcome: when a customer-facing agent gives wrong answers because of its data, the cause is found and fixed through its owner the same morning, and nobody ships a prompt patch that hides it.

The risk story is plain. Data Workers changes nothing in LangSmith. dbt fixes arrive as diffs your engineers merge; cleanups are applied by the data owner; reruns start only after a named approver says yes. Each action leaves a receipt with the trigger, evidence, diff, approver, checks and undo path. Unanswered requests expire and escalate, never auto-grant, and no agent can promote its own work. Autonomy is set per domain from L0 manual to L4 autonomous, and an org-wide stop halts all autonomous dispatch. The agents run in your infrastructure; your data stays in your systems, and the hosted Conductor sees workflow metadata only. Zero migration.

Why now: agents answer customers, sales and finance directly, and Engine and coding agents ship prompt changes within hours. A wrong usage number in front of a customer becomes a billing dispute. The first win is two production agents with checks and baselines on the tables they query, read-only, and a report of every upstream change that would have reached them. What stays the same: LangSmith, your evaluators, your agent code and your owners. Start with a pilot; the path is on the pricing page, and the pilot is credited in full against the first year. For the numbers, see the ROI of agentic data operations.

The sentence to repeat upstairs: "LangSmith tells us when an agent goes wrong; Data Workers makes sure the data it answers from is right, with an approval and a receipt for every change."

Getting started

Start with a pilot. Pick the two agents your customers or executives use most, record the tables their tools query, connect Data Workers read-only to your warehouse, dbt project, pull requests and orchestrator, have your load jobs post rows inserted, and let it report what it would have caught before LangSmith's alerts fired, before you turn on the first fix class. The pilot path and plans are on the pricing page, and the pilot is credited in full against the first year.

FAQ

Does Data Workers read our LangSmith traces, datasets or evals? LangSmith stays the system of record for those, and your client runs the LangSmith Remote MCP side by side with Data Workers' MCP servers, so one assistant sees both. Data Workers works on the tables under the agent.

Our evals passed while users complained. How would Data Workers have helped? Groundedness evaluators check the answer against the tool output, so a faithful agent on a frozen table passes. Data Workers checks the table: nulls, key uniqueness, row floors and the baselines your jobs post, such as rows inserted per run. In the example, that baseline flagged the table eight hours before the first complaint.

Will Data Workers catch every bad row? It catches what its checks and baselines cover. The null check passes up to 10% nulls by default and uniqueness passes at a 99% distinct ratio, so teams set stricter thresholds for tables an agent leans on and post the metrics that matter. The owner confirms the cause in the rows, as the analytics engineer did in the example.

Should we merge Engine's prompt fix or Data Workers' data fix? Often one of each, decided by different owners. Data Workers' receipt shows the table, the diff and the re-check, and the AI engineer judges whether Engine's prompt change still helps. Engine reopens the issue if it resurfaces.

What about RAG agents and agents built in Fleet? The same applies. Record the agent as a reader of the table behind its retrieval index or tool, and that table's checks, baselines, lineage and owners cover it; re-indexing stays with your team.

Where does our data go? The agents run in your infrastructure and hold the credentials. Your data stays in your systems; the hosted Conductor sees workflow metadata only, never rows, credentials or model keys.

Sources

  • •LangChain, LangSmith product page: Engine, Observability, Evaluation, Deployment, Sandboxes, LLM Gateway, Fleet, SmithDB, https://www.langchain.com/langsmith (checked Oct 3, 2026)
  • •LangChain blog, LangSmith Engine v2, agents, fine-tuning and Trajectories (Sep 25, 2026), https://www.langchain.com/blog/langsmith-engine-agents-fine-tuning-trajectories (checked Oct 3, 2026)
  • •LangSmith docs, Engine: issue lifecycle, pull requests, feedback signals, fix validation (private beta), https://docs.langchain.com/langsmith/engine (checked Oct 3, 2026)
  • •LangSmith Cloud changelog: LangGraph Platform is now LangSmith Deployment (October 13-17, 2025); Engine reads tool result contents (September 14-21, 2026), https://docs.langchain.com/langsmith/changelog (checked Oct 3, 2026)
  • •LangSmith docs, LangSmith Deployment, https://docs.langchain.com/langsmith/deployment (checked Oct 3, 2026)
  • •LangSmith docs, LangSmith Fleet ("Agent Builder is now LangSmith Fleet"), https://docs.langchain.com/langsmith/fleet/index (checked Oct 3, 2026)
  • •LangSmith docs, Alerts: run count, cost, errors, feedback score, latency; Slack, PagerDuty, webhook, https://docs.langchain.com/langsmith/alerts (checked Oct 3, 2026)
  • •LangSmith docs, online evaluators (LLM-as-judge, code, multi-turn, composite), https://docs.langchain.com/langsmith/online-evaluations-code (checked Oct 3, 2026)
  • •LangSmith docs, LangSmith Remote MCP (OAuth or API key; same tools as the standalone server), https://docs.langchain.com/langsmith/langsmith-remote-mcp (checked Oct 3, 2026)
  • •LangSmith docs, LangSmith MCP Server (deprecated in favour of the Remote MCP), https://docs.langchain.com/langsmith/langsmith-mcp-server (checked Oct 3, 2026)
  • •LangSmith docs, user management (SAML SSO, SCIM), https://docs.langchain.com/langsmith/user-management (checked Oct 3, 2026)
  • •Data Workers, Client setup (open-source docs), https://dataworkers.io/opensource-docs/client-setup/ (checked Oct 3, 2026)
  • •Data Workers open-source repository, agent tools and start-agent.sh, https://github.com/DataWorkersProject/dataworkers-claw-community (checked Oct 3, 2026)
  • •Data Workers pricing, pilot and hosted Conductor, https://dataworkers.io/pricing/ (checked Oct 3, 2026)