LangSmith

LangChain's observability and evaluation platform: zero-config LangGraph tracing, CI-gated evals, prompt hub; closed-source SaaS.

Best for: Teams building on LangChain/LangGraph that want managed deployment, Studio debugging, and CI-gated evals from one vendor

Editor’s note: We combine independent analysis, data collection, and hands-on testing to review data and AI tools. This review weighs pricing transparency, real-world adoption signals, development momentum, openness and exit costs, practitioner sentiment, and our editorial verdict.

Quick verdict: We recommend LangSmith for teams building on LangChain or LangGraph, where zero-config tracing, Studio debugging, and CI-gated evals repay the premium. The trade-off: bills scale with agent chattiness, and self-hosting is Enterprise-only via sales. This LangSmith review covers pricing, a walkthrough, and rival LLM observability and evaluation tools in our directory.

Key Takeaways 🔍

  • Three tiers (verified September 2026): Developer free (1 seat, 5,000 base traces a month), Plus $39/seat (10,000 included), Enterprise custom
  • Fast backend: the SmithDB Rust data layer (May 2026) loads trace trees at P50 92ms, per the vendor
  • Beyond LangChain: langsmith[otel] adds OpenTelemetry tracing for OpenAI, Anthropic, and Vercel AI SDK stacks; LangGraph stays the only zero-config path
  • Con: the retention trap. Evaluator feedback, run rules, and annotation queues silently upgrade traces to 400-day retention at roughly double the per-trace rate, per pricing trackers
  • Con: closed source, no self-serve self-hosting; secondary trackers put 1M traces a month on Plus at roughly $2,514 for one seat

Pros and Cons

Pros

  • Zero-config LangGraph tracing captures every node, edge, and state transition
  • Evals wire into pytest, Vitest, and GitHub workflows, so a PR fails when scores drop
  • SmithDB backend loads trace trees at a P50 of 92ms, per the vendor
  • SDKs in four languages: Python, TypeScript, Go, and Java
  • SOC 2 Type II, GDPR, and HIPAA compliance on the cloud offering

Cons

  • Closed source; self-hosting requires an Enterprise contract through sales
  • A single agent run emits many traces, so bills track chattiness, not value
  • Evaluator feedback auto-upgrades traces to 400-day retention at roughly double the cost
  • G2 reviewers flag reliability wobbles and evolving documentation
  • ISO 27001 is absent from the compliance list

How Much Does LangSmith Cost?

The free tier is the most misleading thing about LangSmith pricing, and nobody is lying to you. The tiers themselves are simple, verified on langchain.com/pricing in September 2026:

  • Developer ($0): 1 seat, 5,000 base traces a month, community support, and a 5 LCU monthly Fleet allowance
  • Plus ($39/seat/month): unlimited seats at that rate, 10,000 base traces included, email support, 1 free small serverless deployment, and 25 LCU a month
  • Enterprise (custom): self-hosted and hybrid deployment, custom SSO, ABAC and RBAC, a support SLA, and deployed engineers
TierPriceIncluded base tracesRetentionKey adds
Developer$05,000/month14 days1 seat, community support, 5 LCU/month
Plus$39/seat/month10,000/month14 daysEmail support, 1 small serverless deployment, 25 LCU/month
EnterpriseCustomCustomCustomSelf-hosted or hybrid deployment, SSO/ABAC/RBAC, SLA, deployed engineers

The real bill comes from usage metering and retention upgrades, not the seat price. Usage is metered in units: 1 LangChain Compute Unit (LCU) costs $1.50, 1 Storage Unit (LSU) costs $1.00, a Tuned Evaluator run costs 0.01 LCU, Engine runs land around 5 to 30 LCU each, and Sandboxes bill per second at variable rates. The vendor page presents trace overage as dynamic pay-as-you-go rather than a flat rate card; secondary pricing trackers estimate roughly $2.50 per 1,000 extra base traces at 14-day retention and about double that at extended retention, so verify the live rate on langchain.com/pricing before budgeting.

Base traces keep 14-day retention, but evaluator feedback, run rules, and annotation queues automatically upgrade a trace to the extended 400-day class at roughly double the cost, and you can also upgrade traces manually for a fee. Evaluation-heavy workflows therefore cost more than tracing-only ones at identical volume.

Is LangSmith Good Value for Money?

  • Model traces per run, not per request. A single agent run emits many traces, so multiply traces per run by runs per day before trusting any tier’s included volume
  • Volume gets expensive fast. At 1M base traces a month on Plus, secondary trackers put a single seat at roughly $2,514/month
  • Teams pay twice. A 5-seat team at 100,000 monthly traces pays $195 in seats plus roughly $225 in tracker-estimated overage, around $420 all in; Langfuse’s paid Cloud bills usage with unlimited seats, so adding a sixth engineer there adds no seat cost
  • Startups get a cushion. LangChain’s startup program offers up to $10,000 in credits plus discounted pricing

Author’s Testing Notes 📝

Prototype on Developer, but do the arithmetic before anyone standardizes. Count the traces one real agent run emits, multiply by daily runs, and price that number against Plus overage. Then budget evaluation as its own line item: the moment evaluators start attaching feedback, those traces move to the 400-day class at roughly double the rate, so an eval-heavy month costs more than a tracing-only month at the same volume.

— Panoply reviewer

My Experience With LangSmith

I tested LangSmith the way most teams meet it: an existing LangGraph agent, an API key, and a deadline. On LangGraph, there is almost nothing to set up.

Tracing a LangGraph Agent

On a LangGraph app, tracing needs an API key and one tracing flag in the environment. No decorators, no wrappers, no callback plumbing: runs start appearing in the project view as nested, replayable traces. Each step in the tree carries its own latency and token cost, so I could see which node in a multi-step run burned the budget without adding a single line of instrumentation. The 2026 update added AI-generated trace summaries that surface failure patterns across runs, which shortens the scan when a project holds thousands of traces.

The free Developer tier covers this stage: 5,000 base traces a month with 14-day retention is enough to connect an app and inspect real runs. It says little about production cost, though, because one agent run emits many traces, not one.

[Screenshot needed: LangSmith trace tree of a multi-step agent run, with per-step latency and token counts in the right panel] The trace tree breaks a single agent run into its individual steps, each with latency and token cost. Source: Panoply

Debugging a Failure Partway Through

The workflow LangSmith is built around: an agent fails on step four of seven, and you need to know why. I could open the exact trace, replay it, and attach evaluations to it, which is how you separate a prompt regression from a model update. The per-step state capture matters here; because LangSmith records every node, edge, and state transition, the point of divergence is visible rather than inferred.

Mind the clock on failures you want to keep. Base traces expire after 14 days, so a broken run worth studying next month has to go into a dataset or an annotation queue first, and the queue moves it into the 400-day retention class.

Wiring an Eval into CI

This is where LangSmith stops being a debugger and starts being a platform. I built a dataset from captured traces, attached evaluators, and gated a pull request on the result: pytest and Vitest integrations let the PR fail automatically when the evaluation score drops below a threshold. Prompts follow the same discipline. The Prompt Hub versions every prompt by commit hash with prod, qa, and beta environment tags, and CI pulls a pinned version with pull_prompt and publishes with push_prompt, so a prompt change ships like a code change.

Author’s Testing Notes 📝

The non-LangChain path is real but not equal. The langsmith[otel] package accepts OpenTelemetry traces from raw OpenAI, Anthropic, or Vercel AI SDK code, but expect manual instrumentation where LangGraph users get everything free. One friction users report: filtered trace views are not preserved in the URL, so sharing a filtered result with a teammate is awkward on large projects.

— Panoply reviewer

LangGraph Integration: The Moat and the Leash

Why can no rival match LangSmith’s LangGraph tracing? Because no rival owns the runtime being traced. LangChain controls LangGraph itself, so LangSmith captures every node, edge, and state transition from a single environment variable, while framework-agnostic competitors instrument from the outside. Monte Carlo’s engineering team cited this zero-setup LangGraph integration as a key reason it chose LangSmith over alternatives, per a published Langfuse-vs-LangSmith comparison.

The moat extends upward through the stack:

  • LangGraph Studio: visual debugging of agent graphs, stepping through state rather than log lines
  • Managed deployment: LangGraph agents hosted on LangChain’s infrastructure, with human-in-the-loop approval steps built in
  • Prompt Hub and evals: versioning and quality gates that assume LangSmith is the system of record

Each layer a team adopts migrates more of the actual runtime onto the vendor’s platform, and the practical exit cost climbs with every one: leaving a tracing tool is a config change, but leaving a managed agent runtime is a migration project. Practitioners on Hacker News report teams that prototyped on LangChain moving off it as they productionize, which is worth weighing before the deployment layer, not after.

Our read: LangSmith’s returns are proportional to ecosystem commitment. Model your trace volume and your exit tolerance before adopting the deeper layers, because the platform rewards all-in usage and quietly penalizes half-in usage. The workable middle path: adopt tracing and evals freely, since OpenTelemetry keeps traces portable, but treat Studio and managed deployment as separate architectural commitments with their own exit plans. For non-LangChain stacks, the langsmith[otel] package makes raw OpenAI, Anthropic, Vercel AI SDK, and LlamaIndex code workable, with manual instrumentation replacing the zero-config magic.

Evaluation, CI Gates, and the Prompt Hub

The evaluation loop is the most mature surface LangSmith has, and the strongest argument for paying at all. Datasets, evaluators, and experiment tracking live in one place, and the CI integrations turn quality from a dashboard into a gate: pytest, Vitest, and GitHub workflows can fail a pull request the moment an evaluation score dips below threshold. That brings deterministic testing discipline to systems whose outputs are anything but deterministic. Tuned Evaluator runs are metered at 0.01 LCU each, so routine scoring stays cheap at the unit level.

Every prompt in the Prompt Hub carries commit-hash versioning, prod, qa, and beta environment tags route the right version to the right environment, and a community library covers cross-team discovery. Prompts move through CI with push_prompt and pull_prompt, so a prompt rollback works exactly like a code rollback.

The billing coupling from the pricing section hits hardest here: evaluator feedback pushes traces into 400-day extended retention at roughly double the base cost, so the better your evaluation hygiene, the higher your per-trace spend. Budget for it deliberately rather than discovering it on an invoice.

The LangSmith Gateway extends the loop to spend: it adds cost controls, rate limiting, PII redaction, and coding-agent integration. LangChain CEO Harrison Chase has framed cost management as the emerging customer priority in his public posts, citing Uber reportedly capping tokens at $1,500 per developer per month.

For teams that treat prompts like code, this is the most complete quality loop in the category. Eval-first shops comparing on price should look at Braintrust, which bills on evaluation scores rather than traces and avoids the retention coupling entirely. Braintrust frames the split itself: it bills scores because evaluation is its product, while LangSmith bills traces because monitoring is.

Self-Hosting and Compliance: Enterprise Only

If your security team requires on-prem deployment without a sales cycle, LangSmith is disqualified before the demo. Self-hosted LangSmith exists only as an Enterprise add-on: the license key comes from contacting sales@langchain.dev, and no self-serve path exists at any price.

Self-hosted LangSmith deploys on Kubernetes (recommended for production) or Docker Compose, has been tested on GKE, EKS, AKS, OpenShift, and Minikube, and bundles PostgreSQL, ClickHouse, and an internal Redis by default, with external managed Redis or Valkey and external storage services advised for production deployments. This is a real on-prem product, not a checkbox. Enterprise also lists a hybrid deployment option alongside full self-hosting, and the contract bundles a support SLA, Slack support, team trainings, architectural guidance, and deployed engineers.

Compliance on the cloud offering:

  • SOC 2 Type II: listed
  • GDPR: listed
  • HIPAA: listed
  • ISO 27001: not listed, a checklist gap against Langfuse, whose cloud carries all four

For teams where data sovereignty or air-gapped hosting is mandatory and procurement overhead is not an option, the sales gate is the disqualifier, and Langfuse’s free MIT self-hosting, with no user or unit caps, is the obvious detour. For enterprises already buying SLAs, deployed engineers, and architectural support, the gate matters less than the stack quality behind it.

Adoption, Performance, and What Reviewers Say

LangSmith’s adoption numbers point in opposite directions, and both are lying a little. The langsmith package on PyPI records roughly 78 million downloads a month (about 17.3 million a week, 1.5 million a day), but the client ships as a transitive dependency of LangChain, so that figure overstates active platform users. The SDK repository holds roughly 618 to 670 GitHub stars against Langfuse’s 17,900-plus, which understates for the opposite reason: the client is thin and the product lives server-side, closed source, where stars cannot accumulate. Neither number stands alone, which is why our adoption read leans on multiple signals rather than any single one.

SmithDB, the Rust-based data layer launched in May 2026, now handles 100 percent of US Cloud ingestion, and the vendor reports trace trees loading at a P50 of 92 milliseconds. That figure is LangChain’s own, not an independent benchmark, so treat it as a vendor claim until your own projects confirm it.

G2 reviewers rate LangSmith 4.8 out of 5, and aggregated summaries spanning 267 reviews land close by at 4.7. The recurring complaints across both are reliability wobbles and documentation that evolves faster than it completes. Strong capability, imperfect polish is the pattern.

How Does LangSmith Compare to Competitors?

Which LangSmith alternatives actually threaten it? Each rival in this directory batch attacks a different flank:

  • Langfuse is the open-source counterweight: MIT-licensed self-hosting at zero cost with no user or unit caps, unlimited users on paid Cloud plans, and ISO 27001 on the compliance list. Its LLM-as-judge evaluators have been fully MIT open source since June 2025, native integrations cover CrewAI, LlamaIndex, DSPy, and Haystack, and GitHub Actions can gate deploys on eval results. The trade: less native depth for LangGraph-specific agent state inspection, and ClickHouse’s $400M acquisition in January 2026 raises roadmap-independence questions, though the MIT license is unchanged
  • Braintrust is the eval-first alternative. It bills on evaluation scores (10,000 free, then $2.50 per 1,000) rather than traces, claims 1 to 3ms of execution overhead per trace, and enforces CI gates natively; its runtime and agent-deployment features are the weaker flank
  • Arize Phoenix offers open-source, OpenTelemetry-native tracing with vendor-neutral instrumentation, the low-commitment way to get traces without picking a platform; the managed story is thinner, with the enterprise layer living in Arize AX
  • Helicone covers lightweight gateway-style logging rather than a full tracing-plus-eval platform, and its hosted product stopped accepting new signups after the March 2026 Mintlify acquisition
  • W&B Weave suits research-heavy teams already on Weights & Biases, with strong experiment tracking and prompt versioning but SDKs in Python and TypeScript only, against LangSmith’s four languages

On LangGraph, LangSmith earns its premium: no rival matches the zero-config state capture, Studio debugging, or managed deployment. Off LangGraph, the premium mostly buys a nicer eval UI, and Langfuse or Phoenix deserve the first look before any card goes down. Eval-first teams sit in the middle: Braintrust’s score-based billing sidesteps the retention trap, at the cost of the weaker deployment story.

How We Test LLM Observability Tools

We combine independent analysis, data collection, and hands-on testing to review data and AI tools. For LLM observability platforms, that means setting the tool up ourselves, tracing and evaluating a real agent workflow end to end, and weighing it against direct rivals on the same criteria.

We collect public signals from GitHub, PyPI, Docker Hub, Stack Overflow, G2, and Gartner peer reviews, and we hand-check vendor pricing pages and date every price we print. Third-party review scores are deliberately small-weighted, sustained practitioner sentiment counts more than any single rating, and areas we cannot measure are marked not applicable rather than scored zero. Signals refresh monthly, editorial verdicts quarterly, and sponsors and affiliates cannot change a score. Prices current as of September 2026.

LangSmith Review: Should You Run Your Agents Through LangSmith?

Is LangSmith worth it? We recommend it for LangChain and LangGraph-native teams, for organizations that want managed agent deployment, Studio debugging, and CI-gated evals from one vendor, and for enterprises comfortable buying self-hosting through a sales-led contract. On that stack, the zero-config tracing and the eval loop covered above are the best in the category, and SmithDB makes the daily experience fast.

Skip it if you need self-serve on-prem deployment (Langfuse self-hosts free under MIT), if you are eval-first and optimizing cost per score (Braintrust bills scores, not traces), or if your stack is heterogeneous enough that the zero-config advantage never materializes and you would be paying LangGraph prices for OTel-level integration.

The next action costs nothing: start on the free Developer tier, run one real agent workload, and build the cost model from what you observe. Multiply traces per run by runs per day, check which workflows trigger 400-day retention, and verify the current overage rates on langchain.com/pricing before committing a card. Early-stage startups on that stack should also apply for LangChain’s startup program, which offers up to $10,000 in credits, before the first overage invoice lands.

FAQ

Does LangSmith require LangChain?

No. The langsmith[otel] package adds OpenTelemetry support, so raw OpenAI, Anthropic, Vercel AI SDK, and LlamaIndex code can send traces to LangSmith without adopting LangChain. LangGraph remains the only zero-config integration, and the deepest features (Studio, state inspection, managed deployment) stay LangGraph-specific.

Can I self-host LangSmith for free?

No. LangSmith is closed source, and self-hosting is an Enterprise-only add-on requiring a license key from LangChain sales. There is no self-serve or free self-hosting path at any tier. Teams that need free on-prem deployment should look at Langfuse, which self-hosts under an MIT license.

How much does LangSmith cost at production volume?

Developer is free for 5,000 base traces a month; Plus is $39 per seat with 10,000 included, then pay-as-you-go overage. Secondary trackers estimate roughly $2.50 per 1,000 extra base traces and about $2,514 a month for a single seat at 1M traces. Verify current rates on langchain.com/pricing, since the vendor presents overage as dynamic rather than a fixed rate card.

What triggers LangSmith’s expensive 400-day retention?

Three things upgrade a base trace to the extended 400-day class automatically: evaluator feedback attached to a trace, run rules firing, and traces entering annotation queues. Extended retention costs roughly double the base rate, so evaluation-heavy workflows cost more than tracing-only ones at identical trace volume. Manual upgrades are also available for a fee.

Can I migrate away from LangSmith later?

Partially. OpenTelemetry support makes trace pipelines portable to other OTel-compatible platforms. The harder exports are evaluation datasets, annotations, and prompt history, which need extracting before retention windows close. The real exit cost is the runtime: teams using Studio, managed deployment, and the Prompt Hub have moved part of their agent infrastructure onto LangSmith itself, and that is a migration project, not a config change.

Spotted a wrong price or a missing integration? Send a correction. A human reads every one.

Similar tools

Other tools in the same category, with the same card and the same honest pricing.

Weights & Biases

Experiment tracking

The polish leader in experiment tracking: dashboards, Sweeps, reports, and Weave for LLM observability; Pro is $60/month flat for up to 10 seats.

Visit site

MLflow

Experiment tracking

Open-source standard for experiment tracking, model registry, and serving; 3.x adds OpenTelemetry tracing, LLM judges, and a Prompt Registry.

Visit site

Langfuse

Observability

Open-source LLM engineering platform: OpenTelemetry-native tracing, LLM-as-judge evals, prompt management; MIT core with free self-hosting.

Visit site

LangChain

RAG framework

The most-adopted LLM framework, rebuilt on LangGraph in 1.0: durable agent execution, 1,000+ integrations, MIT-licensed; LangSmith is the paid layer.

Visit site

LlamaIndex

RAG framework

RAG-first data framework for LLM apps: 160+ connectors, purpose-built indexes, MIT-licensed; LlamaCloud adds managed parsing and indexing.

Visit site

Weaviate

Dedicated

Open-source vector database with native BM25 + vector hybrid search, per-tenant shard isolation, and multimodal vectorizer modules.

Visit site

Pinecone

Dedicated

Fully managed serverless vector database with zero-ops scaling, deep RAG framework integrations, and enterprise compliance.

Visit site