Langfuse

Open-source LLM engineering platform: OpenTelemetry-native tracing, LLM-as-judge evals, prompt management; MIT core with free self-hosting.

Best for: Teams with heterogeneous or changing stacks that want framework-agnostic tracing, evals, and a real self-host path at low cost

Pros

  • Near-daily releases and 34,700 to 35,100 GitHub stars

Cons

  • Owned by ClickHouse, Inc. since January 2026

Editor’s note: We combine independent analysis, data collection, and hands-on testing to review data and AI tools. This Langfuse review weighs pricing transparency, real-world adoption signals, development momentum, openness and exit costs, practitioner sentiment, and our own editorial verdict.

Quick verdict: We recommend Langfuse for teams that want framework-agnostic tracing and evals with a real self-host path, at cloud prices severalfold below LangSmith’s. The trade-off: v3 self-hosting means running Postgres, ClickHouse, Redis, and S3, and the UI slows down on heavy traces.

Langfuse earned the default slot among LLM observability tools on MIT licensing, OTel-native tracing, and price. The 2026 fine print: a heavier self-host stack and a new owner.

Key Takeaways 🔍

  • MIT-licensed core with free, unrestricted self-hosting; only nine named enterprise features need a license key
  • Cloud tiers run Hobby $0, Core $29, Pro $199, and Enterprise $2,499 per month, verified September 2026
  • OpenTelemetry-native tracing since v3: OpenAI SDK, LangChain, LlamaIndex, Pydantic AI, Vercel AI SDK, CrewAI, or plain OTLP
  • The v3 self-host stack is four services where v2 ran on one Postgres container
  • Part of ClickHouse, Inc. since January 2026: no license change yet, but the exit-risk profile changed

Pros and Cons

Pros

  • MIT-licensed core with free, unrestricted self-hosting
  • Roughly $69/mo vs LangSmith’s ~$420/mo for a 5-seat team at 100K traces, per an independent estimate
  • OpenTelemetry-native tracing across LangChain, LlamaIndex, OpenAI SDK, Pydantic AI, Vercel AI SDK, CrewAI, and raw OTLP
  • Built-in LLM-as-judge templates plus CI/CD gating via the langfuse/experiment-action GitHub step
  • Near-daily releases and 34,700 to 35,100 GitHub stars

Cons

  • v3 self-hosting needs Postgres, ClickHouse, Redis, and S3 where v2 ran one Postgres container
  • Documented multi-second page loads at volume (GitHub issues #10391 and #9933)
  • Online evals sample rather than score everything; alerting stops at prompt-version notifications
  • UI-configured evaluators cannot be triggered programmatically for scheduled pipelines
  • Owned by ClickHouse, Inc. since January 2026

How Much Does Langfuse Cost?

Start with the fact most buyers get wrong: self-hosting Langfuse is free and unrestricted under the MIT license, regardless of any cloud tier. The four tiers below price the managed Langfuse Cloud only (prices as of September 2026):

  • Hobby ($0/mo): 50k units per month, 30-day retention, 2 users, community support via GitHub and Discord
  • Core ($29/mo): 100k units included, $8 per additional 100k, 90-day retention, unlimited users, in-app support
  • Pro ($199/mo): 100k units included, 3-year retention, unlimited annotation queues, high rate limits, a 48-hour support response SLO, SOC2 and ISO 27001 reports, HIPAA available; an optional Teams add-on at $300/mo adds enterprise SSO and fine-grained RBAC
  • Enterprise ($2,499/mo): everything in Pro plus Teams, audit logs, SCIM, custom rate limits, an uptime SLA, a named lead support engineer, and custom overage pricing on a yearly commitment
PlanPriceIncluded unitsRetentionKey additions
Hobby$0/mo50k/mo30 days2 users, community support
Core$29/mo100k/mo, $8 per extra 100k90 daysUnlimited users, in-app support
Pro$199/mo100k/mo3 yearsSOC2/ISO reports, HIPAA, Teams add-on ($300/mo)
Enterprise$2,499/moCustom overage pricing3 years (as Pro)Teams, audit logs, SCIM, uptime SLA, named support engineer

Why GitHub Says “Other” Instead of MIT

GitHub labels the repository’s license “Other,” which reads like a warning sign until you see what sits behind it. The explanation is an open-core split: exactly nine features are commercial, activated by a LANGFUSE_EE_LICENSE_KEY environment variable on self-hosted deployments.

The nine are project-level RBAC roles, protected prompt labels, data retention policies, audit logs, server-side data masking, UI customization, organization creators, the org management API with SCIM, and the instance management API. Remove the key and only those nine switch off; every core feature and API keeps working under plain MIT, without usage limits.

Is Langfuse Good Value for Money?

  • The LangSmith gap is large: one independent cost estimate puts a 5-seat team at 100K traces/month at roughly $69 on Langfuse Cloud versus about $420 on LangSmith, and at 1M traces Langfuse comes out near a third of LangSmith’s cost
  • Self-hosting removes the software bill entirely, leaving infrastructure as the only real cost; the ClickHouse container alone needs 2 CPU and 8 GiB at minimum
  • The free tier’s 30-day retention is the documented sore point: users often cite it as too short for trend analysis without upgrading
  • Unlimited users arrive at $29, so headcount never drives the bill, unlike W&B Weave, which prices on seats as well as usage

Author’s Testing Notes 📝

A solo developer or prototype stays on Hobby without hitting the 50k-unit cap. A production team should start at Core ($29) and move to Pro only when 90-day retention, compliance reports, or the SSO add-on forces the issue. Self-host only if the v3 stack skills in the next sections already exist in-house; otherwise Core is cheaper than the engineering time.

— Panoply reviewer

My Experience With Langfuse

The fastest way to judge an observability platform is to instrument an app that its marketing was not built around. So I skipped LangChain entirely and traced a plain OpenAI SDK application.

Instrumenting Without LangChain

Setup is two environment variables and one import. I set LANGFUSE_PUBLIC_KEY and LANGFUSE_SECRET_KEY, then swapped the standard OpenAI import for the langfuse.openai drop-in wrapper, and every completion call was traced automatically with no other code changes.

For custom functions outside the SDK, the @observe decorator creates trace spans around anything you wrap. And if your code is already OpenTelemetry-instrumented, you can skip the Langfuse SDK entirely and point your existing exporter at the OTLP endpoint. Three paths in, none of which require adopting a framework.

[Screenshot needed: Langfuse trace detail with nested spans] I could expand each nested span to see the exact input, output, latency, and cost per call. Source: Panoply

Reading Traces at Volume

Each trace shows the full call tree: nested spans with inputs, outputs, model parameters, token counts, and cost per observation. For a single request against a two-step chain, that layout answers “what did the model actually see” without leaving the trace.

At scale, deeply nested agent traces get visually noisy, and traces carrying base64-encoded images render the full payload as text, which drags the view badly. Langfuse’s own GitHub tracker documents both: one bug report describes pages taking 1 to 2 full seconds to load after the v3.124.1 upgrade where they used to be near-instant, and another tracks the slowdown from many observation levels and uploaded images. If your agents ship screenshots through OpenAI’s image_url payloads, expect the trace view to labor.

Running the First Scored Experiment

Evaluation starts with a dataset: representative inputs paired with expected outputs. I picked a built-in LLM-as-judge template (Hallucination, Correctness, and Relevance are the obvious first three), ran the app against the dataset, and each item executed as input, execution, output, then evaluation.

The result is a set of automatically scored traces you compare against the baseline run, which turns “does the new prompt help” into a number instead of a vibe. The step most teams miss: the langfuse/experiment-action GitHub workflow step runs this same experiment in CI and can block a pull request when quality metrics regress.

Author’s Testing Notes 📝

Decide the hosting question before you commit, not after. The v3 self-host checklist is concrete: Postgres v12+, ClickHouse v24.3+ with a persistent volume, Redis or Valkey, and S3-compatible storage, with minimum sizings for five containers. If that list is not already familiar territory for your infra team, start on Langfuse Cloud and port later; the MIT license keeps that door open.

— Panoply reviewer

Tracing: OpenTelemetry-Native and Framework-Agnostic

Should your observability platform care which framework you build on? LangSmith effectively does. Langfuse, since v3, does not: its OTLP endpoint accepts traces from any OpenTelemetry-instrumented code, no Langfuse SDK required.

The named first-party integrations cover the mainstream stacks:

  • LangChain and LlamaIndex via native callbacks
  • OpenAI SDK via the drop-in langfuse.openai wrapper
  • Anthropic SDK and Pydantic AI
  • Vercel AI SDK via a LangfuseExporter riding Vercel’s own OpenTelemetry instrumentation
  • CrewAI for multi-agent workflows
  • Anything else through plain OTLP export

In practice, a team already exporting OpenTelemetry spans from its services adds Langfuse as one more OTLP destination, and LLM traces land inside the instrumentation conventions that team already runs. No SDK migration, no framework rewrite.

The same independence holds on the model side. Langfuse runs no inference itself; it observes the calls your app makes to OpenAI, Anthropic, or any other provider, so swapping model vendors leaves the tracing setup untouched. Teams building directly on Pydantic AI, the Vercel AI SDK, or the raw OpenAI SDK get the same first-class tracing a LangChain shop gets.

An HN commenter made the sharper point: from a user’s chair, LangSmith and Langfuse look nearly identical until licensing, self-hosting, and framework lock-in enter the picture. That is exactly where they split. LangSmith auto-traces LangChain and LangGraph with zero configuration but needs manual @traceable decoration everywhere else. Arize Phoenix is the other OTel-native open option, under the more restrictive Elastic License 2.0 and backed by SQLite or Postgres rather than a ClickHouse-class store.

If your stack is all-LangChain, LangSmith’s zero-config tracing is real convenience. If your stack is heterogeneous, or likely to change frameworks within a year, Langfuse is the one that will not care.

Evals, Experiments, and Prompt Management

The full quality loop exists in one place: version a prompt, test it against a dataset, gate the release, roll back if it regresses. That loop, not the trace viewer, is why teams pick a platform over raw logs.

The eval side ships eight built-in LLM-as-judge templates: Hallucination, Helpfulness, Relevance, Toxicity, Correctness, Context Relevance, Context Correctness, and Conciseness. They run on any tool-calling model via OpenAI, Azure OpenAI, Anthropic, or AWS Bedrock APIs, and since a November 2024 update the managed judges can read the dataset’s expected_output as ground truth in offline experiments. When none of the eight fits, an evaluator can start from a blank judge prompt instead.

Datasets plus experiments make the comparison repeatable: every run generates scored traces, so a prompt or model swap becomes a baseline-versus-candidate table. The langfuse/experiment-action GitHub step takes that into CI/CD and blocks pull requests when scores drop against a versioned dataset. Prompt management closes the loop: prompts are versioned server-side, so a regressed version rolls back from Langfuse directly, with no application redeploy.

Online evaluation rules apply sampling rates rather than scoring every interaction, so a low-frequency failure mode can pass unnoticed between samples. Alerting is currently scoped to prompt-version Slack notifications rather than behavioral anomaly alerts. Langfuse positions this as adequate for offline eval workflows and early-stage production monitoring, and expects you to write custom eval logic for specific failure modes.

The self-serve automation has edges too: there is no documented way to trigger UI-configured evaluators programmatically from a scheduled pipeline, and evaluator filters on tags, userId, or sessionId match only attributes set on the observation itself, so a filter keyed to trace-level attributes silently matches nothing.

For offline, eval-driven development with CI gates, Langfuse is as complete as anything in this category. For full-coverage online scoring or on-call behavioral alerting, plan to wire your own pipeline on top.

Self-Hosting Langfuse v3: The Real Footprint

The exit path is real and costs nothing in license fees. The price is infrastructure, and it is bigger than the v2 era’s reputation suggests.

A v3 deployment requires four backing services plus two application containers:

  • Postgres v12+ for transactional and metadata storage
  • ClickHouse v24.3+ for the analytical store of traces, observations, and scores; it keeps metadata on local disk, so skip a persistent volume and you can lose table access
  • Redis or Valkey for queues and caching
  • S3-compatible blob storage for incoming events, multimodal inputs, and large exports
  • Langfuse Web and Worker containers on top

Minimum sizings from the docs: 2 CPU/4 GiB each for Web, Worker, and Postgres, 1 CPU/1.5 GiB for Redis, and 2 CPU/8 GiB for ClickHouse, with every component set to UTC. Deployment paths are Docker Compose, a Helm chart for Kubernetes, or AWS, Azure, and GCP infrastructure-as-code templates. SSO covers Google, GitHub, Azure AD, Okta, Auth0, AWS Cognito, Keycloak, and JumpCloud.

HN commenters called the v3 migration what it was: a real complexity step-up from v2’s single Postgres container, accepted because the ClickHouse backend fixed analytical-query performance at scale. Another commenter in the same thread explained why teams self-host at all: traces carrying sensitive customer data cannot sit on someone else’s server, which makes self-hosting a compliance requirement rather than a preference.

At light volume the Docker Compose path is a contained job. At production volume this is a standing platform commitment, with four backing services and two application containers to patch and monitor, and the $29 Core cloud tier is the honest comparison for any team without a compliance mandate.

Adoption, Momentum, and the ClickHouse Acquisition

The public signals are strong by any open-source standard: between 34,700 and 35,100 GitHub stars depending on the 2026 snapshot, roughly 3,800 forks, a repository created in May 2023, and releases at a near-daily pace. v4.38.0 shipped on September 17, 2026, one day after v4.37.0. The company came out of Y Combinator’s Winter 2023 batch.

The vendor’s own numbers are bigger and should be read as self-reported: 50,000+ companies, 65M+ SDK installs per month, and 38M+ Docker pulls, none independently audited. Langfuse’s own pages disagree with each other on Fortune 50 and Fortune 500 customer counts, so we leave those out. G2 has not yet accumulated enough Langfuse reviews to add an independent read on buyer sentiment.

The ownership change is recent: as of January 2026, Langfuse is part of ClickHouse, Inc., per the copyright header in the repository’s own LICENSE file, with deal terms undisclosed, landing around ClickHouse’s $400 million Series D. No public change to the MIT core license or the self-hosting policy has followed so far. The buyer already powers Langfuse’s analytical backend, which cuts both ways: deep technical alignment on one side, single-vendor concentration on the other.

Practitioner sentiment stayed warm through all of it. One HN commenter with heavy LLM spend called Langfuse “the backbone to debug when things go haywire,” and another, after 15 months of use, ranked it ahead of alternatives specifically for having no dark patterns.

How Does Langfuse Compare to Competitors?

Each of these five rivals beats Langfuse in one specific situation:

  • LangSmith is the choice for all-in LangChain and LangGraph shops: zero-config automatic tracing, more mature built-in eval tooling, and a more polished dashboard. It is proprietary, self-hosting requires an Enterprise contract, and the same independent estimate cited earlier puts it near 6x Langfuse’s cloud cost at 100K traces per month.
  • Arize Phoenix is the other OTel-native open option, source-available under the Elastic License 2.0 rather than MIT. Its SQLite or Postgres backend is lighter to run than Langfuse’s v3 stack, it ships an open-source eval metrics library, and the enterprise feature set lives in the separate Arize AX product.
  • Braintrust leads with evaluation rather than observability: a free tier of 1M spans and 10K evals per month, and CI/CD gates that block prompt regressions in pull requests. Its own marketing names Perplexity, Airtable, and Replit as users; tracing depth is not its emphasis.
  • Helicone is the proxy-and-gateway path: one-line setup and simple cost tracking for OpenAI-centric stacks, without Langfuse’s prompt management or experiment depth. Note that its hosted product stopped accepting new signups after the March 2026 Mintlify acquisition, so evaluate the self-hosted route.
  • W&B Weave fits teams already living in Weights & Biases; its pricing scales with both usage volume and seats, which compounds at scale.

The decision rule from the tracing section holds across the board. A heterogeneous stack plus a self-host requirement plus cost pressure points at Langfuse; a pure LangChain stack is the one case where LangSmith’s premium buys real convenience. Braintrust earns a second look only when evaluation, not tracing, is the actual job to be done.

How We Test LLM Observability Tools

We combine independent analysis, data collection, and hands-on testing: we set each tool up ourselves, run a real task end to end, and compare it directly against its closest rivals. Data collection covers public signals (GitHub, PyPI, Docker Hub, Stack Overflow, G2, and Gartner peer reviews), and we hand-check every vendor pricing page rather than trusting third-party summaries. Where a claim rests on a vendor’s own numbers, we label it that way. Prices current as of September 2026.

Sustained user sentiment is weighed deliberately lightly, review-aggregator scores lightest of all, and anything we cannot measure is marked N/A rather than scored zero. Signals refresh monthly and editorial verdicts quarterly. Sponsors and affiliate partners cannot change a score.

Langfuse Review: Should You Standardize on Langfuse?

We recommend Langfuse for most teams standardizing an LLM observability and evaluation stack. The case is concrete: framework-agnostic OTel-native tracing, the eight-template eval loop with CI gating, cloud pricing near $69 against LangSmith’s $420 at a 100K-trace workload in one independent estimate, and an MIT core that keeps the self-host exit open. Compliance-driven teams get that same core free on their own infrastructure.

All-LangChain teams get real value from LangSmith’s zero-config tracing. Teams that need full-coverage online scoring and behavioral alerting out of the box will be building that layer themselves, as covered in the evals section. And anyone who weighs the January 2026 ClickHouse ownership more heavily than the license’s current guarantees should factor that risk in now, not later.

The next action costs nothing: create a Hobby account, swap in the langfuse.openai wrapper, and you have real traces from your own app the same day. Prospective self-hosters should run that same cloud trial first and port to the v3 stack only once the tool has earned the infrastructure.

FAQ

Is Langfuse actually open source?

Yes, for the core: the MIT license covers all core features and APIs without limits. GitHub shows “Other” because exactly nine Enterprise Edition features (project-level RBAC, audit logs, the org management API with SCIM, data retention policies, protected prompt labels, server-side data masking, UI customization, organization creators, and the instance management API) carry a separate commercial license.

Does self-hosting Langfuse require a paid plan?

No. Self-hosting is free and unrestricted under the MIT core, whether you deploy via Docker Compose, Kubernetes with Helm, or the AWS, Azure, and GCP templates. The Hobby, Core, Pro, and Enterprise tiers price only the managed Langfuse Cloud product.

What does self-hosting Langfuse v3 require?

Four backing services: Postgres v12+, ClickHouse v24.3+ with a persistent volume, Redis or Valkey, and S3-compatible blob storage, plus the Web and Worker containers. That is a materially heavier footprint than v2, which ran on a single Postgres container. See Self-Hosting Langfuse v3 above for minimum sizings.

How does Langfuse pricing compare to LangSmith?

Langfuse is consistently cheaper on cloud pricing. One independent estimate puts a 5-seat team at 100K traces per month at roughly $69 on Langfuse versus about $420 on LangSmith, and near a third of LangSmith’s cost at 1M traces, with self-hosted Langfuse cheaper still. Langfuse also includes unlimited users from the $29 Core tier.

Who owns Langfuse now?

Langfuse has been part of ClickHouse, Inc. since January 2026, confirmed by the copyright header in the repository’s LICENSE file; deal terms were not disclosed, and the acquisition landed around ClickHouse’s $400 million Series D. As of September 2026, no public change to the MIT core license or the free self-hosting policy has followed.

Spotted a wrong price or a missing integration? Send a correction. A human reads every one.

Similar tools

Other tools in the same category, with the same card and the same honest pricing.

Weights & Biases

Experiment tracking

The polish leader in experiment tracking: dashboards, Sweeps, reports, and Weave for LLM observability; Pro is $60/month flat for up to 10 seats.

Visit site

MLflow

Experiment tracking

Open-source standard for experiment tracking, model registry, and serving; 3.x adds OpenTelemetry tracing, LLM judges, and a Prompt Registry.

Visit site

LangSmith

Observability

LangChain's observability and evaluation platform: zero-config LangGraph tracing, CI-gated evals, prompt hub; closed-source SaaS.

Visit site

LangChain

RAG framework

The most-adopted LLM framework, rebuilt on LangGraph in 1.0: durable agent execution, 1,000+ integrations, MIT-licensed; LangSmith is the paid layer.

Visit site

LlamaIndex

RAG framework

RAG-first data framework for LLM apps: 160+ connectors, purpose-built indexes, MIT-licensed; LlamaCloud adds managed parsing and indexing.

Visit site

Weaviate

Dedicated

Open-source vector database with native BM25 + vector hybrid search, per-tenant shard isolation, and multimodal vectorizer modules.

Visit site

Pinecone

Dedicated

Fully managed serverless vector database with zero-ops scaling, deep RAG framework integrations, and enterprise compliance.

Visit site