Braintrust

Eval-first platform for AI products: datasets, experiments, CI-gated scoring, and a hybrid VPC data plane; bills scores, not traces.

Best for: Production AI teams with a release cadence who want regressions blocked in CI and failures converted into permanent tests

Pros

  • Datasets, experiments, traces, and online scoring run in one loop instead of stitched-together tools
  • One-click conversion of a production failure into a permanent regression test
  • Free tier includes unlimited users and 10,000 scores per month
  • Enterprise data plane runs in your own AWS, GCP, or Azure VPC via an official Terraform module
  • Autoevals scoring library is open source and works without the paid platform

Cons

  • No plan between $0 and $249 per month
  • Data ($3/GB) and score ($1.50 per 1,000) overages stack with no hard spending cap
  • Self-hosting and SSO/SAML are Enterprise-only
  • Engineers drive it; product managers ride along
  • Confirmed May 2026 breach came with a key-rotation advisory

Editor’s note: We combine independent analysis, data collection, and hands-on testing to review data and AI tools, weighing pricing transparency, real-world adoption signals, development momentum, openness and exit costs, practitioner sentiment, and our editorial verdict.

Quick verdict: We recommend Braintrust for teams with a real release process who want regressions blocked in CI and production failures converted into permanent test cases. The trade-off: nothing between free and $249 per month, and no hard cap on either usage meter.

This Braintrust review tests the evals-first bet, since rival vendors write most of its coverage. Among the LLM observability and evaluation tools in our directory, Braintrust puts evaluation, not tracing, at the center.

Key Takeaways 🔍

  • Free Starter tier is genuinely usable: unlimited users, 10,000 scores a month, 1 GB of data (verified September 2026)
  • Notion’s ~70-engineer AI team spends ~80% of its work evaluating in Braintrust and ships new frontier models within 24 hours (vendor case study)
  • Enterprise data plane runs in your own VPC via Terraform; Braintrust’s servers do not store those traces
  • Pro jumps straight to $249 a month, and moderate usage lands near $339 (third-party estimate)
  • A confirmed May 2026 security incident came with a key-rotation advisory

Pros and Cons

Pros

  • Datasets, experiments, traces, and online scoring run in one loop instead of stitched-together tools
  • One-click conversion of a production failure into a permanent regression test
  • Free tier includes unlimited users and 10,000 scores per month
  • Enterprise data plane runs in your own AWS, GCP, or Azure VPC via an official Terraform module
  • Autoevals scoring library is open source and works without the paid platform

Cons

  • No plan between $0 and $249 per month
  • Data ($3/GB) and score ($1.50 per 1,000) overages stack with no hard spending cap
  • Self-hosting and SSO/SAML are Enterprise-only
  • Engineers drive it; product managers ride along
  • Confirmed May 2026 breach came with a key-rotation advisory

How Much Does Braintrust Cost?

Pro’s $249 headline sits on top of two usage meters, and those meters decide the bill. Braintrust sells three tiers, which I verified on braintrust.dev/pricing in September 2026:

  • Starter ($0): $10 per month in model credits, 1 GB of processed data ($4/GB over), 10,000 scores ($2.50 per 1,000 over), 14-day retention, unlimited users, projects, datasets, and experiments, 1 human review score per project; no RBAC, no SSO
  • Pro ($249/month): $100 in model credits, 5 GB of data ($3/GB over), 50,000 scores ($1.50 per 1,000 over), 30-day retention ($0.50/GB/month to extend), basic RBAC, custom charts, and Environments; startups may qualify for 6 to 12 months free
  • Enterprise (custom): custom data and score volumes, custom retention with export options, custom RBAC, SSO/SAML, self-hosted or hybrid deployment, premium support, HIPAA BAA
PlanPriceData includedScores includedRetentionKey adds
Starter$01 GB/month ($4/GB over)10,000/month ($2.50/1k over)14 daysUnlimited users and projects, $10 model credits
Pro$249/month5 GB/month ($3/GB over)50,000/month ($1.50/1k over)30 days ($0.50/GB/month to extend)Basic RBAC, custom charts, Environments, $100 model credits
EnterpriseCustomCustomCustomCustom, with export optionsSSO/SAML, custom RBAC, self-hosting, HIPAA BAA, premium support

“Processed data” is broader than traces: it counts the bytes ingested across logs, experiments, datasets, and attachments, so an eval-heavy team fills the data meter even with modest production traffic.

On Pro, the sticker is not the bill. A team logging 10 GB of trace data and running 100,000 scores a month pays roughly $339: the $249 base, plus $15 for the 5 extra GB at $3/GB, plus $75 for the 50,000 extra scores at $1.50 per 1,000. That figure is a third-party estimate, but I cross-checked it against the vendor’s own per-unit rates and the arithmetic holds.

Both meters run open-ended. Braintrust has no hard spending cap, so a traffic spike or a verbose agent can move the bill without asking you first.

Is Braintrust Good Value for Money?

  • The free tier funds a real eval practice: unlimited seats, 10,000 scores, and 1 GB a month are enough to run evaluations on a small team’s releases before anyone pays
  • The missing mid-tier is the real gap: a team outgrowing Starter has nowhere to land except $249 plus two open meters
  • Rivals meter differently: LangSmith Plus costs $39 per seat a month with 10,000 base traces, and Langfuse Cloud’s Core plan starts at $29 a month, but neither maps onto Braintrust’s score meter, so compare on your own usage shape rather than sticker prices

Author’s Testing Notes 📝

Exhaust the Starter tier before you spend anything: with unlimited users, the only real constraints are the 1 GB data meter and the 14-day retention. Watch both meters from day one, because they, not the $249 base, decide your actual bill. And if you are early-stage, apply for the startup program (6 to 12 months of Pro free) before paying list price.

— Panoply reviewer

My Experience With Braintrust

The whole platform is organized around one loop: dataset, experiment, score, ship. I ran that loop the way Braintrust documents it, from an empty project to a CI gate.

Building the First Eval

I started with a dataset: rows of inputs paired with expected outputs, and later rows converted from logged traces. From there the Playground let me run prompts against multiple models side by side, so a prompt tweak shows its effect across, say, two frontier models on the same rows before anything ships. The dataset, the scorers, and the resulting experiments all live inside the same project, so nothing gets exported between tools along the way.

Scoring is the step most tools leave to you, and Braintrust gives it three concrete forms: heuristic checks, custom code scorers, and LLM-as-judge graders. I could assemble any of them in the UI without writing code, or pull the same scorer types from the autoevals library in Python or TypeScript for a code-first setup. Results save as an experiment, and every experiment compares against a baseline run so a score change is visible as a diff, not a vibe.

[Screenshot needed: Braintrust experiment comparison against baseline] Each experiment run diffs against the baseline, so I could see exactly which rows regressed after a prompt change. Source: Panoply

Turning a Production Failure into a Test

One click turns a production failure into a permanent test, and that workflow is the product’s best argument. A failing interaction arrives as a production trace, captured through the SDK or an OpenTelemetry integration. Braintrust converts that trace into a dataset entry. Attach or reuse a scorer that defines what passing looks like for that case, and the failure joins the standard experiment suite.

From then on, every future experiment re-runs that case, so the same failure cannot silently ship again. Bug reports become tests instead of tickets.

Gating CI

Braintrust runs evals inside CI, with GitHub Actions supported directly and statistical significance testing built in. A prompt or model change that regresses scores against the baseline dataset fails the merge, the same way a broken unit test would. LangSmith expects you to write your own eval runner for this; here the gate is the native workflow.

A passing gate still has limits: Braintrust’s Notion case study describes models that pass a regression eval at 100% yet behave differently in ways that matter for the product, which is why that team pairs regression suites with frontier evals.

Author’s Testing Notes 📝

The fit test is blunt: if your team has no eval dataset, no recurring release process, and no appetite for scorer design, Braintrust is process overhead, not tooling. Adopt it when scoring outputs has become routine, not before. Prototypes get nothing from a CI gate they never run.

— Panoply reviewer

Scorers, Autoevals, and Loop

An eval feature does nothing until someone writes the scorers, and scorer design is exactly the appetite a team new to evals lacks. Braintrust attacks that bottleneck from three sides: by hand in the UI, from an open library, or generated by its Loop assistant.

The library side is genuinely open. Autoevals (github.com/braintrustdata/autoevals, installable via pip or npm) ships heuristic, statistical, and model-graded scorers, adapted in part from OpenAI’s evals project. It reads an OPENAI_BASE_URL environment variable or falls back to the Braintrust AI Gateway, and it runs entirely without the paid platform. That is the honest boundary between open and proprietary here: the scoring primitives are free to take with you; the datasets, experiments, tracing, and CI gating around them are the paid product.

Loop is the generated side. You describe the measurement in natural language, such as tool-selection accuracy, conversation quality, or a domain business rule, and Loop produces production-ready scorers. It also reads your production trace data directly and proposes improved prompts and datasets from what it finds.

Founder Ankur Goyal laid out the thesis on the Latent Space podcast: “Evals are the centerpiece of systematic AI Engineering,” and the design goal was “making evals something that software engineers, product managers can immediately do.” One of his datapoints stuck with me: open-source models account for less than 5% of usage among customers running Braintrust in production, so most production traffic still runs on closed frontier models.

The verdict on this cluster: Loop lowers the scorer-design bar that leaves eval features idle, and autoevals makes the exit path real rather than rhetorical.

Hybrid Deployment: Your VPC, Their Control Plane

Braintrust’s Enterprise self-hosting splits the product in two: a vendor-hosted control plane and a data plane you run yourself. The control plane covers the UI, metadata, and auth. The data plane, deployed inside your own AWS, GCP, or Azure VPC through an official Terraform module (on AWS it packages as Lambda functions plus a Postgres database), holds everything sensitive: traces, datasets, prompts, completions, and customer inputs.

SDKs and the browser UI talk directly to your data plane, and Braintrust’s servers do not store or proxy that data. An on-prem variant exists for teams that cannot use a cloud VPC at all. The AWS module is public on GitHub (braintrustdata/terraform-aws-braintrust-data-plane), so your platform team can read exactly what gets deployed before signing. Neither Starter nor Pro offers any self-hosting; the whole arrangement is Enterprise-only, packaged alongside SSO/SAML and a HIPAA BAA.

The security record belongs in the same section. Braintrust confirmed a breach around May 6, 2026, and advised customers to rotate sensitive API keys as a precaution. No detailed public postmortem has been published, so the technical scope beyond that advisory is not publicly documented.

Among rivals, LangSmith is closed source and also gates self-hosting behind its Enterprise tier, while Langfuse and Arize Phoenix are open source and self-host free. Braintrust’s differentiator is not that self-hosting exists but that the Terraform module and the direct SDK-to-data-plane path are documented concretely.

Who Actually Runs Braintrust

The customer evidence starts with Notion, though it comes from Braintrust’s own case study, so read it as vendor-sourced. Notion’s roughly 70-engineer AI team spends about 80% of its work evaluating feedback and traces inside the platform, and deploys a newly released frontier model within 24 hours by running it through existing regression and frontier evals. One multilingual evaluation across Japanese, Korean, and English became one of the top quality improvements of the year for APAC customers. Sarah Sachs, Notion’s AI modeling lead, describes the practice plainly: “I sat down in Braintrust and looked at some of the worst experiences our customers had and tried to understand how we can be better.”

In Forbes coverage, Notion cofounder Simon Last credits the measurement layer for ambition: “It lets us build much more complex stuff, more confidently.” Investor Elad Gil, in the same coverage, frames the pitch: “This allows you to take the vibes out of AI development.”

The least polished datapoint is the most independent: Zapier cofounder Bryan Helmig, in his own Hacker News comment, wrote “We’ve been using Braintrust for evals at Zapier and it’s been really great.” It is one forum comment, but no vendor wrote it.

Founded in August 2023 by Figma’s former AI lead Ankur Goyal, Braintrust raised a $36 million Series A led by Martin Casado at a16z, with Databricks and Datadog investing, at an estimated $150 million valuation. In February 2026 it added an $80 million Series B led by ICONIQ at a reported $800 million valuation.

How Does Braintrust Compare to Competitors?

Every rival here solves an overlapping problem with a different center of gravity:

  • LangSmith is trace-first, with the deepest developer experience for LangChain and LangGraph pipelines and OpenTelemetry support for everything else. It meters traces, not scores, and its free tier caps at one user and 5,000 traces a month. CI evaluation means writing your own runner, where Braintrust gates merges natively.
  • Langfuse is open source and self-hostable free at any tier, which makes it the cost-efficient choice at scale, and the YC W23 company has strong community adoption. Its Core cloud plan starts at $29 a month. What it lacks is Braintrust’s opinionated eval methodology; you assemble your own practice on top.
  • Arize Phoenix offers free, ungated open-source tracing and evals with more than 10,000 GitHub stars and ML-monitoring heritage, so drift and classic model observability run deeper. Pre-production simulation tooling is minimal.
  • Helicone is one-line proxy logging: swap the OpenAI base URL and get cost, latency, and time-to-first-token analytics. It is not built for structured evals or regression gating, and its hosted product stopped new signups after the March 2026 Mintlify acquisition.
  • W&B Weave fits teams already running Weights & Biases for experiment tracking and model registry, keeping ML experiments and LLM traces in one environment. Its training-paradigm framing suits research workflows more than production LLM reliability.

The decision rule is short. Pick Braintrust when regressions blocking merges is the feature you are buying. Pick a logging-first tool when you just need to see traces, because you would be paying score-meter prices for a workflow you do not run. And when free self-hosting is the requirement that outranks everything else, Langfuse and Phoenix are the two names on the list.

How We Test LLM Observability Tools

We combine independent analysis, data collection, and hands-on testing: we set the tool up ourselves, run a real eval task end to end, and compare it directly against the rivals in its category. Alongside that, we collect public signals (GitHub activity, PyPI and Docker Hub downloads, Stack Overflow volume, G2 and Gartner peer reviews) and hand-check every vendor pricing page rather than repeating figures other sites publish. Prices current as of September 2026.

Aggregate review scores get deliberately small weight, because they lag and they game easily. Where an area cannot be measured for a tool, we mark it N/A rather than scoring it zero. Signals refresh monthly, editorial verdicts quarterly, and sponsors and affiliates cannot change a score.

Braintrust Review: Should You Gate Your AI Releases on Braintrust?

We recommend Braintrust for production AI teams that ship on a cadence, want regressions blocked in CI rather than caught by users, and can absorb $249 a month plus meters or qualify for the startup program. Compliance-sensitive enterprises get a second reason to shortlist it: the VPC data plane covered above is documented down to a public Terraform module.

Skip it in three cases. Prototypes without a release process will feel it as pure overhead, as covered in My Experience With Braintrust. Teams that only want trace viewing have cheaper paths in Langfuse or Phoenix. And budget-capped teams should weigh the uncapped data and score meters seriously, because the $339 worked example (a third-party estimate) reflects moderate usage, not an edge case.

The next action costs nothing: run a real eval practice on the free Starter tier, with its unlimited seats and 10,000 scores a month. The moment both meters trend toward their caps is the moment the $249 conversation is worth having, ideally with the 6 to 12 month startup program on the table.

FAQ

Is Braintrust worth it for a small team?

Yes, once you are scoring outputs systematically and shipping to production on a cadence; the free Starter tier supports that with unlimited users and 10,000 scores a month. Before that point it is overkill: a prototype with no eval dataset and no release process gains nothing from regression gating. Startups that outgrow Starter may qualify for 6 to 12 months of Pro free.

How does Braintrust pricing scale past the $249 sticker price?

Two meters run on top of the base: processed data (bytes ingested across logs, experiments, datasets, and attachments) at $3/GB beyond Pro’s 5 GB, and scores at $1.50 per 1,000 beyond 50,000. Neither has a hard cap. A team logging 10 GB and running 100,000 scores a month lands near $339, per the third-party worked example in the pricing section above.

Can our data stay in our own cloud?

Yes, on the Enterprise tier. The data plane (traces, datasets, prompts, completions) runs inside your own AWS, GCP, or Azure VPC via an official Terraform module, and SDKs talk to it directly. The control plane (UI, metadata, auth) stays vendor-hosted. HIPAA BAA and SAML SSO are also Enterprise-only.

Is autoevals the same as the Braintrust platform?

No. Autoevals is a standalone open-source scoring library with heuristic, statistical, and LLM-as-judge scorers, installable via pip or npm and usable without a Braintrust account. It reads an OPENAI_BASE_URL environment variable or falls back to the Braintrust AI Gateway. The paid platform is everything built around it: datasets, experiments, tracing, Loop, and CI gating.

Has Braintrust had security incidents?

Yes. Braintrust confirmed a breach around May 6, 2026, and advised customers to rotate sensitive API keys as a precaution. No detailed public postmortem has been published, so the technical scope beyond that advisory is not publicly documented.

Spotted a wrong price or a missing integration? Send a correction. A human reads every one.

Similar tools

Other tools in the same category, with the same card and the same honest pricing.

Comet

Experiment tracking

Experiment tracking with bundled production drift monitoring, plus Opik, its Apache-2.0 open-source LLM evaluation product.

Visit site

Weights & Biases

Experiment tracking

The polish leader in experiment tracking: dashboards, Sweeps, reports, and Weave for LLM observability; Pro is $60/month flat for up to 10 seats.

Visit site

MLflow

Experiment tracking

Open-source standard for experiment tracking, model registry, and serving; 3.x adds OpenTelemetry tracing, LLM judges, and a Prompt Registry.

Visit site

LangSmith

Observability

LangChain's observability and evaluation platform: zero-config LangGraph tracing, CI-gated evals, prompt hub; closed-source SaaS.

Visit site

Langfuse

Observability

Open-source LLM engineering platform: OpenTelemetry-native tracing, LLM-as-judge evals, prompt management; MIT core with free self-hosting.

Visit site

LangChain

RAG framework

The most-adopted LLM framework, rebuilt on LangGraph in 1.0: durable agent execution, 1,000+ integrations, MIT-licensed; LangSmith is the paid layer.

Visit site

LlamaIndex

RAG framework

RAG-first data framework for LLM apps: 160+ connectors, purpose-built indexes, MIT-licensed; LlamaCloud adds managed parsing and indexing.

Visit site

Weaviate

Dedicated

Open-source vector database with native BM25 + vector hybrid search, per-tenant shard isolation, and multimodal vectorizer modules.

Visit site

Pinecone

Dedicated

Fully managed serverless vector database with zero-ops scaling, deep RAG framework integrations, and enterprise compliance.

Visit site