Editor’s note: We combine independent analysis, data collection, and hands-on testing to review data and AI tools. This review weighs pricing transparency, real-world adoption signals, development momentum, openness and exit costs, practitioner sentiment, and our editorial verdict.
Quick verdict: We recommend Arize for engineering teams that want the deepest evaluation and drift tooling among LLM observability tools, plus a free, ungated Phoenix core. The trade-offs: the platform assumes engineers, and unpublished Enterprise pricing that third-party trackers put at $60,000-100,000+ per year. This Arize AI review covers pricing, the Phoenix boundary, and the Dynatrace deal.
Key Takeaways 🔍
- Dynatrace is buying Arize: a $915M cash-and-stock agreement announced August 13, 2026; both companies state Arize stays independent after close
- AX tiers: Free ($0, 25k spans/month), Pro ($50/month, 50k spans), Enterprise custom; unlimited users, evals, and experiments on every tier (verified September 2026)
- Phoenix, the open-source core, self-hosts under Elastic License 2.0 with zero feature gating and 11,600+ GitHub stars
- Con: engineering-centric. Product managers report needing an engineer to drive it
- Con: trace ingestion once lagged rivals. The May 2026 realtime rollout fixed it on every tier, but older comparisons still cite it

Pros and Cons
Pros
- Phoenix self-hosts free under Elastic License 2.0 with zero feature gating or caps
- Unlimited users, evals, and experiments on every AX tier, including the $0 Free tier
- OpenInference rides on OpenTelemetry, so instrumentation ports to any OTel-compatible backend
- Eval and drift depth built on six years of classical ML observability heritage
- Realtime trace ingestion on all tiers since the May 2026 rollout
Cons
- Enterprise pricing unpublished; third-party trackers estimate $60,000-100,000+ per year
- Engineering-centric platform that non-engineers struggle to drive alone
- Span-and-volume billing gets costly once workloads pass Pro’s 50k spans and 10GB
- Modest G2 review base: 4.2/5 from roughly 28 reviews as of 2026
- Pending $915M Dynatrace acquisition adds roadmap uncertainty until close
How Much Does Arize Cost?

Arize AX sells three tiers, and the two published ones are cheap; the budget risk all lives in the third. Here is the lineup as of September 2026:
- Free ($0): 25,000 spans/month, 1GB storage, 15-day retention, 10 issues/month
- Pro ($50/month): 50,000 spans/month, 10GB storage, 30-day retention, 25 issues/month
- Enterprise (custom): unlimited signals and issues, custom span and storage volumes, SaaS or self-hosted deployment
Every tier, including Free, comes with unlimited users, evals, and experiments. Against seat-billed rivals that is a real differentiator: LangSmith Plus runs about $39 per seat per month, so a 20-person team pays roughly $780 in seat fees there versus a flat $50 on AX Pro.
| Tier | Price | Spans/month | Storage | Retention | Issues/month |
|---|---|---|---|---|---|
| Free | $0 | 25,000 | 1GB | 15 days | 10 |
| Pro | $50/month | 50,000 | 10GB | 30 days | 25 |
| Enterprise | Custom | Custom | Custom | Custom | Unlimited |
AX bills on span counts and data volume, not seats, so a data-heavy workload that outgrows Pro’s 50,000 spans and 10GB has nowhere to go but an Enterprise quote.
That Enterprise quote is also where the real feature boundary sits. A custom deal adds:
- Self-hosted deployment of the full AX platform
- Enterprise SSO and audit logs
- SOC 2 Type II, GDPR, and HIPAA compliance packaging
- Multi-organization and multi-space RBAC
- Data Fabric
- Custom managed agents with repository access
- Swarm observability across managed and third-party agents
- Dedicated support with custom SLAs
Arize publishes no Enterprise price. Third-party pricing trackers report a median around $60,000 per year, scaling past $100,000 for larger deployments. That range is tracker-sourced, not vendor-confirmed, so treat it as a planning number and get the real quote in procurement.
Phoenix, the open-source core, self-hosts free with no caps, and Arize states that a later migration to AX carries over your projects, datasets, and experiments, so starting at $0 does not strand your eval history.
Is Arize Good Value for Money?
- Free and Pro are low-risk. $0 and $50 flat, with unlimited seats, so a small team knows its ceiling before it commits
- Unlimited seats beat seat billing. The LangSmith Plus math above flips in Arize’s favor at two seats ($78 versus $50)
- Braintrust’s ladder is steeper: its free Starter tier jumps straight to $249/month with no mid-tier, plus $3/GB for tracing
- The value question is Enterprise. Pay the custom quote for the governance list above (self-hosting, RBAC, SSO, audit logs, SLAs), and not before you need it
Author’s Testing Notes 📝
Develop on self-hosted Phoenix, where the eval tooling costs nothing. Buy Pro at $50 when managed dashboards and hosted alerting start saving real engineer time. Price Enterprise against the governance features alone; if you do not need self-hosted AX, multi-org RBAC, or compliance packaging, the tracker-estimated $60,000+ spend has nothing to buy.
— Panoply reviewer
My Experience With Arize
You can evaluate Arize’s entire core stack without talking to sales, because the open-source Phoenix product runs in your own environment with no account and no license key. That is where I started, and it is where I would tell any team to start.
🐳 Standing Up Phoenix
Phoenix runs locally, in a Docker container, on Kubernetes via Helm, or inside your own cloud account. Development needs no signup at all: I could launch the server and start sending traces without creating an Arize account. Everything stays in your environment too, with no phone-home telemetry, which matters if your traces carry customer prompts.
Author’s Testing Notes 📝
Budget time for documentation spelunking on the self-host path. On Arize’s own community forum, one self-hoster evaluating Phoenix as an alternative to Braintrust and Humanloop reported checking the website, YouTube, and GitHub without finding setup docs that just work; the Arize team and community resolved it, but the docs did not get them there alone. The happy path is smooth, the edges are not.
— Panoply reviewer
🔌 Instrumenting With OpenInference
Instrumentation runs through the openinference-instrumentation package family, with per-framework packages for OpenAI, LangChain, DSPy, Vertex AI, and others. Once installed, the spans capture the LLM-specific context that generic APM tracing misses: full prompts, retrievals from vector stores, and tool or API calls.
The design decision that impressed me most is that OpenInference is a convention layered on OpenTelemetry, not a proprietary format. Any OTel-compatible backend can receive the same spans, so the hours you spend instrumenting are portable even if you later leave Arize entirely. Compare that with LangSmith, whose deepest integration assumes you live in LangChain.
📈 Moving to AX
The managed AX layer is where the same traces gain operational machinery: online evaluations that run against production traffic, custom alerting and dashboards, and Alyx, the platform’s agent copilot. AX reads the same OpenInference trace format Phoenix already uses, so there is no re-instrumentation, and Arize states that projects, datasets, and experiments carry over. The migration is an upgrade rather than a rebuild.
Phoenix answered my development questions (what did this chain retrieve, which prompt version scored better), while AX exists for the questions that only appear in production: is quality drifting this week, and who gets paged when it does. Arize’s own case studies claim Handshake deployed and scaled 15+ LLM use cases on AX in under six months; that is a vendor-published figure, but it matches the product’s pitch as the production layer.
Phoenix vs AX: What Free Actually Includes
Phoenix, the free core, includes tracing, evals, datasets, experiments, and a prompt playground, self-hosted under Elastic License 2.0 with no feature gating.
What AX adds sits almost entirely in operations and governance:
- Online evaluations against live production traffic
- Custom alerting and dashboards
- Alyx, the agent copilot
- Multi-organization RBAC, enterprise SSO, and audit logs
- Dedicated support, plus an optional self-hosted enterprise deployment mode
Realtime trace ingestion is no longer on that list either: it reached every AX tier in May 2026, and self-hosted Phoenix had it before then.
That makes the verdict unusual for this category: teams willing to self-host and instrument their own code can genuinely skip AX. Pay for AX when you want zero-setup managed dashboards, realtime alerting on production traffic without running the infrastructure yourself, or RBAC and SSO for a larger organization. Compare Braintrust, whose free Starter tier caps you at 1GB of processed data and 10K eval runs before a $249/month jump; Phoenix’s limits are whatever your own hardware imposes.
Adoption signals back the free path as a safe bet. The Phoenix repo sits at 11,600+ GitHub stars with 1,200+ forks, up from the 8,600 to 10,300 stars cited in earlier 2026 coverage of the same repo. The arize-phoenix package sits at version 17.2.0 on PyPI, with the openinference-instrumentation packages shipping alongside it. Arize itself reports 2.4M+ monthly Phoenix installs, 6M+ monthly OpenInference downloads, and 32M+ all-time downloads. Those are company figures rather than audited numbers, but an independent review separately cites 3M+ monthly downloads, and the public star count points the same direction.
Evals and Drift: Where the ML Heritage Shows
Why does a platform that turned to LLMs relatively late lead the category on evaluation? Because Arize spent its first years on the harder statistical problem. The company was founded in 2020 as a machine learning model monitoring platform for tabular and computer-vision models, has raised $131 million, and carried that drift-detection and embeddings toolkit into the LLM era intact.
An independent comparison of six observability platforms rates Phoenix first on eval rigor for exactly this reason: drift detection and embeddings analysis run on deeper statistics than LLM-native rivals ship. The same comparison puts Arize behind LangSmith and Langfuse on LLM-specific UI ergonomics. That split fits the platform’s engineering-centric reputation: the primitives go further, and the dashboards make you work harder for them.
The historical knock on Arize was ingestion speed. Before May 2026, complete traces could take noticeably longer to appear in the AX dashboard than in LangSmith, where traces land in seconds, and a June 2025 enterprise reviewer on G2 named latency and custom instrumentation as the main downside. A May 2026 release closed that gap by extending realtime trace ingestion to every AX tier, including Free; it had previously been limited to Enterprise customers and self-hosted Phoenix.
Gartner named Arize a Cool Vendor in its Enterprise AI Operationalization and Engineering report, per Arize’s announcement of the recognition. And AWS’s machine learning blog documents evaluating agent workflows built on its own Strands Agents SDK with Arize AX, evidence the tracing depth extends beyond the LangChain ecosystem.
If your evals stop at LLM-as-judge scoring, several rivals do the job. If they extend into drift, embedding drift, and production regression, Arize is the strongest pick in the category.
The Dynatrace Acquisition: What Buyers Should Price In
Six weeks before this review, Arize agreed to be acquired. On August 13, 2026, Dynatrace announced a definitive agreement to acquire Arize in a cash-and-stock deal valued at $915 million. Co-founders Jason Lopatecki and Aparna Dhinakaran join Dynatrace, reporting to CEO Rick McConnell. Both companies state that Arize will continue operating independently after close.
At announcement the deal was pending regulatory approval and other closing conditions, and the companies said they expected it to close later that quarter or early the next. Dynatrace is an APM incumbent buying an LLM observability layer; Datadog took the build route with its own LLM Observability product at $31+ per host per month plus a volume add-on.
The open questions are the ones no press release answers yet. Does AX pricing stay flat once it sits inside a Dynatrace price book? Does Phoenix keep shipping at its current pace under ELv2? Does “operating independently” survive the first joint roadmap cycle? Phoenix’s license is unchanged so far, and nothing in the announcement alters what self-hosters run today.
For anyone negotiating an AX Enterprise contract, put the acquisition questions (pricing continuity, support terms, roadmap commitments) directly into procurement before signing.
How Does Arize Compare to Competitors?
Which rival wins depends on what your stack already looks like. Among LLM observability and evaluation tools, each of these owns a niche Arize does not:
- LangSmith has the deepest LangChain and LangGraph integration in the category, with node-by-node state diffs, traces that appear in seconds, and a workflow for replaying evals against new models. It is cloud-first, seat-billed at about $39 per seat on Plus, and closed source; VPC deployment only arrives on its enterprise tier.
- Langfuse is the MIT-licensed open-core alternative: fully self-hostable on Postgres plus ClickHouse at $0 platform cost, framework-agnostic via OpenTelemetry, with a cloud tier from roughly $29-59 per seat. Its per-framework depth runs shallower than LangSmith’s.
- Braintrust is eval-first with strong experiment tooling, but its pricing ladder jumps from a free Starter tier (1GB processed data, 10K eval runs) straight to $249/month with no mid-tier, plus $3/GB for tracing.
- Helicone installs as a drop-in proxy in about five minutes with no SDK changes and captures API-level logs, not agent-execution tracing. Its hosted product stopped taking new signups after Mintlify acquired it in March 2026.
- W&B Weave suits research teams already on Weights & Biases, with automatic call tracking (inputs, outputs, cost, latency) through a decorator and strong prompt versioning. It leans toward training experimentation over production reliability and has no self-hosted option; Phoenix is the standard answer when that requirement appears.
- Datadog LLM Observability fits teams consolidating on existing APM, since LLM traces sit next to their infrastructure data, priced at $31+ per host per month plus an LLM-volume add-on rather than LLM-native usage pricing.
The decision rule: pick Arize for eval and drift depth plus the most complete ungated free core; pick LangSmith for LangGraph-native ergonomics; pick Langfuse for an MIT-licensed core and per-seat cloud pricing far below tracker-estimated AX Enterprise deals.
How We Test LLM Observability Tools
We combine independent analysis, data collection, and hands-on testing to review data and AI tools. For LLM observability that means standing the tool up ourselves, instrumenting a real application, and running a traced task end to end before weighing it against direct rivals on the same criteria.
We collect public signals from GitHub, PyPI, Docker Hub, Stack Overflow, G2, and Gartner peer reviews, and we hand-check every vendor pricing page rather than relying on republished numbers. Third-party review scores are deliberately small-weighted, and sustained practitioner sentiment counts more than any single rating. Areas we cannot measure are marked not applicable rather than scored zero. Signals are refreshed monthly and editorial verdicts quarterly, and sponsors and affiliates cannot change a score. Prices current as of September 2026.
Arize Review: Should You Build Your Eval Stack on Arize?
We recommend Arize for engineering teams whose evaluation needs run past LLM-as-judge into drift and embeddings analysis, for self-hosters (Phoenix is the most complete ungated open-source core in this category), and for teams tired of seat billing, since every AX tier includes unlimited users, evals, and experiments.
Look harder before committing if your team is not engineering-led; the spans-and-drift vocabulary that makes Arize powerful is also its learning curve. Arize watches agents and LLM apps but does not build them, so flow design, knowledge bases, and agent deployment stay in a separate tool. The ingestion-delay complaint in older comparisons is no longer a reason to pass, since realtime landed on every tier in May 2026. LangGraph-native shops get faster ergonomics from LangSmith, and procurement teams that cannot tolerate an unpublished Enterprise price (third-party trackers say $60,000-100,000+ per year) stacked on a pending acquisition should get answers in writing first.
The evaluation costs nothing and can start today: run Phoenix in Docker this afternoon, instrument one service with OpenInference, and upgrade only when managed dashboards or governance features force the question. Your projects, datasets, and experiments carry over when they do.
FAQ
What’s the difference between Arize AX and Phoenix?
Phoenix is the free, self-hostable open-source core: tracing, evals, datasets, experiments, and a prompt playground under Elastic License 2.0 with no feature gating. Arize AX is the managed and enterprise layer on top, adding online evaluations, custom alerts and dashboards, the Alyx agent copilot, multi-organization RBAC, SSO and audit logs, dedicated support, and an optional self-hosted enterprise deployment mode.
Is Arize Phoenix good enough to skip AX entirely?
Yes, for teams willing to self-host and instrument their own code, because Phoenix carries no feature gating versus the core platform. You pay for AX to get zero-setup managed dashboards, online evaluations and alerting against production traffic, and enterprise controls like RBAC, SSO, and audit logs. See Phoenix vs AX above for the full boundary.
Does the Dynatrace acquisition change anything for Arize customers?
Not yet. Dynatrace announced a definitive agreement to acquire Arize for $915 million in cash and stock on August 13, 2026, with the founders joining Dynatrace and both companies stating Arize continues operating independently. The deal was pending regulatory approval at announcement, so product and roadmap impact is unconfirmed; Enterprise buyers should raise it in procurement.
Is Arize AX worth it?
The Free ($0) and Pro ($50/month) tiers are low-risk, and self-hosted Phoenix removes cost from development entirely. Enterprise is custom-quoted; third-party trackers put typical deals around $60,000-100,000+ per year, a range Arize does not confirm. It earns that spend only if you need self-hosted deployment, multi-org RBAC, SSO and audit logs, or dedicated SLAs.
Can I use Arize tracing without getting locked in?
Yes. Both Phoenix and AX trace through OpenInference, a convention built on OpenTelemetry and designed to work with any OpenTelemetry-compatible backend. Instrumentation you write for Arize ports to another OTel-compatible observability tool without rewriting your spans.