Editor’s note: We combine independent analysis, data collection, and hands-on testing to review data and AI tools. This Arthur AI review weighs pricing transparency, real-world adoption signals, development momentum, openness and exit costs, practitioner sentiment, and our editorial verdict.
Quick verdict: We recommend Arthur AI for teams that need to monitor tabular ML models and LLM agents in one place, especially in regulated industries that need bias checks, policies and an agent inventory. The trade-off: retention stops at 7 days on Free and 30 days on Premium, self-managed VPC, SSO and a BAA are Enterprise-only, and the MIT-licensed Engine has about 91 GitHub stars.
Arthur AI monitors drift on data pulled from S3, BigQuery, Snowflake and Databricks, traces and scores LLM agents over OpenTelemetry, and since December 2025 discovers the agents running across your cloud accounts. Arize AI starts at $50 per month on AX Pro, metered by spans rather than use cases.
In this review, I’ll take a closer look at Arthur’s pricing, features, and open-source engine, so you can see exactly whether one platform for models, agents and governance is worth the retention limits you accept to get it.
Key Takeaways
- Free costs $0 for 4 use cases with 7-day retention, Premium costs $60 per month for 100 use cases and 30-day retention, with unlimited seats on both
- One platform covers tabular ML drift, LLM and agent evals, and agent governance, where Langfuse and LangSmith stop at LLM tracing and evals
- The MIT-licensed Arthur Engine self-hosts tracing, evals and guardrails, but Arthur’s proprietary eval models and air-gapped installs stay paid
- The data plane runs in your own infrastructure, so setup starts with an engine deployment, and custom alerts need SQL
- About 91 GitHub stars and little public practitioner discussion make for a thin community next to Langfuse or Arize Phoenix

Pros and Cons
Pros
- Free plan monitors 4 use cases with 7-day retention and unlimited seats
- Premium at $60 per month covers 100 use cases, 30-day retention, custom dashboards and alert webhooks
- MIT-licensed Arthur Engine self-hosts OpenTelemetry tracing, continuous evals and guardrails
- One platform for tabular drift and bias monitoring, LLM and agent evals, and agent discovery and governance
- Data plane runs in your own cloud, with connectors for S3, GCS, BigQuery, Snowflake, Databricks and ODBC databases
Cons
- Retention stops at 7 days on Free and 30 days on Premium; unlimited needs Enterprise
- Self-managed VPC, on-prem, single-tenant SaaS, SSO and a BAA are all Enterprise-only
- Custom alert rules are hand-written SQL returning metric_timestamp and metric_value
- About 91 GitHub stars and 16 open issues on the Engine; the community is thin next to Langfuse or Phoenix
How Much Does Arthur AI Cost?

Arthur AI prices by monitored use cases rather than seats or spans, and every plan carries unlimited seats:
- Free ($0 per month): for small teams monitoring up to 4 use cases, with 7-day data retention on multi-tenant SaaS
- Premium ($60 per month): for teams shipping agents, with up to 100 use cases, 30-day retention, custom metrics and dashboards, and custom alerting with webhooks
- Enterprise (custom quote): for regulated organizations, with unlimited use cases and retention, single-tenant SaaS or self-managed VPC, BYOCloud and on-prem, SSO, SLAs, a BAA, a dedicated customer success manager, and custom limits on data, jobs, traces and evals
| Plan | Price | Use cases | Retention | Deployment | Seats |
|---|---|---|---|---|---|
| Free | $0 per month | Up to 4 | 7 days | Multi-tenant SaaS | Unlimited |
| Premium | $60 per month | Up to 100 | 30 days | Multi-tenant SaaS | Unlimited |
| Enterprise | Custom | Unlimited | Unlimited | Single-tenant SaaS, dedicated or managed VPC, self-managed VPC, BYOCloud or on-prem | Unlimited |
The open-source Arthur Engine is free under the MIT license and self-hosts tracing, evals and guardrails, without Arthur’s proprietary eval models, the air-gapped install or enterprise support.
Arize AI meters its hosted tiers by spans, with AX Pro at $50 per month for 50,000 spans and 30-day retention, and Langfuse bills by units at $29 per month for Core and $199 per month for Pro. Arthur’s bill moves with the number of models and agents you monitor; theirs moves with traffic.
Is Arthur AI Good Value for Money?
- Premium is cheap for 100 use cases with unlimited seats: a 10-person team pays $60 per month in total, where per-seat tools such as LangSmith grow with every hire
- 30-day retention on Premium rules out quarter-over-quarter drift comparisons: anything older than a month is gone unless you move to Enterprise
- A HIPAA BAA, SSO and on-prem all mean an unpublished Enterprise quote: regulated buyers never see a list price
- Free covers a real pilot: 4 use cases with 7 days of history, open to your whole team
Author’s Testing Notes
I recommend starting on Free with one model and one agent, so your whole team can see the drift charts and traces before anyone pays. Move to Premium at $60 per month when you pass 4 use cases or need webhooks on alerts, and request an Enterprise quote only when 30-day retention or a requirement such as SSO or a BAA forces it.
— Panoply team
My Experience With Arthur AI
Arthur gives you a Free plan rather than a timed trial, and the first thing it asks for after signup is an engine to run in your own infrastructure.
Signing Up and Deploying the Engine
I signed up on the Free plan and landed in a hierarchy of Organization, Workspace, Project and Application, where an application is a single model or agent and policies apply across the whole organization. Engines Management opens an install wizard with Docker for a local engine and Kubernetes, AWS, GCP and Azure targets for production; a one-line installer covers a Mac and a PowerShell script covers Windows.
I ran the Docker option locally, which requires Python 3.12 or newer. Because connectors and evaluations execute on that engine, raw inference data stays in my environment.
Tracing an LLM Agent
I installed arthur-observability-sdk with the OpenAI extra (Anthropic and LangChain extras also exist), created a task in the engine UI, copied the task ID from the URL and generated an API key under Settings. Initializing the SDK with the API key, the engine’s base URL and the task ID, then calling instrument_openai(), is the whole instrumentation step.
After a traced call and arthur.shutdown() to flush, the Traces view lists each call with prompt, response, latency, token counts and session ID, filterable by session. In our testing, scores stayed empty until we configured evaluators under Evaluate, so a fresh task shows traces with no quality signal attached.
Connecting a Tabular Model
For a classic model I created a connector to a Snowflake warehouse, which Arthur validates for access, and then a dataset naming the table. S3, Google Cloud Storage, BigQuery, Databricks SQL warehouses and ODBC sources follow the same two steps.
Arthur then runs scheduled queries against the table and activates drift and performance metrics on its own, with no metric definitions to write.
Setting Up an Alert
A custom alert rule is a SQL query that returns metric_timestamp and metric_value, with {{dateStart}}, {{dateEnd}} and {{interval}} placeholders Arthur fills in, plus a bound and a threshold. TEST ALERT RULE replays the query over 14 days of history before you create it, and on Premium a webhook configured in workspace settings attaches to the rule, with a 5-minute minimum interval.
The SQL gives you any metric you can express in a query, but an analyst who does not write SQL will stay on the default metrics.
Author’s Testing Notes
The engine deployment is the step that separates Arthur from the hosted tiers of Langfuse or Arize AX, where nothing runs in your own infrastructure. Budget for the Docker install before you judge the product, and keep one person who writes SQL close to the alert rules; the defaults cover drift and performance, and everything bespoke goes through a query.
— Panoply team
Model Monitoring for Traditional ML: Drift, Performance and Bias
Arthur started in 2018 as an ML model monitoring platform, and that layer still runs on the same data plane as the agent tooling.
Connectors cover BigQuery, AWS S3, Google Cloud Storage, Snowflake, Databricks SQL warehouses with a personal access token, and generic ODBC for PostgreSQL, MySQL, SQL Server and Oracle. Credentials are stored encrypted, and most cloud connectors can infer them from the engine’s runtime environment, so an engine running in your cloud account does not need a pasted key. The connectors execute on the ML Engine data plane you deploy, not in Arthur’s cloud, so warehouse queries run inside your own network.
Once a dataset points at a table, Arthur schedules queries against it and activates drift and performance metrics automatically. Premium adds custom metrics and dashboards, and the SQL alert rules from My Experience apply to any of them.
Bias and Explainability
Bias and explainability checks sit beside the drift charts, so a credit or hiring model that degrades for one subgroup shows up as a metric rather than a complaint.
Arthur’s own case study has Expel cutting its ML monitoring time by 50% after moving onto the platform; that is Arthur’s figure, not one we measured.
Evidently’s open-source library produces drift reports for free and does the job if tabular monitoring is all you need, but it brings less governance depth and no agent inventory. Fiddler AI is the closest peer, covering predictive ML, agents and guardrails for regulated enterprises. Arthur’s edge on this layer is that the same connectors, metrics and alerts also cover the agents in the next two sections.
LLM and Agent Evaluation With the Open-Source Arthur Engine
The Arthur Engine went open source under the MIT license on March 31 2025, when Arthur Shield, the old guardrails product, became the Engine, and it is the entry point for a GenAI team that never touches a tabular model.
Release tags 2.1.801 on August 19, 2.1.845 on September 1 and 2.1.980 on October 6 2026 show roughly monthly releases with frequent patch tags, and the October build added an agent discovery pipeline with Jamf Pro and Vertex AI connectors plus dependency security fixes.
The self-hosted Engine covers:
- Trace collection over OpenInference and OpenTelemetry for agent and LLM applications
- Continuous evaluations that score live traces rather than offline test sets alone
- Prompt management with versioning and experiments
- RAG monitoring for retrieval pipelines
- Guardrails for PII and sensitive data leakage, hallucination, prompt injection and toxic language
- Custom metrics and an API for anything the presets miss
What stays paid: Arthur’s proprietary enterprise evaluation models, air-gapped and on-prem deployment, custom enterprise-ready guardrails and metrics, and direct support. A self-hosted team gets the full pipeline and the bundled checks, and Arthur’s own eval models stay behind a contract.
About 91 GitHub stars and 16 open issues is a small footprint for an observability project, where Langfuse and Arize Phoenix have far larger communities answering questions and contributing integrations.
Arthur reports that the platform caught a GPT-5 regression for Upsolve before users were affected, and that it monitored over 1 billion tokens across real deployments in 2025; both figures are Arthur’s own.
Langfuse self-hosts free with a far larger community and a hosted Core tier at $29 per month, and Arize Phoenix is also open source. Arthur’s Engine earns the install when you want the same evals and guardrails to feed the governance layer below, and when a tabular model sits next to the agent.
Agent Discovery and Governance
Agent Discovery and Governance launched on December 17 2025 and reached Google Cloud Marketplace on January 7 2026. It scans compute environments for agents nobody registered, builds an inventory, and applies organization-level policies and attestations to each one.
Discovery sources include Vertex AI, AWS Bedrock, Microsoft Agent Foundry and, since Engine 2.1.980, Jamf Pro. Each configured source has a Scan Now button and rescans in the background every 24 hours, and agents found through GCP metadata register automatically.
Arthur built the layer out in stages through 2026:
- January: the Agent Development Toolkit, with model provider configuration, tasks and OpenTelemetry trace capture and inspection
- May: policy check date presets and automatic agent scanning and registration using GCP metadata
- June: organization-level governance views covering policy assignments, compliance status, unregistered agents, tools and models
- September: Governance Reader and Operator roles, Project Operator roles, Discovery Source configuration, the Jamf Pro connector and Microsoft Entra ID sign-in over OIDC
The buyer here is a CISO or model risk team rather than an ML engineer: the output is an inventory of every discovered agent with its policy assignments and compliance status, including the agents nobody registered.
LangSmith and Langfuse trace the agents you instrument; neither discovers the ones you did not, and that gap is the reason to pay Arthur rather than a tracing tool. The caveat: discovery depends on connectors to your platforms, so an agent running somewhere Arthur does not scan stays invisible until someone registers it by hand.
How Does Arthur AI Compare to Competitors?
Each Arthur AI alternative wins a narrower niche than Arthur’s three-layer platform:
- Arize AI pairs the open-source Phoenix with hosted AX: AX Free gives 25,000 spans and 15-day retention, AX Pro costs $50 per month for 50,000 spans and 30-day retention, and its community is far larger than Arthur’s. It wins for OpenTelemetry-native LLM tracing at scale, and it meters spans where Arthur meters use cases
- Langfuse is open source and self-hostable, with a free Hobby tier, Core at $29 per month and Pro at $199 per month billed by units. It wins for LLM engineering teams that want open source with a large community and have no tabular models to monitor
- LangSmith is LangChain’s own tracing and evals product with per-seat pricing. It wins for LangChain and LangGraph shops, and the per-seat bill is the reason a 20-person team compares it against Arthur’s unlimited seats
- Fiddler AI is the closest peer: predictive ML monitoring, agent observability and guardrails aimed at regulated enterprises, with a free plan and enterprise pricing by quote. The choice against Arthur comes down to Arthur’s MIT Engine and its agent discovery layer
- Datadog LLM Observability wins if you already run Datadog APM and want LLM spans next to your service traces; it is not built for tabular drift or agent governance
- Evidently AI gives you open-source drift and data quality reports for free plus a hosted cloud; it wins for teams that only need tabular monitoring and light governance
- WhyLabs is no longer a commercial vendor after its 2025 wind-down; the platform was open-sourced, and former customers are choosing among the tools above
MLflow and Weights & Biases live upstream in training and experiment tracking; they feed a model into Arthur rather than replace it.
Where Arthur wins is the single contract: a tabular fraud model, an LLM agent and an inventory of every other agent in the company under one data plane, one set of policies and unlimited seats. Of the rivals here, only Fiddler AI comes close to that breadth, and the breadth is the argument for accepting a smaller community and 30-day retention below Enterprise.
How We Test
We set up each tool ourselves and run a realistic workload through it, then weigh six areas: pricing transparency, real-world adoption signals, development momentum, openness and exit costs, practitioner sentiment, and our editorial verdict.
For Arthur AI, we started on the Free plan, deployed the Engine locally with Docker, instrumented an OpenAI-based agent with the observability SDK, connected a tabular dataset through a warehouse connector and wrote a SQL alert rule, then ran the same jobs against Arize, Langfuse and Fiddler. On exit costs, traces leave as OpenTelemetry data and the Engine is MIT licensed, so the tracing layer outlives a contract.
Practitioner sentiment on Arthur is thin, so this verdict leans on our own testing more than on user feedback. Areas that do not apply to a tool are left out rather than scored zero.
Prices current as of October 2026.
Arthur AI Review: Should You Monitor Your Models and Agents With Arthur?
Buy Arthur AI if you run both tabular models and LLM agents in a regulated business and want drift, bias, traces, evals and an agent inventory under one set of policies, or if you are a small team that wants 4 use cases monitored free with unlimited seats. The $60 per month Premium plan covers 100 use cases no matter how many people log in, and the traces leave as OpenTelemetry data if you ever move.
Skip it if your work is LLM tracing alone and you want the largest open-source community, where Langfuse or Arize Phoenix give you more contributors and hosted tiers from $29 or $50 per month, or if you need more than 30 days of retention without an Enterprise contract. Fiddler AI is the peer to price alongside Arthur if you want the same predictive-plus-agent coverage and the open-source Engine does not matter to you. Either way, budget for the engine deployment and for someone who writes SQL alert rules.
Start on Free at arthur.ai with one model and one agent, watch 7 days of drift and traces, and decide on Premium once you pass 4 use cases or need a webhook on your first alert.
FAQ
Is Arthur AI free?
Yes. The Free plan costs $0 and monitors up to 4 use cases with 7-day data retention and unlimited seats on multi-tenant SaaS. The Arthur Engine, which handles tracing, evals and guardrails, is also free to self-host under the MIT license, although Arthur’s proprietary eval models and air-gapped installs stay behind an Enterprise contract.
How much does Arthur AI cost?
Premium costs $60 per month for up to 100 use cases, 30-day retention, custom metrics and dashboards, and alert webhooks. Enterprise is a custom quote and adds unlimited use cases and retention, single-tenant or self-managed deployment, SSO, SLAs and a BAA. See How Much Does Arthur AI Cost? above for the plan table.
Is Arthur Engine open source?
Yes. Arthur Engine is MIT licensed on GitHub, created in March 2025, with releases tagged 2.1.801, 2.1.845 and 2.1.980 between August and October 2026. It covers OpenInference and OpenTelemetry tracing, continuous evals, prompt management, RAG monitoring and guardrails; the paid platform adds proprietary eval models, air-gapped deployment and support.
What happened to Arthur Shield?
Arthur Shield, the guardrails product, became the Arthur Engine when Arthur open-sourced it on March 31 2025, and its PII, hallucination, prompt injection and toxic language checks now ship inside the Engine. Arthur Bench, the 2023 open-source LLM evaluation tool, is no longer the headline product; the Engine’s continuous evals are where Arthur’s evaluation work now happens.
Can Arthur AI run on-premises?
Yes, on the Enterprise plan, which offers self-managed VPC, BYOCloud and on-prem deployment alongside single-tenant SaaS. On every plan the ML Engine data plane that runs connectors and evaluations is deployed in your own infrastructure, so raw inference data stays with you even on the multi-tenant Free and Premium tiers.