Comet

Experiment tracking with bundled production drift monitoring, plus Opik, its Apache-2.0 open-source LLM evaluation product.

Best for: Teams that want experiment tracking with bundled drift monitoring under W&B pricing, and LLM teams via the open-source Opik

Editor’s note: We combine independent analysis, data collection, and hands-on testing to review data and AI tools. This review weighs pricing transparency, real-world adoption signals, development momentum, openness and exit costs, practitioner sentiment, and our editorial verdict.

Quick verdict: We recommend Comet for teams that want drift monitoring bundled with their experiment tracker at $19 per user per month, and Opik for anyone who wants a genuinely open LLM eval stack. The trade-offs: a closed-source core platform, Enterprise-only self-hosting, and cost as the top user complaint.

Two numbers frame this Comet ML review: Opik, Comet’s LLM eval product, holds 22.3k GitHub stars; the repo named for the company holds 87. Among MLOps platforms in our directory, Comet is two products under one brand that earn separate verdicts.

Key Takeaways 🔍

  • MLOps Pro is $19/user/month with 1,500 training hours and 500GB included, verified September 2026; the circulating $39 figure is stale
  • Opik is Apache-2.0, has 22.3k GitHub stars, and self-hosts free
  • Production monitoring runs real statistics: KS test, Population Stability Index, Jensen-Shannon divergence
  • The core platform is closed source; self-hosting it requires the custom-priced Enterprise tier
  • No bulk export API: leaving Comet means parallel logging, not a one-time migration
Comet homepage
Comet’s homepage. Source: Panoply.

Pros and Cons

Comet’s edge is drift monitoring bundled with tracking; its costs are a closed core and metered overages.

Pros

  • Drift monitoring (KS test, PSI, Jensen-Shannon divergence) bundled with experiment tracking
  • Opik is Apache-2.0, holds 22.3k GitHub stars, and self-hosts free
  • Pro at $19/user/month undercuts the $50-60 range cited for Weights & Biases
  • 40+ framework integrations that start from a two-line setup
  • Enterprise on-prem with SSO/SAML/LDAP, audit logging, and SOC 2

Cons

  • Core MLOps SDK is closed source
  • Self-hosting the core platform requires the custom-priced Enterprise tier
  • Overages stack on the seat fee: $1 per training hour past 1,500, $3 per extra 100GB
  • UI slows as experiment volume grows, per Capterra reviewers
  • No bulk export API for historical runs

How Much Does Comet Cost?

Comet pricing
Comet’s pricing plans. Source: Panoply.

Comet’s pricing page carries two separate tables, one for the MLOps Platform and one for Opik, both verified September 2026. The $39 per month Pro price that two secondary sources still print matches neither table, so treat it as stale.

MLOps Platform:

  • Free ($0): 1 team member, fair usage policy, 100GB storage, with experiment tracking, dataset management, Python visualizations, and the model registry included
  • Pro ($19/user/month): up to 10 team members, 1,500 training hours included ($1 per additional hour), 500GB storage ($3 per extra 100GB per month), email support
  • Enterprise (custom): unlimited members, hours, and storage, plus self-hosted or VPC deployment, SSO, service accounts, view-only users, production monitoring, and SLA-backed support

Opik:

  • Open Source ($0): the full backend self-hosted, with no commercial license required
  • Free Cloud ($0): up to 10 members, 25,000 spans per month, 60-day retention
  • Pro Cloud ($19/month): up to 50 members, 100,000 spans per month, customizable retention
  • Enterprise (custom): unlimited scale, flexible deployment, and a SOC 2, ISO 27001, HIPAA, and GDPR compliance set
ProductPlanPriceWhat’s includedOverages
MLOps PlatformFree$01 seat, 100GB, tracking, datasets, registryFair usage policy
MLOps PlatformPro$19/user/month10 members, 1,500 training hours, 500GB$1 per training hour, $3 per 100GB/month
MLOps PlatformEnterpriseCustomUnlimited usage, self-hosting, SSO, production monitoringNone listed
OpikOpen Source$0Full self-hosted backendNone
OpikFree Cloud$010 members, 25k spans/month, 60-day retentionNone listed
OpikPro Cloud$19/month50 members, 100k spans/monthNone listed
OpikEnterpriseCustomUnlimited, SOC 2, ISO 27001, HIPAA, GDPRNone listed

Neither free tier asks for a credit card, so both tracks can be trialed before anyone signs a purchase order.

Is Comet Good Value for Money?

  • $19/user undercuts Weights & Biases. W&B team pricing is cited across sources at roughly $50-60 per user per month, so a 10-seat team saves $310 to $410 a month on seats alone
  • The overages compound. A 10-seat Pro team pays $190 per month before a single extra training hour. Log 2,000 training hours against the 1,500 included and the meter adds $500; storing 1TB adds another $15 for the five extra 100GB blocks
  • Cost is a documented churn risk. One Capterra reviewer reported their business “thought about moving away from the product due to the high cost”
  • The differentiator costs extra. Production monitoring and self-hosting sit only in the custom Enterprise tier, so the $19 seat buys a tracker, not the drift layer; if monitoring drives the purchase, budget for a sales conversation from day one
  • Cheaper exists at both ends: ClearML’s Pro tier is cited around $15/user/month, and MLflow is free if you run the infrastructure yourself

Author’s Testing Notes 📝

Solo work belongs on the Free tier; the single seat is the real limit, not the 100GB. Teams should model their training hours before committing to Pro, because the $1 per hour meter past 1,500 is where budgets slip quietly. And LLM-only teams can skip this pricing table entirely: self-host Opik free and never open the MLOps track.

— Panoply reviewer

My Experience With Comet

Comet’s pitch is that instrumenting an existing training script takes two lines of code. The documented minimal path really is that short, and it is where I started.

The Two-Line Setup

The whole quickstart is import comet_ml, then comet_ml.login(), then exp = comet_ml.start(). From there, exp.log_parameters() and exp.log_metrics() push hyperparameters and metrics into the workspace, and the run shows up in the web UI with built-in visualizations attached.

That minimal path extends across the stack most teams already run. Comet integrates with 40+ frameworks and providers, and user reviews consistently describe PyTorch, TensorFlow, Keras, scikit-learn, XGBoost, LightGBM, and Hugging Face Transformers as working with minimal boilerplate. The two lines are the honest baseline, though Capterra reviewers note that getting past logging into the platform’s full feature set took real onboarding effort.

Comparing Runs and Building Panels

Once runs are logged, the workspace is built around selecting experiments and comparing them: parameters side by side and metric charts overlaid across runs. The comparison view is the product’s center of gravity, and the Free tier gives a solo user tracking, Python visualizations, and comparison with no feature held back.

Beyond the stock charts, Comet Panels let you build custom visualizations against your logged data. Pento.ai, in a vendor case study, uses custom panels for computer vision application development, where loss curves alone say little about model quality.

[Screenshot needed: Comet UI comparing multiple runs with metric charts and custom panels] Comparing logged runs is where Comet earns its seat fee; custom panels extend the stock charts. Source: Panoply

Where the UI Strains

Capterra reviewers report the interface slowing as logged experiment volume grows, and custom view management creating duplicate work. Several hit errors uploading larger files against size restrictions. A G2-derived review summary adds inconsistencies between what the web UI exposes and what the Python SDK can do, observed when automatically uploading and tagging models in the registry, plus recurring gripes about documentation quality and limited UI customization.

Test for this before you buy: log a realistic run history on the Free tier, including your largest artifacts, and check whether the comparison view still responds at that volume.

Author’s Testing Notes 📝

The finding I would price in before standardizing: Comet has no bulk export API for historical runs. If you ever switch trackers, the practical path is running parallel logging in the new tool until you trust it, not a one-time data migration. That is the lock-in, and it costs more the longer you wait to notice it.

— Panoply reviewer

From Experiments to Production: Registry and Drift Monitoring

What happens after training is where most trackers quietly stop. Comet keeps going: the model registry versions your trained models, and Model Production Monitoring watches what those models do once deployed.

The registry is included from the Free tier, so the lineage from logged run to registered model version survives without a second tool. One caveat carries over from the previous section: the G2-derived summary reports UI-versus-SDK inconsistencies precisely here, when uploading and tagging registry models automatically, so script that workflow against the SDK before you depend on it.

The monitoring compares training data distributions against live production distributions using the Kolmogorov-Smirnov test, Population Stability Index, and Jensen-Shannon divergence, with configurable alert thresholds. Those are the same drift tests a platform team would otherwise implement by hand.

Know where it sits, though: production monitoring is listed under the Enterprise tier on the pricing page, not Free or Pro. The budget version of Comet is a tracker; the monitored version is a sales conversation.

A sane MLOps stack has one development record, one deployment path, and one monitoring system, and Comet’s pitch is that its monitoring connects deployment behavior back to the development record inside the same product. Deployment orchestration stays your problem, and the connection only works if SDK instrumentation is consistent across teams.

Against the field: MLflow and base Weights & Biases do not bundle drift detection natively, while Arize built its business on model and generative-AI observability. Comet sits between them, a tracker that ships its own drift layer.

Opik: The Open-Source Half

Search GitHub for Comet and the repository named after the company is an 87-star issue tracker, now inactive, redirecting questions to the Opik repo or Slack. The core MLOps SDK is closed source. Comet’s entire open-source reputation lives one repo over, in comet-ml/opik: Apache-2.0, 22.3k stars, 1.8k forks.

Opik is Comet’s LLM observability and evaluation product, covering agent tracing, automated evals, and production monitoring for LLM applications. The part that matters for buyers: the full backend self-hosts without a commercial license. That is the server, the web app, tracing, and evaluation, not a crippled community edition.

It also has its own pricing track, separate from the MLOps Platform:

  • Open Source ($0): the full feature set, self-hosted
  • Free Cloud ($0): 10 members, 25,000 spans per month, 60-day retention
  • Pro Cloud ($19/month): 50 members, 100,000 spans per month

Note the unit on that Pro price: $19 per month, listed without a per-user multiplier. A 50-member team on Opik Pro Cloud pays what a single seat costs on the MLOps track.

Vendor case studies show Opik adopted on its own merits: Zencoder uses it to build and test fully agentic software pipelines, and Pattern runs data-driven LLM evaluation on it. Neither is a core-platform tracking customer.

Opik competes with Langfuse and Arize Phoenix in LLM observability, not with Weights & Biases in experiment tracking. Judged in that category, it is a top-tier open option regardless of what you think of the closed platform attached to it. If your workload is LLM evals, you can adopt Opik and never touch the rest of this review.

Adoption and Who Runs It

The comet_ml package pulls 79,535 downloads per week on PyPI and 346,851 per month, in line with the roughly 79,500 weekly figure in independent reporting. That is a real, currently active user base, not a logo wall.

Comet reports 150,000+ developers across roughly 10,000 teams, and third-party coverage repeats a claim of more than 500 million experiments tracked. Treat both as marketing figures rather than measured ones.

Comet was founded in 2017 by Gideon Mendels and Nimrod Shabtay, is based in New York, and has raised roughly $70 million, including a Series B led by OpenView Venture Partners with Scale Venture Partners and Trilogy Equity Partners participating. A $63 million figure still in circulation likely reflects an earlier funding snapshot rather than a contradiction.

The customer list, all vendor case studies: Uber manages AI experiments across its ML organization on Comet, Etsy took image search from prototype to production with it, AssemblyAI builds end-to-end speech recognition in PyTorch on it, and NatWest Group uses it to standardize its AI platform, the regulated-industry reference point. NTT DATA Group, Ancestry, and Netflix also appear on the roster.

How Does Comet Compare to Competitors?

A Comet evaluation runs into four live rivals, each winning a different niche, plus one that is no longer available at all:

  • Weights & Biases: the polish and collaboration leader, with real-time dashboards, shareable reports, and Sweeps for Bayesian hyperparameter tuning. It costs more, at the $50-60 per user per month range cited across sources, it is cloud-first with self-hosting gated to its own Enterprise tier, and Weave covers its LLM work
  • MLflow: free, open source, 27,000+ GitHub stars, and zero vendor lock-in, with a scope that now spans tracking, registry, tracing, evals, and a prompt registry. You run the infrastructure, accept a plainer UI, and get no native collaboration; it is the default inside the Databricks ecosystem
  • ClearML: broader end-to-end scope than a pure tracker, folding in pipeline orchestration, dataset versioning, and serving at a cited $15/user/month Pro tier; less LLM-eval depth than Opik
  • Neptune.ai: acquired by OpenAI (announced December 2025), with its hosted service shut down in March 2026; a discontinued rival, not a live alternative
  • Arize Phoenix: the observability-first open alternative, with open-source Phoenix as a local-first entry point, covering tracing, evaluations, drift, and human feedback loops for both predictive and generative models; strongest at the production layer where Comet’s monitoring is Enterprise-gated

Tracking plus bundled drift monitoring on a budget: Comet. Maximum dashboard polish and team collaboration: Weights & Biases. Zero vendor risk at the cost of running it yourself: MLflow. LLM evals alone: Opik or Langfuse, and skip the platform question. Regulated enterprises that want on-prem tracking and monitoring in one contract are the specific case where Comet’s Enterprise tier beats assembling the stack from parts.

How We Test MLOps Platforms

We combine independent analysis, data collection, and hands-on testing to review data and AI tools. For MLOps platforms that means collecting public signals (GitHub activity, PyPI downloads, Docker Hub pulls, Stack Overflow volume, G2 and Gartner peer reviews), hand-checking vendor pricing pages rather than repeating secondhand figures, and setting the tool up ourselves to run a real task end to end against its named rivals.

Review-site sentiment is deliberately small-weighted; a dozen reviews should color a verdict, not decide it. Where an area cannot be measured, we mark it N/A rather than scoring it zero. Signals are refreshed monthly and editorial verdicts quarterly. Sponsors and affiliates cannot change a score. Prices current as of September 2026.

Comet Review: Should You Track Your Experiments on Comet?

We recommend Comet for teams that want drift monitoring attached to the same product as their experiment tracker without paying Weights & Biases prices: $19 per user per month against a cited $50-60. Regulated enterprises needing on-prem deployment with SSO and audit logging have a real path here, priced at the Enterprise tier. And LLM teams should take Opik seriously on its own: Apache-2.0, 22.3k stars, free to self-host.

Skip it if you require open source end to end (that is MLflow, with the DevOps bill attached), if you want the most polished dashboards and collaboration in the category (that is Weights & Biases), or if the no-bulk-export switching cost is unacceptable for a tool you might outgrow. And budget for the Enterprise conversation if drift monitoring is why you came; it does not live in the $19 tier.

Next action: run the Free tier’s two-line setup on one real training script and see whether the comparison view earns a rollout. LLM-only teams should docker-compose Opik instead and start there.

FAQ

Is Comet ML free?

Both product lines have a real free tier. The MLOps Platform Free plan costs $0 and includes 100GB of storage, experiment tracking, datasets, and the model registry, but its single seat makes it a solo or evaluation tier. Opik is free two ways: fully self-hosted, or a Free Cloud tier with 10 members and 25,000 spans per month.

What is Opik, and is it the same product as Comet?

No. Opik is Comet’s open-source (Apache-2.0) LLM observability and evaluation product, covering agent tracing, automated evals, and LLM production monitoring. It has its own pricing track, a free self-hosted deployment, and a 22.3k-star community the closed-source core platform does not share. See Opik: The Open-Source Half above.

Can I self-host Comet?

Opik, yes, free: the full backend self-hosts with no commercial license. The core MLOps Platform, only at the custom-priced Enterprise tier, which requires a license token from Comet sales and a Linux server meeting the vendor’s hardware requirements, with SSO/SAML/LDAP available for compliance-heavy teams.

How does Comet pricing compare to Weights & Biases and Neptune?

Comet MLOps Pro is $19 per user per month plus overages ($1 per training hour past 1,500, $3 per extra 100GB), against a $50-60 per user range cited for Weights & Biases, which charges flat per seat. Neptune is no longer a comparison: OpenAI’s acquisition was announced in December 2025 and the hosted service shut down in March 2026.

Is Comet worth it compared to free MLflow?

It is a build-vs-buy call. MLflow costs nothing and locks you into nothing, but you run the infrastructure and accept a plainer UI with no native collaboration. Comet trades $19 per user per month for a managed UI, bundled drift monitoring, and, through Opik, an open LLM eval layer MLflow does not match at the same depth.

Spotted a wrong price or a missing integration? Send a correction. A human reads every one.

Similar tools

Other tools in the same category, with the same card and the same honest pricing.

Weights & Biases

Experiment tracking

The polish leader in experiment tracking: dashboards, Sweeps, reports, and Weave for LLM observability; Pro is $60/month flat for up to 10 seats.

Visit site

MLflow

Experiment tracking

Open-source standard for experiment tracking, model registry, and serving; 3.x adds OpenTelemetry tracing, LLM judges, and a Prompt Registry.

Visit site

LangSmith

Observability

LangChain's observability and evaluation platform: zero-config LangGraph tracing, CI-gated evals, prompt hub; closed-source SaaS.

Visit site

Langfuse

Observability

Open-source LLM engineering platform: OpenTelemetry-native tracing, LLM-as-judge evals, prompt management; MIT core with free self-hosting.

Visit site

LangChain

RAG framework

The most-adopted LLM framework, rebuilt on LangGraph in 1.0: durable agent execution, 1,000+ integrations, MIT-licensed; LangSmith is the paid layer.

Visit site

LlamaIndex

RAG framework

RAG-first data framework for LLM apps: 160+ connectors, purpose-built indexes, MIT-licensed; LlamaCloud adds managed parsing and indexing.

Visit site

Weaviate

Dedicated

Open-source vector database with native BM25 + vector hybrid search, per-tenant shard isolation, and multimodal vectorizer modules.

Visit site

Pinecone

Dedicated

Fully managed serverless vector database with zero-ops scaling, deep RAG framework integrations, and enterprise compliance.

Visit site