Milvus

Open-source, billion-scale vector database under LF AI & Data; 3.0 indexes lakehouse data in place, with DiskANN and GPU CAGRA indexes.

Best for: Teams scaling past tens of millions of vectors, lakehouse-native stacks, and GPU-index workloads with platform engineering on staff

Editor’s note: We combine independent analysis, data collection, and hands-on testing to review data and AI tools. This review weighs pricing transparency, real-world adoption signals, development momentum, openness and exit costs, practitioner sentiment, and our editorial verdict.

Quick verdict: We recommend Milvus for teams scaling past tens of millions of vectors, for lakehouse stacks using 3.0’s External Collections, and for GPU or DiskANN index workloads. The trade-off: distributed mode means Kubernetes, etcd, object storage, and a write-ahead log before the first query, and under roughly 10 million vectors that stack is over-engineered for the job.

This Milvus review is current on 3.0, the July 2026 release most published advice misses.

Key Takeaways 🔍

  • Milvus 3.0 went GA July 29, 2026: External Collections index Parquet, Lance, and Iceberg files in place
  • Widest index menu in the category: HNSW, IVF variants, FLAT, SCANN, DiskANN, and GPU CAGRA
  • 46,271 GitHub stars as of September 28, 2026; Apache 2.0, LF AI & Data governed
  • Distributed operations are real work: documented rolling-upgrade failure modes, and one team reports outspending a managed service on ops within six months
  • Self-hosters accept a performance ceiling: Cardinal, Zilliz’s fastest search core, stays proprietary
Milvus homepage
Milvus’s homepage. Source: Panoply.

Pros and Cons

Every pro below is about scale, and every con is an operations cost:

Pros

  • Scales to billions of vectors on disaggregated storage and compute
  • Widest index selection in the category, including DiskANN and GPU CAGRA
  • 3.0 indexes Parquet, Lance, and Iceberg data in place, no ETL copy
  • Native BM25 hybrid search since version 2.6
  • Apache 2.0 under LF AI & Data governance, no field-of-use clauses

Cons

  • Distributed mode needs Kubernetes, etcd, object storage, and a WAL before the first query
  • Over-engineered for collections under roughly 10M vectors
  • Cardinal, Zilliz’s fastest search core, is withheld from open source
  • Rolling upgrades have documented edge-case failures in distributed clusters
  • A self-hosting operator describes the monitoring as spotty

The pricing section puts dollar and hour figures on those cons, and the upgrade risks get their own walkthrough below.

How Much Does Milvus Really Cost?

Milvus costs $0 in license fees under Apache 2.0. The real bill is infrastructure plus engineering time, and it varies by deployment tier:

  • Milvus Lite ($0): embedded in your Python process via pip, zero infrastructure, built for prototyping
  • Standalone (a VPS bill): one Docker container with persistent storage, the tier for small-to-mid production workloads
  • Distributed (real money): Kubernetes plus etcd for metadata, S3 or MinIO for objects, and the Woodpecker write-ahead log that replaced the Kafka/Pulsar dependency in 2.6

For distributed Milvus pricing, one independent 2026 cost model sized a 100-million-vector cluster on AWS EKS: an r6g.xlarge Spot query node, an r6g.large Spot index node, an m6g.large on-demand coordinator (kept off Spot because it must stay stable), the EKS control plane, and about 1.2TB of S3 storage at roughly $0.02 per GB-month. Treat the figures as third-party estimates, not quotes:

Setup at 100M vectorsMonthly estimate (third party)
Self-hosted Milvus on EKS, single replica~$370
Self-hosted Milvus on EKS, with HA second query node~$442
Zilliz Cloud Dedicated, equivalent capacity$600-1,200
Pinecone Pods, equivalent capacity~$5,760

High availability is the cheapest line on that bill: the second query node adds roughly $72 a month.

The line item no model captures is engineering time. One operator running roughly 50 million vectors on a single instance reports about 4 hours a month of maintenance: backups, occasional restarts, and monitoring alerts that sometimes fire wrong, plus one incident in eight months that took two hours to resolve. The counter-testimony from the same forum thread is a team that spent more on operations in six months than a managed service would have cost. The same cost model puts the self-hosting break-even near $500 a month of managed spend; below that, ops hours eat the savings.

Index choice is the biggest cost lever at scale. DiskANN, which splits the index between memory and disk, was the only option that cost model found economically viable under $1,000 a month at 100 million vectors, where full in-memory HNSW is not. RaBitQ 1-bit quantization, added in 2.6, compresses indexes to roughly 1/32 of their original size at about 95% recall, which cuts the RAM a billion-vector collection needs.

Is Milvus Good Value for Money?

  • Far cheaper at scale: ~$370-442 a month self-hosted covers capacity the same estimate prices near $5,760 on Pinecone and $600-1,200 on Zilliz Cloud Dedicated
  • A money-loser below break-even: under roughly $500 a month of managed spend, engineering time outweighs the savings
  • Free all the way down: Lite and Standalone cost nothing beyond the hardware they sit on
  • Budget hours, not just dollars: even a smooth 50M-vector deployment ran 4 hours a month of maintenance

Author’s Testing Notes 📝

Prototype on Milvus Lite, then run Standalone in Docker until you pass roughly 10 million vectors or need high availability. Go Distributed only with a platform engineer on staff who already owns Kubernetes; the etcd, object storage, and WAL stack is a job responsibility, not a side quest.

— Panoply reviewer

My Experience With Milvus

Milvus meets you at three levels of commitment, and the sane path runs through all three in order.

Starting with Milvus Lite

The whole install is a single pip install through the pymilvus client. Lite runs inside the Python process, persists to a local file, and needs no server, no Docker, and no cluster. I treat it as the honest on-ramp: the client API is the same one you keep at every tier, so the collection and search code from a notebook carries forward with only the connection URI changed when you outgrow it.

Standalone in Docker

Standalone is a single Docker container with persistent volumes, and it is the tier most teams should actually run. It gives you a real server, real backups, and none of the distributed dependency stack. I could browse collections, inspect indexes, and check segment state through the web console instead of writing client code for every look.

[Screenshot needed: Attu or Milvus WebUI collection view] I used the web console to inspect collections and index state on a Standalone deployment. Source: Panoply

Picking a Consistency Level

Milvus exposes four consistency levels, settable per collection or per query, and the choice is a real workload decision:

  • Strong: always reads the latest data, at a search-speed cost; use it for financial or compliance-sensitive lookups
  • Bounded (the default): tolerates brief staleness; right for most recommendation-style workloads
  • Session: a client immediately sees its own writes; the setting for upload-then-search flows, where a user adds a document and queries it in the next request
  • Eventually: fastest, no freshness guarantee; use it when raw query speed wins

Where Distributed Gets Real

A distributed cluster is Milvus plus three stateful services: etcd for metadata, S3 or MinIO for segment storage, and the Woodpecker WAL, which runs as its own service in 3.0. Before 2.6, that slot was a Kafka or Pulsar cluster, so the stack is one system lighter than older guides describe. Kubernetes coordinates all of it, which is why the tier starts at a platform team, not a weekend.

Plan maintenance windows around two known rolling-upgrade failure modes. Replica rebalancing can start moving segments inside the replica still serving traffic before its replacement begins loading, taxing that replica’s CPU, memory, and I/O mid-upgrade. A QueryNode shutdown can also hang with no timeout when object storage returns errors for segment files. Don’t trust a green post-upgrade readiness check on its own either; open issues describe checks reporting healthy when they are not.

One more trap: copying volume directories to a new server can restart clean while collections stay invisible, so migrate with backup tooling, not file copies.

Author’s Testing Notes 📝

Treat version upgrades as releases in their own right. Given the documented rebalancing and shutdown-hang cases, run every rolling upgrade against a staging replica first, and keep a tested restore point before touching the serving cluster.

— Panoply reviewer

Milvus 3.0: Lake-Native Changes the Deployment Math

Most Milvus advice you will find describes 2.x. Version 3.0 reached general availability on July 29, 2026, and its headline feature deletes a whole category of pipeline work.

External Collections build vector, full-text, JSON, and scalar indexes directly over data sitting in Parquet, Lance, Iceberg, or Vortex files in object storage. Nothing is copied into Milvus first. If your embeddings already live in an Iceberg table on S3, Milvus indexes them where they are, and the ETL job that used to shovel them into the database stops existing.

Under the hood, 3.0 ships a new storage engine called Loon. It is manifest-based, designed to cut read amplification for low-latency point access on object storage, and it uses Vortex, an Arrow-compatible columnar format, as the default. The LF AI & Data Foundation frames the shift plainly: index vectors where they already live instead of moving them to the index.

Adopting 2.x meant provisioning a cluster and holding the data twice, once in the lake and once in Milvus segments, with an ingestion pipeline keeping the copies in sync. With External Collections, the object store you already pay for is the storage layer, and Milvus’s job shrinks to compute and indexing. The boundary is the format list: Parquet, Lance, Iceberg, and Vortex are covered, and data anywhere else still needs an export step first.

For teams on a lakehouse stack, this collapses the strongest argument against running a separate vector database, because the duplication and sync pipeline was the argument. For everyone else, 3.0 changes little today; a standalone RAG app with 5 million vectors gains nothing from lake-native indexing it is not doing. Decide on 3.0 either way, because the 2.x-era articles you are comparing against are reviewing a smaller product.

No rival ships a wider index menu, and that menu is the practical reason Milvus wins at scale:

  • HNSW: the graph-index default for low-latency, in-memory search
  • IVF_FLAT and IVF_SQ8: cluster-based indexes; IVF_FLAT trades a little recall for speed, and IVF_SQ8 adds scalar quantization to cut memory
  • FLAT: brute force, exact results, small collections only
  • SCANN: Google’s quantization-heavy design for high throughput
  • DiskANN: hybrid disk and memory, built for 100-million-vector-class collections that can tolerate around 100ms latency
  • GPU CAGRA: the CUDA graph index added in 2.4

The CAGRA numbers come from Zilliz’s own engineering write-up, so weigh them accordingly: index builds 10 to 50x faster than CPU HNSW, and sub-millisecond p50 query latency when the whole index fits in GPU VRAM. That VRAM condition is the catch, and the GPU path is CUDA-only, with no Huawei Ascend support.

An independent 2026 Salt Technologies AI benchmark at 1 million vectors and 1,536 dimensions measured Milvus with GPU at 6ms p50 and 12-18ms p99, against 15-25ms CPU-only. That trails self-hosted Qdrant’s 4ms p50 but beats Pinecone Serverless (20-30ms p50, with cold-query p99 at 40-80ms) and Weaviate Cloud (50-70ms p50). Benchmarks shift with hardware and dataset, so read them as ordering, not gospel.

Hybrid search went native in 2.6 with built-in BM25, so keyword and vector retrieval run in one system instead of a Milvus cluster beside an Elasticsearch cluster. The one published throughput figure is a single comparison author’s own benchmark: 400% higher hybrid-search throughput than Elasticsearch on equivalent hardware. Treat it as one data point until your own workload confirms it. Pair native BM25 with the RaBitQ quantization covered in the pricing section and one cluster covers both dense and keyword retrieval at billion scale.

Milvus vs Zilliz Cloud: What the Open Source Version Holds Back

Zilliz, the company that created Milvus, sells Zilliz Cloud, and that managed service runs Cardinal, a proprietary search core that does not ship in open-source Milvus. Self-hosting means accepting a performance ceiling on the identical API surface.

What the money buys is mostly subtraction: no etcd, no MinIO, no WAL service, no Kubernetes, no rolling-upgrade maintenance windows. Backups, replication, and security hardening, which self-hosters own outright, move to Zilliz as well. Pricing, per independent price trackers:

  • Serverless: roughly $0.096 per compute-unit-hour, so one CU running all month (720 CU-hours) comes to about $69
  • Dedicated: listed as starting somewhere between $70 and $99 a month
  • Free tier: 100 compute-unit-hours and 5GB of storage, enough to try the managed API before committing

Zilliz also announced a storage price of $0.04 per GB-month from January 2026, an 87% cut from $0.30, alongside a 25% compute price cut. Confirm current rates on the live Zilliz Cloud pricing page before budgeting.

The decision rule from the pricing section carries over: below about $500 a month of managed spend, Zilliz Cloud is usually the better buy; above it, self-hosting starts paying for its own ops. Zilliz Cloud’s own plans and features are a separate evaluation; this review covers the open-source database.

Who Actually Runs Milvus

The GitHub repository stands at 46,271 stars, 4,275 forks, and 1,422 open issues as of September 28, 2026; older articles still citing 30,000 or 44,000 stars are dating themselves.

Shopee runs Milvus for instant video recall in its recommendation systems, leaning on multi-replica mode for low-latency, high-availability retrieval. eBay uses it to power ad recommendations, matching users to sponsored listings by semantic similarity rather than keywords alone.

Beyond those two, the logo list (Salesforce, PayPal, NVIDIA, Airbnb, IBM, AT&T, Roblox, Bosch, and others), spanning search, recommendation, agentic-pipeline, and multimodal workloads, traces back largely to vendor material, so read it as reported users rather than audited deployments.

The tempering signal: PeerSpot practitioners rate Milvus 7.4 out of 10, ranking it #13 among vector databases and #10 among open source databases on that platform. They praise stability and scalability once it is running and name a steep learning curve, especially for distributed setups, as the recurring complaint. Good, not adored, and the gap is operations, not search quality.

How Does Milvus Compare to Competitors?

Cross-shopping Milvus alternatives comes down to one number: how many vectors you will hold in two years.

  • Qdrant: the lowest published latency of any purpose-built vector database (4ms p50 self-hosted in the independent Salt Technologies benchmark) and a single binary with no Kubernetes requirement. Self-hosted cost is the lowest of the group too: roughly $20-40 a month of VPS handles 10 million vectors with binary quantization. Its ecosystem at billion scale is thinner than Milvus’s, and it is our default suggestion for most RAG teams below that scale.
  • Pinecone: zero-ops serverless and the fastest path from prototype to production. The cost model above estimates ~$5,760 a month for capacity a $370-442 self-hosted Milvus cluster covers, and cold queries spike to 40-80ms p99.
  • Weaviate: the most mature turnkey hybrid search, built-in vectorizer modules for OpenAI and Cohere, and native MCP server support since v1.37. Dimension-based billing multiplies with replication, which makes binary quantization near-mandatory above 1 million vectors, and Weaviate Cloud ran 50-70ms p50 in the same independent benchmark.
  • pgvector: the right answer up to roughly 10-50 million vectors if you already run Postgres, with joins and transactions for free. There is no DiskANN and no GPU path beyond that scale.
  • Elasticsearch: the one-stop shop when you need logs, keyword search, and vectors in one system. A benchmark cited on Zilliz’s own comparison page puts Elasticsearch exact kNN at 1,600ms against Milvus’s 7ms p99; it is vendor-published, and exact kNN is brute force, so read it as a gap in approach rather than a like-for-like race.

Our rule: above 100 million vectors, Milvus is the only option in this group built for the job. Below 10 million it is usually the wrong first choice, and Qdrant, pgvector, or Chroma will get you there with far less machinery.

How We Test Vector Databases

Our reviews combine independent analysis, data collection, and hands-on testing. For vector databases, that means setting the tool up ourselves, running a real ingestion-and-search task end to end, and comparing the workflow and the results directly against rivals in the same category.

We track public signals (GitHub, PyPI, Docker Hub, Stack Overflow, G2, and Gartner peer reviews) and date every price we quote. Sustained user sentiment is weighed, deliberately lightly. Where an area cannot be measured we mark it N/A rather than counting it against the tool. Signals refresh monthly, editorial verdicts quarterly, and neither sponsors nor affiliate relationships can change a verdict. Prices current as of September 2026.

Milvus Review: Should You Build Your Vector Infrastructure on Milvus?

Among the vector databases in our directory, Milvus is the scale pick: disaggregated storage and compute built for billion-vector collections, governed by the LF AI & Data Foundation rather than a single vendor. We recommend it for teams scaling past tens of millions of vectors, for lakehouse stacks that can put 3.0’s External Collections to work, for GPU-index workloads, and for organizations that want open source with no license risk on the exit path.

Skip it in three cases. Under roughly 10 million vectors, Qdrant, pgvector, or Chroma deliver the same outcome with a fraction of the machinery. Without platform engineering on staff, buy Zilliz Cloud or another managed service instead of learning etcd on production traffic. And if single-digit-millisecond latency on a small deployment is the whole game, Qdrant’s benchmark numbers win it.

The next action costs nothing: pip install Milvus Lite in a notebook this afternoon. Promote it to Standalone in Docker when the prototype graduates, and take on Distributed only when the vector count, and the staffing, both say so.

FAQ

Is Milvus really free?

The software is, under Apache 2.0, with no seat limits or field-of-use clauses. The bill is infrastructure and engineering time: a third-party model priced a 100-million-vector AWS EKS cluster at $370-442 a month before staff hours. Milvus Lite and Standalone run on hardware you already have. Below roughly $500 a month of equivalent managed spend, the ops hours usually cost more than they save.

Do I need Kubernetes to run Milvus?

Only for the Distributed tier. Milvus Lite runs embedded in a Python process with zero infrastructure, and Standalone runs as a single Docker container that handles small-to-mid production workloads. Kubernetes, etcd, object storage, and the Woodpecker write-ahead log enter the picture only when you need horizontal scale or high availability.

What is the difference between Milvus and Zilliz Cloud?

Zilliz Cloud is the managed service from Milvus’s commercial steward, built on the same API but running Cardinal, a proprietary search core open-source Milvus does not get. It removes the etcd, MinIO, and WAL operational burden for usage-based fees. Below about $500 a month of managed spend, managed usually wins; see the Zilliz Cloud section above.

Is Milvus overkill for small projects?

Usually, yes. Under roughly 10 million vectors, the multi-service architecture adds operations without adding performance you can feel. Milvus Lite and Standalone soften this, but Chroma, Qdrant, or pgvector are simpler starting points at that scale, and Qdrant self-hosts 10 million vectors on a $20-40 VPS.

How does Milvus handle consistency?

Through four tunable levels, set per collection or per query: Strong always reads the latest data, Bounded (the default) tolerates brief staleness, Session guarantees a client sees its own writes, and Eventually maximizes speed. See Picking a Consistency Level above for which workload fits each.

Spotted a wrong price or a missing integration? Send a correction. A human reads every one.

Similar tools

Other tools in the same category, with the same card and the same honest pricing.

ClearML

ML platform

End-to-end open-source MLOps: experiment tracking, GPU orchestration with fractional GPUs, dataset versioning, and pipelines at $15/user.

Visit site

Zilliz Cloud

Dedicated

Fully managed Milvus from its commercial steward, with AutoIndex tuning, the Cardinal engine, and post-2026 storage pricing at $0.04/GB/month.

Visit site

Qdrant

Dedicated

Open-source Rust vector database with in-graph payload filtering, three quantization families, and a free forever cloud tier.

Visit site

Haystack

RAG framework

deepset's open-source RAG framework: typed component pipelines, YAML serialization, and enterprise connectors under Apache-2.0.

Visit site

Braintrust

Evals

Eval-first platform for AI products: datasets, experiments, CI-gated scoring, and a hybrid VPC data plane; bills scores, not traces.

Visit site

Comet

Experiment tracking

Experiment tracking with bundled production drift monitoring, plus Opik, its Apache-2.0 open-source LLM evaluation product.

Visit site

Weights & Biases

Experiment tracking

The polish leader in experiment tracking: dashboards, Sweeps, reports, and Weave for LLM observability; Pro is $60/month flat for up to 10 seats.

Visit site

MLflow

Experiment tracking

Open-source standard for experiment tracking, model registry, and serving; 3.x adds OpenTelemetry tracing, LLM judges, and a Prompt Registry.

Visit site

LangSmith

Observability

LangChain's observability and evaluation platform: zero-config LangGraph tracing, CI-gated evals, prompt hub; closed-source SaaS.

Visit site