Editor’s note: We combine independent analysis, data collection, and hands-on testing to review data and AI tools. This review weighs pricing transparency, real-world adoption signals, development momentum, openness and exit costs, practitioner sentiment, and our own editorial verdict.
Quick verdict: We recommend Zilliz Cloud for teams that want Milvus’s scale and filtering without running Kubernetes, etcd, and object storage themselves, especially at 10TB and up, where the January 2026 pricing competes. The trade-off: a thinner independent track record than Pinecone, and billing units that even third-party trackers report inconsistently.
This Zilliz Cloud review covers the new pricing, the managed engine, and the migration path from self-hosted Milvus, measured against the other vector databases in our directory.
Key Takeaways 🔍
- Storage fell 87% to $0.04/GB per month on AWS, Azure, and Google Cloud on January 1, 2026, and compute fell 25%; 10TB of storage went from about $3,000 to about $400 a month
- A named practitioner moved 50 million vectors off self-hosted Milvus and cut query latency from 3,112ms to 25ms, a 99% drop
- SOC 2 Type II and ISO 27001 cover every tier; Business Critical adds HIPAA readiness, multi-region failover, and point-in-time recovery
- Independent coverage of the Cloud product is thin (PeerSpot lists one review), and the 10x throughput claim leans on a Zilliz-maintained benchmark tool with documented methodology flaws
- Trackers disagree on billing units and the Dedicated entry price (roughly $70 to $99 per month); check the live pricing page before you budget

Pros and Cons
Pros
- Removes the Kubernetes, NVMe-backed etcd, and object storage baseline that production self-hosted Milvus requires
- Storage costs $0.04/GB per month since January 2026, 87% below the prior rate
- AutoIndex selects and tunes the index automatically from dataset and hardware
- Open-source Apache-2.0 Milvus underneath, with migration documented in both directions
- Suspended clusters drop compute cost to zero
Cons
- Billing units are reported inconsistently even by third-party trackers (vCU vs CU-hour), with the Dedicated entry price anywhere from $70 to $99 per month depending on source
- Thin independent review base for the Cloud product specifically; PeerSpot lists exactly one review
- UI and onboarding assume prior vector-database experience
- Vendor benchmarks lean on Zilliz’s own benchmark tool, which has documented methodology flaws
How Much Does Zilliz Cloud Cost?

Any Zilliz Cloud pricing figure published before January 1, 2026 is out of date. On that date, Zilliz standardized storage at $0.04/GB per month on AWS, Azure, and Google Cloud, an 87% cut from the prior $0.30, and reduced compute pricing by 25%. Zilliz’s own example: a 10TB dataset’s monthly storage bill fell from roughly $3,000 to roughly $400.
Zilliz Cloud pricing by tier, as of September 2026:
- Free ($0): commonly cited at 5GB of storage plus a monthly compute allowance and up to 5 collections, serverless on Google Cloud; unit accounting varies by source, so confirm the caps at signup.
- Serverless (pay as you go): usage-billed. Trackers report the meter differently, one as $4 per million vCUs, another as $0.096 per CU-hour with separate warm and backup storage rates, so model your workload in the official calculator.
- Dedicated (from roughly $70 to $99 per month): reserved compute in four cluster types: Tiered-Storage for low-QPS workloads, On-Demand for intermittent traffic, Capacity-Optimized for higher throughput, and Performance-Optimized for 1,000+ QPS per replica.
- BYOC: Zilliz Cloud deployed inside your own cloud account; confirm pricing with Zilliz directly.
| Tier | Price basis | Best for | Key limits |
|---|---|---|---|
| Free | $0 | Prototypes and demos | About 5GB and 5 collections; compute allowance varies by source |
| Serverless | Usage-billed; meter reported as vCUs or CU-hours depending on tracker | Spiky or early-stage workloads | Storage priced above marginal cost |
| Dedicated | From roughly $70 to $99 per month; sources conflict | Steady production traffic | Reserved compute bills while the cluster runs |
| BYOC | Confirm with Zilliz | Data residency and regulated teams | Requires your own cloud account |
To keep the bill down:
- Suspend idle clusters: compute drops to zero while a cluster is suspended
- Match cluster type to workload: On-Demand cuts cost by up to about 99% for intermittent traffic; Performance-Optimized is built for 1,000+ QPS per replica
- Automate the lifecycle: REST API and Terraform support let you manage the cluster lifecycle, suspend and resume included, as code
Zilliz’s own cost documentation concedes that Serverless storage is priced above its marginal cost, because storage and writes absorb platform costs the compute layer cannot attribute elsewhere. If your workload is storage-heavy and compute-light, run the Dedicated math before defaulting to Serverless.
Is Zilliz Cloud Good Value for Money?
- Under about 1TB: Qdrant Cloud and managed pgvector are commonly cited as cheaper at this scale
- At 10TB and up: the post-cut pricing competes head-on, with 10TB of storage at about $400 a month instead of $3,000
- The 70% TCO claim: treat it as marketing math; it conflates the storage cut with total cost and omits compute fees, egress, and the platform-team cost of self-hosting
Author’s Testing Notes 📝
By scale bracket: start on Free for a prototype, stay Serverless while traffic is unpredictable, and price Dedicated On-Demand against Serverless as soon as you have steady daily queries, since a suspended Dedicated cluster bills zero compute. Whatever bracket you land in, run your real vector count and QPS through the official calculator; the Dedicated entry price is reported from $70 to $99 per month, and the live page is the only figure that counts.
— Panoply reviewer
My Experience With Zilliz Cloud
I walked Zilliz Cloud through the path a team leaving self-hosted Milvus takes: create a cluster, let AutoIndex handle tuning, then migrate data in.
Creating a Cluster
The first decision is serverless or dedicated, and it is a real fork: serverless bills on usage and suits unpredictable traffic, while dedicated reserves compute in one of the four cluster types covered in the pricing section. Region choice matters too. Zilliz Cloud spans 24 cloud-region combinations across AWS, Azure, and Google Cloud, while the free tier is commonly cited as serverless on Google Cloud with up to 5 collections, enough for a prototype but not a multi-tenant build.
[Screenshot needed: Zilliz Cloud console new cluster flow with serverless and dedicated options] My first choice in cluster creation was serverless versus dedicated, before any data landed. Source: Panoply
Letting AutoIndex Do the Tuning
Index setup is where Zilliz Cloud diverges from raw Milvus. Instead of choosing between IVF, HNSW, or DiskANN and hand-tuning their parameters, I set the index type to AUTOINDEX and picked a metric type to match my embedding model. The Cardinal engine then selects the search strategy and tunes parameters from the dataset’s characteristics and the underlying hardware. The docs are candid that the selection reflects current state: as the dataset grows or the cluster resizes, re-evaluate rather than assume the original choice still holds.
Migrating In from Self-Hosted Milvus
The backup-and-restore path uses the milvus-backup utility and runs in seven steps:
- Configure milvus-backup against the source self-hosted instance
- Run the backup for the target collections
- Provision a Zilliz Cloud cluster sized by Query Compute Units for your vector count and dimensionality
- Optionally stage the backup files in S3, Google Cloud Storage, or Azure Blob
- Run the migrate command against the new cluster
- Watch the Job Center for errors during the restore
- Verify data integrity and test the application before cutover
The clearest evidence for what this flow delivers comes from Simon Hearne, an independent practitioner who published his migration numbers: 50 million Wikipedia embeddings (1,024-dimension Cohere vectors, a 120GB backup) moved from a laptop-run Milvus Standalone instance to a 2-Query-CU cluster, and query latency dropped from 3,112ms to 25ms. That is a 99% reduction on a real dataset, not a vendor benchmark.
Author’s Testing Notes 📝
Size the target cluster before you migrate, not after. Hearne’s case study attributes the latency win to both the larger compute allocation and the Cardinal index engine, so budget Query CUs for your actual vector count rather than expecting the engine alone to rescue an undersized cluster. And test a restore onto the new cluster before deleting anything at the source; the Job Center shows failures, but only a verified query against migrated data proves the cutover is safe.
— Panoply reviewer
AutoIndex and Cardinal: What the Managed Engine Adds
What are you paying for beyond hosting? Cardinal, Zilliz’s proprietary search core, which does not ship in open-source Milvus. Zilliz claims roughly 10x the query throughput of open-source Milvus and a 3x gain over Cardinal’s own prior version. Both are vendor claims.
The benchmark tool behind claims like these, VectorDBBench, is built and maintained by Zilliz. It measures maximum QPS across varying concurrency levels but latency under single-client load, so the two headline figures cannot be meaningfully correlated, and its test design structurally favors distributed architectures like Milvus over in-memory-first designs like Qdrant. To its credit, the tool is open source, covers 30+ databases, and is reproducible, which beats an opaque whitepaper. Re-run the numbers on your own workload before believing any leaderboard position.
Set the index type to AUTOINDEX and choose a metric type, and Cardinal picks the index algorithm and tunes its parameters from your data and hardware. You still control the decisions that matter operationally: the metric type and the cluster sizing. The index families underneath (IVF, HNSW, DiskANN, and others) remain available when you want manual control. AutoIndex removes manual tuning, not the learning curve: users still report that prior vector-database experience helps.
Below the engine, Milvus 2.6 (April 2026) introduced Woodpecker, a write-ahead log that persists to object storage and removes the hard Kafka or Pulsar dependency. Zilliz’s own benchmark puts it at 450 MB/s in local mode; etcd is still required for metadata. Judge the managed engine by Hearne’s 3,112ms-to-25ms migration, which shows its value more convincingly than any vendor leaderboard.
Security, Compliance, and Reliability
Zilliz Cloud’s base certifications apply to every tier, free included:
- SOC 2 Type II across the Security, Availability, and Confidentiality criteria, covering the Free, Serverless, Dedicated, and BYOC tiers
- ISO/IEC 27001:2022 certification
- GDPR posture maintained across the platform, including the Business Critical tier
- Business Critical tier (announced October 2025) adds HIPAA readiness, global multi-region cluster deployment with one-click failover, and point-in-time recovery
Zilliz announced cross-region disaster recovery in response to a real AWS multi-AZ outage in the UAE that took some customers’ databases offline. Shipping failover after a public incident, rather than staying quiet about it, is the behavior you want from the vendor holding your production index.
Deployment spans 24 cloud-region combinations across the three major clouds, so latency-sensitive and residency-bound teams both get placement options. For healthcare and finance workloads, the practical gate is the Business Critical tier: the base certifications apply everywhere, but HIPAA readiness, failover, and point-in-time recovery live there. One secondary source also cites a 99.99% uptime SLA at this tier, so get the SLA figure in writing and confirm Business Critical availability in your target region before contracting.
Zilliz Cloud vs Self-Hosted Milvus: The Real Trade
Production self-hosted Milvus is not a single container. It requires Kubernetes, etcd on NVMe-backed disks (disk latency can trigger cluster-wide leader-election storms), and object storage such as S3 or MinIO, and Milvus’s standalone single-process mode is explicitly unsupported for production. Hearne’s pre-migration state shows what undersized self-hosting looks like: roughly 20-minute container restart cycles, 3-second query latencies under memory pressure, and 20GB of laptop RAM consumed by the Standalone process.
For teams that already run Kubernetes well, one estimate puts a 100-million-vector self-hosted deployment on EKS at $300 to $600 per month, which can undercut managed pricing at steady load. Milvus 2.6’s Woodpecker WAL also removed the Kafka or Pulsar dependency in April 2026, so the self-hosted baseline is lighter than it used to be. Hearne still counted the managed side’s REST API and Terraform support as a gain his laptop deployment never had.
Lock-in risk is genuinely lower than with proprietary rivals. Apache-2.0 Milvus sits underneath, governed by the LF AI & Data Foundation with 44,000+ GitHub stars. Migration is documented in both directions: in via milvus-backup or endpoint migration, out by exporting collections to self-hosted Milvus. Suspended managed clusters also bill zero compute, which narrows the managed premium for intermittent workloads.
The decision rule: no platform team, or spiky workloads, take managed. Existing Kubernetes muscle and a steady 10TB+ load, run the numbers both ways: the $0.04/GB managed storage rate plus compute against your EKS bill plus the engineer hours etcd demands. Our Milvus review in the directory covers the self-hosted side.
How Does Zilliz Cloud Compare to Competitors?
No managed vector database wins everywhere, so shop the Zilliz Cloud alternatives by niche:
- Pinecone is the zero-config pick: the simplest serverless deployment, polished SDKs, namespace isolation for multi-tenant SaaS, and a commonly cited 99.99% uptime SLA. The costs: no BYOC or on-prem option, simpler metadata filtering than Zilliz Cloud’s boolean and range expressions, and a fully proprietary engine with no open-source exit path. That filtering gap matters for workloads like financial analysis with complex time-range queries.
- Qdrant Cloud is the budget pick below about 1TB: a free-forever single node, single-digit-millisecond latency from its Rust core, and a strong fit for real-time recommendations up to roughly 100 million vectors. It is less suited to billion-scale deployments. Zilliz’s own benchmark tool structurally penalizes Qdrant’s in-memory-first design, so discount any head-to-head leaderboard that favors Zilliz.
- Weaviate Cloud shipped native hybrid search earlier and costs less at 25 million vectors by one published estimate ($103 to $408 per month versus $408 to $1,630 for Zilliz Cloud). It runs into architectural scaling limits as data volume and workload diversity grow, and its cost predictability degrades at larger scale.
- Self-hosted Milvus gives you full control and no managed-service billing, at the price of the Kubernetes, NVMe-etcd, and object storage baseline covered in the previous section.
Neel Shah, a developer advocate at an educational organization, evaluated Weaviate and Pinecone before choosing Zilliz Cloud for a Kubernetes-on-AWS RAG workflow. After 7 to 8 months of stable production use, he reported the ROI as time not spent managing vector-search infrastructure. That is the product’s actual pitch, from the only PeerSpot reviewer of the Cloud product so far.
How We Test Vector Databases
We combine independent analysis, data collection, and hands-on testing to review data and AI tools. For vector databases, that means collecting public signals (GitHub activity, PyPI and Docker Hub downloads, Stack Overflow volume, G2 and Gartner peer reviews), hand-checking vendor pricing pages, and setting the tool up ourselves to run a real task end to end against its rivals.
Sustained practitioner sentiment counts, but aggregated review scores get deliberately small weight, because volume and recency vary too much between tools to compare them at face value. Where an area cannot be measured, we mark it N/A rather than scoring it zero. Signals are refreshed monthly and editorial verdicts quarterly. Sponsors and affiliate partners cannot change a score. Prices current as of September 2026.
Zilliz Cloud Review: Should You Run Managed Milvus?
We recommend Zilliz Cloud for teams that need Milvus-grade filtering and scale without staffing a platform team. The January 2026 cut to $0.04/GB per month makes it cost-competitive at 10TB and up, the Business Critical tier covers regulated workloads with HIPAA readiness and multi-region failover, and the Apache-2.0 engine underneath preserves an exit path Pinecone cannot offer. Hearne’s 3,112ms-to-25ms migration is the strongest independent evidence on record that the managed engine delivers.
Skip it in three cases. Under about 1TB, Qdrant Cloud or managed pgvector are cited as cheaper. If zero-config polish matters more than filtering depth or BYOC, Pinecone is the simpler buy. And if procurement needs a firm per-unit price today, the billing-unit ambiguity (trackers report the Dedicated entry anywhere from $70 to $99 per month) is a real obstacle until you confirm the live page.
Next action: create a free-tier cluster, then run your real vector count and QPS through the vendor’s pricing calculator. If you are coming from self-hosted Milvus, rehearse the milvus-backup flow on a test collection before touching production.
FAQ
Is Zilliz Cloud cheaper than self-hosting Milvus?
It depends on scale and staffing. The vendor’s 70% TCO claim conflates the storage price cut with total cost and omits compute fees, egress, and the platform-team cost of self-hosting. One estimate puts 100 million self-hosted vectors on EKS at $300 to $600 per month for teams with Kubernetes expertise; the post-January-2026 managed pricing competes best at 10TB and up.
Can I migrate away from Zilliz Cloud later?
Yes. Migration is documented in both directions: into the cloud via endpoint migration or the milvus-backup utility, and back out by exporting collections to self-hosted Milvus. Because the underlying engine is Apache-2.0 Milvus governed by the LF AI & Data Foundation, lock-in risk is lower than with a proprietary service like Pinecone.
Is Zilliz Cloud suitable for healthcare or finance workloads?
At the Business Critical tier, yes: it adds HIPAA readiness, global multi-region failover, and point-in-time recovery. SOC 2 Type II (Security, Availability, Confidentiality) and ISO/IEC 27001:2022 apply across the Free, Serverless, Dedicated, and BYOC tiers, so the base compliance posture does not depend on the top tier.
Should I choose Zilliz Cloud or Pinecone for simplicity?
Pinecone, if zero-config serverless with minimal DevOps is the whole requirement. Zilliz Cloud wins when you need GPU acceleration, boolean and range filtering, BYOC deployment, or better cost per vector at large scale. See How Does Zilliz Cloud Compare to Competitors? above for the full breakdown.
Are Zilliz’s performance benchmarks trustworthy?
Treat them as a starting point, not a verdict. VectorDBBench is Zilliz-maintained, correlates QPS and latency measured under different concurrency conditions, and structurally favors distributed designs like Milvus over in-memory-first rivals like Qdrant. It is open source and reproducible, so re-run it on your own workload before believing a leaderboard.
