Vector Database Comparison 2026: Pinecone vs Weaviate vs Milvus vs Qdrant
What You'll Learn
- How a vector database stores embeddings and supports retrieval augmented generation.
- Where Pinecone, Weaviate, Qdrant, Milvus, Chroma, and pgvector differ in deployment and search.
- How metadata filters, hybrid retrieval, indexes, backups, and operations affect the real cost.
- How to compare databases fairly with a labeled corpus instead of relying on vendor benchmarks.
1. What a Vector Database Does
A vector database stores numerical embeddings alongside identifiers, metadata, and often the source text or a pointer to it. An application converts a user query into a vector, searches for nearby records, applies filters, and sends selected context to a generation model. This pattern supports semantic search, retrieval augmented generation, recommendations, duplicate detection, and multimodal retrieval.
The database is not the same as the embedding model. A model controls how content becomes a vector. The database controls storage, indexing, filtering, querying, replication, backups, access, and operations. A change in embedding model can require re-embedding the corpus and rebuilding or migrating the index.
Distance and index settings also matter. Exact nearest-neighbor search can provide a reference result, while approximate indexes trade some recall for speed and lower resource use. A production decision should record the embedding model, vector dimension, distance function, index type, metadata schema, top-k, reranker, and region.
The AI systems explainer discusses a separate application category. Here the focus is the storage and retrieval layer that supports those applications.
2. Why a Single “Best” Database Claim Fails
The legacy article declared Pinecone the lowest-latency managed option, Qdrant the most cost-efficient, Weaviate the best hybrid-search product, Milvus the choice for billion-scale data, Chroma suitable up to a fixed vector count, and pgvector suitable below a hard threshold. Those statements were not supported by a reproducible workload and should not be treated as current facts.
Latency depends on vector dimension, index type, filters, top-k, concurrency, region, cache state, replication, result size, and whether a reranker is used. Cost depends on memory, storage, reads, writes, egress, backups, support, staff, and migration work. Scale depends on the architecture and the team operating it.
A small product with an existing Postgres service may gain more from pgvector than from a new managed database. A regulated team may value a private deployment. A prototyping team may prefer a local library. A large distributed system may require a database with a deliberate sharding and operations plan. These are decision constraints, not leaderboard positions.
For adjacent technology context, see the AI search optimization guide. Search visibility and vector retrieval are related engineering concerns, but neither source proves that one database is right for every site.
3. Vector Database Comparison 2026 at a Glance
The following matrix uses the providers’ own documentation to describe architecture and features. “Open source” refers to the project or software availability. It does not mean the managed cloud service is free, and it does not remove the cost of hardware, monitoring, backups, or engineering.
| Database | Documented form | Useful starting point | Primary question |
|---|---|---|---|
| Pinecone | Managed serverless and other hosted plans | Teams that prefer usage-based hosted operations | Do read, write, storage, and egress costs fit the workload? |
| Weaviate | Open-source project and managed cloud plans | Teams needing hybrid search, vector compression, and cloud options | Which plan, region, index type, and service usage apply? |
| Qdrant | Open-source, managed cloud, hybrid cloud, and private cloud | Teams needing dense plus sparse or multi-stage retrieval | Is the required deployment and security tier affordable? |
| Milvus | Lite, standalone, distributed, and managed cloud paths | Teams with a reason to operate a distributed vector system | Who will manage cluster, storage, scaling, and recovery? |
| Chroma | Open-source local or self-hosted and cloud service | Applications needing a simple retrieval stack or prototype | Which cloud or self-hosted operating limits apply? |
| pgvector | Postgres extension | Teams that already store application data in Postgres | Can Postgres handle the index, filters, tenants, and query load? |
4. Pinecone: Managed Serverless Cost Model
Pinecone’s official cost documentation says serverless is usage-based. It identifies read units, write units, storage, and egress as the main serverless usage metrics. The documentation says the Starter plan has no monthly minimum, while the listed Builder, Standard, and Enterprise plans have minimum usage commitments of $20, $50, and $500 per month respectively. The plan details and current rates should be checked on the live pricing page.
Pinecone describes a query as using 1 read unit per 1 GB of namespace size with a minimum of 0.25 read units per query. It describes fetch as 1 read unit per 10 records with a minimum of 1 and List as 1 read unit per call with up to 100 records per call. These unit rules show why request volume and namespace size should be included in a budget.
Storage is also part of the calculation. Pinecone’s documentation describes dense-vector index size using record count, identifier size, metadata size, and vector dimensions multiplied by 4 bytes. Egress is metered for data returned by in-scope reads. A system that returns large metadata or vector values may therefore have a different cost profile from one that returns only identifiers and selected fields.
Pinecone documents dense and sparse hybrid search and metadata filters. Its hybrid-search guidance warns that dense and sparse scores are not automatically normalized and recommends explicit weighting or a document-schema approach. This is an important engineering detail because an uncalibrated hybrid query can favor the sparse component for reasons unrelated to relevance.
Read the Pinecone cost documentation and the Pinecone hybrid-search guide. Do not reuse the legacy 33 ms p99 claim because no workload, test harness, region, or independent result was provided.
5. Weaviate: Managed Cloud and Vector Services
Weaviate’s current pricing page lists an always-free plan with 1 cluster per user, up to 100,000 objects, 1 GB of memory, 10 GB of disk, 1 collection, and up to 3 tenants. It lists 2,000 embedding requests per day and a 1,000-request-per-month Query Agent allowance on that plan. The page lists Flex starting at $45 per month and Premium starting at $400 per month.
The same page describes hybrid search on the listed plans and shows that vector-dimension and storage rates vary by plan, region, cloud provider, index type, and compression method. It lists example hosted embedding prices of $0.025, $0.040, and $0.065 per 1M tokens for named models. Those are page examples, not a complete cost quote for every configuration.
Weaviate’s pricing page describes shared and dedicated deployment, replication, backups, multi-tenancy, RBAC, and enterprise features. The practical choice is not simply open-source versus cloud. It is whether the team needs managed upgrades, availability targets, support, security controls, compression, and a particular cloud region.
Weaviate can be a useful candidate for hybrid search and a managed application service. It should still be tested with the intended metadata filters and query distribution. A feature checkbox does not predict recall or latency for a particular corpus.
The official Weaviate pricing page should be used for current plan and rate checks. The commercial AI safety guide is a separate topic and is not evidence for vector-database pricing.
6. Qdrant: Hybrid and Multi-Stage Retrieval
Qdrant’s official documentation describes hybrid and multi-stage queries using the Query API and a prefetch parameter. Prefetch queries run first, and the main query is applied over their results. This supports retrieval pipelines that combine dense and sparse vectors or add later-stage scoring.
Qdrant documents result fusion through Reciprocal Rank Fusion and DBSF. It says the RRF constant k can be parameterized and notes that the feature is available as of v1.16.0. These controls are useful when separate representations of the same content need to be combined, but the weights and limits still need evaluation on real queries.
Qdrant’s pricing page lists a Free Tier with 0.5 vCPU, 1 GB RAM, and 4 GB disk for tests and prototypes. It describes Standard as usage-based with dedicated resources, backups, flexible scaling, and a 99.5% uptime SLA. Premium requires minimum spend and adds features such as SSO and private VPC links. The page also describes hybrid cloud and private cloud options for data-residency and regulated workloads.
Qdrant is a candidate for teams that want dense and sparse retrieval, deployment flexibility, or a path from local software to managed infrastructure. The article removes the legacy claim that a March 2026 funding round proves product cost efficiency or search quality.
Consult the official Qdrant hybrid-query documentation and Qdrant pricing page before making an architecture or budget decision.
| Deployment need | Plausible candidates | Documented reason | Operational question |
|---|---|---|---|
| Hosted with usage metering | Pinecone, Weaviate, Qdrant Cloud | Each documents managed plans and usage or resource-based billing | What are the minimums, storage, read, write, and egress charges? |
| Hybrid or private cloud | Qdrant, Weaviate, Milvus ecosystem | Official pages describe hybrid, private, or managed deployment paths | Who controls the network, keys, backups, and upgrades? |
| Single-machine service | Milvus Standalone, Chroma, pgvector, Qdrant self-hosted | Each has a documented local or single-service path | How will recovery and capacity be handled? |
| Distributed system | Milvus Distributed or a managed service | Milvus documents a Kubernetes-oriented distributed mode | Does the team need distributed operations or only a larger index? |
| Existing relational stack | pgvector | Vectors remain alongside application records in Postgres | Can the current database handle search load and maintenance? |
7. Milvus: Multiple Deployment Modes and Indexes
The official Milvus overview describes Milvus as an open-source vector database under the Apache 2.0 license and says it is available as software and as a cloud service. It identifies Milvus Lite as a Python library for prototyping and edge devices, Milvus Standalone as a single-machine server deployment, and Milvus Distributed as a Kubernetes-oriented architecture for very large scenarios.
Milvus documentation lists ANN search, filtering search, range search, hybrid search over multiple vector fields, full-text search based on BM25, reranking, fetch, and query. It also lists index types such as HNSW, IVF, FLAT, SCANN, and DiskANN, along with sparse vectors, JSON, arrays, and multi-tenancy choices.
These options bring flexibility and operational decisions. The team must choose storage, index parameters, hardware, replication, backups, monitoring, and upgrade procedures. Milvus publishes performance statements based on its own benchmark references, but those results do not establish a universal 2x or 5x advantage for every workload.
Milvus is a reasonable candidate when a team has a reason to operate a distributed vector system or needs its index and data-model choices. It may be unnecessary for a small application that already has a capable relational database. The correct choice depends on the operational boundary, not on a scale label alone.
See the official Milvus overview for deployment modes and search capabilities. Treat performance numbers in vendor pages as claims to reproduce against the target corpus.
8. Chroma: Simple Retrieval and Local Development
Chroma’s official documentation describes Chroma as open-source data infrastructure for AI. It lists document and metadata storage, dense and sparse search, metadata filtering, full-text and regex search, and multimodal retrieval. It also documents integrations with embedding models from OpenAI, Cohere, Hugging Face, and sentence-transformers.
The documentation describes a self-hosted or cloud database path. That makes Chroma a practical candidate for a developer who wants a direct retrieval workflow without immediately operating a distributed cluster. The application still needs to define data persistence, access control, backups, resource limits, and migration procedures.
Chroma should not be assigned a hard vector-count boundary from the legacy article. Performance depends on the data type, index, hardware, query mix, and service mode. A prototype result does not establish production reliability, and cloud billing should be checked separately from local software use.
Chroma can fit early retrieval development or an application whose search requirements are moderate and whose team values a simple SDK. If the product later requires strict availability, advanced tenancy, or large distributed operations, include migration cost in the initial design.
The official Chroma documentation provides the current feature and deployment overview. For broader AI infrastructure context, the AI coding cost analysis illustrates why operational effort belongs in a technology comparison.
9. pgvector: Vector Search Inside Postgres
The official pgvector repository describes pgvector as open-source vector similarity search for Postgres. It supports exact and approximate nearest-neighbor search, single-precision, half-precision, binary, and sparse vectors, multiple distance functions, ACID compliance, joins, and point-in-time recovery through Postgres.
Exact search provides a useful recall reference. Approximate indexes support HNSW and IVFFlat and trade some recall for speed. The README says HNSW generally offers a stronger speed and recall trade-off than IVFFlat but uses more memory and takes longer to build. The correct setting depends on the data and query load.
The pgvector project documents vector up to 2,000 dimensions, halfvec up to 4,000, bit up to 64,000, and sparsevec up to 1,000 non-zero elements. It also documents filtering with Postgres indexes, iterative index scans, partitioning for tenant isolation, and hybrid search with Postgres full-text search and result fusion.
Filtering deserves attention. The README warns that with approximate indexes, filtering is applied after the index scan. This can reduce the number of qualifying results for selective filters unless iterative scans, partial indexes, partitioning, or another strategy is used. This is a reason to test tenant and metadata queries rather than measuring only unfiltered nearest-neighbor speed.
Read the official pgvector repository. pgvector is often attractive when application data and embeddings already belong in Postgres, but the database team must size memory, maintenance, indexes, replicas, and backups.
| Cost area | Managed service example | Self-hosted or Postgres concern | Measurement |
|---|---|---|---|
| Storage | Vector dimensions, object storage, backup, and region rates | Disk, memory, replication, and backup capacity | Bytes per vector and total records |
| Queries | Read units, request units, egress, or resource usage | CPU, memory, cache, concurrency, and network | Queries per second and result size |
| Writes | Write units, imports, or service usage | Index build, compaction, WAL, and maintenance | Upsert volume and re-embedding rate |
| Operations | Plan minimums, support, availability, and private networking | Staff time, alerts, upgrades, and incident response | Monthly run cost and recovery time |
| Migration | Export, import, API, and model compatibility | Schema changes and index rebuilds | Time and compute for a full reindex |
10. Hybrid Search, Metadata, and Reranking
Dense search is useful for meaning and paraphrase. Sparse or lexical search is useful for exact terms, identifiers, product codes, names, and rare words. Hybrid search combines these signals, but each database exposes different representations and score-combination methods.
Pinecone documents dense and sparse vectors in one index, separate indexes, and a document-schema pattern. It warns that dense and sparse scores have different ranges and should be weighted or normalized. Qdrant documents prefetch with RRF or DBSF. Milvus documents multiple vector fields and BM25 full-text search. Chroma documents dense, sparse, full-text, and regex search. pgvector documents hybrid search with Postgres full-text search and RRF or a cross-encoder.
Metadata filters can decide whether a search result is usable. Tenant ID, language, publication state, region, date, permissions, and document type should be applied consistently. An approximate index can return fewer qualifying records when filters are applied after candidate generation. Record filter selectivity in the evaluation.
Reranking adds another stage. It can improve ordering but adds latency and cost. Measure the complete pipeline, not only the first vector query. A database with a slightly slower initial search may produce a lower end-to-end cost if its filters and reranking path reduce unnecessary generation calls.
For an Indian news or finance application, test named entities, ticker symbols, scheme names, dates, transliteration, and exact numbers. The search optimization article is relevant to information retrieval planning, but the vector database must still be measured directly.
11. How to Run a Fair Vector Database Test
Create a labeled query set from real application traffic or a carefully sampled development corpus. Mark relevant documents, expected filters, language, and answer-support passages. Keep the embedding model and vector dimension fixed while comparing databases. If the embedding model changes, run a separate experiment because it changes the retrieval input.
Use the same chunking, vector values, distance function, top-k, metadata filters, reranking model, and region where the comparison permits. Record warm and cold behavior separately. Measure recall against exact search where available, precision, latency percentiles, throughput, failure rate, storage, write time, backup time, and recovery behavior.
Do not report a p99 number without its dataset size, vector dimension, hardware or plan, index parameters, concurrency, filters, result count, and test date. Do not use a funding announcement as evidence of quality or cost. Do not use a vendor’s own “faster” statement as a result for an unrelated corpus.
Run at least one filtered and one hybrid test. For tenant systems, include tenants of different sizes and check whether one tenant affects another. For multilingual products, report results by language. For a Postgres design, compare exact and approximate search with the actual filter and transaction mix.
| Test stage | Keep constant | Measure | Record with the result |
|---|---|---|---|
| Data preparation | Documents, chunks, metadata, embedding model, and dimension | Record count, token count, and vector bytes | Corpus version and checksum |
| Retrieval | Distance, top-k, filters, reranker, and query set | Recall, precision, latency, and duplicate rate | Index type and parameters |
| Operations | Region, concurrency, warm state, and retry policy | Throughput, failure rate, recovery, and p99 | Plan, hardware, and test date |
| Cost | Same query, write, storage, egress, and backup assumptions | Monthly service and operating cost | Pricing page and rate date |
| Migration | Same model change and reindex procedure | Rebuild time, downtime, and rollback path | Export format and compatibility notes |
12. Practical Decision Checklist
Choose a starting database by answering six questions. Is the application already on Postgres? Does it need a managed service or private deployment? Will search combine dense, sparse, and lexical signals? How selective are the metadata filters? What are the recovery and compliance requirements? What is the total cost at the expected record and query volume?
Pinecone is a candidate for teams that prefer a hosted usage model and want documented serverless operations. Weaviate is a candidate when its managed plans, hybrid search, compression, and cloud features fit the service boundary. Qdrant is a candidate for dense and sparse retrieval with cloud, private, or self-managed paths.
Milvus is a candidate when the team needs its deployment modes, index set, or distributed architecture. Chroma is a candidate for simple local or cloud retrieval workflows. pgvector is a candidate when application records, filters, transactions, and vectors should remain in Postgres. None of these descriptions is a quality guarantee.
Keep a decision record with the model, dimension, database version, index, filters, query set, region, price date, cost assumptions, and evaluation result. Re-run the comparison when a provider changes a model, pricing plan, API, index behavior, or service region.
The AI jobs analysis is a separate article. For this database decision, the strongest evidence is a reproducible test on the application’s own documents and queries.
Frequently Asked Questions
SK Jabedul Haque
Building India's most trusted finance education platform — simplifying news, schemes and market trends so anyone can understand and invest confidently.
Read full bioNever miss an update
Get our clearest explainers on schemes, markets and money — read what matters, without the noise.
Explore more articles