The Context Window

Vector Database Selection for RAG Pipelines

Match the workload to the database, not the benchmark, to build RAG systems that actually work.

Senior Writer · · 9 min read
Cover illustration for “Vector Database Selection for RAG Pipelines”
Context Delivery · September 25, 2026 · 9 min read · 1,966 words

Choosing a vector database for a RAG pipeline gets treated like picking a phone: check the review sites, find the top spot, buy it. That instinct fails here. The right database depends on how many vectors you're storing, how complex your filters get, whether you need hybrid search, and where the data legally has to sit. There's no single winner in 2026. The right pick depends entirely on the shape of the workload sitting in front of it, and most teams pick wrong because they benchmark the database instead of the workload.

What the RAG pipeline asks of a vector database

Before a single query hits the database, most of the important decisions have already been made upstream. Chunking strategy, embedding model, and dimension count all get locked in during document preparation, and the database has to live with whatever those choices produce.

At query time, the sequence is mechanical: embed the incoming query, run a similarity search to pull the nearest chunks, apply whatever filters and ranking logic the application specifies, then hand the top-k results to the LLM as context. That's the whole job. The database doesn't grade the model's homework. It has no way to know whether the LLM actually used the chunks it received, used them correctly, or ignored them and hallucinated anyway. That's a separate evaluation problem, handled by different tooling.

Treating a vector database benchmark, recall@K, queries per second, p99 latency, as a stand-in for RAG quality is a mistake most teams make once and regret. Chunking strategy, embedding model choice, query rewriting, reranking, and hardware allocation all interact with each other, so a database that tops a synthetic benchmark can still sit inside a mediocre RAG system. A database that wins in isolation proves it's good at its narrow job. It says nothing about the pipeline built around it, and treating the two as the same question is how teams end up over-engineering the wrong layer.

The five decision dimensions that separate good fits from bad ones

Four of these carry real weight in practice, though the literature likes to frame a fifth around latency targets on its own. Below roughly 50 to 100 million vectors, pgvector and several other options work fine for most RAG workloads, and the gaps that appear at that scale involve cost and convenience, not raw capability. Past that line, index rebuild time, memory headroom, and distributed architecture start making the decision for you whether you've planned for it or not. HNSW indexes are not free: a 5-million-chunk corpus can eat 8 to 15 GB of RAM just holding the index. A team planning for real growth needs to account for that memory curve now, not wait until it appears in an incident at 2 a.m.

Filtering complexity is the second dimension, and it's the one most teams underestimate until it bites them. Filtered vector search, narrowing by tenant ID, date, permission level, category, is structurally harder than plain nearest-neighbor search, because the database has to reconcile two different query strategies at once. Some databases pre-filter, narrowing the candidate set before running the vector search. Others post-filter, running the vector search first and throwing away results after the fact. At low selectivity, where a filter only matches a thin slice of the corpus, that architectural difference can quietly wreck recall without anyone noticing until a customer complains. Cardinality estimation for filtered vector queries, essentially guessing how selective a filter will be before running it, remains an open engineering problem across the industry (arXiv 2512.09695), and it directly shapes how good a database's query plan turns out to be.

Third: hybrid search, and this one is not optional anymore. Dense vector search and BM25 keyword search fail in different places. Dense embeddings miss exact-match terms, product identifiers, and legal citations, strings that need to match character for character. BM25 misses semantic synonymy, so it won't connect "car" and "automobile" unless someone tells it to. Reciprocal Rank Fusion combines the two without the score-incompatibility problems that break naive weighted averaging, since BM25 scores and cosine similarity scores don't live on comparable scales. Stacking a cross-encoder reranker on top, Cohere Rerank 3, Voyage rerank-2, or a self-hosted bge-reranker, adds 50 to 150 milliseconds of latency but usually buys a 10 to 15 point jump in precision@3. Whether that trade is worth it depends on the latency budget of the application, not on which reranker is fashionable this quarter. Eight of the ten most-used vector databases now ship hybrid search by default, so this stopped being a capability question a while ago and became a tuning question instead.

Deployment model is the fourth axis, and it's frequently decided by legal before engineering ever gets a vote. Managed and serverless options remove infrastructure work and get a team running fast, but the cost structure grows with query volume and storage, and data sovereignty guarantees may not exist. Self-hosting hands back full control and usually lowers unit cost at scale, at the price of needing in-house expertise for provisioning, upgrades, backups, and monitoring. Healthcare and fintech teams frequently can't touch a purely managed cloud service no matter how good the technology is, because data residency rules eliminate the option before performance even enters the room.

pgvector: the right default for teams already on PostgreSQL

pgvector is a PostgreSQL extension that adds vector storage plus HNSW and IVF indexing directly to a Postgres instance a team is probably already running. No new service to deploy, patch, or monitor. That alone should settle the question for most teams before a single benchmark gets run.

It holds up for RAG systems up to roughly 50 to 100 million vectors, and past that ceiling, the architecture begins to show meaningful constraints. Within that range, the performance case is stronger than most people expect: pgvector paired with pgvectorscale hit 28 times lower p95 latency and 16 times higher query throughput than Pinecone's storage-optimized index, at 99% recall on 50 million 1536-dimensional embeddings, self-hosted. The bigger structural win, though, is that vectors sit next to the relational data. Joins, access controls, and transactional consistency come along for free, with no separate metadata store to keep in sync with the vector index. For a team already running Postgres, pgvector removes an entire class of synchronization bugs before those bugs get a chance to exist. Reaching for a dedicated vector database before hitting that 50 to 100 million ceiling is, in most cases, solving a problem the team doesn't have yet.

Pinecone: managed convenience with a cost curve that needs watching

Pinecone is fully managed and closed-source, reachable only through its cloud API. There's no general self-hosted version of the core product, though Pinecone Nexus, the enterprise tier, can run through a self-hosted Kubernetes installer, and a bring-your-own-cloud option lets the data plane run inside a customer's own cloud account. Neither of those is a real self-hosted deployment in the way Qdrant or Milvus offer it.

The default shape in 2026 is serverless: pay for storage, reads, and writes, with no idle charges sitting on the bill when the index isn't being queried. Published pricing runs $0.33 per GB per month for storage, $8.25 for reads, and $2.00 for writes. That's convenient for teams that don't want to run infrastructure, but the convenience has a price attached that scales with usage, not with value delivered. A team with heavy query volume needs to model that curve before signing up, not after the first invoice lands and someone asks why the bill tripled.

Qdrant: the performance and cost efficiency case for self-hosting

Qdrant is open-source and written in Rust, available either self-hosted or through Qdrant Cloud as a managed option. In standardized benchmarks among purpose-built vector databases, it posted 4 ms p50 latency at a large scale of vectors and 1536 dimensions, near the front of the pack on raw speed.

The cost story is where Qdrant separates itself from the managed alternatives. A self-hosted deployment on a single VPS runs somewhere around $50 to $100 a month for a large number of vectors, with solid performance at that price point. For mid-scale workloads, Pinecone's per-query pricing simply can't compete with that math. Qdrant's payload filtering narrows results by metadata condition during the search itself, and payload indexes accelerate those filtered queries rather than treating them as an afterthought bolted onto vector search after the fact. That makes Qdrant the natural fit for filter-heavy RAG: systems with per-tenant isolation, permission boundaries, date-range constraints, or document-type segmentation, where the filter is doing as much work as the vector search itself.

Weaviate: when hybrid search and extensibility are the primary requirements

Weaviate is open-source, available for self-hosting or as a managed cloud service. Its defining feature is hybrid search built into the core query layer: vector search and BM25 keyword search run together, with tunable fusion between the two, rather than hybrid being an add-on feature stapled onto a vector-only engine as an afterthought.

Benchmarked latency runs 30 to 70 ms overall, and turning on hybrid search adds another 10 to 20 ms on top of vector-only queries, a reasonable trade for teams that need both exact-match and semantic recall in the same request. Weaviate's architecture is also modular: embedding models, vectorizers, and rerankers can be swapped in without rebuilding the surrounding application. Teams that expect their retrieval stack to keep changing, new embedding models, new rerankers, over the next year or two get real value out of that flexibility, more than teams chasing the lowest latency number on a spec sheet.

Milvus and Zilliz Cloud: when the dataset outgrows everything else

Milvus is the most widely adopted open-source vector database as of 2026, with a large and active GitHub following and a cloud-native, Kubernetes-compatible design. Zilliz Cloud is its managed counterpart, run by the company behind the open-source project.

Both are built for billion-scale indexing, with GPU-accelerated index construction and tiered storage. That architecture exists specifically because flat single-node designs stop working once a corpus crosses into that range. Community consensus treats Milvus as overkill for anything under 50 million vectors, and that consensus is right: past that line it becomes the obvious choice, precisely because distributed scale and tiered storage stop being nice-to-haves and start being requirements nobody can skip. Zilliz Cloud's Cardinal engine gets cited in benchmark discussions as a real step up over open-source HNSW, and it addresses the most common complaint leveled at self-hosted Milvus: the operational complexity of running a distributed system at that scale without a managed layer underneath it holding things together.

Vespa, Chroma, and other candidates worth knowing

Vespa combines vector search, structured search, and machine-learned ranking inside a single system. It's the choice for teams where retrieval is only the first stage of a more elaborate ranking pipeline, supporting custom ranking models, learning-to-rank, and cascading retrieval stages. That's substantial overkill for a standard RAG setup. But for search products where ranking quality is the actual competitive edge, Vespa's depth is the point, not a cost to tolerate.

Chroma sits at the other end of that spectrum. It's built for prototyping, local development, and MVPs, and it now ships with an object-storage backend plus collection forking, enough to support lightweight production use. It's not a serious candidate for high-throughput or large-scale production RAG, and pretending otherwise just delays the eventual migration. Its value is speed of setup and how little friction it puts between a developer and a working prototype.

Teams already running Elasticsearch, OpenSearch, or Redis can often extend what's already there instead of adopting a new system. When scale and retrieval needs stay within reason, that route avoids the operational cost of bringing in another piece of infrastructure to monitor, patch, and staff for, which is often the more expensive decision than any per-query pricing line on a vendor's website.

Sources

  1. Top 15 Vector Databases in 2026: A Production Guide | Medium
  2. Exqutor: Extended Query Optimizer for Vector-augmented Analytical Queries
  3. weaviate.io
Filed underContext Delivery

More in Context Delivery