Pick the vector database that fits your scale, your team and the database you already run.
You built a chatbot. It works on your machine until you feed it 50,000 internal documents. Then it hallucinates, it slows down, and it pulls a report from three years ago when you asked for last week’s.
Nine times out of ten, the weak link is the RAG stack, and usually the vector database underneath it.
The database wall
A Retrieval-Augmented Generation (RAG) system sounds like a weekend project. Past the “hello world” stage, you hit the database wall.
Choosing the wrong vector store early leads to high latency, rising cloud costs and a painful migration six months later when your data outgrows the setup.
Pick for your scale
I’ve seen teams freeze in front of the number of options in the AI ecosystem. Pick the right tool for your scale and your team.
Here is the 2026 landscape, so you can stop scrolling and start shipping.
Why the database matters in RAG
An LLM like Claude or GPT knows a lot but remembers nothing about your data. RAG gives it that memory, and your vector database is the librarian. If the librarian is slow or loses books, the model can’t do its job.
If you’re still deciding whether RAG is the right tool at all, versus baking knowledge into the model, start with my breakdown of RAG vs fine-tuning.
Three things to look for
- Latency: can it find the right chunk in tens of milliseconds? When it can’t, a Redis semantic cache in front of the store often helps more than swapping databases.
- Hybrid search: can it search by meaning (vectors) and by exact keywords (full-text)?
- Developer experience: how much time will you spend on DevOps?
The contenders
pgvector: “I already have a database”
If you already run Postgres for your web applications, pgvector is usually the first stop. It’s an extension that adds vector types and indexes to the database you already trust.
It fits well under about 10 million vectors. You keep ACID guarantees and simple backups, and your relational data stays next to your embeddings. No new infra, no new security review.
- Pros: zero new infrastructure on Postgres, easy joins with user metadata, wide ecosystem support (Laravel, Django, Node.js).
- Cons: 100M+ vectors needs serious hardware, and hybrid search means combining it with Postgres full-text search yourself.
Pinecone: “I want zero ops”
Pinecone is a managed, serverless vector database. You don’t manage clusters or tune indexes. You send vectors and get results, and you pay for what you use.
It’s the usual pick for teams that want to grow toward a billion vectors without hiring a dedicated DevOps engineer.
- Pros: fully managed (pick a region and go), low latency, enterprise features like SOC 2.
- Cons: there is no open-source or self-hosted version. Since August 2026, Enterprise customers can use Bring Your Own Cloud (BYOC) to run the data plane in their own AWS, GCP or Azure account. Costs also climb quickly with heavy read and write volume.
Weaviate and Qdrant: hybrid search first
If your app needs to blend semantic search with classic keyword search, these two lead. Both are built for vector retrieval from the ground up, and both offer open-source and managed cloud versions.
Weaviate does hybrid search out of the box. Qdrant, written in Rust, is fast and light on memory.
- Pros: strong hybrid search (BM25 plus vectors), flexible hosting (self-hosted Docker or managed cloud), good filtering (“documents from 2025 that mention security”).
- Cons: more operational overhead than Pinecone, and a new database API to learn.
| Option | Best for | Hosting | Hybrid search |
|---|---|---|---|
| pgvector | Teams already on Postgres | Your Postgres | With Postgres full-text |
| Pinecone | Zero ops at large scale | Managed (BYOC on Enterprise) | Supported |
| Weaviate | Hybrid search and filtering | Self-hosted or managed | Built in |
| Qdrant | Fast, memory-efficient search | Self-hosted or managed | Built in |
How to choose: the engineering trade-offs
Pick the database that matches your constraints.
Factor 1: the “billions” problem
Most startups don’t have a billion vectors. They have a few thousand PDFs. In the sub-1M range, pgvector is almost always the right answer.
If you’re building a global legal search engine or a large e-commerce recommendation system, you need the distributed design of Milvus or Pinecone. Don’t build a large distributed system for a small amount of data.
Factor 2: hybrid search
Pure vector search is weak at exact technical terms. Search for “PHP 8.4 features” and you might get general PHP articles. Hybrid search blends the meaning of the vector with the precision of a keyword match.
If search quality is your top metric, look at Weaviate or Qdrant. If your data is highly relational, weigh vector search vs graph search before you commit to pure vectors.
Factor 3: the DevOps tax
Every new piece of infrastructure is another thing that can break at 3 AM.
With a small team, lean on managed services like Pinecone or Zilliz. With a strong infra team that wants to save on cloud margins at high scale, self-hosting Qdrant on Coolify or Kubernetes is the move.
pgvector with Laravel
Since I work a lot with custom web development using Laravel, here is what this looks like in practice. Laravel 13 with the AI SDK (laravel/ai) has native vector columns and a whereVectorSimilarTo query method on Postgres with pgvector.
// migration: an HNSW index with cosine distance
$table->vector('embedding', dimensions: 1536)->index();
// query: pass the user's question as a string and Laravel embeds it for you
$results = Document::query()
->select('content')
->whereVectorSimilarTo('embedding', $query, minSimilarity: 0.4)
->limit(5)
->get();
Under the hood this is pgvector’s cosine distance operator, <=>. That snippet is the core of a RAG system: find the content, send it to the LLM and get a grounded answer.
On older Laravel versions, the pgvector/pgvector-php package gives you a Vector cast and a nearestNeighbors scope that do the same job.
Fat vector index
Lean vector index
Three practical tips
These will save you weeks of refactoring.
Index after the initial load
Vector indexes like HNSW are fast to search but slow to insert into. For a large initial load, insert your vectors first, then create the index. The pgvector docs say the same: it’s faster to build the index after loading your data.
Normalize your vectors
Make sure your embedding model and your vector database agree. If you use cosine similarity, normalize your vectors. It keeps scores consistent and prevents odd ranking bugs.
Keep the metadata lean
Don’t store the whole JSON document inside the vector store. Store the vector and an ID, and keep the heavy data in your primary database. The index stays small and fast.
These are the cheap wins. For the deeper traps (bad chunking, missing reranking, stale embeddings), see the 7 mistakes you’re making with your production RAG stack. To test chunk sizes on your own documents, try my RAG chunk splitter.
My rule of thumb
I’ve built these systems for startups and established businesses, and this is how I usually guide them:
- Default to pgvector. It’s the path of least resistance for most web apps.
- Move to Pinecone if you need high performance and don’t want to manage servers.
- Choose Weaviate or Qdrant if your app depends on hybrid search and metadata filtering.
Key takeaways
- pgvector is the default for teams already on Postgres.
- Pinecone is the zero-ops option for scale, with BYOC for Enterprise teams that need data in their own cloud.
- Hybrid search (keyword plus vector) usually beats vector search alone.
- Keep it simple. Don’t design for billions of vectors when you have thousands.
The right stack is the one that lets you ship your AI features today. If you’re choosing one for production, here’s how I help teams design and ship RAG systems, or reach out with your war stories.
What’s the hardest problem you’ve hit with retrieval in your RAG system so far?