Retrieval-Augmented Generation104 · Module II · Lesson 04 of 13
Article · 12 min

Vector databases

Pinecone, Weaviate, Qdrant, pgvector — compared honestly.

Summary

A vector database stores embeddings and returns nearest neighbors by similarity. The category is crowded: Pinecone, Weaviate, Qdrant, Milvus, Chroma, pgvector. The choice matters less than the discipline of using one as a database — with backups, migrations, and access control.

Objectives
  • 01Name the four axes on which vector databases differ.
  • 02Recognize when pgvector suffices.
  • 03State one reason to pick each of the leading options.
The Lesson

The axes

1. Managed vs. self-hosted. 2. Standalone vs. embedded (in Postgres). 3. Index algorithm (HNSW, IVF, DiskANN). 4. Hybrid search support (BM25 alongside vectors). Every serious DB supports HNSW; the differences are operational.

Choosing

pgvector if you already run Postgres and have <10M vectors. Qdrant for self-hosted with strong metadata filtering. Pinecone for zero-ops managed. Weaviate for hybrid search and modular design. Milvus for scale beyond 100M vectors.

Key Ideas
  • Vector DBs mostly do the same thing; operations differ.
  • Start with pgvector; move only when it hurts.