Vector search, embeddings and RAG
Embeddings and similarity, approximate nearest-neighbour indexes (HNSW, IVF, product quantisation), vector databases versus adding vectors to your existing store, hybrid search, and retrieval-augmented generation (RAG) pipelines end to end.
Reading is half of it. See this used in a real interview: walk through Design an LLM Inference Service →
Embeddings turn text, images and users into vectors so that "similar meaning" becomes "nearby in space". They power semantic search, recommendations, deduplication and, most visibly now, retrieval-augmented generation (RAG), where an LLM answers questions using documents retrieved by similarity. Interviews increasingly include a vector search or RAG component, and the questions are about scale, freshness and quality.
Embeddings and similarity
An embedding model maps an input to a fixed-length vector (commonly 384 to 3,072 dimensions). Inputs with similar meaning get vectors pointing in similar directions. Similarity is usually cosine similarity (or dot product on normalised vectors).
Sizing: 100 M documents × 768 dimensions × 4 bytes (float32) ≈ 300 GB of raw vectors, before index overhead. That number drives most decisions below.
Exact versus approximate search
Exact nearest neighbours means comparing the query with every vector: 100 M comparisons per query, far too slow at scale. Approximate nearest-neighbour (ANN) indexes return almost the best matches (say 95 to 99 % recall) in milliseconds.
| Index | How it works | Strengths | Costs |
|---|---|---|---|
| HNSW (graph) | a layered graph; greedy search hops toward the query | excellent recall and latency | memory-hungry (vectors + graph in RAM); slower to build |
| IVF (clustering) | cluster vectors; search only the nearest clusters | smaller, faster to build | recall depends on how many clusters you probe |
| Product quantisation (PQ) | compress vectors into short codes | 10–30× less memory | lower accuracy; often re-rank with full vectors |
| DiskANN-style | graph index on SSD | billions of vectors per machine | higher latency than RAM |
Common production combination: IVF or HNSW with quantised vectors in memory, then re-rank the top few hundred candidates with full-precision vectors.
Where to store vectors
- A dedicated vector database (Pinecone, Weaviate, Milvus, Qdrant): built for ANN at scale, with filtering, sharding and replication.
- Your existing store with a vector extension (pgvector for Postgres, vector fields in Elasticsearch or OpenSearch, Redis): one system, transactional updates, easy joins with metadata. Excellent up to tens of millions of vectors.
- A library (FAISS) inside your own service, for full control.
Start with the extension in the store you already run unless scale or latency clearly needs a dedicated system.
Filtering and hybrid search
Real queries combine meaning with constraints: "similar products, in stock, under $50, in this region". Filters can be applied before the ANN search (accurate but can break the index’s efficiency for very selective filters) or after (fast but may return too few results). Good vector stores support filtered ANN natively.
Hybrid search combines vector similarity with keyword search (BM25): keywords catch exact terms, product codes and names that embeddings blur; vectors catch paraphrases. Merge the two result lists (for example reciprocal rank fusion) and optionally re-rank with a cross-encoder model. See search, indexing and autocomplete.
RAG end to end
Ingestion (offline, continuous):
- Collect documents and keep their permissions.
- Chunk them into passages of a few hundred tokens, with some overlap, along natural boundaries (headings, paragraphs).
- Embed each chunk and store vector, text and metadata (source, permissions, updated time).
- Re-process on change, through change events, so answers do not cite deleted or outdated text.
Query (online):
- Embed the question (optionally rewrite it first).
- Retrieve the top 20 to 100 chunks with hybrid search, filtered by what this user may see.
- Re-rank to the best 5 to 10.
- Build a prompt with those chunks and the question; the LLM answers and cites sources.
Most RAG quality problems are retrieval problems: bad chunking, missing keyword matching, no re-ranking, stale indexes. Evaluate retrieval separately (did the right chunk appear in the top k?) before blaming the model.
Scale, freshness and cost
- Updates: HNSW handles inserts but deletes are awkward; many systems mark deletions and rebuild periodically.
- Sharding: split vectors across nodes by id; each query fans out to all shards and merges, so per-shard latency matters.
- Re-embedding: changing the embedding model means re-embedding everything; plan it as a batch job with a new index swapped in.
- Permissions: enforce access control in the retrieval filter, never by trusting the LLM to withhold text it was given.
Checklist
- Embedding model, dimensions and total vector memory.
- ANN index type, expected recall and latency; quantisation and re-ranking.
- Store choice: extension in the existing database, or a dedicated vector store.
- Metadata filters and hybrid keyword search.
- For RAG: chunking, permission-aware retrieval, re-ranking, citations.
- Freshness on document changes, and evaluation of retrieval quality.