SysDesignPrep.com
Study guide 76 of 183

Vector search, embeddings and RAG

Embeddings and similarity, approximate nearest-neighbour indexes (HNSW, IVF, product quantisation), vector databases versus adding vectors to your existing store, hybrid search, and retrieval-augmented generation (RAG) pipelines end to end.

Reading is half of it. See this used in a real interview: walk through Design an LLM Inference Service →

Embeddings turn text, images and users into vectors so that "similar meaning" becomes "nearby in space". They power semantic search, recommendations, deduplication and, most visibly now, retrieval-augmented generation (RAG), where an LLM answers questions using documents retrieved by similarity. Interviews increasingly include a vector search or RAG component, and the questions are about scale, freshness and quality.

Embeddings and similarity

An embedding model maps an input to a fixed-length vector (commonly 384 to 3,072 dimensions). Inputs with similar meaning get vectors pointing in similar directions. Similarity is usually cosine similarity (or dot product on normalised vectors).

Sizing: 100 M documents × 768 dimensions × 4 bytes (float32) ≈ 300 GB of raw vectors, before index overhead. That number drives most decisions below.

Exact nearest neighbours means comparing the query with every vector: 100 M comparisons per query, far too slow at scale. Approximate nearest-neighbour (ANN) indexes return almost the best matches (say 95 to 99 % recall) in milliseconds.

IndexHow it worksStrengthsCosts
HNSW (graph)a layered graph; greedy search hops toward the queryexcellent recall and latencymemory-hungry (vectors + graph in RAM); slower to build
IVF (clustering)cluster vectors; search only the nearest clusterssmaller, faster to buildrecall depends on how many clusters you probe
Product quantisation (PQ)compress vectors into short codes10–30× less memorylower accuracy; often re-rank with full vectors
DiskANN-stylegraph index on SSDbillions of vectors per machinehigher latency than RAM

Common production combination: IVF or HNSW with quantised vectors in memory, then re-rank the top few hundred candidates with full-precision vectors.

Where to store vectors

  • A dedicated vector database (Pinecone, Weaviate, Milvus, Qdrant): built for ANN at scale, with filtering, sharding and replication.
  • Your existing store with a vector extension (pgvector for Postgres, vector fields in Elasticsearch or OpenSearch, Redis): one system, transactional updates, easy joins with metadata. Excellent up to tens of millions of vectors.
  • A library (FAISS) inside your own service, for full control.

Start with the extension in the store you already run unless scale or latency clearly needs a dedicated system.

Real queries combine meaning with constraints: "similar products, in stock, under $50, in this region". Filters can be applied before the ANN search (accurate but can break the index’s efficiency for very selective filters) or after (fast but may return too few results). Good vector stores support filtered ANN natively.

Hybrid search combines vector similarity with keyword search (BM25): keywords catch exact terms, product codes and names that embeddings blur; vectors catch paraphrases. Merge the two result lists (for example reciprocal rank fusion) and optionally re-rank with a cross-encoder model. See search, indexing and autocomplete.

RAG end to end

Documents are chunked, embedded and indexed; a question is embedded, similar chunks are retrieved and re-ranked, and the LLM answers using them
Retrieval-augmented generation. Quality is decided by chunking, retrieval and re-ranking long before the LLM sees anything.

Ingestion (offline, continuous):

  1. Collect documents and keep their permissions.
  2. Chunk them into passages of a few hundred tokens, with some overlap, along natural boundaries (headings, paragraphs).
  3. Embed each chunk and store vector, text and metadata (source, permissions, updated time).
  4. Re-process on change, through change events, so answers do not cite deleted or outdated text.

Query (online):

  1. Embed the question (optionally rewrite it first).
  2. Retrieve the top 20 to 100 chunks with hybrid search, filtered by what this user may see.
  3. Re-rank to the best 5 to 10.
  4. Build a prompt with those chunks and the question; the LLM answers and cites sources.

Most RAG quality problems are retrieval problems: bad chunking, missing keyword matching, no re-ranking, stale indexes. Evaluate retrieval separately (did the right chunk appear in the top k?) before blaming the model.

Scale, freshness and cost

  • Updates: HNSW handles inserts but deletes are awkward; many systems mark deletions and rebuild periodically.
  • Sharding: split vectors across nodes by id; each query fans out to all shards and merges, so per-shard latency matters.
  • Re-embedding: changing the embedding model means re-embedding everything; plan it as a batch job with a new index swapped in.
  • Permissions: enforce access control in the retrieval filter, never by trusting the LLM to withhold text it was given.

Checklist

  • Embedding model, dimensions and total vector memory.
  • ANN index type, expected recall and latency; quantisation and re-ranking.
  • Store choice: extension in the existing database, or a dedicated vector store.
  • Metadata filters and hybrid keyword search.
  • For RAG: chunking, permission-aware retrieval, re-ranking, citations.
  • Freshness on document changes, and evaluation of retrieval quality.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.