SysDesignPrep.com
Study guide 40 of 183

Document databases and MongoDB-style modelling

When document stores fit and how to model them: documents and collections, embedding versus referencing, designing for access patterns, unbounded arrays, indexes, sharding and shard keys, transactions and consistency options, schema validation and common pitfalls.

Reading is half of it. See this used in a real interview: walk through Design Yelp →

Document databases (MongoDB, Couchbase, Firestore, Cosmos DB and others) store JSON-like documents rather than rows. They are popular for product catalogues, user profiles, content and any data whose shape varies or nests naturally. In interviews they are a reasonable choice when you justify them from access patterns and model documents sensibly. This guide covers the modelling decisions that matter.

The model

  • A document is a nested structure (objects and arrays) with an id.
  • Documents live in collections, without a fixed schema (though validation can be added).
  • Reads and writes of a single document are atomic.
  • Queries can filter on any field, including nested ones, using indexes.

The central design question is what goes inside one document.

Embedding versus referencing

Embed related data inside the parent document when:

  • It is read together almost every time (a restaurant with its address, hours and photos list).
  • It belongs to one parent only (a listing's amenities).
  • It is bounded in size.

Reference (store ids and fetch separately) when:

  • The related data is large, unbounded or grows forever (all reviews of a restaurant).
  • It is shared by many parents (a user referenced by many posts).
  • It is updated independently and often.

A common middle ground: embed a summary (the latest 5 reviews, the author's display name) and reference the full set. See data modelling and denormalisation.

Model for the access patterns

As with any NoSQL store, start from the queries:

  • "Show a listing page" fetches one listing document with embedded details: one read.
  • "Show reviews for a listing, newest first, paged" reads a reviews collection indexed by (listing_id, created_at).
  • "Show a user's bookings" reads a bookings collection indexed by user_id.

Duplicated fields (names, prices at booking time) are fine if you know how they are kept in sync or why they should not be. See Design Airbnb.

Pitfalls

  • Unbounded arrays: embedding every comment or follower in one document eventually hits the document size limit (16 MB in MongoDB) and makes every update rewrite a growing document. Use a separate collection.
  • Over-normalizing: modelling exactly like relational tables, then doing many round trips or $lookup joins on every request.
  • Missing indexes: unindexed queries scan the collection. Index fields used for filtering and sorting, in the right compound order. See database indexing.
  • Schema drift: without validation, documents accumulate inconsistent shapes. Add schema validation and version fields, and migrate lazily on read or with backfills. See schema evolution and serialization.

Sharding

Document stores shard collections by a shard key:

  • Choose a high-cardinality key that matches the main query filter, so queries go to one shard (user_id for user data, listing_id for reviews).
  • Avoid monotonically increasing keys (timestamps, ObjectIds at the front) that send all inserts to one shard; use hashed sharding when range queries on the key are not needed.
  • Queries without the shard key scatter to all shards. See sharding and partitioning and hot keys and skew.

Consistency and transactions

  • Single-document operations are atomic, which is why embedding data that must change together is valuable.
  • Multi-document ACID transactions exist in modern MongoDB, at a performance cost; use them for occasional cross-document invariants, not as the default.
  • Replica sets offer write concerns (acknowledged by a majority, for durability) and read concerns and preferences (read from primary for fresh data, or secondaries for scale with staleness). See consistency models.

When document stores fit

Good fits:

  • Catalogues and content with varied attributes (products with different specs, listings with optional fields).
  • User profiles and settings.
  • Aggregates read and written as a unit (a cart, an order with its items). See shopping cart and checkout.
  • Rapidly evolving products where schemas change often.

Less good:

  • Highly relational data with many-to-many queries and complex joins (consider relational or graph stores). See graph data.
  • Heavy analytical queries (use a warehouse). See OLTP versus OLAP.
  • Strict multi-entity transactional workloads like ledgers (relational is simpler). See payments and ledgers.

Postgres with JSONB columns covers many document use cases while keeping relational features; it is a fair alternative to mention. See SQL vs NoSQL.

In the interview

"Businesses are documents with embedded address, hours, categories and a summary of recent reviews; full reviews live in a separate collection indexed by business id and date; the businesses collection is sharded by hashed business id; geo queries use a geospatial index; writes use majority write concern." See Design Yelp.

Checklist

  • Model from access patterns; one read per main screen where possible.
  • Embed bounded, co-read, owned data; reference unbounded or shared data.
  • No unbounded arrays; summaries embedded, details referenced.
  • Compound indexes for filters and sorts; schema validation and versions.
  • Shard key matching the main filter, not monotonic.
  • Write and read concerns chosen per operation; transactions sparingly.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.