SysDesignPrep.com
Study guide 142 of 183

Designing an email inbox (Gmail-style)

The backend of a webmail service: receiving mail over SMTP, spam filtering, storing messages and attachments, threading into conversations, labels and folders, per-user search, sync with IMAP and mobile clients, quotas, deduplication and retention.

Reading is half of it. See this used in a real interview: walk through Design a Notification System →

"Design Gmail" is a classic interview question that combines inbound message processing, per-user storage at enormous scale, search, and sync across devices. Unlike chat, email arrives from the outside world over an old protocol, much of it spam, and users keep it for decades. This guide covers the receiving and storage side; for sending, see sending email at scale.

Receiving mail

  1. Other mail servers find yours via MX records in DNS and connect over SMTP.
  2. Inbound MTAs (mail transfer agents) accept the connection, check sender reputation and authentication (SPF, DKIM, DMARC), and apply rate limits.
  3. Accepted messages are queued durably (once accepted, you are responsible for them) and processed asynchronously: parsing, spam and malware scanning, classification, then delivery to each recipient's mailbox.

The SMTP edge must reject abusive senders cheaply and absorb spikes. See DDoS protection and abusive traffic.

Spam and safety

A large share of inbound mail is spam or malicious. Filtering combines reputation (IPs, domains), authentication results, content classifiers, link and attachment scanning, and user feedback ("report spam" and "not spam" are training labels). Messages are routed to the spam folder rather than dropped, so mistakes can be recovered. See trust and safety.

Storage model

Per user, the system stores:

  • Message metadata: id, thread id, sender, recipients, subject, date, labels, read and starred flags, size, a short snippet.
  • Message bodies and attachments: the raw MIME message, often large.

Split them:

  • Metadata in a database partitioned by user id (wide-column or sharded relational), indexed by (user, label, date) for folder views. See data modelling for Cassandra and DynamoDB.
  • Bodies and attachments in object storage, referenced by id, compressed and encrypted at rest. See object storage and files.
  • Deduplicate identical attachments (the same PDF sent to 500 people) by content hash, with reference counting. See hashing and encoding.

Threads (conversations)

Group messages into threads using the Message-ID, In-Reply-To and References headers, falling back to subject similarity and participants. Each message gets a thread id at delivery time; the inbox view lists threads by their latest message, with the thread's labels being the union of its messages' labels.

Labels and folders

Gmail-style labels are many-to-many: a message can have several. Store labels as a set on the message metadata and maintain per-label indexes (user, label, date) for fast listing. Unread counts per label are counters updated on changes and corrected periodically. See counting at scale. Folders are a special case (one label per message).

Users search their own mail, so the index is partitioned by user:

  • Each user's messages are indexed (sender, recipients, subject, body text, attachment names and sometimes contents) in a per-user or user-sharded inverted index. See how Elasticsearch works.
  • Queries always include the user id, so they hit one shard, unlike web search's scatter-gather. See web search engine architecture.
  • Indexing is asynchronous after delivery; new mail becomes searchable within seconds.

Sync with clients

  • The web and mobile apps sync through an API with a per-user change log (history id): "what changed since history 9876?" returns new messages, label changes and deletions. See offline-first apps and sync.
  • Push notifications or a persistent connection signal new mail. See push notifications.
  • IMAP support for third-party clients maps folders to labels and keeps per-folder message sequence numbers, which is notoriously awkward with labels.

Quotas and retention

  • Per-user storage quotas (counting deduplicated attachments in some way that is fair to users).
  • Trash and spam auto-delete after 30 days; deletion removes metadata and decrements attachment references, with garbage collection of unreferenced blobs. See privacy and data deletion.
  • Enterprise retention and legal holds override user deletion for business accounts. See audit logs.

Scale

Billions of mailboxes, hundreds of billions of messages per day (most of it spam), and petabytes of storage. Everything is partitioned by user, which keeps most operations single-shard; the heavy work (spam filtering, indexing) is asynchronous and horizontally scalable. Cold mail moves to cheaper storage tiers. See cost-aware system design.

In the interview

"Inbound SMTP servers authenticate and rate-limit senders, accept into a durable queue, then a pipeline scans for spam and malware, threads the message and delivers it to each recipient: metadata into a user-partitioned store with label indexes, bodies and deduplicated attachments into object storage, and an entry into the user's change log. A per-user search index is updated asynchronously; clients sync via the change log and get pushes for new mail."

Checklist

  • MX and SMTP edge with authentication checks and rate limits; durable acceptance.
  • Spam and malware pipeline with user feedback.
  • Metadata partitioned by user; bodies and attachments in object storage with deduplication.
  • Threading by headers; labels as sets with per-label indexes and counters.
  • Per-user search shards updated asynchronously.
  • Change-log sync, push, IMAP compatibility.
  • Quotas, retention, legal holds and garbage collection.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.