SysDesignPrep.com
Interviewer kit

Design Instagram

Run this for someone else. You hold the answers; they do not. Read the prompt, keep the clock, and use the probes below when an answer is thin. Do not show them this page.

The candidate should have a blank page and a whiteboard, not this.

Open with this

Upload photos, publish posts and stories, and load a feed of them for 500 M daily users, with every image served from a CDN in milliseconds. Take a couple of minutes on requirements, then we will do some numbers, then the design. I will interrupt to keep us moving.

The clock

  • 4 min: functional requirements and scope
  • 4 min: non-functional requirements, with numbers
  • 5 min: back-of-envelope estimates
  • 16 min: high-level design and one or two flows
  • 16 min: deep dives and the close

Move them on out loud when a section overruns. The commonest failure is spending twenty minutes on requirements and never reaching a deep dive, and preventing that is your job as much as theirs.

Requirements · 8 min

Listen for: a scoped set of capabilities, an explicit out-of-scope list, and numeric targets rather than adjectives. Prompt with “what are you not building?” if they never scope, and “what number would make that requirement real?” if they say “fast” or “highly available”.

Functional (7)
  • Upload a photo and publish a post: One to ten photos with a caption. The post appears only once every image is processed, never half-rendered.
  • Home feed: Posts from accounts you follow, ranked, paginated. The first screen must load fast on a phone on a weak connection.
  • Stories: Photos that disappear after 24 hours, shown in a tray at the top of the app, with "seen" state per viewer.
  • Likes and comments: Like counts shown on every post. Exact for small accounts; may lag by seconds on a post getting a million likes an hour.
  • Follow and unfollow: Changes what the feed shows from the next refresh. The graph is asymmetric: following does not need approval for public accounts.
  • Profile grid: Every post by one user, newest first, as thumbnails.
  • Out of scope: Video and Reels transcoding (see Design YouTube), direct messages, search and Explore, ads, and content moderation beyond a hook in the pipeline.
Non-functional (7)
  • Scale (500 M DAU · 100 M photos/day): Reads dominate: about a hundred feed loads for every photo uploaded.
  • Feed latency (p99 < 300 ms for the first page): The feed is the app’s home screen. Images come from the CDN separately, so the API returns ids and URLs, not bytes.
  • Image delivery (p95 < 100 ms per image, CDN hit > 95 %): This is the requirement that drives the hardest trade-off: petabytes a day of egress have to come from the edge, so every image needs stable, cacheable URLs in several sizes.
  • Durability (no lost photos): People’s memories. Originals are stored with 11 nines durability; derived sizes can always be regenerated.
  • Availability (99.95 % for reads): A feed that loads slightly stale beats an error. Uploads can queue and retry on the client.
  • Freshness (followers see a post within seconds): Eventual consistency is fine; your own post must appear in your own feed and grid immediately (read your writes).
  • Privacy: Private accounts: only approved followers can see posts, and image URLs must not be guessable or usable forever once shared.

Estimates · 5 min

Ask for two or three numbers, not all of them. What matters is whether they state assumptions, round sensibly, and say what the number implies. Push once with “where did that come from?”

The numbers (7)
  • Photo uploads per second: ~1.2 k avg · ~3.5 k peak. 100 M/day ÷ 86 400 s ≈ 1.2 k/s; × 3 for the evening peak ≈ 3.5 k/s. Each upload fans into several resized variants, so the processing fleet sees about 5× that in image operations.
  • Feed loads per second: ~120 k avg · ~350 k peak. 500 M DAU × 20 feed loads a day = 10 B/day ÷ 86 400 ≈ 116 k/s; × 3 at peak ≈ 350 k/s. About 100 reads for every upload.
  • New image storage per day: ~300 TB. Original ~2 MB plus four variants (1080, 640, 320 px and a thumbnail) adding ~1 MB ≈ 3 MB per photo. 100 M × 3 MB = 300 TB/day, about 110 PB a year before replication.
  • Image egress: ~10 PB/day. Each feed load shows ~10 images at ~100 KB (the phone-sized variant) ≈ 1 MB. 10 B loads × 1 MB = 10 PB/day ≈ 115 GB/s on average. Only a CDN can serve this; at a 95 % hit rate the origin still serves ~6 GB/s.
  • Likes per second: ~45 k avg · ~150 k peak. 4 B likes/day ÷ 86 400 ≈ 46 k/s. The total is easy; the problem is concentration: one celebrity post can take 5 k likes a second on a single counter.
  • Timeline cache size: ~4 TB. Keep the newest 500 post ids per active user: 500 × 16 B (id + score) = 8 KB × 500 M users = 4 TB of RAM across a Redis cluster. Large but ordinary: ~60 nodes with 64 GB each.
  • Stories alive at any moment: ~250 M. 250 M stories posted a day, each living 24 hours, so about 250 M exist at once. Metadata ~200 B each is 50 GB: small enough to keep entirely in memory with a TTL.

High-level design · 16 min

Let them draw. Interrupt only to ask what backs a component or what a box actually does. Then pick one flow below and ask them to walk it end to end.

Components (17)
  • Mobile app: Uploads photos straight to object storage with a presigned URL, calls the API for feeds and posts, and loads every image from the CDN. Holds a resumable upload queue so a dropped connection does not lose a post.
  • Image CDN (edge cache, signed URLs): Serves every image byte. URLs are content-addressed and immutable (a new size or edit is a new URL), so they cache forever. Private-account images use signed URLs that expire.
  • API gateway: Authenticates, rate limits and routes API calls. Returns JSON only: post ids, captions and image URLs, never image bytes.
  • Upload service: Hands out presigned upload URLs and tracks each upload’s state. Keeps the gateway out of the byte path, so a 4 MB photo never passes through an API server.
  • Object storage (S3 · originals + variants): Originals and derived sizes, keyed by content hash. Lifecycle rules move originals of old posts to cheaper tiers; variants stay hot because the CDN refills from them.
  • Media workers (resize · strip EXIF · moderate): Triggered when an original lands: validate, strip location metadata, resize into four variants, run the moderation classifier, then emit media.ready. Idempotent per content hash, so retries are free.
  • Post service: Creates posts in a pending state, publishes them once all media is ready, and serves post metadata for feeds and grids. Owns read-your-writes for the author.
  • Post store (sharded Postgres · key=user_id): Posts, captions, media references. Sharded by author, with post ids that embed the shard so any id can be routed without a lookup (Instagram’s published scheme). A profile grid is one shard, one index scan.
  • Event bus (Kafka): post.published, media.ready, like and follow events. Decouples the write path from fan-out, counters and notifications, and lets them lag without slowing a publish.
  • Feed service: Builds a page of the home feed: read the precomputed timeline, merge in recent posts from followed celebrities, hydrate, rank, and return ids with image URLs.
  • Timeline cache (Redis sorted sets): Per user, the newest ~500 post ids pushed by fan-out. About 4 TB across the cluster. Rebuilt from the post store on a miss, so losing it costs latency, not data.
  • Fan-out workers: On post.published, look up the author’s followers and push the post id into each follower’s timeline. Skipped for accounts above the celebrity threshold, whose posts are pulled at read time instead.
  • Social graph (follower lists, sharded): Who follows whom, stored in both directions so fan-out can page through followers and the feed can list followed celebrities. Follower counts decide which side of the celebrity threshold an account is on.
  • Stories service: Publishes stories with a 24-hour expiry, builds the tray (accounts you follow with unseen stories first), and records views.
  • Stories store (Redis + TTL · seen bitmaps): Live stories per author with a 24-hour TTL, so expiry is the store’s job, not a cleanup cron. Per-viewer seen state is a compact set per tray.
  • Like service: Records who liked what (exactly once per user and post) and maintains the displayed counts through sharded counters that absorb hot posts.
  • Counter store (sharded counters): Like and comment counts split across N sub-counters per hot post and summed on read, with the total snapshotted to the post store every few seconds.
Flows to ask them to walk (5)
  1. Upload photos and publish a post: Bytes go straight from the phone to object storage, processing is asynchronous, and the post only becomes visible once every image is ready. The API never touches a byte of image data.
    1. App asks for upload URLs, one per photo, and creates the post in a pending state.
    2. App uploads each original directly to object storage.
    3. Storage emits an object-created event; a media worker validates, strips EXIF location, resizes and moderates.
    4. Worker publishes media.ready; the post service marks that photo done and publishes the post when all are ready.
    5. Post service emits post.published; the author sees it at once in their own grid and feed.
  2. Load the home feed: A precomputed timeline makes the common case one cache read. Celebrities are merged in at read time, and images arrive from the CDN in parallel with the JSON.
    1. App requests the first page of the feed.
    2. Feed service reads the user’s timeline of pushed post ids.
    3. It merges recent posts from celebrities the user follows, which were never pushed.
    4. Candidates are hydrated and ranked, and the page is returned with image URLs.
    5. App fetches the images from the CDN; most are edge hits.
  3. A celebrity with 300 M followers posts: The scale-breaking case on both the write side (300 M timeline inserts) and the engagement side (a like storm on one counter).
    1. The post is published like any other.
    2. Fan-out sees the author is above the threshold and skips pushing.
    3. Followers’ feed loads pull the post from the celebrity’s recent-posts cache.
    4. Likes pour in at thousands per second on one post; each goes to a random shard of its counter.
    5. Like events stream onward for notifications and for the periodic count snapshot.
  4. An upload dies halfway, or a worker crashes: Mobile networks drop constantly. Every step is retryable and idempotent, and nothing half-done ever becomes visible.
    1. The connection drops during the second photo’s upload.
    2. The app retries the create call; the idempotency key returns the same pending post.
    3. A media worker crashes mid-resize; the event is redelivered and another worker redoes the job.
    4. The post stays pending until every media.ready has arrived, so followers never see a broken post.
    5. A sweeper deletes posts still pending after 24 hours and their orphaned originals.
  5. Post a story, view the tray, expire after 24 hours: Stories reuse the media pipeline but live in a store with a TTL, and the tray is ordered by what each viewer has not yet seen.
    1. The story image is uploaded and processed exactly like a post photo.
    2. Stories service records the story with a 24-hour TTL.
    3. A viewer opens the app; the tray lists followed accounts with live stories, unseen first.
    4. Viewing a story marks it seen for that viewer.

Deep dives · 16 min

Pick two. Ask the headline question, let them answer, then use the follow-ups. The follow-ups are where the level gets decided, so leave time for at least three of them.

Getting photos in and out

Ask: Why not just POST the photo to your API servers and resize it there?

Good answers name: Presigned direct upload + event-driven resize into fixed variants, Upload through the API, resize synchronously, Store only originals; resize on the fly at the CDN.

Our pick: Create the post first with an idempotency key, return a presigned PUT URL per photo bound to the key, content type and a size cap, and let the app upload directly with multipart resume. An object-created event triggers a media worker that validates the image, strips EXIF location, writes four variants under content-hashed keys, runs the moderation classifier and emits media.ready. The post publishes when every photo is ready. Common sizes are precomputed because they are requested billions of times; an on-demand resizer behind the CDN handles anything unusual and its output is cached. Instagram has written about moving uploads off its web tier for exactly these reasons.

  1. Someone uploads a 200 MB file, or a file that is not an image. Where is that stopped?
    Twice. The presigned URL fixes Content-Type and a maximum Content-Length, so S3 rejects an oversized upload outright. Then the worker decodes the file with a hardened image library in a sandbox and rejects anything that does not parse as an image within limits on pixel count, which also stops decompression bombs. The post stays pending and the app is told the upload failed.
  2. Why strip EXIF data?
    Phone photos carry GPS coordinates, the device model and sometimes the owner’s name. Publishing those on every photo leaks where people live. The worker strips everything except orientation, which it applies to the pixels first so the image does not appear rotated.
  3. At 10× uploads, what breaks first?
    The media workers, because resize is CPU-bound: 35 k uploads a second at peak means ~175 k resize operations a second. They autoscale on queue depth, and the queue absorbs bursts so users see a slightly longer pending state rather than errors. Object storage and presigning scale on their own.
  4. How do you know the pipeline is healthy?
    Measure time from upload complete to post published at p50 and p99, queue age (the oldest unprocessed event), worker failure rate by error type, and the count of posts pending longer than a minute. A rising queue age is the earliest warning; it moves before users notice.
Building the home feed

Ask: Do you precompute every user’s feed when someone posts, or assemble it when they open the app?

Good answers name: Hybrid: push for normal accounts, pull for celebrities above a threshold, Push to everyone (fan-out on write), Pull for everyone (fan-out on read).

Our pick: Fan out on write to followers’ Redis timelines for accounts below ~1 M followers, keeping the newest 500 ids per user. Skip fan-out for accounts above the threshold and keep their last few days of posts in a replicated per-author cache; at read time the feed service merges the viewer’s timeline with the recent posts of the celebrities they follow (usually under ten). Skip pushing to followers inactive for 30 days and rebuild their timeline on return. Rank the merged ~300 candidates with a light model and paginate with a score-and-id cursor. This is the same hybrid as Design a News Feed, which goes deeper on ranking and cursor stability.

  1. A user unfollows someone. Do their posts vanish from the feed?
    At read time, yes: hydration filters out authors the viewer no longer follows, which is cheap because the follow set is cached. A background job also removes that author’s ids from the timeline so they stop taking slots. Filtering on read makes the change instant without waiting for the cleanup.
  2. The timeline cache cluster loses a node. What do users see?
    Users on that shard get timeline misses, and the feed service rebuilds their timelines from the post store and graph, which is slower (a few hundred ms) but correct. Rebuilds are rate limited so a node loss does not become a database stampede; some users briefly see a feed assembled from fewer sources.
  3. How do you choose the celebrity threshold?
    From cost: compare the write cost of pushing (followers × cache write) against the read cost of pulling (followers’ feed loads per day × merge cost). The crossover for Instagram-like ratios is in the hundreds of thousands to low millions. Measure fan-out lag and feed p99 and move the threshold to balance them.
  4. How does a private account change fan-out?
    Only approved followers are in the follower list, so fan-out is naturally limited to them. The extra rule is on read: hydration checks the viewer is still an approved follower, so a post pushed before a follower was removed is not shown afterwards.
Like counts on hot posts

Ask: Why not just UPDATE posts SET likes = likes + 1?

Good answers name: Idempotent like row + sharded counter, summed on read, snapshotted periodically, Single counter column, incremented in place, Count asynchronously from the event stream only.

Our pick: A like is an insert of (post_id, user_id) with a uniqueness constraint, which gives exactly-once semantics and the "did I like this" check. Only if the insert was new do we increment the counter. Counters start as one key per post; when a post’s like rate crosses a threshold it is promoted to 64 sub-counters and each increment goes to a random one. Reads sum the sub-counters and cache the result for a second. A job snapshots totals to the post row every few seconds so cold posts need no counter read. The viewer’s own like is applied optimistically in the app, so they always see their action reflected.

  1. Someone likes and unlikes rapidly a hundred times. What happens to the count?
    The like row is the source of truth: insert on like, delete on unlike, and only a successful insert or delete moves the counter. Rapid toggling produces matched increments and decrements. The client also debounces, sending only the final state after a short pause, which cuts the write rate.
  2. A bot farm adds a million likes. How do you take them back?
    Like events go through the stream to an abuse pipeline. When accounts are judged fake, their like rows are deleted in bulk and the counters recomputed from the like table for affected posts, rather than decremented one by one. Recomputing from the source of truth avoids drift.
  3. How do you show "liked by alice and 4,812 others" quickly?
    The "alice" part is a lookup of which of the viewer’s followees liked the post, which is expensive in general. Precompute it lazily: when hydrating, check the viewer’s top followees against a per-post Bloom filter or small set of recent likers, and fall back to just the count if nothing is found quickly.
Stories that disappear

Ask: Stories expire after 24 hours. Do you run a job that deletes them?

Good answers name: Expiry as data: store with TTL, filter by expires_at on read; seen state per (viewer, author), Stories as normal posts with a deletion cron, Precompute each viewer’s tray on write.

Our pick: Each author has a Redis sorted set of live story ids scored by expires_at, and the key itself expires when the newest story does. Reads ignore entries whose score is in the past, so a story is never served late even if eviction lags. The tray is assembled on read from the followed authors who have live stories (a set kept per viewer and updated by story events), ordered by unseen first and by recency, and cached for a minute. Seen state is the newest story timestamp seen per (viewer, author). Media originals carry a lifecycle rule that deletes them after the window plus the author’s private archive copy.

  1. The author saves the story to their archive. How does that interact with the TTL?
    The archive is a separate, private record in the post store pointing to the same media, created when the story is posted if archiving is on. The public story still expires from the stories store; only the media lifecycle rule must respect the archive reference, so archived media is moved to a non-expiring prefix instead of being deleted.
  2. How do you show the author who viewed their story?
    Each view appends the viewer id to a per-story set, deduplicated. The list is read rarely (only the author, a few times) and dropped with the story, so it lives in the same TTL store. For accounts with millions of viewers, store only the first N names and an approximate total count with HyperLogLog.
  3. What happens if the stories cache loses data?
    Stories metadata is also written to a durable log, so the cache is rebuilt from the last 24 hours of events. The cost of loss is a brief gap in the tray, not lost content. For a feature whose content expires in a day, that is the right durability trade.
Sharding the post store

Ask: Billions of posts: how do you shard them, and how do you find a post from just its id?

Good answers name: Shard by author; ids embed timestamp + logical shard + sequence, Shard by post id, secondary index by author, A wide-column store keyed by (author, post id).

Our pick: Thousands of logical shards mapped onto a smaller number of Postgres clusters, with posts placed on the shard of their author. Post ids are 64 bits: 41 bits of milliseconds since a custom epoch, 13 bits of logical shard id and 10 bits of per-shard sequence, generated inside the database so no separate id service is needed. Hydration groups ids by shard and issues one multi-get per shard. Moving a logical shard to a new physical cluster is a copy plus a map update, so growth never re-hashes data. Likes and comments, which are append-heavy and huge, live in a wide-column store keyed by post id.

  1. One author posts constantly and their shard is hot. What do you do?
    Logical shards are small, so the hot logical shard can be moved to its own physical cluster. Reads of a celebrity’s posts are absorbed by the per-author cache anyway; the store mostly sees writes, which even for a prolific account are a few a minute.
  2. Why not use UUIDs for post ids?
    They are 128 bits, do not sort by time (except v7), and carry no routing information, so you would need a lookup from id to shard for every hydration. A 64-bit time-ordered id with an embedded shard is smaller in every index and cache and routes for free.
  3. How do you run a schema migration across thousands of shards?
    Make changes backwards compatible (add nullable columns, backfill, then switch reads), and roll them out shard by shard with automation that can pause and resume. The application must tolerate both schemas during the rollout. Never make a change that requires all shards to flip at the same instant.

Close · 5 min

Ask what breaks first at ten times the load, and what they would build next. Then give them your read: one thing that was strong, one thing that was missing, one thing to practise. Be specific; “good job” helps nobody.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.