SysDesignPrep.com
Study guide 120 of 183

Microservices, service discovery and API gateways

Monolith versus microservices and when to split, service discovery, API gateways and backends for frontends, synchronous versus asynchronous communication, service meshes, and the failure modes that come with a distributed system.

Reading is half of it. See this used in a real interview: walk through Design DoorDash →

Almost every system design diagram is drawn as a set of services, so interviewers sometimes ask why. "Microservices are more scalable" is not an answer. The real reasons are organisational and operational, and the costs (network calls, partial failure, distributed data) are real. This guide covers how to decide where to draw service boundaries and how services find and talk to each other.

Monolith or microservices

A monolith is one deployable application, usually with one database. A microservice architecture splits it into services that are deployed, scaled and owned independently, each with its own data.

MonolithMicroservices
Calls between partsfunction calls: nanoseconds, never partially failnetwork calls: milliseconds, can time out or half-succeed
Transactionsone database transactionsagas and eventual consistency across services
Deployseverything togethereach team ships independently
Scalingscale the whole thingscale the hot parts (search, not settings)
Failure isolationa memory leak takes everything downone service degrades; others keep working (if designed to)
Operational costone thing to runservice discovery, tracing, many pipelines and dashboards

Split when different parts have very different scaling or reliability needs (the read path of a URL shortener versus link creation), when many teams are blocked on one deploy, or when a part needs a different technology (a GPU inference service). Do not split a small product with one team: a well-structured modular monolith is faster to build and easier to run.

In an interview, drawing separate services is fine, but justify the ones that matter: "redirects and link creation are separate because reads outnumber writes 100 to 1 and must stay up even if creation fails".

Drawing boundaries

Good boundaries follow business capabilities and data ownership: an order service owns orders, a payment service owns payments, and nobody else writes their tables. Signs of a bad boundary:

  • Two services that must always be deployed together.
  • A call chain of five services to serve one request (a "distributed monolith").
  • Services that share a database and write each other’s tables.

Each service owns its data and exposes it through an API or events. Other services that need it either call that API or keep their own read-optimised copy from events.

Service discovery

Instances come and go with deploys and autoscaling, so callers cannot use fixed addresses.

  • DNS-based (Kubernetes services): a name resolves to a stable virtual IP that load-balances across healthy pods. Simple and the default today.
  • Registry-based (Consul, etcd, Eureka): instances register with a health check; clients or proxies look up the current list.
  • Client-side load balancing: the caller holds the list and picks an instance (gRPC does this well), avoiding an extra hop.
  • Server-side: the caller talks to a load balancer, which picks.

Health checks decide who is in the list: liveness (restart me if this fails) and readiness (do not send me traffic yet).

API gateways and backends for frontends

An API gateway is the single public entry point: it terminates TLS, authenticates, rate limits, routes to services and can aggregate responses. It keeps cross-cutting concerns out of every service and hides internal topology from clients. See API design and rate limiting.

A backend for frontend (BFF) is a gateway layer per client type (mobile, web, partners) that shapes responses for that client: one call from the phone returns exactly the screen’s data, fetched from several services server-side, instead of six calls over a slow mobile network.

Synchronous or asynchronous

  • Synchronous (REST, gRPC): the caller waits. Right when the user needs the answer now (is this seat free?). Every hop adds latency and a failure mode.
  • Asynchronous (events through Kafka or a queue): the caller publishes and moves on. Right when the work can happen later (send a receipt, update search, recompute recommendations). It decouples availability: the publisher keeps working even if consumers are down. See message queues and streams.

A healthy design is mostly synchronous on the user’s critical path and asynchronous everywhere else. Long synchronous chains multiply failure: five services each at 99.9 % availability give 99.5 % for the chain.

Living with partial failure

Every network call can be slow, fail, or succeed without you hearing back. The standard defences:

  • Timeouts on every call, shorter than the caller’s own deadline.
  • Retries with backoff and jitter, only for idempotent operations, with a retry budget so retries do not multiply load.
  • Circuit breakers that stop calling a failing dependency and fail fast.
  • Bulkheads: separate pools per dependency so one slow service cannot consume all threads.
  • Fallbacks: serve cached or default data when a non-critical dependency fails.
  • Idempotency keys so retries of writes are safe. See distributed transactions and idempotency.

Service mesh

A service mesh (Istio, Linkerd) runs a proxy next to every service instance (a sidecar) that handles mutual TLS, retries, timeouts, load balancing and telemetry, configured centrally. It removes that code from every service, at the cost of more moving parts and a little latency per hop. Worth mentioning for large fleets with many languages; overkill for a handful of services.

Observability is not optional

With dozens of services, "why is this request slow" needs distributed tracing (a trace id propagated through every call), structured logs with that id, and per-service metrics on latency, errors and saturation. See observability, operations and rollouts.

Checklist

  • Why each service boundary exists (scaling, reliability, ownership).
  • Each service owns its data; no shared tables.
  • Discovery and load balancing between services.
  • A gateway at the edge; a BFF if clients differ a lot.
  • Synchronous only on the critical path; events elsewhere.
  • Timeouts, retries with idempotency, circuit breakers.
  • Tracing across services.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.