SysDesignPrep.com
Study guide 122 of 183

Service discovery and service mesh

How services find and talk to each other in a dynamic fleet: DNS and registries, client-side versus server-side discovery, health checks, sidecar proxies, what a service mesh adds (mTLS, retries, traffic shifting, telemetry) and when it is worth the cost.

Reading is half of it. See this used in a real interview: walk through Design DoorDash →

In a fleet of containers that scale up and down, move between hosts and get replaced on every deploy, "call the payments service" raises questions: at which address, which instance, over what protocol, with which retries and timeouts, and is it really the payments service? Service discovery answers where; a service mesh standardises how. Both come up when a design has many microservices.

Service discovery

Instances register themselves (or are registered by the platform) in a registry, and callers look them up.

MechanismHow it worksNotes
DNSthe service name resolves to instance IPssimple and universal; caching and TTLs make updates slow
Kubernetes Servicesa stable virtual IP and DNS name, kept current by the platformthe default inside Kubernetes
Registry (Consul, etcd, Eureka)instances register with health checks; clients query or watchfast updates, richer metadata
Cloud load balancera managed load balancer per servicesimple; an extra hop and cost

Health checks keep the list honest: instances failing liveness or readiness checks are removed, and instances shutting down deregister first and drain in-flight requests.

Client-side versus server-side

  • Server-side discovery: the caller sends to a load balancer, which picks an instance. Clients stay simple; the balancer is an extra hop. See load balancing.
  • Client-side discovery: the caller gets the instance list and picks one itself (round robin, least requests, zone-aware). One fewer hop and smarter balancing, but every language needs a library that does it well.

A sidecar proxy gives you the client-side benefits without per-language libraries.

What a service mesh adds

A service mesh (Istio, Linkerd, Consul Connect, or Envoy-based setups) puts a proxy next to every service instance (a sidecar, or a per-node proxy in newer designs). All traffic flows through these proxies, configured by a central control plane. Features, without changing application code:

  • mTLS everywhere: encrypted, mutually authenticated traffic with automatically rotated certificates, and policies like "only checkout may call payments". See security and auth.
  • Resilience: timeouts, retries with budgets, circuit breaking and outlier ejection, applied consistently. See rate limiting and resilience.
  • Traffic management: canary releases ("5 % to v2"), header-based routing, mirroring traffic to a new version, failover across zones.
  • Telemetry: request rates, errors, latency and traces for every call, uniformly. See observability and operations.

The costs

  • Latency: each hop through a proxy adds a little, typically well under a millisecond, but it adds up across deep call chains.
  • Resources: a proxy per instance costs CPU and memory across the fleet.
  • Complexity: another distributed system to run, upgrade and debug; misconfigured retries in the mesh can amplify outages.

A mesh pays off with many services, several languages, and strong security or compliance needs. With a handful of services, a library plus a load balancer is simpler.

Retries done right

Retries at every layer multiply: three layers each retrying three times turn one failure into 27 requests. Retry at one layer, only idempotent requests, with exponential backoff, jitter and a retry budget (for example, at most 10 % extra traffic). Meshes make the budget easy to enforce.

Edge versus internal traffic

The API gateway handles north-south traffic (clients to the system): authentication, rate limiting, request shaping. The mesh handles east-west traffic (service to service). See API gateways.

In the interview

You rarely need to design a mesh. One sentence is usually enough: "services discover each other through Kubernetes DNS; a mesh provides mTLS, timeouts, retries with budgets, and per-call metrics, and lets us canary new versions." In large microservice designs like Design DoorDash or Design Netflix it shows you know how the pieces actually connect.

Checklist

  • A registry or platform DNS with health checks and graceful deregistration.
  • Client-side balancing (via library or sidecar) or server-side load balancers.
  • mTLS and service-to-service authorization.
  • Timeouts, budgeted retries and circuit breakers in one place.
  • Canary and traffic shifting for safe deploys.
  • A mesh only when the number of services justifies its cost.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.