Service discovery and service mesh
How services find and talk to each other in a dynamic fleet: DNS and registries, client-side versus server-side discovery, health checks, sidecar proxies, what a service mesh adds (mTLS, retries, traffic shifting, telemetry) and when it is worth the cost.
Reading is half of it. See this used in a real interview: walk through Design DoorDash →
In a fleet of containers that scale up and down, move between hosts and get replaced on every deploy, "call the payments service" raises questions: at which address, which instance, over what protocol, with which retries and timeouts, and is it really the payments service? Service discovery answers where; a service mesh standardises how. Both come up when a design has many microservices.
Service discovery
Instances register themselves (or are registered by the platform) in a registry, and callers look them up.
| Mechanism | How it works | Notes |
|---|---|---|
| DNS | the service name resolves to instance IPs | simple and universal; caching and TTLs make updates slow |
| Kubernetes Services | a stable virtual IP and DNS name, kept current by the platform | the default inside Kubernetes |
| Registry (Consul, etcd, Eureka) | instances register with health checks; clients query or watch | fast updates, richer metadata |
| Cloud load balancer | a managed load balancer per service | simple; an extra hop and cost |
Health checks keep the list honest: instances failing liveness or readiness checks are removed, and instances shutting down deregister first and drain in-flight requests.
Client-side versus server-side
- Server-side discovery: the caller sends to a load balancer, which picks an instance. Clients stay simple; the balancer is an extra hop. See load balancing.
- Client-side discovery: the caller gets the instance list and picks one itself (round robin, least requests, zone-aware). One fewer hop and smarter balancing, but every language needs a library that does it well.
A sidecar proxy gives you the client-side benefits without per-language libraries.
What a service mesh adds
A service mesh (Istio, Linkerd, Consul Connect, or Envoy-based setups) puts a proxy next to every service instance (a sidecar, or a per-node proxy in newer designs). All traffic flows through these proxies, configured by a central control plane. Features, without changing application code:
- mTLS everywhere: encrypted, mutually authenticated traffic with automatically rotated certificates, and policies like "only checkout may call payments". See security and auth.
- Resilience: timeouts, retries with budgets, circuit breaking and outlier ejection, applied consistently. See rate limiting and resilience.
- Traffic management: canary releases ("5 % to v2"), header-based routing, mirroring traffic to a new version, failover across zones.
- Telemetry: request rates, errors, latency and traces for every call, uniformly. See observability and operations.
The costs
- Latency: each hop through a proxy adds a little, typically well under a millisecond, but it adds up across deep call chains.
- Resources: a proxy per instance costs CPU and memory across the fleet.
- Complexity: another distributed system to run, upgrade and debug; misconfigured retries in the mesh can amplify outages.
A mesh pays off with many services, several languages, and strong security or compliance needs. With a handful of services, a library plus a load balancer is simpler.
Retries done right
Retries at every layer multiply: three layers each retrying three times turn one failure into 27 requests. Retry at one layer, only idempotent requests, with exponential backoff, jitter and a retry budget (for example, at most 10 % extra traffic). Meshes make the budget easy to enforce.
Edge versus internal traffic
The API gateway handles north-south traffic (clients to the system): authentication, rate limiting, request shaping. The mesh handles east-west traffic (service to service). See API gateways.
In the interview
You rarely need to design a mesh. One sentence is usually enough: "services discover each other through Kubernetes DNS; a mesh provides mTLS, timeouts, retries with budgets, and per-call metrics, and lets us canary new versions." In large microservice designs like Design DoorDash or Design Netflix it shows you know how the pieces actually connect.
Checklist
- A registry or platform DNS with health checks and graceful deregistration.
- Client-side balancing (via library or sidecar) or server-side load balancers.
- mTLS and service-to-service authorization.
- Timeouts, budgeted retries and circuit breakers in one place.
- Canary and traffic shifting for safe deploys.
- A mesh only when the number of services justifies its cost.