SysDesignPrep.com
Study guide 20 of 183

Kubernetes for system design interviews

The Kubernetes concepts that come up in system design: pods, deployments, services and ingress, autoscaling of pods and nodes, health probes, resource requests and limits, stateful sets, jobs, rolling updates, multi-zone scheduling, and what Kubernetes does not solve for you.

Reading is half of it. See this used in a real interview: walk through Design a URL Shortener →

Most companies run their services on Kubernetes, so interviewers sometimes ask how a design would be deployed, scaled or kept healthy on it. You do not need to be a cluster administrator; you need to know the main building blocks, how they map to the boxes in your diagram, and what Kubernetes handles versus what your design still must handle.

The building blocks

ObjectWhat it isIn your design
Podone or more containers scheduled together, with shared network and storageone instance of a service
Deploymentkeeps N identical pods running; handles rolling updatesa stateless service tier
Servicestable virtual IP and DNS name load-balancing across matching podshow other services find this one
Ingress or GatewayHTTP routing from outside into servicesthe edge entry point
ConfigMap and Secretconfiguration and credentials injected into podssettings and keys
StatefulSetpods with stable identities and their own persistent volumesdatabases, brokers, stateful gateways
Job and CronJobrun-to-completion work, once or on a schedulebatch tasks, periodic jobs
DaemonSetone pod per nodelog and metrics agents
Namespacea scope for names, quotas and access policiesteams or environments

Scheduling and resources

Each container declares requests (guaranteed CPU and memory, used for placement) and limits (the maximum; exceeding the memory limit kills the container, exceeding CPU throttles it). The scheduler places pods on nodes with enough free requested capacity.

  • Set requests from measured usage; too low causes noisy neighbours, too high wastes nodes.
  • Anti-affinity and topology spread place replicas across nodes and zones, so one failure does not take out a whole service.
  • Taints and node pools separate workloads (GPU nodes for inference, isolated nodes for untrusted code). See running untrusted code safely.

Health probes

  • Liveness probe: is the container stuck? Failing restarts it.
  • Readiness probe: can it serve traffic now? Failing removes it from the service's endpoints without restarting.
  • Startup probe: gives slow-starting apps time before liveness checks begin.

Readiness should reflect real ability to serve (warm caches, connections ready), while liveness should be simple, so a slow dependency does not cause restart loops.

Autoscaling

  • Horizontal Pod Autoscaler: adjusts replica count based on CPU, memory or custom metrics (queue depth, requests per second).
  • Cluster autoscaler or node provisioners: add nodes when pods cannot be scheduled, remove idle ones.
  • Event-driven autoscalers scale workers on queue length, including to zero.
  • Node provisioning takes minutes, so keep headroom for spikes. See autoscaling.

Rolling updates

A Deployment replaces pods gradually (maxUnavailable, maxSurge), waiting for readiness before continuing. Combined with preStop hooks and graceful shutdown (stop accepting work, finish in-flight requests), deploys cause no errors. Canary and blue-green need extra tooling (a service mesh or progressive delivery controllers). See deployment strategies.

Pod disruption budgets limit how many replicas can be down during voluntary disruptions such as node upgrades.

Networking

  • Every pod gets an IP; Services give stable names (payments.prod.svc) via cluster DNS.
  • Network policies restrict which pods can talk to which.
  • A service mesh adds mTLS, retries and traffic shifting. See service discovery and service mesh.

State

Kubernetes can run databases with StatefulSets and persistent volumes, but operating them well (backups, failover, upgrades) is hard. Many teams run stateless services on Kubernetes and use managed databases, caches and queues. If you run stateful systems in the cluster, use mature operators that automate their lifecycle. See stateful vs stateless services.

Jobs and batch work

Jobs run to completion with retries; CronJobs schedule them. Useful for periodic tasks and batch processing, though large-scale scheduling and workflows usually need a dedicated system on top. See delayed jobs and distributed cron and Design a Job Scheduler.

What Kubernetes does not solve

  • Your application's correctness under failure: retries, idempotency, timeouts, data consistency.
  • Database scaling and replication design.
  • Multi-region architecture (clusters are usually per region; global routing is separate). See multi-region architecture.
  • Capacity planning and cost.
  • Observability (it provides hooks; you still need metrics, logs and traces).

In the interview

One or two sentences are usually enough: "Each stateless service is a Deployment spread across three zones, behind a Service; the HPA scales on requests per second, with the cluster autoscaler adding nodes; readiness probes and graceful shutdown make rolling deploys safe; databases are managed services outside the cluster." For GPU inference (Design LLM Inference) or sandboxed execution (Design LeetCode), mention dedicated node pools.

Checklist

  • Deployments for stateless tiers; Services for discovery; Ingress at the edge.
  • Requests and limits set from measurements; spread across zones.
  • Liveness, readiness and startup probes used correctly.
  • HPA on meaningful metrics plus node autoscaling with headroom.
  • Rolling updates with graceful shutdown and disruption budgets.
  • Managed or operator-run stateful systems; application resilience still your job.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.