SysDesignPrep.com
Study guide 173 of 183

Distributed tracing and structured logging

Following one request through dozens of services: traces, spans and context propagation, OpenTelemetry, sampling strategies including tail-based sampling, correlating logs, metrics and traces, structured logging, log pipelines, cardinality and cost.

Reading is half of it. See this used in a real interview: walk through Design a Metrics and Monitoring System →

In a monolith, a slow request leaves a stack trace in one log. In a system of fifty services, the same request touches a gateway, an auth service, three backends, two databases, a cache and a queue, each with its own logs. Without tracing, "why was this checkout slow?" becomes guesswork. Distributed tracing and well-structured logs are how teams debug modern systems, and they come up in monitoring designs and in any discussion of operating microservices.

Traces and spans

  • A trace represents one request's journey through the system, identified by a trace id.
  • A span is one operation within it (an HTTP call, a database query, a queue publish), with a start time, duration, service name, attributes and a parent span id.
  • Spans form a tree, displayed as a timeline (waterfall) that shows exactly where time went: which calls ran in sequence, which in parallel, and which were slow or failed.

Context propagation

The trace id must travel with the request across every boundary:

  • HTTP and gRPC calls carry it in headers (the W3C traceparent standard).
  • Messages put on queues carry it in message headers, so consumers continue the trace (often as a linked span, since processing happens later).
  • Thread pools and async code must pass the context along; losing it breaks the trace.

OpenTelemetry provides vendor-neutral SDKs that instrument common frameworks automatically and export spans to any backend (Jaeger, Tempo, Zipkin or commercial tools).

Sampling

Tracing every request at high traffic is expensive to send and store. Sampling options:

StrategyHowTrade-off
Head-baseddecide at the start (keep 1 %), propagate the decisioncheap; misses most rare errors and slow requests
Tail-basedbuffer all spans, decide after the trace completeskeeps every error and slow trace; needs a collector tier with memory
Rate-limited per endpointat most N traces per second per routecovers low-traffic endpoints fairly
Always for some100 % for errors, specific users or debug flagstargeted investigation

Tail-based sampling ("keep all traces with errors or above p99 latency, plus 1 % of the rest") gives the most useful data per dollar. See cost-aware system design.

Structured logging

Logs should be machine-parseable events, not free-form text:

{"ts":"2026-10-05T12:00:01Z","level":"error","service":"payments","trace_id":"4bf9...","order_id":"o_981","msg":"card declined","code":"insufficient_funds"}
  • Consistent field names across services.
  • Trace id in every log line, so you can jump from a trace to the logs of that request, and back.
  • Log levels used consistently; no debug noise in production by default.
  • No personal data or secrets in logs. See privacy and data deletion.

Log pipelines

Agents on each host ship logs to a pipeline (often via Kafka) that parses, enriches, samples or drops noisy logs, and writes to search storage for recent data and cheap object storage for long retention. Volume is the main cost: control it with levels, sampling of repetitive logs and retention tiers. See how Elasticsearch works.

Connecting the three signals

  • Metrics say something is wrong (error rate up, p99 up). See SLIs, SLOs and error budgets.
  • Traces show where in the call graph it is wrong. Exemplars attach example trace ids to metric data points, linking a latency spike directly to slow traces.
  • Logs explain why, for that specific request.

Shared attributes (service, version, region, trace id) make jumping between them one click. See observability and operations.

Cardinality

Metrics with high-cardinality labels (user id, order id) explode storage costs. Put high-cardinality details in traces and logs, which are designed for per-request data, and keep metric labels bounded (service, endpoint, status, region). See Design a Monitoring System.

Using traces beyond debugging

  • Service dependency maps built automatically from spans.
  • Critical path analysis: which calls actually determine end-to-end latency. See tail latency.
  • Capacity insight: which callers drive load on a service.

In the interview

For a microservice design such as Design DoorDash: "every request gets a trace id at the gateway, propagated through RPC headers and queue messages via OpenTelemetry; we use tail-based sampling to keep errors and slow traces; structured logs carry the trace id; metrics link to traces via exemplars." For Design a Monitoring System, discuss trace storage and sampling as part of the design.

Checklist

  • Trace ids created at the edge and propagated through every hop, including queues.
  • OpenTelemetry instrumentation and a neutral export pipeline.
  • Tail-based or hybrid sampling that keeps errors and slow requests.
  • Structured logs with consistent fields and trace ids; no personal data.
  • Log pipelines with sampling, retention tiers and cost control.
  • Bounded metric labels; high-cardinality detail in traces and logs.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.