"Lambda vs Kappa architecture"
Two designs for combining real-time and historical data processing: Lambda with separate batch and speed layers merged at query time, Kappa with a single replayable stream, their trade-offs in correctness, complexity and cost, and how modern stream processors and lakehouses changed the choice.
Reading is half of it. See this used in a real interview: walk through Design an Ad Click Aggregator →
Many systems need numbers that are both fresh (clicks in the last minute) and correct (billing totals for the month). Batch jobs are accurate but slow; stream processing is fast but historically less reliable. Lambda and Kappa are two well-known architectures for getting both. The names come up often in analytics and aggregation interviews, so it helps to know what each means and when each makes sense today.
Lambda architecture
Three layers:
- Batch layer: stores the immutable master dataset (all raw events) and periodically recomputes views from scratch (for example, nightly).
- Speed layer: processes new events in real time to produce incremental views covering the gap since the last batch run.
- Serving layer: answers queries by merging the batch view (accurate, up to the last run) with the speed view (approximate, recent).
When the next batch run completes, its results replace the speed layer's for that period, correcting any errors.
Strengths: the batch layer is the source of truth, so bugs or approximations in the speed layer are corrected automatically; reprocessing is just rerunning the batch.
Weaknesses: two codebases implementing the same logic (often in different frameworks), which drift apart; merging views adds complexity; operating two pipelines costs more.
Kappa architecture
One pipeline:
- All data flows through a durable, replayable log (Kafka with long retention or tiered storage). See how Kafka works.
- A stream processor computes all views.
- To fix a bug or change logic, deploy a new version of the job that replays the log from the beginning into a new output table; switch readers to it when it catches up; delete the old one.
Strengths: one codebase and one processing model; reprocessing uses the same code as real-time.
Weaknesses: replaying months or years of events through a stream job can be slow and expensive; requires long log retention; complex historical analyses can be awkward in streaming form.
Comparison
| Lambda | Kappa | |
|---|---|---|
| Pipelines | batch plus speed | stream only |
| Codebases | two (same logic twice) | one |
| Source of truth | batch over the master dataset | the replayable log |
| Correcting errors | next batch run | replay with a fixed job |
| Freshness | real time (approximate) plus batch (exact) | real time |
| Operational cost | higher | lower, except during replays |
| Best when | exact batch results are required (billing), complex historical jobs | logic fits streaming, replay is affordable |
What changed
Modern tools blurred the line:
- Stream processors such as Flink offer exactly-once state, event-time windows and watermarks, making streaming results trustworthy. See windowing and watermarks.
- Unified APIs (Spark Structured Streaming, Beam, Flink SQL) run the same logic in batch and streaming modes, removing the two-codebase problem.
- Lakehouse tables (Iceberg, Delta) accept streaming writes and batch reads on the same data, with time travel for reprocessing. See data lakes and lakehouses.
Many teams now run a mostly Kappa-style streaming pipeline, writing to lakehouse tables, plus targeted batch jobs for reconciliation and heavy historical computation: a pragmatic hybrid.
Examples in interview questions
- Ad click aggregation: stream aggregates per minute for dashboards; a daily batch over raw clicks reconciles and produces billing numbers (Lambda-like for money). See Design an Ad Click Aggregator.
- Metrics: streaming ingestion and rollups; reprocessing rarely needed (Kappa-like). See Design a Monitoring System.
- Autocomplete rankings: batch rebuilds of top suggestions from logs plus a streaming layer for trending queries, merged at serving time (classic Lambda). See Design Typeahead.
- Leaderboards: real-time updates from the stream, with periodic recomputation from the score history to fix drift. See Design a Leaderboard.
Choosing in an interview
- If correctness must be audited (billing, payouts), keep a batch reconciliation over immutable raw data, whatever the real-time path is.
- If logic is the same for real-time and history, prefer one engine with unified APIs to avoid duplicate code.
- Keep raw events immutable and long-retained; that is what makes either architecture recoverable.
Say it simply: "Streaming for freshness, batch over raw events for the numbers we bill on; same logic via a unified API where possible." See batch and stream processing.
Checklist
- Lambda: batch for truth, speed layer for freshness, merged at serving.
- Kappa: one replayable log and stream jobs; reprocess by replay.
- Two-codebase drift is Lambda's main cost; replay time is Kappa's.
- Unified engines and lakehouse tables enable hybrids.
- Immutable, retained raw data in either case.
- Batch reconciliation wherever money is involved.