Message ordering and sequence numbers
Keeping messages and events in the right order: why global order is expensive, per-conversation and per-key ordering, server-assigned sequence numbers, client ids and timestamps, detecting gaps, reordering buffers, ordering with partitions and retries, and how chat and trading systems do it.
Reading is half of it. See this used in a real interview: walk through Design WhatsApp →
Users notice when order breaks: a reply appears above the question, a "deleted" message comes back, a cancel arrives before the order it cancels. Ordering is cheap within one machine and expensive across many, so good designs choose what must be ordered (usually per conversation, per account or per instrument) and enforce exactly that. Chat, trading and collaboration questions all probe this.
Choose the scope of order
- Global total order: every message in the system in one sequence. Requires a single sequencer or consensus on every write; limits throughput. Rarely needed.
- Per key order: messages within one conversation, account, document or order book are ordered; different keys are independent. Scales horizontally by partitioning on that key. Almost always what products need.
- Causal order: a reply comes after what it replies to, even across keys. Achieved with references or logical clocks. See logical clocks.
Server-assigned sequence numbers
The simplest robust approach for chat:
- All messages for a conversation go through the conversation's owner (a partition of the message service, chosen by hashing the conversation id).
- The owner assigns a monotonically increasing sequence number per conversation when it persists the message.
- Clients display messages sorted by sequence number.
Because one partition owns each conversation, assigning numbers is a local increment, not a distributed agreement. Group chats and channels work the same way. See Design WhatsApp and Design Slack.
Store messages keyed by (conversation_id, sequence), which also makes pagination and sync trivial: "give me messages after 1042". See data modelling for Cassandra and DynamoDB.
Client ids and timestamps
- Client-generated message ids (UUIDs) make sends idempotent: a retried send is recognised and not stored twice. See idempotency keys in practice.
- Client timestamps are for display ("sent at 10:42"), not ordering; device clocks are often wrong.
- Optimistic display: the sender's client shows the message immediately in a pending state, then places it at its server-assigned position when acknowledged.
Gaps and reordering
Messages can arrive at a client out of order (push notification first, then the socket catch-up, retries, multiple devices):
- Sequence numbers make gaps visible: if a client has 1040 and receives 1043, it fetches 1041 and 1042 before displaying, or displays with a placeholder.
- A small reordering buffer holds early messages briefly.
- On reconnect, the client asks for everything after its last contiguous sequence number. See presence and connection management.
Ordering in queues and streams
- Kafka orders records within a partition; partition by the key whose order matters (conversation id, account id). See how Kafka works.
- Consumers must process a partition sequentially for ordered keys; parallelising inside a partition breaks order unless you shard by key within the consumer.
- Retries break order: if message 5 fails and 6 proceeds, 6 overtakes 5. Either block the key until 5 succeeds, or make handlers tolerant (versions, sequence checks that discard older updates). See delivery semantics.
- Repartitioning (adding partitions) changes key placement and can reorder in-flight keys; plan it carefully.
Edits, deletes and reactions
Edits and deletes are new events referencing the original message id, with their own sequence numbers. Clients apply them in sequence order; a delete arriving before its message (rare, but possible across channels) is held until the message appears, or recorded as a tombstone so the message never shows. Using versions per message ensures the latest edit wins regardless of arrival order.
Trading and other strict systems
Exchanges need a total order per instrument for fairness: a sequencer stamps every order with a sequence number before matching, and the matching engine processes strictly in that order. Recovery replays the sequenced log. See low-latency systems and matching engines and Design a Stock Exchange.
Collaborative editing
Edits to one document are ordered by the document's session server (OT) or merged without global order (CRDTs). Either way, per-document scope keeps ordering cheap. See CRDTs vs operational transformation and Design Google Docs.
In the interview
"Each conversation is owned by one partition of the message service, which assigns a per-conversation sequence number when persisting; clients send messages with client-generated ids for idempotency, display by sequence number, detect gaps and fetch missing ranges on reconnect. Kafka topics downstream are partitioned by conversation id so consumers see each conversation in order."
Checklist
- Order only within the scope that matters (conversation, account, instrument).
- A single owner per key assigns sequence numbers.
- Client ids for idempotency; client timestamps only for display.
- Gap detection, reordering buffers and catch-up by sequence.
- Partition streams by the ordering key; handle retries without overtaking.
- Edits and deletes as sequenced events referencing the original.