hard general · part of Practice Questions · Senior SWE Roadmap · topic form: Message Queues & Event-Driven Architecture

Requirements to clarify

  • Functional: producers publish messages to topics/queues, consumers subscribe and process them, in order (at least per-partition) and without loss.
  • Non-functional: high throughput, durability (messages survive broker crashes), horizontal scalability, configurable delivery guarantees.

Core components

  • Partitioning: a topic is split into partitions, each an ordered, append-only log — ordering is only guaranteed within a partition, which is what allows parallel consumption across partitions.
  • Replication: each partition is replicated across multiple brokers (leader + followers) for durability; a broker failure triggers leader election among the in-sync replicas.
  • Consumer groups: multiple consumers in a group split the partitions among themselves, each partition consumed by exactly one consumer in the group at a time — this is how horizontal consumer scaling works.
  • Offset tracking: consumers track how far they’ve read (an offset) so they can resume after a crash without reprocessing everything (or, for at-least-once, resume slightly behind and reprocess a few messages).

Key tradeoffs

  • At-least-once (simple, occasional duplicates, requires idempotent consumers) vs exactly-once (needs transactional writes/dedup, meaningfully more complex) — most systems default to at-least-once and push idempotency onto the consumer.

Approach / Notes

The best interview answer starts by naming the contract the queue must provide, then showing how the internal pieces preserve that contract under failure.

High-level architecture

flowchart LR
	P[Producers] --> B[(Broker cluster)]
	B --> R1[Partition replicas]
	B --> C[Consumer groups]
	B --> D[(DLQ)]

Producers append to topics. The broker shards each topic into partitions, replicates those partitions across brokers, and lets consumer groups read them independently. The queue’s job is to preserve ordering where promised, survive broker failure, and let throughput scale horizontally.

Core data flow

  1. A producer writes a message to a topic with a partition key.
  2. The broker chooses the partition for that key and appends the message to the partition log.
  3. The leader replica replicates the record to followers before acknowledging it.
  4. A consumer group reads from the partition, processes the message, and commits the offset.
  5. If processing fails repeatedly, the message moves to a dead-letter queue.
sequenceDiagram
	participant Prod as Producer
	participant Lead as Partition leader
	participant F1 as Follower 1
	participant Cons as Consumer

	Prod->>Lead: publish event(key=user42)
	Lead->>F1: replicate record
	F1-->>Lead: ack
	Lead-->>Prod: ack write
	Cons->>Lead: fetch next offset
	Lead-->>Cons: message
	Cons->>Lead: commit offset after success

What must be solved

  • Durability: a committed message survives broker crashes.
  • Ordering: messages with the same key stay in order within a partition.
  • Scalability: more partitions and more consumers increase throughput.
  • Failure handling: retries, redelivery, and DLQs keep poison messages from blocking the stream.

Partitioning and ordering example

If you partition by order_id, all events for one order stay ordered even if unrelated orders are spread across partitions.

flowchart TD
	O1[order-101 events] --> P0[Partition 0]
	O2[order-102 events] --> P1[Partition 1]
	O3[order-103 events] --> P0

That is the practical compromise: per-key ordering is usually enough, global ordering is too expensive.

Consumer-group scaling example

If a topic has 6 partitions and a consumer group has 3 workers, each worker gets 2 partitions. Add a 4th worker and it may still sit partially idle if partitions cannot be split further.

flowchart LR
	P0[Partition 0] --> C1[Worker 1]
	P1[Partition 1] --> C1
	P2[Partition 2] --> C2[Worker 2]
	P3[Partition 3] --> C2
	P4[Partition 4] --> C3[Worker 3]
	P5[Partition 5] --> C3

Delivery guarantees

  • At-most-once: fast, but messages can be lost.
  • At-least-once: default for real systems; duplicates are possible, so consumers must be idempotent.
  • Exactly-once: approximated with deduplication and transactional writes; true end-to-end exactly-once is not realistic across external side effects.

Practical answer shape

If asked live, say: I would build a partitioned, replicated log with producer acks, consumer-group offsets, retries, and a DLQ. I would accept at-least-once semantics and make consumers idempotent, because that is the simplest design that survives real failures without losing messages.