medium general · part of Practice Questions · Senior SWE Roadmap

Requirements to clarify

  • Functional: send notifications across multiple channels (push, email, SMS), triggered by events, possibly with user preferences/opt-outs.
  • Non-functional: must not lose notifications, must not spam (dedup/rate-limit), must handle bursty demand (e.g., a breaking-news push to millions at once), third-party delivery providers can be slow/rate-limited themselves.

Core components

  • Event ingestion: upstream services publish “notify user X about Y” events onto a message queue, decoupling notification sending from the triggering action (see Message Queues & Event-Driven Architecture).
  • Notification service: consumes events, checks user preferences (which channels, opted in/out), applies dedup/rate-limiting rules, and routes to the appropriate channel-specific sender.
  • Channel-specific senders: separate workers/queues per channel (push via APNs/FCM, email via an ESP, SMS via a provider) — isolates one channel’s slowness/failures from the others.
  • Retry & backoff: failed sends go through exponential backoff retry, with a dead-letter queue for persistent failures.

Key tradeoffs

  • At-least-once delivery (risk of duplicate notifications) is usually accepted over risking lost notifications — idempotency keys on the client side handle the rare duplicate.

Approach / Notes

Think of the system as a fan-out pipeline with reliability baked in at every step. The hard part is not “sending an email” or “calling FCM”; it’s turning one business event into many channel-specific deliveries without losing messages, spamming users, or letting a slow provider back up the whole app.

High-level architecture

flowchart LR
	E[Upstream business event
(order shipped, comment replied, etc.)] --> Q[(Event queue / bus)]
	Q --> N[Notification orchestrator]
	N --> P[(Preferences store/cache)]
	N --> D[(Dedup / delivery log)]
	N --> CP[(Channel priority queues)]
	CP --> E1[Email workers]
	CP --> P1[Push workers]
	CP --> S1[SMS workers]
	E1 --> ESP[Email provider]
	P1 --> APNS[APNs / FCM]
	S1 --> SMSP[SMS provider]

The orchestrator owns policy: which users should be notified, on which channels, at what priority, and whether this event has already been handled. Channel workers own transport: format the payload, call the provider, retry on transient failures, and record the final outcome.

Core data model

  • Notification event: the incoming business event, usually keyed by event_id, user_id, type, payload, priority, and created_at.
  • User preferences: per-user channel opt-ins, quiet hours, locale, device tokens, email addresses, SMS numbers, and per-type overrides.
  • Delivery record: one row per (event_id, user_id, channel) with status pending, sent, failed, or suppressed.
  • Provider attempt log: retry metadata, provider response codes, backoff state, and last error for observability/debugging.

Persisting the delivery record before or alongside enqueueing is what prevents “we accepted the event, then crashed before sending” from becoming a lost notification.

End-to-end flow

  1. A source service emits an event like OrderShipped or MentionCreated.
  2. The notification service consumes it, normalizes the payload, and writes a durable delivery record.
  3. The orchestrator checks preferences, quiet hours, and dedup rules.
  4. Eligible deliveries are fanned out into per-channel queues.
  5. Channel workers send through the external provider and store the result.
  6. Failures that look transient are retried with exponential backoff; permanent failures go to a dead-letter queue.
sequenceDiagram
	participant O as Order Service
	participant Q as Event Queue
	participant N as Notification Orchestrator
	participant P as Preferences Store
	participant C as Channel Queue
	participant W as Worker
	participant X as Provider

	O->>Q: OrderShipped(event_id=123)
	Q->>N: consume event
	N->>P: load user preferences
	P-->>N: push+email enabled, SMS off
	N->>N: dedup check event_id=123
	N->>C: enqueue email + push jobs
	C->>W: deliver job
	W->>X: send notification
	X-->>W: 200 OK
	W->>N: mark sent

Reliability choices

  • At-least-once processing is the safe default. It is better to send twice than not at all for most notifications.
  • Idempotency keys on (event_id, user_id, channel) stop duplicate delivery when retries happen.
  • Transactional outbox fits well if the notification event originates from a service that also writes a database row. Write the business change and outbox entry in one DB transaction, then relay the outbox to the queue.
  • Channel isolation matters. A slow SMS provider should not delay push notifications; separate queues and worker pools keep one failure domain from contaminating the others.

Scaling and burst handling

  • Use a broker/queue to absorb spikes, then let workers drain at provider-safe rates.
  • Partition by user_id or channel depending on which ordering property matters more.
  • Scale workers horizontally off queue depth and provider latency, not CPU alone.
  • Use provider-specific throttles so you stay under APNs/FCM/ESP/SMS quotas.

The right mental model is “load leveling with policy in the middle”: the queue protects the backend from bursts, while the orchestrator decides who should actually receive the notification.

Edge cases and tradeoffs

  • Duplicates vs loss: duplicates are usually cheaper than loss, but some categories (password reset, payment confirmation) need stricter idempotency and sometimes one-time tokens.
  • User preferences vs compliance: users can opt out of marketing, but transactional alerts may still be required.
  • Realtime vs digest: some notifications should be immediate, while others can be batched into hourly or daily digests.
  • Ordering: per-user ordering is usually enough; global ordering is not worth the throughput cost.
  • Retention: keep delivery history long enough for audit/support, but archive old attempts so the hot path stays small.

Good interview follow-ups

  • How do you prevent the same event from notifying the same user twice?
  • What happens if APNs or the email provider is down for 30 minutes?
  • How do you support quiet hours and time zones?
  • How do you ensure one noisy channel cannot starve the others?
  • What do you do with undeliverable addresses or permanently bounced emails?

Practical answer shape

If asked to design this live, the clean answer is: ingest events into a durable queue, persist a delivery record, evaluate user preferences and dedup keys, fan out into separate per-channel queues, and have isolated workers retry with backoff and DLQs. That gives you reliability, burst absorption, and provider isolation without forcing the user-facing request to wait on third-party delivery.