medium general · part of Practice Questions · Senior SWE Roadmap
Requirements to clarify
- Functional: send notifications across multiple channels (push, email, SMS), triggered by events, possibly with user preferences/opt-outs.
- Non-functional: must not lose notifications, must not spam (dedup/rate-limit), must handle bursty demand (e.g., a breaking-news push to millions at once), third-party delivery providers can be slow/rate-limited themselves.
Core components
- Event ingestion: upstream services publish “notify user X about Y” events onto a message queue, decoupling notification sending from the triggering action (see Message Queues & Event-Driven Architecture).
- Notification service: consumes events, checks user preferences (which channels, opted in/out), applies dedup/rate-limiting rules, and routes to the appropriate channel-specific sender.
- Channel-specific senders: separate workers/queues per channel (push via APNs/FCM, email via an ESP, SMS via a provider) — isolates one channel’s slowness/failures from the others.
- Retry & backoff: failed sends go through exponential backoff retry, with a dead-letter queue for persistent failures.
Key tradeoffs
- At-least-once delivery (risk of duplicate notifications) is usually accepted over risking lost notifications — idempotency keys on the client side handle the rare duplicate.
Approach / Notes
Think of the system as a fan-out pipeline with reliability baked in at every step. The hard part is not “sending an email” or “calling FCM”; it’s turning one business event into many channel-specific deliveries without losing messages, spamming users, or letting a slow provider back up the whole app.
High-level architecture
flowchart LR E[Upstream business event (order shipped, comment replied, etc.)] --> Q[(Event queue / bus)] Q --> N[Notification orchestrator] N --> P[(Preferences store/cache)] N --> D[(Dedup / delivery log)] N --> CP[(Channel priority queues)] CP --> E1[Email workers] CP --> P1[Push workers] CP --> S1[SMS workers] E1 --> ESP[Email provider] P1 --> APNS[APNs / FCM] S1 --> SMSP[SMS provider]
The orchestrator owns policy: which users should be notified, on which channels, at what priority, and whether this event has already been handled. Channel workers own transport: format the payload, call the provider, retry on transient failures, and record the final outcome.
Core data model
- Notification event: the incoming business event, usually keyed by
event_id,user_id,type,payload,priority, andcreated_at. - User preferences: per-user channel opt-ins, quiet hours, locale, device tokens, email addresses, SMS numbers, and per-type overrides.
- Delivery record: one row per
(event_id, user_id, channel)with statuspending,sent,failed, orsuppressed. - Provider attempt log: retry metadata, provider response codes, backoff state, and last error for observability/debugging.
Persisting the delivery record before or alongside enqueueing is what prevents “we accepted the event, then crashed before sending” from becoming a lost notification.
End-to-end flow
- A source service emits an event like
OrderShippedorMentionCreated. - The notification service consumes it, normalizes the payload, and writes a durable delivery record.
- The orchestrator checks preferences, quiet hours, and dedup rules.
- Eligible deliveries are fanned out into per-channel queues.
- Channel workers send through the external provider and store the result.
- Failures that look transient are retried with exponential backoff; permanent failures go to a dead-letter queue.
sequenceDiagram participant O as Order Service participant Q as Event Queue participant N as Notification Orchestrator participant P as Preferences Store participant C as Channel Queue participant W as Worker participant X as Provider O->>Q: OrderShipped(event_id=123) Q->>N: consume event N->>P: load user preferences P-->>N: push+email enabled, SMS off N->>N: dedup check event_id=123 N->>C: enqueue email + push jobs C->>W: deliver job W->>X: send notification X-->>W: 200 OK W->>N: mark sent
Reliability choices
- At-least-once processing is the safe default. It is better to send twice than not at all for most notifications.
- Idempotency keys on
(event_id, user_id, channel)stop duplicate delivery when retries happen. - Transactional outbox fits well if the notification event originates from a service that also writes a database row. Write the business change and outbox entry in one DB transaction, then relay the outbox to the queue.
- Channel isolation matters. A slow SMS provider should not delay push notifications; separate queues and worker pools keep one failure domain from contaminating the others.
Scaling and burst handling
- Use a broker/queue to absorb spikes, then let workers drain at provider-safe rates.
- Partition by
user_idorchanneldepending on which ordering property matters more. - Scale workers horizontally off queue depth and provider latency, not CPU alone.
- Use provider-specific throttles so you stay under APNs/FCM/ESP/SMS quotas.
The right mental model is “load leveling with policy in the middle”: the queue protects the backend from bursts, while the orchestrator decides who should actually receive the notification.
Edge cases and tradeoffs
- Duplicates vs loss: duplicates are usually cheaper than loss, but some categories (password reset, payment confirmation) need stricter idempotency and sometimes one-time tokens.
- User preferences vs compliance: users can opt out of marketing, but transactional alerts may still be required.
- Realtime vs digest: some notifications should be immediate, while others can be batched into hourly or daily digests.
- Ordering: per-user ordering is usually enough; global ordering is not worth the throughput cost.
- Retention: keep delivery history long enough for audit/support, but archive old attempts so the hot path stays small.
Good interview follow-ups
- How do you prevent the same event from notifying the same user twice?
- What happens if APNs or the email provider is down for 30 minutes?
- How do you support quiet hours and time zones?
- How do you ensure one noisy channel cannot starve the others?
- What do you do with undeliverable addresses or permanently bounced emails?
Practical answer shape
If asked to design this live, the clean answer is: ingest events into a durable queue, persist a delivery record, evaluate user preferences and dedup keys, fan out into separate per-channel queues, and have isolated workers retry with backoff and DLQs. That gives you reliability, burst absorption, and provider isolation without forcing the user-facing request to wait on third-party delivery.