Javathoughts Logo
Javathoughts
Published on
Views

Design a Notification System Design Interview Guide

Authors
  • avatar
    Name
    Javed Shaikh
    Twitter

← System Design Interview Preparation

This guide walks through Design a Notification System the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.


1. Problem

A notification system tells a user that something happened, using email, SMS, or mobile/web push, without slowing down the product API that caused the event.

Example:

  • An order ships
  • The order service publishes OrderShipped
  • The user gets an email, and a push if the app is installed
  • They do not get SMS if they turned SMS off

Chat unread messages, news-feed mentions, OTPs, and marketing campaigns all share this backbone. The product service should not call Twilio or Gmail itself. That is slow, flaky, and hard to retry.

In this design we are building a system that:

  1. Accepts notification events from other services
  2. Respects user preferences
  3. Renders a template
  4. Sends through email / push / SMS providers
  5. Records status and retries failures

We are not building a full marketing automation suite. We are designing the reliable send pipeline.


2. Functional Requirements / FR

RequirementWhat it means
ChannelsEmail, SMS, push at minimum.
TemplatesOrder {{orderId}} shipped with locale.
PreferencesUser can mute a channel or a category (marketing vs transactional).
PriorityOTP and security > chat > marketing.
DeduplicationThe same event should not email twice.
RetryProvider 500s get retried with backoff.
Statusqueued, sent, failed, skipped (muted).

Out of scope for a 45-minute interview:

  • Full in-app notification inbox UI (mention a table)
  • A/B testing of subject lines
  • Fancy customer-data platform
  • Voice calls

Ask whether transactional (OTP, receipts) and marketing share one pipeline. They can share workers but should have separate queues and rate limits.


3. Non-Functional Requirements / NFR

RequirementWhy it matters
AsyncProduct APIs must not wait on Gmail or SMS.
At-least-once send with idempotencyDuplicates annoy users and cost money (SMS).
High throughputNews-feed or chat can create bursts.
Provider isolationTwilio down should not block email.
Low latency for OTPSeconds, not minutes.
Auditability“Did we send that reset email?” must be answerable.

A good interview sentence: the notification system is a worker pipeline with a queue in front, not a library call inside checkout.


4. Back-of-the-Envelope Calculation

Traffic assumptions

  • 100 million notifications per day across all channels
  • Mix: 70% push, 25% email, 5% SMS (SMS is expensive)
  • 80% transactional, 20% marketing
  • Upstream product writes that cause notifications are lower; one order may create 2–3 notifications
  • Read/write: this system is write/send heavy. Status reads are admin/debug.

QPS

Average send QPS:

100,000,000 / 86,400 ≈ 1,160 QPS

Peak 5× → around 6,000 QPS.

Chat-style bursts can be higher for a minute. Design workers and queues for 10,000 QPS peak if the interviewer mentions chat or live sports. Keep 6,000 as the baseline story.

Channel split at peak 6,000:

  • Push ≈ 4,200 / sec
  • Email ≈ 1,500 / sec
  • SMS ≈ 300 / sec

Storage

Each notification log row ≈ 300 bytes (ids, channel, status, provider id, timestamps).

100,000,000 × 300 bytes ≈ 30 GB/day

Keep 30 days hot:

30 × 30 GB ≈ 900 GB

Archive the rest. Templates and preferences are tiny.

Cache / memory estimate

Preference objects: 50 million users × 200 bytes ≈ 10 GB if all cached. Cache active users only, e.g. 5 million × 200 bytes ≈ 1 GB.

Template cache: a few thousand templates in each worker’s memory. Almost free.

Server estimate

A worker that calls a provider might do 200–500 sends/sec depending on I/O.

Use 300 / sec per worker conservative:

6,000 / 300 = 20 workers

Run ~30 across three pools (email, push, SMS) so one provider slowness does not stall OTP. Queue brokers extra.

These are interview estimates, not exact production numbers.


5. APIs

Product services call a small internal API or publish an event. Users call preference APIs.

Enqueue a notification

POST /internal/notifications

{
  "idempotencyKey": "order-shipped:ord_91:email",
  "userId": "u_42",
  "category": "transactional.shipping",
  "templateId": "order_shipped_v3",
  "channels": ["email", "push"],
  "priority": "high",
  "data": { "orderId": "ord_91", "carrier": "Delhivery" }
}

202 Accepted:

{
  "notificationId": "n_77",
  "status": "queued"
}

Duplicate idempotencyKey returns the original id, does not send twice.

User preferences

GET /api/notifications/preferences

PUT /api/notifications/preferences

{
  "email": { "transactional": true, "marketing": false },
  "sms": { "transactional": true, "marketing": false },
  "push": { "transactional": true, "chat": true, "marketing": true }
}

OTP / security may be non-disableable. Say that out loud.

Status

GET /internal/notifications/{id}

{
  "notificationId": "n_77",
  "status": "sent",
  "attempts": 1,
  "provider": "ses",
  "sentAt": "2026-09-19T08:00:02Z"
}

6. Data Model

notification_templates

FieldNotes
template_id
channelemail / sms / push
localeen-IN
subjectemail only
bodywith placeholders
version

user_preferences

user_id, channel, category, enabled, quiet_hours.

notifications (send log)

FieldNotes
notification_id
idempotency_keyunique
user_id
channel
template_id
priority
statusqueued, sent, failed, skipped
attempts
provider
provider_message_id
last_error
created_at

Index (user_id, created_at) and unique idempotency_key.

Queue payload

Do not put huge blobs in Kafka. Put ids + template data. Render in the worker.


7. High-Level Design

Product services drop events on a queue. Workers check preferences and templates, then talk to providers. Status goes to the notification DB.

Notification System architecture

Notification System architectureProduct events go on a queue. A worker loads user preferences and a template, then sends email, push, or SMS and stores delivery status.checkstatus⚙️Product Service📩Notification Queue👷Notification Worker👤User Preferences📝Template Service📧Email📱Push💬SMS🗄️Notification DB
Product events go on a queue. A worker loads user preferences and a template, then sends email, push, or SMS and stores delivery status.

Components:

  • Product Service: order, chat, auth. Emits an event and moves on.
  • Notification Queue: buffers bursts. Separate topics for high vs low priority. Kafka is a common choice.
  • Notification Worker: the brain of this design.
  • User Preferences: should this user get this channel + category?
  • Template Service: render subject/body with user locale.
  • Email / Push / SMS providers: SES, FCM/APNs, Twilio, etc.
  • Notification DB: audit log and idempotency.

Send flow:

  1. Product calls enqueue or publishes OrderShipped.
  2. API writes an idempotent row queued and puts a queue message.
  3. Worker loads preferences. If muted → skipped.
  4. Worker renders template.
  5. Worker calls the provider.
  6. On success → sent. On retryable error → backoff and retry. On hard error → failed.

Why a queue? Providers are slow and fail. Checkout should not wait. This is the same idea as event-driven architecture.


8. Deep Dives

Email, SMS, push

  • Email: cheap, not instant, good for receipts. Needs unsubscribe for marketing.
  • SMS: expensive, good for OTP, strict rate limits and country rules.
  • Push: cheap, needs device tokens, can fail if the user disabled OS notifications.

One event can fan out to several channels. That is several worker jobs, not one giant transaction.

OTP: dedicated high-priority queue, more workers, no mixing with “20% off” campaigns.

Template service

Keep templates out of Java code. A template id + JSON data lets non-engineers edit copy.

Render in the worker:

  • Missing variable → fail the send, do not send Hello {{name}}
  • Locale fallback: hi-IN → en
  • Version templates so you can roll back a bad email

For email, send HTML + text. For SMS, hard cap length. For push, title + body + deep link.

Preferences

Check before calling a paid provider.

Order of checks:

  1. User exists and is not banned
  2. Category allowed on that channel
  3. Quiet hours (skip or delay marketing; still send OTP)
  4. Device token present for push
  5. Dedup key

Cache preferences. Invalidate on PUT.

Retry logic

Retry provider timeouts and 5xx. Do not retry 4xx like invalid email or bounced address (mark failed, maybe suppress future email).

Backoff: 1s, 5s, 25s, then dead-letter queue. Max attempts (say 5).

Workers must be idempotent: if the process dies after the provider accepted the message but before you wrote sent, retry might double-send. Use:

  • idempotency key to the provider when they support it
  • or store provider_message_id with a unique constraint and accept a rare duplicate on crash (say this tradeoff)

SMS duplicates are costly. Be extra careful there.

Deduplication

Keys like order-shipped:ord_91:email or chat-push:u_42:msg_15.

Window: for chat, collapse “5 messages in 10 seconds” into one push: debounce in Redis (SETNX notify:chat:u_42 with 10s TTL). That is different from idempotency of a single event.

Priority

Three queues:

  • P0: OTP, security, password reset
  • P1: shipping, chat, mentions
  • P2: marketing, reminders

Workers steal from P0 first. Marketing has its own provider rate limit so it cannot eat the SES quota for OTP.

Provider failure

One SMS vendor down: fail over to a second vendor for P0. Email: SES to another provider. Push: FCM vs APNs are already split by platform.

Circuit-breaker around a sick provider so threads are not stuck. See circuit breaker.


9. Bottlenecks

BottleneckWhat happensWhat you do
Provider rate limits429 from Twilio/FCMPer-channel queues, token bucket, extra vendors
Poison messageBad template crashes the worker loopDLQ, skip, alert
Preference DBEvery send hits itCache. Batch.
Hot userOne user gets 10k marketing emailsPer-user rate limit
Kafka lagBurst from feed/chatAutoscale workers. Separate P0 topic.
Status table writes100M rows/dayPartition by date. Async batch.

The first bottleneck to name is slow or limited third-party providers.


10. Tradeoffs

Sync send vs async queue

OTP product managers want sync so they can show errors. Still, async with a fast P0 queue is usually enough (send in 1–2 seconds). True sync couples you to Twilio latency.

One pipeline vs two (transactional vs marketing)

One codebase, two queues. Mixing them is how marketing delays a password reset.

Exactly-once vs at-least-once

Exactly-once through Gmail is not something you control. Idempotency keys + at-least-once is the honest answer.

Render at enqueue vs in worker

Enqueue: product waits on templates. Worker: better. Product only sends data.

Store full body vs template + data

Storing rendered body helps audits. Costs space. Compromise: store template id + data, render on demand for the first 30 days.


11. Failure Modes

Queue down
Product API should still succeed. Local outbox on the product DB, replay later. Do not lose OTP events; fail the product call if outbox cannot write.

Worker crash
Message returns to the queue. Idempotency prevents most double sends.

Provider outage
Retry, then fallback vendor for P0. Pause P2 automatically.

Bad template deploy
Version pin. Feature-flag new template id. Render errors go to DLQ, not to users.

Preference cache stale
User turns off email, still gets one. Short TTL + invalidate on write. Accept rare extra send, or read DB for P2.

Fan-out storm from news feed
A celebrity post should not create 50 million emails. Notification policy: in-app + sampled push, never email for that event. Rate limit at enqueue.

Duplicate chat pushes
Debounce. The chat system should send one “unread” event, not one per message, when the user is offline.


12. Interview Answer in 10 Minutes

Here is a version you can speak out loud.

"I would design notifications as an async pipeline. Product services do not call email or SMS directly. They publish an event or call an enqueue API with an idempotency key, then return to the user.

Assume 100 million notifications a day. That is about 1,160 QPS, around 6,000 at peak, mostly push, some email, a little SMS. Storage for send logs is tens of gigabytes a day, so I would partition the log table and archive old rows.

A worker picks the job, checks user preferences, renders a template, and calls the right provider. I would split queues by priority so OTP never sits behind marketing. Email, push, and SMS have separate worker pools because each provider fails differently.

Retries use backoff and a dead-letter queue. Dedup uses a unique idempotency key. Chat can also debounce many messages into one push. If a provider is down, P0 can fail over to a second vendor. Status is stored so we can answer whether a reset email was sent.

That is the design: queue, worker, preferences, templates, providers, and an audit table."

Practice this until it is under 10 minutes.


13. Interview Talking Points

  • Never block checkout/chat on a provider.
  • Idempotency keys save money and trust.
  • Preferences before send.
  • Priority queues.
  • Separate worker pools per channel.
  • Retry + DLQ + circuit breaker.
  • Debounce chat/feed bursts.
  • Metrics: queue lag, send success, provider latency, skip (muted) rate.
  • This is a classic microservices + queue story.

14. Follow-up Questions

How do you send 10 million marketing emails?
Dedicated P2 queue, warmup IPs, provider batch APIs, unsubscribe headers, slower rate. Do not use the OTP SES quota.

How do you handle quiet hours?
Delay marketing on a delayed queue. Send transactional anyway.

What if the user has three devices?
Push to all device tokens. Dedup in the app. Remove dead tokens on provider errors.

How is this different from in-app notifications?
In-app is a table + feed API, no third party. You can write in-app in the same worker and skip providers.

Java implementation?
Spring events or Kafka consumers, template engine, per-channel clients, Redis for debounce and preference cache. See Pulse-Guard Kafka for idempotent consumers.

Other probes: GDPR deletion, PII in logs, and multi-tenant SaaS templates.


Related JavaThoughts reading:


This is the last of the first six system design guides. More questions stay on the System Design landing page as they are written.