- Published on
- Views
Design a Notification System Design Interview Guide
- Authors

- Name
- Javed Shaikh
← System Design Interview Preparation
This guide walks through Design a Notification System the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.
1. Problem
A notification system tells a user that something happened, using email, SMS, or mobile/web push, without slowing down the product API that caused the event.
Example:
- An order ships
- The order service publishes
OrderShipped - The user gets an email, and a push if the app is installed
- They do not get SMS if they turned SMS off
Chat unread messages, news-feed mentions, OTPs, and marketing campaigns all share this backbone. The product service should not call Twilio or Gmail itself. That is slow, flaky, and hard to retry.
In this design we are building a system that:
- Accepts notification events from other services
- Respects user preferences
- Renders a template
- Sends through email / push / SMS providers
- Records status and retries failures
We are not building a full marketing automation suite. We are designing the reliable send pipeline.
2. Functional Requirements / FR
| Requirement | What it means |
|---|---|
| Channels | Email, SMS, push at minimum. |
| Templates | Order {{orderId}} shipped with locale. |
| Preferences | User can mute a channel or a category (marketing vs transactional). |
| Priority | OTP and security > chat > marketing. |
| Deduplication | The same event should not email twice. |
| Retry | Provider 500s get retried with backoff. |
| Status | queued, sent, failed, skipped (muted). |
Out of scope for a 45-minute interview:
- Full in-app notification inbox UI (mention a table)
- A/B testing of subject lines
- Fancy customer-data platform
- Voice calls
Ask whether transactional (OTP, receipts) and marketing share one pipeline. They can share workers but should have separate queues and rate limits.
3. Non-Functional Requirements / NFR
| Requirement | Why it matters |
|---|---|
| Async | Product APIs must not wait on Gmail or SMS. |
| At-least-once send with idempotency | Duplicates annoy users and cost money (SMS). |
| High throughput | News-feed or chat can create bursts. |
| Provider isolation | Twilio down should not block email. |
| Low latency for OTP | Seconds, not minutes. |
| Auditability | “Did we send that reset email?” must be answerable. |
A good interview sentence: the notification system is a worker pipeline with a queue in front, not a library call inside checkout.
4. Back-of-the-Envelope Calculation
Traffic assumptions
- 100 million notifications per day across all channels
- Mix: 70% push, 25% email, 5% SMS (SMS is expensive)
- 80% transactional, 20% marketing
- Upstream product writes that cause notifications are lower; one order may create 2–3 notifications
- Read/write: this system is write/send heavy. Status reads are admin/debug.
QPS
Average send QPS:
100,000,000 / 86,400 ≈ 1,160 QPS
Peak 5× → around 6,000 QPS.
Chat-style bursts can be higher for a minute. Design workers and queues for 10,000 QPS peak if the interviewer mentions chat or live sports. Keep 6,000 as the baseline story.
Channel split at peak 6,000:
- Push ≈ 4,200 / sec
- Email ≈ 1,500 / sec
- SMS ≈ 300 / sec
Storage
Each notification log row ≈ 300 bytes (ids, channel, status, provider id, timestamps).
100,000,000 × 300 bytes ≈ 30 GB/day
Keep 30 days hot:
30 × 30 GB ≈ 900 GB
Archive the rest. Templates and preferences are tiny.
Cache / memory estimate
Preference objects: 50 million users × 200 bytes ≈ 10 GB if all cached. Cache active users only, e.g. 5 million × 200 bytes ≈ 1 GB.
Template cache: a few thousand templates in each worker’s memory. Almost free.
Server estimate
A worker that calls a provider might do 200–500 sends/sec depending on I/O.
Use 300 / sec per worker conservative:
6,000 / 300 = 20 workers
Run ~30 across three pools (email, push, SMS) so one provider slowness does not stall OTP. Queue brokers extra.
These are interview estimates, not exact production numbers.
5. APIs
Product services call a small internal API or publish an event. Users call preference APIs.
Enqueue a notification
POST /internal/notifications
{
"idempotencyKey": "order-shipped:ord_91:email",
"userId": "u_42",
"category": "transactional.shipping",
"templateId": "order_shipped_v3",
"channels": ["email", "push"],
"priority": "high",
"data": { "orderId": "ord_91", "carrier": "Delhivery" }
}
202 Accepted:
{
"notificationId": "n_77",
"status": "queued"
}
Duplicate idempotencyKey returns the original id, does not send twice.
User preferences
GET /api/notifications/preferences
PUT /api/notifications/preferences
{
"email": { "transactional": true, "marketing": false },
"sms": { "transactional": true, "marketing": false },
"push": { "transactional": true, "chat": true, "marketing": true }
}
OTP / security may be non-disableable. Say that out loud.
Status
GET /internal/notifications/{id}
{
"notificationId": "n_77",
"status": "sent",
"attempts": 1,
"provider": "ses",
"sentAt": "2026-09-19T08:00:02Z"
}
6. Data Model
notification_templates
| Field | Notes |
|---|---|
template_id | |
channel | email / sms / push |
locale | en-IN |
subject | email only |
body | with placeholders |
version |
user_preferences
user_id, channel, category, enabled, quiet_hours.
notifications (send log)
| Field | Notes |
|---|---|
notification_id | |
idempotency_key | unique |
user_id | |
channel | |
template_id | |
priority | |
status | queued, sent, failed, skipped |
attempts | |
provider | |
provider_message_id | |
last_error | |
created_at |
Index (user_id, created_at) and unique idempotency_key.
Queue payload
Do not put huge blobs in Kafka. Put ids + template data. Render in the worker.
7. High-Level Design
Product services drop events on a queue. Workers check preferences and templates, then talk to providers. Status goes to the notification DB.
Notification System architecture
Components:
- Product Service: order, chat, auth. Emits an event and moves on.
- Notification Queue: buffers bursts. Separate topics for high vs low priority. Kafka is a common choice.
- Notification Worker: the brain of this design.
- User Preferences: should this user get this channel + category?
- Template Service: render subject/body with user locale.
- Email / Push / SMS providers: SES, FCM/APNs, Twilio, etc.
- Notification DB: audit log and idempotency.
Send flow:
- Product calls enqueue or publishes
OrderShipped. - API writes an idempotent row
queuedand puts a queue message. - Worker loads preferences. If muted →
skipped. - Worker renders template.
- Worker calls the provider.
- On success →
sent. On retryable error → backoff and retry. On hard error →failed.
Why a queue? Providers are slow and fail. Checkout should not wait. This is the same idea as event-driven architecture.
8. Deep Dives
Email, SMS, push
- Email: cheap, not instant, good for receipts. Needs unsubscribe for marketing.
- SMS: expensive, good for OTP, strict rate limits and country rules.
- Push: cheap, needs device tokens, can fail if the user disabled OS notifications.
One event can fan out to several channels. That is several worker jobs, not one giant transaction.
OTP: dedicated high-priority queue, more workers, no mixing with “20% off” campaigns.
Template service
Keep templates out of Java code. A template id + JSON data lets non-engineers edit copy.
Render in the worker:
- Missing variable → fail the send, do not send
Hello {{name}} - Locale fallback:
hi-IN→en - Version templates so you can roll back a bad email
For email, send HTML + text. For SMS, hard cap length. For push, title + body + deep link.
Preferences
Check before calling a paid provider.
Order of checks:
- User exists and is not banned
- Category allowed on that channel
- Quiet hours (skip or delay marketing; still send OTP)
- Device token present for push
- Dedup key
Cache preferences. Invalidate on PUT.
Retry logic
Retry provider timeouts and 5xx. Do not retry 4xx like invalid email or bounced address (mark failed, maybe suppress future email).
Backoff: 1s, 5s, 25s, then dead-letter queue. Max attempts (say 5).
Workers must be idempotent: if the process dies after the provider accepted the message but before you wrote sent, retry might double-send. Use:
- idempotency key to the provider when they support it
- or store
provider_message_idwith a unique constraint and accept a rare duplicate on crash (say this tradeoff)
SMS duplicates are costly. Be extra careful there.
Deduplication
Keys like order-shipped:ord_91:email or chat-push:u_42:msg_15.
Window: for chat, collapse “5 messages in 10 seconds” into one push: debounce in Redis (SETNX notify:chat:u_42 with 10s TTL). That is different from idempotency of a single event.
Priority
Three queues:
- P0: OTP, security, password reset
- P1: shipping, chat, mentions
- P2: marketing, reminders
Workers steal from P0 first. Marketing has its own provider rate limit so it cannot eat the SES quota for OTP.
Provider failure
One SMS vendor down: fail over to a second vendor for P0. Email: SES to another provider. Push: FCM vs APNs are already split by platform.
Circuit-breaker around a sick provider so threads are not stuck. See circuit breaker.
9. Bottlenecks
| Bottleneck | What happens | What you do |
|---|---|---|
| Provider rate limits | 429 from Twilio/FCM | Per-channel queues, token bucket, extra vendors |
| Poison message | Bad template crashes the worker loop | DLQ, skip, alert |
| Preference DB | Every send hits it | Cache. Batch. |
| Hot user | One user gets 10k marketing emails | Per-user rate limit |
| Kafka lag | Burst from feed/chat | Autoscale workers. Separate P0 topic. |
| Status table writes | 100M rows/day | Partition by date. Async batch. |
The first bottleneck to name is slow or limited third-party providers.
10. Tradeoffs
Sync send vs async queue
OTP product managers want sync so they can show errors. Still, async with a fast P0 queue is usually enough (send in 1–2 seconds). True sync couples you to Twilio latency.
One pipeline vs two (transactional vs marketing)
One codebase, two queues. Mixing them is how marketing delays a password reset.
Exactly-once vs at-least-once
Exactly-once through Gmail is not something you control. Idempotency keys + at-least-once is the honest answer.
Render at enqueue vs in worker
Enqueue: product waits on templates. Worker: better. Product only sends data.
Store full body vs template + data
Storing rendered body helps audits. Costs space. Compromise: store template id + data, render on demand for the first 30 days.
11. Failure Modes
Queue down
Product API should still succeed. Local outbox on the product DB, replay later. Do not lose OTP events; fail the product call if outbox cannot write.
Worker crash
Message returns to the queue. Idempotency prevents most double sends.
Provider outage
Retry, then fallback vendor for P0. Pause P2 automatically.
Bad template deploy
Version pin. Feature-flag new template id. Render errors go to DLQ, not to users.
Preference cache stale
User turns off email, still gets one. Short TTL + invalidate on write. Accept rare extra send, or read DB for P2.
Fan-out storm from news feed
A celebrity post should not create 50 million emails. Notification policy: in-app + sampled push, never email for that event. Rate limit at enqueue.
Duplicate chat pushes
Debounce. The chat system should send one “unread” event, not one per message, when the user is offline.
12. Interview Answer in 10 Minutes
Here is a version you can speak out loud.
"I would design notifications as an async pipeline. Product services do not call email or SMS directly. They publish an event or call an enqueue API with an idempotency key, then return to the user.
Assume 100 million notifications a day. That is about 1,160 QPS, around 6,000 at peak, mostly push, some email, a little SMS. Storage for send logs is tens of gigabytes a day, so I would partition the log table and archive old rows.
A worker picks the job, checks user preferences, renders a template, and calls the right provider. I would split queues by priority so OTP never sits behind marketing. Email, push, and SMS have separate worker pools because each provider fails differently.
Retries use backoff and a dead-letter queue. Dedup uses a unique idempotency key. Chat can also debounce many messages into one push. If a provider is down, P0 can fail over to a second vendor. Status is stored so we can answer whether a reset email was sent.
That is the design: queue, worker, preferences, templates, providers, and an audit table."
Practice this until it is under 10 minutes.
13. Interview Talking Points
- Never block checkout/chat on a provider.
- Idempotency keys save money and trust.
- Preferences before send.
- Priority queues.
- Separate worker pools per channel.
- Retry + DLQ + circuit breaker.
- Debounce chat/feed bursts.
- Metrics: queue lag, send success, provider latency, skip (muted) rate.
- This is a classic microservices + queue story.
14. Follow-up Questions
How do you send 10 million marketing emails?
Dedicated P2 queue, warmup IPs, provider batch APIs, unsubscribe headers, slower rate. Do not use the OTP SES quota.
How do you handle quiet hours?
Delay marketing on a delayed queue. Send transactional anyway.
What if the user has three devices?
Push to all device tokens. Dedup in the app. Remove dead tokens on provider errors.
How is this different from in-app notifications?
In-app is a table + feed API, no third party. You can write in-app in the same worker and skip providers.
Java implementation?
Spring events or Kafka consumers, template engine, per-channel clients, Redis for debounce and preference cache. See Pulse-Guard Kafka for idempotent consumers.
Other probes: GDPR deletion, PII in logs, and multi-tenant SaaS templates.
15. Internal Links
Related JavaThoughts reading:
- System Design Interview Preparation
- Design a Chat Messaging System
- Design a News Feed
- Design a Rate Limiter
- Kafka
- Event-driven architecture
- Message queues
- Java
- Spring Boot
- Microservices
- Distributed Systems
- Scaling Java event-driven systems
- 20 system design concepts
This is the last of the first six system design guides. More questions stay on the System Design landing page as they are written.
