Javathoughts Logo
Javathoughts
Published on
Views

Design a Chat Messaging System Design Interview Guide

Authors
  • avatar
    Name
    Javed Shaikh
    Twitter

← System Design Interview Preparation

This guide walks through Design a Chat / Messaging System the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.


1. Problem

A chat system lets people send messages to each other in real time, keep history, and know whether the other person is online.

Example:

  • User A types “on my way”
  • If User B is online, the message appears in a second
  • If User B is offline, the message is stored and a push notification is sent
  • Later both can scroll older messages

This is WhatsApp / Slack / Messenger at a small slice: 1:1 chat, group chat, delivery status, online/offline.

In this design we are building:

  1. One-to-one conversations
  2. Group conversations
  3. Real-time delivery over WebSockets
  4. Message persistence
  5. Push for offline users

We are not building voice calls, end-to-end encryption math, or a full Slack workspace admin panel. Mention E2E encryption as a follow-up if asked.


2. Functional Requirements / FR

RequirementWhat it means
1:1 chatTwo users, one conversation.
Group chatMany users, one conversation. Start with a cap (e.g. 200 members).
Send / receiveText first. Media ids later.
HistoryMessages are stored and can be paged.
Delivery statusSent, delivered, read (optional ticks).
Online / offlinePresence: last seen or a green dot.
PushIf the user has no open socket, send a mobile/web push.

Out of scope for a 45-minute interview:

  • Voice and video
  • Full search across all history (mention later)
  • Disappearing messages product details
  • Cross-region multi-device sync edge cases beyond “each device is a connection”

Confirm message size (4 KB text) and whether exactly-once is required. At-least-once + idempotent message ids is the practical answer.


3. Non-Functional Requirements / NFR

RequirementWhy it matters
Low latencyChat should feel instant. Aim for under a few hundred ms in-region.
High connection countEach online user holds a WebSocket. Servers must hold many idle connections.
Ordered messagesIn one conversation, order should be stable.
Durable historyA dropped socket cannot lose a sent message.
Fanout for groupsOne group message becomes N deliveries.
AvailabilityPeople treat chat as infrastructure.

A good interview sentence: the WebSocket path is for speed. The database is for truth. The push path is for users who are not connected.


4. Back-of-the-Envelope Calculation

Traffic assumptions

  • 50 million Daily Active Users
  • 25% online at peak → 12.5 million concurrent WebSockets
  • Each DAU sends 20 messages/day → 1 billion messages/day
  • 90% 1:1, 10% group
  • Average group size 8
  • Read/write: each message is written once and read by 1–8 clients, plus history scrolls. Treat send QPS as the core write number.

QPS

Average send QPS:

1,000,000,000 / 86,400 ≈ 11,600 messages/sec

Peak 5× → around 58,000 messages/sec.

Deliveries are higher because of groups:

90% × 1 delivery + 10% × 8 deliveries ≈ 1.7 deliveries per message
Peak deliveries ≈ 100,000 / sec

History reads: assume each DAU also loads 5 history pages/day.

250,000,000 / 86,400 ≈ 2,900 QPS average, ~15,000 peak

Storage

Message row ≈ 200 bytes (ids, text pointer or short text, timestamps, status). Long text in object storage or a text column.

1,000,000,000 × 200 bytes ≈ 200 GB/day

One year ≈ 70 TB. Partition by conversation_id. Keep recent messages on fast storage, archive old ones.

Media is separate: not in the message table.

Cache / memory estimate

Presence is a natural cache:

12.5 million online users × 50 bytes ≈ 625 MB

Recent conversation inbox (last message per chat) for online users:

12.5 million × 20 chats × 100 bytes ≈ 25 GB

Fits a Redis cluster. Full history does not belong in Redis.

Server estimate

WebSocket connections are the special number.

If one gateway process holds 50,000 idle connections (conservative):

12,500,000 / 50,000 = 250 gateway instances

Some teams hold 100k+ per box. Use 250 in the interview and say you would load-test.

Chat service workers for 58k messages/sec: if one node handles 5,000 msgs/sec,

58,000 / 5,000 ≈ 12 write nodes

Use ~20 plus gateways. Gateways dominate hardware.

These are interview estimates, not exact production numbers.


5. APIs

REST for history and send-when-needed. WebSocket for live traffic.

REST

POST /api/conversations — create 1:1 or group

GET /api/conversations — inbox list

GET /api/conversations/{id}/messages?cursor=&limit=50

POST /api/conversations/{id}/messages

{
  "clientMsgId": "c_88",
  "text": "on my way"
}

clientMsgId makes retries safe.

WebSocket

Connect: wss://chat.example.com/ws?token=...

Client → server:

{ "type": "send", "conversationId": "cv_1", "clientMsgId": "c_88", "text": "on my way" }
{ "type": "ack", "messageId": "m_15", "status": "read" }
{ "type": "typing", "conversationId": "cv_1" }

Server → client:

{ "type": "message", "messageId": "m_15", "conversationId": "cv_1", "from": "u_a", "text": "on my way" }
{ "type": "status", "messageId": "m_15", "status": "delivered" }
{ "type": "presence", "userId": "u_b", "online": false }

If there is no socket, the chat service enqueues a push notification.


6. Data Model

conversations

FieldNotes
conversation_id
typedirect or group
created_at
last_message_idinbox preview
last_message_atsort inbox

conversation_members

conversation_id, user_id, role, joined_at, last_read_message_id

Index: (user_id, last_message_at) for inbox. Index: (conversation_id) for group fanout.

messages

FieldNotes
message_idserver id
client_msg_idunique per sender
conversation_idpartition key
sender_id
text
created_at
seqmonotonic per conversation

Unique (sender_id, client_msg_id). Index (conversation_id, seq) or (conversation_id, created_at).

message_status (optional table or columns)

Per user: delivered_at, read_at. For 1:1 you can keep two flags on the message. For groups, a separate table or “read up to seq”.

Presence

Do not put heartbeat rows in SQL. Use Redis: presence:{userId} → {gatewayId, lastSeen} with a short TTL. Heartbeats refresh TTL.


7. High-Level Design

Online users stay on a WebSocket gateway. The chat service persists every message. Offline users get a push from a queue.

Chat / Messaging architecture

Chat / Messaging architectureOnline users talk through a WebSocket gateway and chat service. Messages are saved. Offline users get a push from the queue.persistoffline👤User A🌐WebSocket Gateway⚙️Chat Service👤User B🗄️Message DB📩Queue🔔Push Notification
Online users talk through a WebSocket gateway and chat service. Messages are saved. Offline users get a push from the queue.

Components:

  • User A / User B: apps holding a socket when open.
  • WebSocket Gateway: holds connections, authenticates, heartbeats. Sticky by user id so you know where a user is connected.
  • Chat Service: permission check, persist, fanout to member sockets, emit offline events.
  • Message DB: source of truth for history.
  • Queue: offline / push events, and sometimes group fanout.
  • Push Notification Service: APNs / FCM / web push. Reuse the notification system if it exists.

Send flow (B online):

  1. A sends on the socket (or REST).
  2. Chat service writes the message (idempotent on clientMsgId).
  3. Lookup B’s connection on the gateway (Redis: userId → gatewayId).
  4. Push the payload to B’s socket.
  5. Mark delivered when B’s client acks.

Send flow (B offline):

  1. Same persist.
  2. No live connection.
  3. Enqueue push: “A: on my way”.
  4. When B opens the app, history API + new socket.

Group flow: persist once, then fanout to N member connections. Large groups: fanout workers on a queue so one request does not send to 200 sockets inline.


8. Deep Dives

One-to-one chat

Canonical conversation id: min(userA, userB) + max(userA, userB) so you do not create two DMs.

Sequence number per conversation keeps order even if timestamps collide.

Group chat

Members table is the fanout list. Check membership on every send.

For small groups (≤ 200), inline fanout is fine. For bigger rooms, that becomes a news-feed problem: persist, then workers deliver. Slack-style channels at huge scale often use a log per channel plus per-user cursors.

Start the interview with small groups, persist once, fanout to online members, push to offline members.

WebSocket connection

Why WebSocket? HTTP request/response is a bad chat loop. The server must push. See WebSockets.

Gateway duties:

  • TLS and auth
  • Heartbeat / ping
  • Map userId → connection
  • Forward to chat service (gRPC/HTTP) so gateways stay thin

If User A is on gateway 7 and User B on gateway 12, chat service publishes “deliver to user B” and gateway 12 writes to the socket. A pub/sub (Redis, NATS, Kafka) between gateways works.

Java note: thread-per-connection will not hold 50k sockets. Use non-blocking I/O (Netty, or Java virtual threads with care). This is a good talking point.

Message persistence

Write to the DB before you tell the sender “sent”. If you ack first and then crash, the message vanished.

Order:

  1. Persist
  2. Ack sender (sent)
  3. Fanout
  4. Delivered / read as extra events

Retries from the client use the same clientMsgId. Unique constraint stops duplicates.

Delivery status

  • Sent: stored on server
  • Delivered: recipient’s device got it (socket ack or app opened and fetched)
  • Read: recipient’s client sent a read receipt

Do not block sending on read receipts. They are extra messages. Users can disable read receipts: still store delivered, skip broadcasting read.

Online / offline users

Presence:

  • On connect: SET Redis key with TTL 45s, subscribe heartbeats
  • On disconnect: delete key, broadcast offline to interested users (contacts / open conversation — not to the whole user base)
  • Last seen: write to DB occasionally, not on every heartbeat

“Who is online” for 12 million users cannot be a SQL SELECT WHERE online=true on every page.

Push notifications

If presence is missing, chat service emits OfflineMessage to a queue. Notification worker checks mute settings and quiet hours, then sends push. When the user opens the socket, suppress duplicate noisy banners if they already fetched the message.


9. Bottlenecks

BottleneckWhat happensWhat you do
Connection countGateways run out of RAM/FDsMore gateways. Efficient event loops.
Group fanoutOne send becomes thousands of socket writesAsync workers. Batch.
Hot conversationA group chat storms one DB partitionShard by conversation id. Cache last messages.
Inbox queriesOpening the app loads 50 chatslast_message_at index + cache
Presence broadcastsA celebrity going online notifies millionsOnly notify relevant conversations
History scansDeep scrollCursor on seq. Archive cold data.

The first bottleneck to name is millions of long-lived WebSocket connections.


10. Tradeoffs

WebSocket vs long polling vs SSE

  • WebSocket: bidirectional, best default for chat
  • SSE: server-to-client only, client still POSTs messages
  • Long poll: simple, more overhead

Pick WebSocket.

Sync persist vs async persist

Async persist is faster and can lose messages. Do not do that for chat. Persist first.

SQL vs wide-column for messages

SQL is fine early. At 200 GB/day, a partitioned store (Cassandra, Dynamo, sharded Postgres) is expected. Partition key = conversation id.

Exactly-once vs at-least-once

Network retries happen. At-least-once + idempotent ids.

Push vs socket only

Socket only fails when the app is killed. Push is required on mobile.


11. Failure Modes

Gateway crash
Connections drop. Clients reconnect to another gateway. Messages already persisted are not lost. Replay from last seq.

Chat service crash after persist, before fanout
Sender has “sent”. Receiver might not get live push. Client B will fetch on reconnect. A delivery worker can retry fanout from an outbox table.

Duplicate message
Same clientMsgId. Unique index. Return the original message id.

Out-of-order delivery
Assign seq in the chat service per conversation (single writer per conversation, or ordered log). Clients sort by seq.

Offline user never gets push
Queue retry. If push provider fails, they still get history on next open. See notification failure modes.

Split brain presence
User shows online but socket is dead. Short TTL + heartbeat fixes it within a minute. Do not trust presence for security checks. Always check membership on send.


12. Interview Answer in 10 Minutes

Here is a version you can speak out loud.

"I would design chat with three paths: WebSocket for online users, a message database for history, and a push queue for offline users.

Assume 50 million DAU, 1 billion messages a day. That is about 12,000 messages per second, around 60,000 at peak. Peak online users might be 12 million sockets. That is the number that scares you, so I would put a dedicated WebSocket gateway tier, maybe a few hundred instances, each holding tens of thousands of connections.

A send writes to the message store first with a client-generated id so retries are idempotent. Then the chat service looks up where the other user is connected and pushes over the socket. If they are offline, it enqueues a notification. Groups persist once and fanout to members, asynchronously if the group is large.

Inbox and last-message previews can live in Redis. Full history is paged from the DB by conversation id and sequence number. Presence is a Redis key with a TTL, not a SQL flag.

If a gateway dies, clients reconnect and load any missed seq. We never ack the sender before the write succeeds.

That is the system: gateway, chat service, message DB, offline queue, push."

Practice this until it is under 10 minutes.


13. Interview Talking Points

  • Persist then fanout. Never the opposite.
  • WebSocket gateways are a separate tier because of connection count.
  • Idempotent client message ids.
  • Seq per conversation for order.
  • Presence in Redis with TTL.
  • Push is a fallback, not the main live path.
  • Group chat is fanout. Keep group size explicit.
  • Java gateways need non-blocking I/O.
  • Monitoring: connect count, send latency, persist failures, push lag.

14. Follow-up Questions

How do you support multiple devices?
Each device has a connection. Fanout to all of a user’s sockets. Last-read seq can be per device or synced.

How would you add end-to-end encryption?
Server stores ciphertext. Keys on devices. Server still routes. Search and push previews become harder. Mention, do not design Signal in 10 minutes.

How do you scale a 10,000 user group?
Treat it like a channel log. Do not write 10,000 inbox rows inline. Consumers pull, or workers fanout in batches.

What if the receiver is online on web but mobile is asleep?
Deliver to the web socket. Optionally still push mobile, or suppress if “active elsewhere”.

REST send vs socket send?
Both call the same chat service. Socket is faster for typing. REST helps when sockets are blocked.

Other probes: media upload, spam, and message deletion.


Related JavaThoughts reading:


Next in this series: Design a Notification System.