Javathoughts Logo
Javathoughts
Published on
Views

Design a Collaborative Document Editor System Design Interview Guide

Authors
  • avatar
    Name
    Javed Shaikh
    Twitter

← System Design Interview Preparation

This guide walks through Design a Collaborative Document Editor the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.


1. Problem

Many people type in the same document at once. Everyone should see edits quickly. The file should not corrupt.

Example:

  • You and a teammate open a spec
  • Cursors and names show presence
  • You insert a sentence while they delete a word
  • The server merges operations
  • History can restore an hour ago
  • Offline, you type on a plane and sync later (simple version)

This is closer to chat (WebSockets) plus a merge algorithm, not a Google Docs clone of every feature.


2. Functional Requirements / FR

RequirementWhat it means
Real-time editingOps in ~100ms on a good network.
Multi-user same docN editors per document.
WebSocketsPush ops and presence.
OT or CRDTExplain one simply.
SnapshotsPeriodic full document.
Version historyRestore snapshot.
Presence / cursorsEphemeral, not in the doc history.
Conflict resolutionAlgorithm, not “last write wins” on the whole file.
Offline (brief)Queue local ops, sync on reconnect.
PermissionsOwner / editor / viewer.

Out of scope: comments, suggestions mode, Excel formulas, 100k-user live lecture (say sharding by doc).


3. Non-Functional Requirements / NFR

RequirementWhy it matters
Low latency opsTyping should feel local.
CorrectnessSame sequence → same doc on all clients.
Doc affinityAll ops for a doc hit one collab process (or a CRDT that can merge anywhere).
DurabilitySnapshots on disk; ops buffered.
Scale by documentHot docs are rare; millions of cold docs.

Interview line: sticky collaboration service per doc, OT or CRDT, snapshots, WebSockets, presence as a side channel.


4. Back-of-the-Envelope Calculation

Traffic assumptions

  • 10 million documents
  • 1 million DAU
  • 5% editing at once → 50,000 concurrent editors
  • Average 1 op/sec while typing (bursts higher)
  • Read/write: live ops are writes; open-doc is snapshot read

QPS

Live ops:

50,000 editors × 1 op/sec = 50,000 ops/sec global
If peak is 5x (many people in a meeting doc), design for around 250,000 ops/sec globally.
Per popular doc: maybe 20 editors × 5 ops = 100 ops/sec — small for one process.

Opens:

Say 20 million opens/day / 86,400 ≈ 230 QPS average, ~1,200 at 5× peak.

Storage

Doc average 50 KB. 10M × 50 KB = 500 GB current.

Ops: 50k ops/sec × 100 bytes × 86,400 ≈ 400 GB/day if you kept every op forever. You will not. Keep ops for 24h then snapshot and compact. Interview: tens of GB/day after compaction.

Cache / memory

Active docs in RAM on the collab node. 50k editors / 5 per doc ≈ 10,000 hot docs × 50 KB ≈ 500 MB plus OT state. Fine.

Servers

Shard by docId. 10k hot docs / 500 per box ≈ 20 collab nodes plus WS gateways. Interview estimates, not exact production numbers.


5. APIs

HTTP:

POST /v1/docs
GET  /v1/docs/{id}          (snapshot + revision)
GET  /v1/docs/{id}/history
POST /v1/docs/{id}/acl

WebSocket:

{ "type": "op", "rev": 1842, "payload": { ... } }
{ "type": "ack", "rev": 1843 }
{ "type": "presence", "cursor": 120, "userId": "..." }
{ "type": "snapshot", "rev": 1900, "text": "..." }

REST is for load and permissions. Live edits are WS.


6. Data Model

documents

doc_id, title, revision, snapshot_blob or S3 key, updated_at.

acl

doc_id, user_id, role.

op_log

doc_id, rev, op, user_id, created_at. Truncate after snapshot.

snapshots

doc_id, rev, storage_key.

Presence is not in SQL. It lives in memory / Redis pubsub.


7. High-Level Design

One document = one in-memory state machine.

Collaborative Document Editor architecture

Collaborative Document Editor architectureEditors talk over WebSockets. The collaboration service applies OT or CRDT ops in memory, then persists snapshots. Presence is broadcast to everyone in the room.ops / snapshots👤Users🌐WebSocket Gateway✍️Collaboration Service🧠Memory / Cache🗄️Document DB📩Queue👷Persistence Worker👀Presence Broadcast
Editors talk over WebSockets. The collaboration service applies OT or CRDT ops in memory, then persists snapshots. Presence is broadcast to everyone in the room.

Components:

  • Users → WebSocket Gateway → Collaboration Service
  • OT/CRDT applied in memory
  • Queue → Persistence Worker → Document DB
  • Presence broadcast back to users

Join flow:

  1. Auth (sessions).
  2. Load snapshot + recent ops.
  3. Subscribe to doc:{id} fanout.
  4. Local typing produces ops; send to server; server transforms; broadcast.

Why a worker? Do not fsync every keystroke on the WS thread. Batch ops, snapshot every N seconds or N revs.


8. Deep Dives

OT vs CRDT (simple)

OT (Operational Transform): server orders ops. If you insert at 5 and they delete 3, the server rewrites your index. Needs a single sequencer per doc (the collab process).

CRDT: each character has an id. Merges commute. Easier multi-master / offline. Bigger metadata.

Interview pick: OT with a single leader per document is enough. Mention CRDT for offline and multi-region.

Snapshots

Every 50 ops or 5 seconds: write snapshot, drop old ops. History restore = old snapshot.

Presence

Throttle cursor messages (10 Hz). Do not persist.

Offline

Client queues ops with client ids. On reconnect, send them; OT/CRDT merges. Last-write-wins on the whole doc is the wrong answer.

Permissions

Viewer: WS receive only. Editor: send ops. Check ACL on connect and on share changes.

Sticky routing

Gateway hashes docId to a collab pod. If the pod dies, reload snapshot elsewhere (ops in Kafka help).

WebSockets intro: WebSockets concept.


9. Bottlenecks

  • One celebrity doc (keynote) — split not possible for one string; scale vertically, or CRDT with care
  • WS connection count on gateway
  • Snapshot storm
  • Huge paste (cap op size)

10. Tradeoffs

ChoiceUpsideDownside
OT + leaderSimple reasoningLeader failover
CRDTOffline, multi-masterMemory, complexity
Persist every opPerfect replayCost
Snapshot oftenCheapCoarser history

Pick: per-doc leader, OT, batched persist, presence separate, ACL on connect.


11. Failure Modes

FailureHandling
Collab node diesClients reconnect, load last snapshot + log
Duplicate opClient op id idempotency
Split brain two leadersShould not: lock docId in Redis
Offline merge messCRDT or reject and show diff
Permission revoked mid-sessionDisconnect WS

12. Interview Answer in 10 Minutes

"I would design a collab editor around a per-document in-memory replica and WebSockets.

Globally you might see 50,000 ops/sec, but one document is tens of ops/sec. I shard by doc id. Clients load a snapshot, then send small operations. A collaboration service sequences them with operational transform so two inserts do not clobber the file. I would mention CRDTs for offline and multi-region.

Presence is a side channel. Persistence is async: a worker writes snapshots and a short op log. History is old snapshots.

If the node dies, clients reconnect and replay from storage. Viewers cannot send ops. We never last-write-wins the whole document."


13. Interview Talking Points

  • Shard by document.
  • OT or CRDT, not LWW file.
  • Snapshot + op log.
  • Presence ≠ history.
  • Sticky WebSocket.
  • Metrics: op latency, WS disconnects, persist lag, snapshot size.
  • Same real-time pattern as chat.

14. Follow-up Questions

Million users in one doc?
That is broadcast, not editing. Fanout like a news feed; freeze editors.

Rich text?
Ops on a tree (insert node), same sequencing.

Mobile flaky network?
Client buffer, backoff, CRDT nicer.

Encryption?
E2E makes server merge hard. Usually encrypt at rest; collaboration needs plaintext on server unless CRDT-in-client only.


Related JavaThoughts reading:


That completes the top 20 questions on the System Design landing page.