- Published on
- Views
Design a Collaborative Document Editor System Design Interview Guide
- Authors

- Name
- Javed Shaikh
← System Design Interview Preparation
This guide walks through Design a Collaborative Document Editor the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.
1. Problem
Many people type in the same document at once. Everyone should see edits quickly. The file should not corrupt.
Example:
- You and a teammate open a spec
- Cursors and names show presence
- You insert a sentence while they delete a word
- The server merges operations
- History can restore an hour ago
- Offline, you type on a plane and sync later (simple version)
This is closer to chat (WebSockets) plus a merge algorithm, not a Google Docs clone of every feature.
2. Functional Requirements / FR
| Requirement | What it means |
|---|---|
| Real-time editing | Ops in ~100ms on a good network. |
| Multi-user same doc | N editors per document. |
| WebSockets | Push ops and presence. |
| OT or CRDT | Explain one simply. |
| Snapshots | Periodic full document. |
| Version history | Restore snapshot. |
| Presence / cursors | Ephemeral, not in the doc history. |
| Conflict resolution | Algorithm, not “last write wins” on the whole file. |
| Offline (brief) | Queue local ops, sync on reconnect. |
| Permissions | Owner / editor / viewer. |
Out of scope: comments, suggestions mode, Excel formulas, 100k-user live lecture (say sharding by doc).
3. Non-Functional Requirements / NFR
| Requirement | Why it matters |
|---|---|
| Low latency ops | Typing should feel local. |
| Correctness | Same sequence → same doc on all clients. |
| Doc affinity | All ops for a doc hit one collab process (or a CRDT that can merge anywhere). |
| Durability | Snapshots on disk; ops buffered. |
| Scale by document | Hot docs are rare; millions of cold docs. |
Interview line: sticky collaboration service per doc, OT or CRDT, snapshots, WebSockets, presence as a side channel.
4. Back-of-the-Envelope Calculation
Traffic assumptions
- 10 million documents
- 1 million DAU
- 5% editing at once → 50,000 concurrent editors
- Average 1 op/sec while typing (bursts higher)
- Read/write: live ops are writes; open-doc is snapshot read
QPS
Live ops:
50,000 editors × 1 op/sec = 50,000 ops/sec global
If peak is 5x (many people in a meeting doc), design for around 250,000 ops/sec globally.
Per popular doc: maybe 20 editors × 5 ops = 100 ops/sec — small for one process.
Opens:
Say 20 million opens/day / 86,400 ≈ 230 QPS average, ~1,200 at 5× peak.
Storage
Doc average 50 KB. 10M × 50 KB = 500 GB current.
Ops: 50k ops/sec × 100 bytes × 86,400 ≈ 400 GB/day if you kept every op forever. You will not. Keep ops for 24h then snapshot and compact. Interview: tens of GB/day after compaction.
Cache / memory
Active docs in RAM on the collab node. 50k editors / 5 per doc ≈ 10,000 hot docs × 50 KB ≈ 500 MB plus OT state. Fine.
Servers
Shard by docId. 10k hot docs / 500 per box ≈ 20 collab nodes plus WS gateways. Interview estimates, not exact production numbers.
5. APIs
HTTP:
POST /v1/docs
GET /v1/docs/{id} (snapshot + revision)
GET /v1/docs/{id}/history
POST /v1/docs/{id}/acl
WebSocket:
{ "type": "op", "rev": 1842, "payload": { ... } }
{ "type": "ack", "rev": 1843 }
{ "type": "presence", "cursor": 120, "userId": "..." }
{ "type": "snapshot", "rev": 1900, "text": "..." }
REST is for load and permissions. Live edits are WS.
6. Data Model
documents
doc_id, title, revision, snapshot_blob or S3 key, updated_at.
acl
doc_id, user_id, role.
op_log
doc_id, rev, op, user_id, created_at. Truncate after snapshot.
snapshots
doc_id, rev, storage_key.
Presence is not in SQL. It lives in memory / Redis pubsub.
7. High-Level Design
One document = one in-memory state machine.
Collaborative Document Editor architecture
Components:
- Users → WebSocket Gateway → Collaboration Service
- OT/CRDT applied in memory
- Queue → Persistence Worker → Document DB
- Presence broadcast back to users
Join flow:
- Auth (sessions).
- Load snapshot + recent ops.
- Subscribe to
doc:{id}fanout. - Local typing produces ops; send to server; server transforms; broadcast.
Why a worker? Do not fsync every keystroke on the WS thread. Batch ops, snapshot every N seconds or N revs.
8. Deep Dives
OT vs CRDT (simple)
OT (Operational Transform): server orders ops. If you insert at 5 and they delete 3, the server rewrites your index. Needs a single sequencer per doc (the collab process).
CRDT: each character has an id. Merges commute. Easier multi-master / offline. Bigger metadata.
Interview pick: OT with a single leader per document is enough. Mention CRDT for offline and multi-region.
Snapshots
Every 50 ops or 5 seconds: write snapshot, drop old ops. History restore = old snapshot.
Presence
Throttle cursor messages (10 Hz). Do not persist.
Offline
Client queues ops with client ids. On reconnect, send them; OT/CRDT merges. Last-write-wins on the whole doc is the wrong answer.
Permissions
Viewer: WS receive only. Editor: send ops. Check ACL on connect and on share changes.
Sticky routing
Gateway hashes docId to a collab pod. If the pod dies, reload snapshot elsewhere (ops in Kafka help).
WebSockets intro: WebSockets concept.
9. Bottlenecks
- One celebrity doc (keynote) — split not possible for one string; scale vertically, or CRDT with care
- WS connection count on gateway
- Snapshot storm
- Huge paste (cap op size)
10. Tradeoffs
| Choice | Upside | Downside |
|---|---|---|
| OT + leader | Simple reasoning | Leader failover |
| CRDT | Offline, multi-master | Memory, complexity |
| Persist every op | Perfect replay | Cost |
| Snapshot often | Cheap | Coarser history |
Pick: per-doc leader, OT, batched persist, presence separate, ACL on connect.
11. Failure Modes
| Failure | Handling |
|---|---|
| Collab node dies | Clients reconnect, load last snapshot + log |
| Duplicate op | Client op id idempotency |
| Split brain two leaders | Should not: lock docId in Redis |
| Offline merge mess | CRDT or reject and show diff |
| Permission revoked mid-session | Disconnect WS |
12. Interview Answer in 10 Minutes
"I would design a collab editor around a per-document in-memory replica and WebSockets.
Globally you might see 50,000 ops/sec, but one document is tens of ops/sec. I shard by doc id. Clients load a snapshot, then send small operations. A collaboration service sequences them with operational transform so two inserts do not clobber the file. I would mention CRDTs for offline and multi-region.
Presence is a side channel. Persistence is async: a worker writes snapshots and a short op log. History is old snapshots.
If the node dies, clients reconnect and replay from storage. Viewers cannot send ops. We never last-write-wins the whole document."
13. Interview Talking Points
- Shard by document.
- OT or CRDT, not LWW file.
- Snapshot + op log.
- Presence ≠ history.
- Sticky WebSocket.
- Metrics: op latency, WS disconnects, persist lag, snapshot size.
- Same real-time pattern as chat.
14. Follow-up Questions
Million users in one doc?
That is broadcast, not editing. Fanout like a news feed; freeze editors.
Rich text?
Ops on a tree (insert node), same sequencing.
Mobile flaky network?
Client buffer, backoff, CRDT nicer.
Encryption?
E2E makes server merge hard. Usually encrypt at rest; collaboration needs plaintext on server unless CRDT-in-client only.
15. Internal Links
Related JavaThoughts reading:
- System Design Interview Preparation
- Design a Chat Messaging System
- Design Authentication and Session System
- Design a Distributed Queue
- WebSockets
- Java
- Spring Boot
- Kafka
- Microservices
- Distributed Systems
- Message queues
- 20 system design concepts
That completes the top 20 questions on the System Design landing page.
