- Published on
- Views
Design File Storage like Dropbox System Design Interview Guide
- Authors

- Name
- Javed Shaikh
← System Design Interview Preparation
This guide walks through Design File Storage like Dropbox the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.
1. Problem
A Dropbox-like product lets a user put files in the cloud, open them from another device, and share a folder with a teammate.
Example:
- You drop
resume.pdfinto/Workon your laptop - The file is split into chunks and uploaded
- Your phone gets a sync event and shows the same file
- A friend with a share link can download it, but cannot delete your originals
The hard part is not “store a file.” The hard part is metadata + sync. Bytes live in object storage. Names, folders, versions, and permissions live in a database. Devices stay in sync with events, not by scanning the whole account every time.
In this design we are building a system that:
- Uploads and downloads files
- Keeps folder structure and file metadata
- Shares files with permission checks
- Syncs changes across devices
- Supports chunked upload, versioning, and basic deduplication
We are not building a full office suite, virus scanning farm, or billing product.
2. Functional Requirements / FR
| Requirement | What it means |
|---|---|
| Upload file | Client sends bytes (or chunks) and gets a file id back. |
| Download file | Owner or permitted user can fetch the latest version. |
| File metadata | Name, size, mime type, owner, folder, timestamps. |
| Folder structure | Nested folders. Move and rename without copying bytes. |
| File sharing | Link share or user share, with view vs edit. |
| Sync across devices | Other devices learn about creates, updates, deletes. |
| Chunked uploads | Large files upload in parts and can resume. |
| Versioning | Keep previous versions for a while. Restore is in scope. |
| Deduplication | Same chunk hash can reuse stored bytes. |
| Permission checks | Every download and share must check access. |
Out of scope for a 45-minute interview:
- Real-time collaborative editing
- Full-text search inside every PDF
- Desktop filesystem driver details
Confirm scope. Then lock it.
3. Non-Functional Requirements / NFR
| Requirement | Why it matters |
|---|---|
| Durable storage | Users treat this as a backup. Losing bytes is a product-killing bug. |
| Upload reliability | Home Wi-Fi drops. Chunks must resume. |
| Download latency | Opening a photo should feel local after the first fetch. Use a CDN. |
| Metadata consistency | Two devices should not disagree on “which version is latest.” |
| Sync freshness | Seconds of delay is fine. Hours is not. |
| Security | Permission checks on metadata and on signed download URLs. |
| Scale | Lots of small files. Metadata QPS will hurt before raw disk does. |
A good interview sentence: object storage holds bytes. The database holds truth about names, versions, and who can see what.
4. Back-of-the-Envelope Calculation
Say these assumptions out loud. Interviewers care more about the method than the exact number.
Traffic assumptions
- 50 million active users
- Each user uploads 2 files/day and downloads 10 files/day
- Average file size 1 MB (mix of docs and photos)
- Read/write ratio ≈ 10 : 2 = 5 : 1 (downloads vs uploads)
Writes are not tiny, because each upload also writes metadata and maybe many chunks.
QPS
Uploads per day: 50,000,000 × 2 = 100 million
100,000,000 requests/day / 86,400 seconds = around 1,160 QPS
Downloads per day: 50,000,000 × 10 = 500 million
500,000,000 / 86,400 ≈ 5,790 QPS
If peak is 5×, design for:
- Peak uploads ≈ 6,000 QPS
- Peak downloads ≈ 29,000 QPS
Metadata reads (list folder, sync cursor) can be even higher than downloads. Mention folder list as a hot API.
Storage
New bytes per day:
100,000,000 files × 1 MB ≈ 100 TB/day
That sounds scary. Dedup, photos that already exist, and version retention policies cut this. In the interview, still show the raw number, then say:
- Keep N versions, not infinite
- Dedup by chunk hash
- Cold versions go to cheaper storage class
Metadata is smaller. Per file ~ 1 KB of DB row + indexes:
100,000,000 × 1 KB ≈ 100 GB/day of metadata growth
Metadata storage is manageable. Object storage cost and CDN bandwidth are the expensive parts.
Cache / memory estimate
Cache:
- Folder listings for active users
- Permission results
- Signed URL targets
Assume 10 million hot folders × 2 KB listing = 20 GB. One Redis cluster is enough to mention.
Server estimate
File metadata service is like a normal Java API. Use 2,000 QPS per server as a conservative number.
Peak metadata ~ 30,000 QPS / 2,000 ≈ 15 servers
Add upload/download proxy capacity, plus workers for sync. About 20–30 app servers plus a worker pool is a fair interview answer. Object storage and CDN are managed services.
These are interview estimates, not exact production numbers.
5. APIs
Keep the API small.
Create upload session
POST /api/files/upload-sessions
{
"folderId": "fld_9a",
"fileName": "resume.pdf",
"sizeBytes": 1048576,
"mimeType": "application/pdf"
}
Response: uploadId, chunkSize (for example 8 MB).
Upload chunk
PUT /api/files/upload-sessions/:uploadId/chunks/:index
Body: raw bytes. Header: Content-MD5 or x-chunk-hash.
Idempotent: same index + same hash is a no-op.
Complete upload
POST /api/files/upload-sessions/:uploadId/complete
Creates or versions the file. Returns fileId and versionId.
List folder
GET /api/folders/:folderId/items?cursor=
Download
GET /api/files/:fileId/download
Returns 302 to a time-limited signed URL on the CDN / object store. The API never streams 2 GB through Java heap if it can avoid it.
Share
POST /api/files/:fileId/shares
{ "mode": "link", "permission": "view" }
Sync
GET /api/sync?cursor=abc
Returns changed file/folder metadata since that cursor.
Errors: 401, 403 not allowed, 409 conflict on rename, 429 upload flood.
6. Data Model
Do not put file bytes in Postgres. Put pointers.
files
| Field | Notes |
|---|---|
file_id | PK |
folder_id | Parent |
owner_id | |
name | Display name |
current_version_id | |
is_deleted | Soft delete for trash |
updated_at |
Unique (folder_id, name) among non-deleted items.
file_versions
| Field | Notes |
|---|---|
version_id | |
file_id | |
size_bytes | |
created_by | |
created_at |
chunks
| Field | Notes |
|---|---|
chunk_hash | PK, content hash |
object_key | Location in object storage |
size_bytes | |
ref_count | For dedup garbage collection |
file_version_chunks
version_id, chunk_index, chunk_hash. Order of chunks rebuilds the file.
folders
folder_id, parent_id, owner_id, name. Root folder per user.
permissions
resource_id, resource_type (file/folder), grantee, role (view/edit).
devices / sync_cursors
Each device stores a cursor. The server keeps an event log per namespace: file created, renamed, version added.
Index folder listings on (folder_id, name). Index sync on (namespace_id, event_id).
7. High-Level Design
Bytes go to object storage. Metadata goes to the database. Sync is async.
File Storage like Dropbox architecture
Components:
- Client: desktop, mobile, or web. Splits large files into chunks.
- API Gateway: auth, TLS, rate limits. See API Gateway.
- File Service: metadata, permissions, upload session state. A Java / Spring Boot service is enough.
- Object Storage: S3-style durable blobs. This is why we do not store 1 MB rows in MySQL.
- Metadata DB: folders, versions, ACLs, chunk maps.
- CDN: download path. Users should not all hit the origin bucket.
- Queue: file-change events. Kafka is a common choice.
- Sync Worker: fans events out to device notifications.
- Notification Service: “your other laptop has a new file.” Same idea as the notification system.
Upload path:
- Client starts a session.
- Chunks go to File Service, then object storage (or direct-to-storage with a signed upload URL).
- Complete upload writes version + chunk map in the DB.
- Event goes on the queue.
Download path: permission check → signed URL → CDN / object storage.
Why a queue? Sync should not sit on the upload HTTP request. That is event-driven architecture.
8. Deep Dives
Chunked uploads
8 MB chunks (example):
- Retry one chunk, not a 2 GB file
- Resume after laptop sleep
- Dedup per chunk: two videos that share a trailer can share those hashes
Direct-to-object-storage uploads with a signed URL keep heavy bytes off your Java servers.
Deduplication
Hash the chunk (SHA-256 is an interview-friendly answer). If chunk_hash exists, bump ref_count and skip the blob write. Privacy note: dedup across users can leak existence of a file. In interviews, say dedup within an account or company first. Cross-tenant dedup is a later debate.
Versioning
Each complete upload creates a new version_id and points files.current_version_id at it. Restore = switch the pointer. Old blobs stay until a GC job sees ref_count = 0.
Permission checks
Walk from file to folder to root, or denormalize an ACL on every item when sharing a folder. Folder share is the painful case. Check before issuing a signed download URL. Short URL TTL (5–15 minutes) limits leaked links.
Sync across devices
Do not ask the phone to list the entire tree. Give an event cursor. Desktop holds a local snapshot. Conflicts: last-writer-wins on versions, or keep both files (resume (conflict).pdf). Mention conflict handling in one sentence.
CDN for downloads
Hot files (shared memes, company templates) belong on the edge. Signed cookies or signed URLs keep the CDN from becoming a public anonymous dump.
9. Bottlenecks
- Metadata DB on
list folderfor huge directories - Upload complete transaction that writes thousands of chunk rows
- Hot shared file download origin if CDN TTL is wrong
- Sync event fan-out to many devices
- Small-file problem: millions of 2 KB files waste object-store request costs
Mitigations: paginate folders, batch chunk rows, CDN, partition the event log by user, pack tiny files into larger blobs if the interviewer wants extra credit.
10. Tradeoffs
| Choice | Upside | Downside |
|---|---|---|
| Direct upload to object storage | Cheap app servers | Harder to scan content on the way in |
| SQL metadata | Strong folder constraints | Big namespaces need sharding |
| Cross-user chunk dedup | Saves storage | Security/privacy risk |
| Infinite versions | User love | Cost explodes |
| Strong consistency on rename | No split-brain names | Higher latency on sync |
Pick: SQL for metadata, object store for bytes, async sync, per-account dedup, limited versions.
11. Failure Modes
| Failure | What users see | What you do |
|---|---|---|
| Object store down | Uploads fail | Fail the session, resume later. Do not mark file complete. |
| Metadata DB down | App looks empty | Serve cached folder lists read-only if you must. Do not accept completes. |
| Queue down | Other devices lag | Uploads still succeed. Sync catches up. |
| Lost chunk | File cannot rebuild | Complete must verify all chunk hashes exist. |
| Stolen share link | Data leak | Expire links, revoke, short signed URL TTL. |
| Two devices edit offline | Conflict copies | Keep both versions, show the user. |
Never mark a file “ready” until every chunk is durable.
12. Interview Answer in 10 Minutes
"I would split Dropbox into bytes and metadata.
Assume 50 million users, 2 uploads and 10 downloads each per day. That is about 100 million uploads a day, around 1,160 upload QPS, and about 6,000 QPS at 5× peak. Downloads are around 6,000 QPS average and about 30,000 at peak. New bytes can look like 100 TB a day before dedup and retention, so object storage plus a CDN matters more than a single disk.
The client uploads chunks, often with signed URLs straight to object storage. A file service records folders, file names, versions, and a chunk list in a database. Completing an upload is the consistency point: all chunks must exist, then we point the file at a new version.
Downloads check permissions, then redirect to a signed CDN URL. We do not stream large files through the API JVM.
Sync is an event log. A worker notifies other devices. Kafka-style queues keep upload HTTP fast.
Dedup is by chunk hash inside an account. Versions are pointers. Sharing is ACL rows plus short-lived URLs.
If the queue is down, uploads still work. If metadata is down, we stop completes. Bytes never live only on one app server."
Practice this until it is under 10 minutes.
13. Interview Talking Points
- Bytes in object storage, names in a database.
- Chunked, resumable upload.
- Signed URLs + CDN for download.
- Version = new pointer, not a rewrite in place.
- Per-account dedup before cross-tenant dedup.
- Cursor-based sync, not full tree scan.
- Permissions before the signed URL.
- Metrics: upload success, complete latency, sync lag, 403 rate, object-store errors.
- This is a distributed systems + microservices story, not a single disk.
14. Follow-up Questions
How do you upload a 10 GB video on bad Wi-Fi?
Chunk it. Persist session. Retry only missing indexes. Complete when the manifest matches.
How do you share a folder with 50 people?
ACL on the folder, inherit on list/download. Cache permission grants. Revoke by bumping a permission version so old signed URLs die.
How do you garbage-collect unused chunks?ref_count. A slow GC deletes blobs at 0 after a grace period, in case a complete is in flight.
Block vs object storage?
Object storage is the interview default: cheap, durable, HTTP-friendly. Block storage is for disks attached to VMs, not user files.
Java details?
Spring Boot file service, Redis for sessions and hot folders, Kafka for sync events. See message queues.
15. Internal Links
Related JavaThoughts reading:
- System Design Interview Preparation
- Design a Notification System
- Design a Distributed Cache
- Design a Rate Limiter
- Java
- Spring Boot
- Kafka
- Microservices
- Distributed Systems
- API Gateway
- Caching
- Event-driven architecture
- 20 system design concepts
Next in this series: Design Video Streaming like YouTube / Netflix.
