Javathoughts Logo
Javathoughts
Published on
Views

Design File Storage like Dropbox System Design Interview Guide

Authors
  • avatar
    Name
    Javed Shaikh
    Twitter

← System Design Interview Preparation

This guide walks through Design File Storage like Dropbox the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.


1. Problem

A Dropbox-like product lets a user put files in the cloud, open them from another device, and share a folder with a teammate.

Example:

  • You drop resume.pdf into /Work on your laptop
  • The file is split into chunks and uploaded
  • Your phone gets a sync event and shows the same file
  • A friend with a share link can download it, but cannot delete your originals

The hard part is not “store a file.” The hard part is metadata + sync. Bytes live in object storage. Names, folders, versions, and permissions live in a database. Devices stay in sync with events, not by scanning the whole account every time.

In this design we are building a system that:

  1. Uploads and downloads files
  2. Keeps folder structure and file metadata
  3. Shares files with permission checks
  4. Syncs changes across devices
  5. Supports chunked upload, versioning, and basic deduplication

We are not building a full office suite, virus scanning farm, or billing product.


2. Functional Requirements / FR

RequirementWhat it means
Upload fileClient sends bytes (or chunks) and gets a file id back.
Download fileOwner or permitted user can fetch the latest version.
File metadataName, size, mime type, owner, folder, timestamps.
Folder structureNested folders. Move and rename without copying bytes.
File sharingLink share or user share, with view vs edit.
Sync across devicesOther devices learn about creates, updates, deletes.
Chunked uploadsLarge files upload in parts and can resume.
VersioningKeep previous versions for a while. Restore is in scope.
DeduplicationSame chunk hash can reuse stored bytes.
Permission checksEvery download and share must check access.

Out of scope for a 45-minute interview:

  • Real-time collaborative editing
  • Full-text search inside every PDF
  • Desktop filesystem driver details

Confirm scope. Then lock it.


3. Non-Functional Requirements / NFR

RequirementWhy it matters
Durable storageUsers treat this as a backup. Losing bytes is a product-killing bug.
Upload reliabilityHome Wi-Fi drops. Chunks must resume.
Download latencyOpening a photo should feel local after the first fetch. Use a CDN.
Metadata consistencyTwo devices should not disagree on “which version is latest.”
Sync freshnessSeconds of delay is fine. Hours is not.
SecurityPermission checks on metadata and on signed download URLs.
ScaleLots of small files. Metadata QPS will hurt before raw disk does.

A good interview sentence: object storage holds bytes. The database holds truth about names, versions, and who can see what.


4. Back-of-the-Envelope Calculation

Say these assumptions out loud. Interviewers care more about the method than the exact number.

Traffic assumptions

  • 50 million active users
  • Each user uploads 2 files/day and downloads 10 files/day
  • Average file size 1 MB (mix of docs and photos)
  • Read/write ratio ≈ 10 : 2 = 5 : 1 (downloads vs uploads)

Writes are not tiny, because each upload also writes metadata and maybe many chunks.

QPS

Uploads per day: 50,000,000 × 2 = 100 million

100,000,000 requests/day / 86,400 seconds = around 1,160 QPS

Downloads per day: 50,000,000 × 10 = 500 million

500,000,000 / 86,400 ≈ 5,790 QPS

If peak is 5×, design for:

  • Peak uploads ≈ 6,000 QPS
  • Peak downloads ≈ 29,000 QPS

Metadata reads (list folder, sync cursor) can be even higher than downloads. Mention folder list as a hot API.

Storage

New bytes per day:

100,000,000 files × 1 MB ≈ 100 TB/day

That sounds scary. Dedup, photos that already exist, and version retention policies cut this. In the interview, still show the raw number, then say:

  • Keep N versions, not infinite
  • Dedup by chunk hash
  • Cold versions go to cheaper storage class

Metadata is smaller. Per file ~ 1 KB of DB row + indexes:

100,000,000 × 1 KB ≈ 100 GB/day of metadata growth

Metadata storage is manageable. Object storage cost and CDN bandwidth are the expensive parts.

Cache / memory estimate

Cache:

  • Folder listings for active users
  • Permission results
  • Signed URL targets

Assume 10 million hot folders × 2 KB listing = 20 GB. One Redis cluster is enough to mention.

Server estimate

File metadata service is like a normal Java API. Use 2,000 QPS per server as a conservative number.

Peak metadata ~ 30,000 QPS / 2,000 ≈ 15 servers

Add upload/download proxy capacity, plus workers for sync. About 20–30 app servers plus a worker pool is a fair interview answer. Object storage and CDN are managed services.

These are interview estimates, not exact production numbers.


5. APIs

Keep the API small.

Create upload session

POST /api/files/upload-sessions

{
  "folderId": "fld_9a",
  "fileName": "resume.pdf",
  "sizeBytes": 1048576,
  "mimeType": "application/pdf"
}

Response: uploadId, chunkSize (for example 8 MB).

Upload chunk

PUT /api/files/upload-sessions/:uploadId/chunks/:index

Body: raw bytes. Header: Content-MD5 or x-chunk-hash.

Idempotent: same index + same hash is a no-op.

Complete upload

POST /api/files/upload-sessions/:uploadId/complete

Creates or versions the file. Returns fileId and versionId.

List folder

GET /api/folders/:folderId/items?cursor=

Download

GET /api/files/:fileId/download

Returns 302 to a time-limited signed URL on the CDN / object store. The API never streams 2 GB through Java heap if it can avoid it.

Share

POST /api/files/:fileId/shares

{ "mode": "link", "permission": "view" }

Sync

GET /api/sync?cursor=abc

Returns changed file/folder metadata since that cursor.

Errors: 401, 403 not allowed, 409 conflict on rename, 429 upload flood.


6. Data Model

Do not put file bytes in Postgres. Put pointers.

files

FieldNotes
file_idPK
folder_idParent
owner_id
nameDisplay name
current_version_id
is_deletedSoft delete for trash
updated_at

Unique (folder_id, name) among non-deleted items.

file_versions

FieldNotes
version_id
file_id
size_bytes
created_by
created_at

chunks

FieldNotes
chunk_hashPK, content hash
object_keyLocation in object storage
size_bytes
ref_countFor dedup garbage collection

file_version_chunks

version_id, chunk_index, chunk_hash. Order of chunks rebuilds the file.

folders

folder_id, parent_id, owner_id, name. Root folder per user.

permissions

resource_id, resource_type (file/folder), grantee, role (view/edit).

devices / sync_cursors

Each device stores a cursor. The server keeps an event log per namespace: file created, renamed, version added.

Index folder listings on (folder_id, name). Index sync on (namespace_id, event_id).


7. High-Level Design

Bytes go to object storage. Metadata goes to the database. Sync is async.

File Storage like Dropbox architecture

File Storage like Dropbox architectureUploads go to object storage. Metadata lives in the database. Sync workers notify other devices. Downloads usually come from a CDN.uploadmetadatadownloadfile change👤Client🌐API Gateway⚙️File Service📦Object Storage🗄️Metadata DB🌍CDN📩Queue👷Sync Worker🔔Device Notification
Uploads go to object storage. Metadata lives in the database. Sync workers notify other devices. Downloads usually come from a CDN.

Components:

  • Client: desktop, mobile, or web. Splits large files into chunks.
  • API Gateway: auth, TLS, rate limits. See API Gateway.
  • File Service: metadata, permissions, upload session state. A Java / Spring Boot service is enough.
  • Object Storage: S3-style durable blobs. This is why we do not store 1 MB rows in MySQL.
  • Metadata DB: folders, versions, ACLs, chunk maps.
  • CDN: download path. Users should not all hit the origin bucket.
  • Queue: file-change events. Kafka is a common choice.
  • Sync Worker: fans events out to device notifications.
  • Notification Service: “your other laptop has a new file.” Same idea as the notification system.

Upload path:

  1. Client starts a session.
  2. Chunks go to File Service, then object storage (or direct-to-storage with a signed upload URL).
  3. Complete upload writes version + chunk map in the DB.
  4. Event goes on the queue.

Download path: permission check → signed URL → CDN / object storage.

Why a queue? Sync should not sit on the upload HTTP request. That is event-driven architecture.


8. Deep Dives

Chunked uploads

8 MB chunks (example):

  • Retry one chunk, not a 2 GB file
  • Resume after laptop sleep
  • Dedup per chunk: two videos that share a trailer can share those hashes

Direct-to-object-storage uploads with a signed URL keep heavy bytes off your Java servers.

Deduplication

Hash the chunk (SHA-256 is an interview-friendly answer). If chunk_hash exists, bump ref_count and skip the blob write. Privacy note: dedup across users can leak existence of a file. In interviews, say dedup within an account or company first. Cross-tenant dedup is a later debate.

Versioning

Each complete upload creates a new version_id and points files.current_version_id at it. Restore = switch the pointer. Old blobs stay until a GC job sees ref_count = 0.

Permission checks

Walk from file to folder to root, or denormalize an ACL on every item when sharing a folder. Folder share is the painful case. Check before issuing a signed download URL. Short URL TTL (5–15 minutes) limits leaked links.

Sync across devices

Do not ask the phone to list the entire tree. Give an event cursor. Desktop holds a local snapshot. Conflicts: last-writer-wins on versions, or keep both files (resume (conflict).pdf). Mention conflict handling in one sentence.

CDN for downloads

Hot files (shared memes, company templates) belong on the edge. Signed cookies or signed URLs keep the CDN from becoming a public anonymous dump.


9. Bottlenecks

  • Metadata DB on list folder for huge directories
  • Upload complete transaction that writes thousands of chunk rows
  • Hot shared file download origin if CDN TTL is wrong
  • Sync event fan-out to many devices
  • Small-file problem: millions of 2 KB files waste object-store request costs

Mitigations: paginate folders, batch chunk rows, CDN, partition the event log by user, pack tiny files into larger blobs if the interviewer wants extra credit.


10. Tradeoffs

ChoiceUpsideDownside
Direct upload to object storageCheap app serversHarder to scan content on the way in
SQL metadataStrong folder constraintsBig namespaces need sharding
Cross-user chunk dedupSaves storageSecurity/privacy risk
Infinite versionsUser loveCost explodes
Strong consistency on renameNo split-brain namesHigher latency on sync

Pick: SQL for metadata, object store for bytes, async sync, per-account dedup, limited versions.


11. Failure Modes

FailureWhat users seeWhat you do
Object store downUploads failFail the session, resume later. Do not mark file complete.
Metadata DB downApp looks emptyServe cached folder lists read-only if you must. Do not accept completes.
Queue downOther devices lagUploads still succeed. Sync catches up.
Lost chunkFile cannot rebuildComplete must verify all chunk hashes exist.
Stolen share linkData leakExpire links, revoke, short signed URL TTL.
Two devices edit offlineConflict copiesKeep both versions, show the user.

Never mark a file “ready” until every chunk is durable.


12. Interview Answer in 10 Minutes

"I would split Dropbox into bytes and metadata.

Assume 50 million users, 2 uploads and 10 downloads each per day. That is about 100 million uploads a day, around 1,160 upload QPS, and about 6,000 QPS at 5× peak. Downloads are around 6,000 QPS average and about 30,000 at peak. New bytes can look like 100 TB a day before dedup and retention, so object storage plus a CDN matters more than a single disk.

The client uploads chunks, often with signed URLs straight to object storage. A file service records folders, file names, versions, and a chunk list in a database. Completing an upload is the consistency point: all chunks must exist, then we point the file at a new version.

Downloads check permissions, then redirect to a signed CDN URL. We do not stream large files through the API JVM.

Sync is an event log. A worker notifies other devices. Kafka-style queues keep upload HTTP fast.

Dedup is by chunk hash inside an account. Versions are pointers. Sharing is ACL rows plus short-lived URLs.

If the queue is down, uploads still work. If metadata is down, we stop completes. Bytes never live only on one app server."

Practice this until it is under 10 minutes.


13. Interview Talking Points

  • Bytes in object storage, names in a database.
  • Chunked, resumable upload.
  • Signed URLs + CDN for download.
  • Version = new pointer, not a rewrite in place.
  • Per-account dedup before cross-tenant dedup.
  • Cursor-based sync, not full tree scan.
  • Permissions before the signed URL.
  • Metrics: upload success, complete latency, sync lag, 403 rate, object-store errors.
  • This is a distributed systems + microservices story, not a single disk.

14. Follow-up Questions

How do you upload a 10 GB video on bad Wi-Fi?
Chunk it. Persist session. Retry only missing indexes. Complete when the manifest matches.

How do you share a folder with 50 people?
ACL on the folder, inherit on list/download. Cache permission grants. Revoke by bumping a permission version so old signed URLs die.

How do you garbage-collect unused chunks?
ref_count. A slow GC deletes blobs at 0 after a grace period, in case a complete is in flight.

Block vs object storage?
Object storage is the interview default: cheap, durable, HTTP-friendly. Block storage is for disks attached to VMs, not user files.

Java details?
Spring Boot file service, Redis for sessions and hot folders, Kafka for sync events. See message queues.


Related JavaThoughts reading:


Next in this series: Design Video Streaming like YouTube / Netflix.