Javathoughts Logo
Javathoughts
Published on
Views

Design Video Streaming like YouTube / Netflix System Design Interview Guide

Authors
  • avatar
    Name
    Javed Shaikh
    Twitter

← System Design Interview Preparation

This guide walks through Design Video Streaming like YouTube / Netflix the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.


1. Problem

A video product has two jobs: creators upload once, and viewers watch many times, on bad networks, without buffering forever.

YouTube is closer to “anyone uploads, anyone watches.” Netflix is closer to “catalog is ingested, then streamed.” The architecture shares the same spine:

  1. Store the raw file
  2. Transcode into many bitrates and resolutions
  3. Package for adaptive streaming
  4. Push playback to a CDN
  5. Keep metadata (title, duration, owner) in a database

We will design that spine. Recommendation can be a short mention. Likes, comments, and copyright are extra, not the first 20 minutes.


2. Functional Requirements / FR

RequirementWhat it means
Upload videoCreator sends a file (or a direct-to-storage upload).
TranscodingTurn one upload into 360p, 720p, 1080p, maybe 4K, plus audio.
Multiple resolutionsPlayer can switch quality.
Video metadataTitle, description, duration, visibility, owner.
Watch / playbackReturn a playlist (HLS/DASH) and stream segments.
CDN deliverySegments served from an edge near the viewer.
Adaptive bitratePlayer picks a bitrate from bandwidth and buffer.
Watch eventsViews, playhead, errors — for analytics.
Copyright / abuseHigh level: scan and take down. Not a full Content ID design.
RecommendationsMention a ranking service. Do not build the ML platform.

Out of scope: live sports latency, DRM license server internals, studio ingest trucks.


3. Non-Functional Requirements / NFR

RequirementWhy it matters
Playback start timeFirst frame in a couple of seconds.
Smooth streamingAvoid rebuffering. CDN + ABR matter more than one fat origin.
Upload durabilityDo not lose a 2-hour file after 99% upload.
Transcode throughputA backlog after a viral day should drain.
Read-heavyWatch QPS dwarfs upload QPS.
AvailabilityIf metadata is down, the catalog looks empty even if CDN is fine.
Cost controlTranscode and egress money can exceed compute money.

Interview line: optimize the watch path. Upload and transcode can be async.


4. Back-of-the-Envelope Calculation

Traffic assumptions

  • 50 million daily active viewers
  • 5 videos watched per user per day
  • Average watch 10 minutes, but we count video start APIs, not every second
  • 100,000 new uploads per day (YouTube-like; Netflix ingest is smaller)
  • Read/write ratio ≈ 250 million watches / 100,000 uploads ≈ 2,500 : 1

This is extremely read-heavy.

QPS

Watch starts per day: 50,000,000 × 5 = 250 million

250,000,000 requests/day / 86,400 seconds = around 2,890 QPS

If peak is 5×, design for around 14,500 QPS of “start playback / fetch playlist.”

Segment fetches are higher (a player may request a small file every few seconds). Say 20× playlist QPS → peak ~300,000 QPS of tiny HTTP GETs, almost all on the CDN, not on your origin.

Uploads:

100,000 / 86,400 ≈ 1.2 QPS average, maybe 10 QPS peak

Upload QPS is tiny. Upload bandwidth is not: if average upload is 500 MB:

100,000 × 500 MB ≈ 50 TB/day ingested

Storage

Raw: 50 TB/day. Transcode often multiplies storage (several renditions). Use 3× as a round number:

50 TB × 3 ≈ 150 TB/day of video objects

Metadata: 100,000 rows/day is nothing. Video DB is not the storage problem. Object storage is.

Cache / memory

  • CDN caches hot segments (the real cache)
  • Origin / Redis: metadata, playlist manifests for hot videos

Cache 1 million hot video metadata records × 2 KB ≈ 2 GB. Easy.

CDN cache size is “as large as the vendor gives you.” In the interview, say hot 10% of catalog gets 90% of watches.

Server estimate

Playlist/metadata API at 15,000 peak QPS / 2,000 QPS per server ≈ 8 servers, plus extras → 12–15.

Transcoding is CPU/GPU heavy. 100,000 videos/day. If one worker does a video in 10 minutes (0.1 hours):

100,000 × 0.1 hour ≈ 10,000 worker-hours/day
10,000 / 24 ≈ 420 workers running 24/7

Peak after evenings needs more. Quote hundreds of transcode workers, autoscaled off a queue.

These are interview estimates, not exact production numbers.


5. APIs

Start upload

POST /api/videos

{
  "title": "System design in 10 minutes",
  "visibility": "public"
}

Returns videoId, uploadUrl (signed).

Finish upload

POST /api/videos/:videoId/complete-upload

Sets status processing and enqueues transcode.

Get watch session

GET /api/videos/:videoId/playback

{
  "videoId": "vid_1",
  "status": "ready",
  "durationSec": 612,
  "manifestUrl": "https://cdn.example/vid_1/master.m3u8"
}

The player then talks to the CDN, not to the video service, for segments.

Update metadata

PATCH /api/videos/:videoId

Analytics (client)

POST /api/events/playback — fire-and-forget. Never block the player.

Errors: 403 private video, 404, 409 still processing, 429.


6. Data Model

videos

FieldNotes
video_id
owner_id
title
statusuploading, processing, ready, failed, blocked
duration_secSet after transcode
visibilitypublic, unlisted, private
created_at

video_assets

FieldNotes
asset_id
video_id
kindraw, video_720p, video_1080p, audio, thumbnail
object_key
bitrate
width, height

playback_manifests

Pointer to master.m3u8 / DASH MPD on the CDN.

moderation_flags

video_id, reason, state. High-level copyright/abuse.

Watch events do not belong in videos. Send them to an event store.


7. High-Level Design

Upload is a pipeline. Playback is CDN.

Video Streaming architecture

Video Streaming architectureCreators upload raw video. Workers transcode many bitrates. Viewers stream processed files from a CDN. Analytics stay off the playback path.transcodemetadatawatch events👤Creator🌐API Gateway⚙️Video Service📦Raw Video Storage📩Queue👷Transcoding Workers📦Processed Video👤Viewer🌍CDN🗄️Video DB📩Event Queue📊Analytics Store
Creators upload raw video. Workers transcode many bitrates. Viewers stream processed files from a CDN. Analytics stay off the playback path.

Components:

  • Creator: uploads to a signed URL.
  • API Gateway: auth and rate limits.
  • Video Service: metadata and job creation. Fine as Spring Boot.
  • Raw Video Storage: original blob. Keep it for re-transcode.
  • Queue: transcode jobs. Kafka or any durable queue.
  • Transcoding Workers: FFmpeg-style jobs. Many bitrates out.
  • Processed Video Storage: segments + manifests.
  • CDN: what the viewer actually hits.
  • Video DB: title, status, pointers.
  • Event Queue + Analytics Store: views and QoE (startup time, rebuffers).

Playback flow:

  1. App asks Video Service for playback.
  2. If ready, return manifestUrl on the CDN.
  3. Player uses adaptive bitrate streaming: if bandwidth drops, request 360p segments.

Why workers? Transcode can take minutes. The upload HTTP call must not wait. Same pattern as event-driven architecture.

Netflix-style catalog ingest is the same pipeline with fewer random creators. For a fuller Netflix picture, see The Netflix Tech Stack.


8. Deep Dives

Transcoding and multiple resolutions

One source file → ladder: 360p, 480p, 720p, 1080p. Each is split into 2–6 second segments. The master playlist lists them.

If you only stored 1080p, a phone on 3G would buffer forever. The extra copies cost storage. They buy watchable video.

Adaptive bitrate streaming

The player measures download speed and buffer. It requests the next segment from a lower or higher ladder rung. Your backend's job is to publish a correct playlist, not to pick the bitrate per request.

CDN delivery

Almost all bytes leave from the edge. Origin fetch happens on a cache miss (new episode drop, or a long-tail video). Pre-warm CDN for a Friday Netflix drop. YouTube long-tail will miss more often; still do not send every byte through the video service.

Recommendations (brief)

A separate ranker takes user history + candidate videos. It is not on the segment path. If ranker is down, show trending/latest.

Likes, comments, views

Views: eventual. A counter service + cache. Do not write videos.view_count += 1 on every play.

Comments: another service. Mention it, do not design it unless asked.

On complete-upload, enqueue a scan job (hash match, audio fingerprint). If it hits, set blocked and stop CDN distribution. Manual review queue for appeals. Keep this as one box on the whiteboard.

Analytics

Client sends play, pause, rebuffer, error. Queue → analytics store. Product uses this for QoE, not for the first frame.


9. Bottlenecks

  • Transcode queue lag after a viral upload day
  • CDN origin when a new episode is watched everywhere at once
  • Hot metadata for a trending video id
  • Analytics write storm if you log every segment request
  • Storage cost if you never delete old renditions

Mitigations: autoscale workers, pre-warm CDN, cache playback API, sample analytics, lifecycle policies on raw files.


10. Tradeoffs

ChoiceUpsideDownside
Many bitratesSmooth ABRStorage and transcode cost
Pre-transcode allSimple playbackSlow “video ready”
Just-in-time transcodeFast publishFirst viewers suffer
Push CDNGreat for Netflix dropsWaste if nobody watches
Pull CDNGood for YouTube long-tailFirst hit is slower
301/immutable segmentsCache foreverHard to fix a bad encode

Interview pick: pre-transcode a standard ladder, pull CDN for UGC, pre-warm for known premieres.


11. Failure Modes

FailureUser impactHandling
Transcode worker crashVideo stays processingJob is retryable; raw blob is source of truth
Bad encodePlayback errorsKeep raw, re-queue, mark failed if repeated
CDN outage in one regionBufferingDNS/failover to another POP or origin (origin will hurt)
Video DB downCannot start new watchesCached playback URLs for hot titles may still work
Event queue downCounts freezePlayback continues
Copyright false positiveAngry creatorAppeal queue, do not delete raw immediately

Playback must not depend on analytics or recommendations.


12. Interview Answer in 10 Minutes

"I would design this as an upload pipeline plus a CDN watch path.

Assume 50 million daily viewers and 5 watches each. That is 250 million starts a day, about 2,900 QPS, around 15,000 QPS at 5× peak. Segment traffic is much higher but belongs on the CDN. Uploads might be 100,000 a day, only a few QPS, but tens of terabytes in. After transcode, storage can be ~150 TB/day. These are interview estimates.

Creator gets a signed upload URL. Raw file lands in object storage. Video service writes metadata as processing and puts a job on a queue. Workers transcode a bitrate ladder and write segments. Status becomes ready.

Viewer calls playback API, gets a manifest URL, and streams from the CDN with adaptive bitrate. Metadata lives in the video DB. Watch events go to a separate queue so they cannot stall the player.

Recommendations and comments are side systems. Copyright is an async scan that can flip status to blocked.

If transcode lags, we scale workers. If CDN misses, origin must be protected. The Java API tier only handles metadata and job control, not 4K bytes."

Practice this until it is under 10 minutes.


13. Interview Talking Points

  • Read-heavy. Watch path ≠ upload path.
  • Never transcode on the request thread.
  • ABR + CDN are the streaming design, not one MP4 URL.
  • Raw blob is sacred so you can re-encode.
  • Analytics are async.
  • Recommendations can fail open.
  • Metrics: time-to-ready, playlist latency, rebuffer ratio, queue lag, origin bandwidth.
  • Related reading: 20 system design concepts, distributed systems.

14. Follow-up Questions

YouTube vs Netflix?
YouTube: huge UGC, unpredictable hot videos, pull CDN, more moderation. Netflix: catalog ingest, premieres, pre-warm CDN, DRM more important.

How do you handle 4K?
Extra ladder rung, more storage, device checks, often restricted to capable clients.

Live streaming?
Different: ingest → packager → low-latency CDN. DVR is extra. Do not pretend VOD workers are live.

How are views counted?
Sampled events, not a SQL increment per segment. Dedup a viewer id for a time window.

Java role?
API + workers with a transcode fleet. See microservices and Kafka.


Related JavaThoughts reading:


Next in this series: Design Search Autocomplete.