- Published on
- Views
Design Video Streaming like YouTube / Netflix System Design Interview Guide
- Authors

- Name
- Javed Shaikh
← System Design Interview Preparation
This guide walks through Design Video Streaming like YouTube / Netflix the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.
1. Problem
A video product has two jobs: creators upload once, and viewers watch many times, on bad networks, without buffering forever.
YouTube is closer to “anyone uploads, anyone watches.” Netflix is closer to “catalog is ingested, then streamed.” The architecture shares the same spine:
- Store the raw file
- Transcode into many bitrates and resolutions
- Package for adaptive streaming
- Push playback to a CDN
- Keep metadata (title, duration, owner) in a database
We will design that spine. Recommendation can be a short mention. Likes, comments, and copyright are extra, not the first 20 minutes.
2. Functional Requirements / FR
| Requirement | What it means |
|---|---|
| Upload video | Creator sends a file (or a direct-to-storage upload). |
| Transcoding | Turn one upload into 360p, 720p, 1080p, maybe 4K, plus audio. |
| Multiple resolutions | Player can switch quality. |
| Video metadata | Title, description, duration, visibility, owner. |
| Watch / playback | Return a playlist (HLS/DASH) and stream segments. |
| CDN delivery | Segments served from an edge near the viewer. |
| Adaptive bitrate | Player picks a bitrate from bandwidth and buffer. |
| Watch events | Views, playhead, errors — for analytics. |
| Copyright / abuse | High level: scan and take down. Not a full Content ID design. |
| Recommendations | Mention a ranking service. Do not build the ML platform. |
Out of scope: live sports latency, DRM license server internals, studio ingest trucks.
3. Non-Functional Requirements / NFR
| Requirement | Why it matters |
|---|---|
| Playback start time | First frame in a couple of seconds. |
| Smooth streaming | Avoid rebuffering. CDN + ABR matter more than one fat origin. |
| Upload durability | Do not lose a 2-hour file after 99% upload. |
| Transcode throughput | A backlog after a viral day should drain. |
| Read-heavy | Watch QPS dwarfs upload QPS. |
| Availability | If metadata is down, the catalog looks empty even if CDN is fine. |
| Cost control | Transcode and egress money can exceed compute money. |
Interview line: optimize the watch path. Upload and transcode can be async.
4. Back-of-the-Envelope Calculation
Traffic assumptions
- 50 million daily active viewers
- 5 videos watched per user per day
- Average watch 10 minutes, but we count video start APIs, not every second
- 100,000 new uploads per day (YouTube-like; Netflix ingest is smaller)
- Read/write ratio ≈ 250 million watches / 100,000 uploads ≈ 2,500 : 1
This is extremely read-heavy.
QPS
Watch starts per day: 50,000,000 × 5 = 250 million
250,000,000 requests/day / 86,400 seconds = around 2,890 QPS
If peak is 5×, design for around 14,500 QPS of “start playback / fetch playlist.”
Segment fetches are higher (a player may request a small file every few seconds). Say 20× playlist QPS → peak ~300,000 QPS of tiny HTTP GETs, almost all on the CDN, not on your origin.
Uploads:
100,000 / 86,400 ≈ 1.2 QPS average, maybe 10 QPS peak
Upload QPS is tiny. Upload bandwidth is not: if average upload is 500 MB:
100,000 × 500 MB ≈ 50 TB/day ingested
Storage
Raw: 50 TB/day. Transcode often multiplies storage (several renditions). Use 3× as a round number:
50 TB × 3 ≈ 150 TB/day of video objects
Metadata: 100,000 rows/day is nothing. Video DB is not the storage problem. Object storage is.
Cache / memory
- CDN caches hot segments (the real cache)
- Origin / Redis: metadata, playlist manifests for hot videos
Cache 1 million hot video metadata records × 2 KB ≈ 2 GB. Easy.
CDN cache size is “as large as the vendor gives you.” In the interview, say hot 10% of catalog gets 90% of watches.
Server estimate
Playlist/metadata API at 15,000 peak QPS / 2,000 QPS per server ≈ 8 servers, plus extras → 12–15.
Transcoding is CPU/GPU heavy. 100,000 videos/day. If one worker does a video in 10 minutes (0.1 hours):
100,000 × 0.1 hour ≈ 10,000 worker-hours/day
10,000 / 24 ≈ 420 workers running 24/7
Peak after evenings needs more. Quote hundreds of transcode workers, autoscaled off a queue.
These are interview estimates, not exact production numbers.
5. APIs
Start upload
POST /api/videos
{
"title": "System design in 10 minutes",
"visibility": "public"
}
Returns videoId, uploadUrl (signed).
Finish upload
POST /api/videos/:videoId/complete-upload
Sets status processing and enqueues transcode.
Get watch session
GET /api/videos/:videoId/playback
{
"videoId": "vid_1",
"status": "ready",
"durationSec": 612,
"manifestUrl": "https://cdn.example/vid_1/master.m3u8"
}
The player then talks to the CDN, not to the video service, for segments.
Update metadata
PATCH /api/videos/:videoId
Analytics (client)
POST /api/events/playback — fire-and-forget. Never block the player.
Errors: 403 private video, 404, 409 still processing, 429.
6. Data Model
videos
| Field | Notes |
|---|---|
video_id | |
owner_id | |
title | |
status | uploading, processing, ready, failed, blocked |
duration_sec | Set after transcode |
visibility | public, unlisted, private |
created_at |
video_assets
| Field | Notes |
|---|---|
asset_id | |
video_id | |
kind | raw, video_720p, video_1080p, audio, thumbnail |
object_key | |
bitrate | |
width, height |
playback_manifests
Pointer to master.m3u8 / DASH MPD on the CDN.
moderation_flags
video_id, reason, state. High-level copyright/abuse.
Watch events do not belong in videos. Send them to an event store.
7. High-Level Design
Upload is a pipeline. Playback is CDN.
Video Streaming architecture
Components:
- Creator: uploads to a signed URL.
- API Gateway: auth and rate limits.
- Video Service: metadata and job creation. Fine as Spring Boot.
- Raw Video Storage: original blob. Keep it for re-transcode.
- Queue: transcode jobs. Kafka or any durable queue.
- Transcoding Workers: FFmpeg-style jobs. Many bitrates out.
- Processed Video Storage: segments + manifests.
- CDN: what the viewer actually hits.
- Video DB: title, status, pointers.
- Event Queue + Analytics Store: views and QoE (startup time, rebuffers).
Playback flow:
- App asks Video Service for playback.
- If
ready, returnmanifestUrlon the CDN. - Player uses adaptive bitrate streaming: if bandwidth drops, request 360p segments.
Why workers? Transcode can take minutes. The upload HTTP call must not wait. Same pattern as event-driven architecture.
Netflix-style catalog ingest is the same pipeline with fewer random creators. For a fuller Netflix picture, see The Netflix Tech Stack.
8. Deep Dives
Transcoding and multiple resolutions
One source file → ladder: 360p, 480p, 720p, 1080p. Each is split into 2–6 second segments. The master playlist lists them.
If you only stored 1080p, a phone on 3G would buffer forever. The extra copies cost storage. They buy watchable video.
Adaptive bitrate streaming
The player measures download speed and buffer. It requests the next segment from a lower or higher ladder rung. Your backend's job is to publish a correct playlist, not to pick the bitrate per request.
CDN delivery
Almost all bytes leave from the edge. Origin fetch happens on a cache miss (new episode drop, or a long-tail video). Pre-warm CDN for a Friday Netflix drop. YouTube long-tail will miss more often; still do not send every byte through the video service.
Recommendations (brief)
A separate ranker takes user history + candidate videos. It is not on the segment path. If ranker is down, show trending/latest.
Likes, comments, views
Views: eventual. A counter service + cache. Do not write videos.view_count += 1 on every play.
Comments: another service. Mention it, do not design it unless asked.
Copyright / abuse (high level)
On complete-upload, enqueue a scan job (hash match, audio fingerprint). If it hits, set blocked and stop CDN distribution. Manual review queue for appeals. Keep this as one box on the whiteboard.
Analytics
Client sends play, pause, rebuffer, error. Queue → analytics store. Product uses this for QoE, not for the first frame.
9. Bottlenecks
- Transcode queue lag after a viral upload day
- CDN origin when a new episode is watched everywhere at once
- Hot metadata for a trending video id
- Analytics write storm if you log every segment request
- Storage cost if you never delete old renditions
Mitigations: autoscale workers, pre-warm CDN, cache playback API, sample analytics, lifecycle policies on raw files.
10. Tradeoffs
| Choice | Upside | Downside |
|---|---|---|
| Many bitrates | Smooth ABR | Storage and transcode cost |
| Pre-transcode all | Simple playback | Slow “video ready” |
| Just-in-time transcode | Fast publish | First viewers suffer |
| Push CDN | Great for Netflix drops | Waste if nobody watches |
| Pull CDN | Good for YouTube long-tail | First hit is slower |
| 301/immutable segments | Cache forever | Hard to fix a bad encode |
Interview pick: pre-transcode a standard ladder, pull CDN for UGC, pre-warm for known premieres.
11. Failure Modes
| Failure | User impact | Handling |
|---|---|---|
| Transcode worker crash | Video stays processing | Job is retryable; raw blob is source of truth |
| Bad encode | Playback errors | Keep raw, re-queue, mark failed if repeated |
| CDN outage in one region | Buffering | DNS/failover to another POP or origin (origin will hurt) |
| Video DB down | Cannot start new watches | Cached playback URLs for hot titles may still work |
| Event queue down | Counts freeze | Playback continues |
| Copyright false positive | Angry creator | Appeal queue, do not delete raw immediately |
Playback must not depend on analytics or recommendations.
12. Interview Answer in 10 Minutes
"I would design this as an upload pipeline plus a CDN watch path.
Assume 50 million daily viewers and 5 watches each. That is 250 million starts a day, about 2,900 QPS, around 15,000 QPS at 5× peak. Segment traffic is much higher but belongs on the CDN. Uploads might be 100,000 a day, only a few QPS, but tens of terabytes in. After transcode, storage can be ~150 TB/day. These are interview estimates.
Creator gets a signed upload URL. Raw file lands in object storage. Video service writes metadata as processing and puts a job on a queue. Workers transcode a bitrate ladder and write segments. Status becomes ready.
Viewer calls playback API, gets a manifest URL, and streams from the CDN with adaptive bitrate. Metadata lives in the video DB. Watch events go to a separate queue so they cannot stall the player.
Recommendations and comments are side systems. Copyright is an async scan that can flip status to blocked.
If transcode lags, we scale workers. If CDN misses, origin must be protected. The Java API tier only handles metadata and job control, not 4K bytes."
Practice this until it is under 10 minutes.
13. Interview Talking Points
- Read-heavy. Watch path ≠ upload path.
- Never transcode on the request thread.
- ABR + CDN are the streaming design, not one MP4 URL.
- Raw blob is sacred so you can re-encode.
- Analytics are async.
- Recommendations can fail open.
- Metrics: time-to-ready, playlist latency, rebuffer ratio, queue lag, origin bandwidth.
- Related reading: 20 system design concepts, distributed systems.
14. Follow-up Questions
YouTube vs Netflix?
YouTube: huge UGC, unpredictable hot videos, pull CDN, more moderation. Netflix: catalog ingest, premieres, pre-warm CDN, DRM more important.
How do you handle 4K?
Extra ladder rung, more storage, device checks, often restricted to capable clients.
Live streaming?
Different: ingest → packager → low-latency CDN. DVR is extra. Do not pretend VOD workers are live.
How are views counted?
Sampled events, not a SQL increment per segment. Dedup a viewer id for a time window.
Java role?
API + workers with a transcode fleet. See microservices and Kafka.
15. Internal Links
Related JavaThoughts reading:
- System Design Interview Preparation
- Design File Storage like Dropbox
- Design a Notification System
- The Netflix Tech Stack
- Java
- Spring Boot
- Kafka
- Microservices
- Distributed Systems
- API Gateway
- Message queues
- Event-driven architecture
- Scaling Java event-driven systems
- 20 system design concepts
Next in this series: Design Search Autocomplete.
