- Published on
- Views
Design a URL Shortener System Design Interview Guide
- Authors

- Name
- Javed Shaikh
← System Design Interview Preparation
This guide walks through Design a URL Shortener the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.
1. Problem
A URL shortener takes a long web address and returns a short one.
Example:
- Long URL:
https://www.javathoughts.com/system-design/url-shortener - Short URL:
https://short.ly/aB3x9K
When a user opens the short URL, the system should send them to the original long URL. Products like Bitly and TinyURL do this. Companies also use short links in SMS, emails, ads, and QR codes because a short link is easier to share.
In this design we are building a service that:
- Creates a short code for a long URL
- Redirects that short code to the original URL
- Stores a few extra features that interviewers often ask about: custom alias, expiry, and click count
We are not building a full marketing platform. We are designing the core create + redirect system.
2. Functional Requirements / FR
| Requirement | What it means |
|---|---|
| Create short URL | A user (or API client) sends a long URL and gets a short URL back. |
| Redirect | Opening GET /:shortCode sends the browser to the original long URL. |
| Custom alias | Optional. The user can ask for https://short.ly/javathoughts if that alias is free. |
| Expiry | Optional. A link can expire after a date or after N days. Expired links should not redirect. |
| Basic analytics | Store a click count. A later version can store referrer and country. |
Out of scope for a 45-minute interview:
- User dashboards and billing
- Full A/B testing
- Custom domains for every customer
- Preview pages before redirect
Confirm scope with the interviewer. Then lock it and move on.
3. Non-Functional Requirements / NFR
| Requirement | Why it matters |
|---|---|
| Low latency redirect | Users click short links and expect an instant jump. Redirects should usually complete in a few milliseconds on the server side. |
| High availability | If the redirect path is down, every shared link looks broken. |
| Read-heavy | People create a link once and click it many times. Redirects will dominate traffic. |
| Scalable storage | New links keep arriving every month. Storage must grow without a redesign. |
| Reliable redirect | The same short code should always map to the same long URL until it expires. |
| Basic abuse protection | Attackers will try to create millions of links, hide malware URLs, or flood redirects. |
A good interview sentence: optimize the redirect path first. Creation can be a bit slower. Redirects cannot.
4. Back-of-the-Envelope Calculation
Say these assumptions out loud. Interviewers care more about the method than the exact number.
Traffic assumptions
- 100 million new URLs per month
- 10 billion redirects per month
- Read/write ratio ≈ 10,000,000,000 / 100,000,000 = 100 : 1
This is clearly a read-heavy system.
Seconds in a month ≈ 30 × 24 × 3600 = 2,592,000
QPS
Write QPS (create short URL):
100,000,000 / 2,592,000 ≈ 39 writes/sec
Read QPS (redirect):
10,000,000,000 / 2,592,000 ≈ 3,860 reads/sec
Peak is often 3× to 5× average. Use 5× to be safe:
- Peak writes ≈ 200 QPS
- Peak redirects ≈ 20,000 QPS
Storage
Average long URL size ≈ 100 bytes
Short code ≈ 7 bytes
Metadata (id, timestamps, expiry, click count, user id) ≈ 400 bytes
Per URL ≈ 500 bytes
New storage per month:
100,000,000 × 500 bytes ≈ 50 GB/month
Five years, ignoring deletes:
50 GB × 12 × 5 ≈ 3 TB
That fits in a normal database cluster. Storage is not the scary part. Redirect QPS is.
Click events are heavier if you store every click as a row:
10 billion clicks/month × 50 bytes ≈ 500 GB/month of event data
Do not write every click into the same table used for redirects. Send events to a queue and store analytics separately.
Cache estimate
If 20% of links get 80% of clicks, cache the hot keys.
Assume we cache 100 million hot mappings. Each cache entry is short code + long URL + a little metadata ≈ 200 bytes.
100,000,000 × 200 bytes ≈ 20 GB
One Redis instance with 20–32 GB RAM can hold this. In production you would shard Redis, but in the interview one cache cluster is enough to mention.
Server estimate
A small Java/Spring Boot instance that mostly hits cache can handle a few thousand redirects per second. Use a conservative 2,000 QPS per server.
Peak 20,000 QPS / 2,000 ≈ 10 redirect servers
Add a few more for create APIs, deploys, and failure. About 12–15 app servers is a reasonable interview answer.
These are rough numbers. The point is: cache-first redirects, async analytics, and horizontal scaling.
5. APIs
Keep the API small.
Create a short URL
POST /api/urls
Request:
{
"longUrl": "https://www.javathoughts.com/system-design/url-shortener",
"customAlias": "jt-url",
"expiresAt": "2027-09-19T00:00:00Z"
}
customAlias and expiresAt are optional.
Success response (201 Created):
{
"shortCode": "jt-url",
"shortUrl": "https://short.ly/jt-url",
"longUrl": "https://www.javathoughts.com/system-design/url-shortener",
"expiresAt": "2027-09-19T00:00:00Z"
}
Error examples:
400invalid URL409custom alias already taken429too many create requests
Redirect
GET /:shortCode
Success: 302 Found with header:
Location: https://www.javathoughts.com/system-design/url-shortener
Use 302 if links can change or expire. Use 301 only if you promise the mapping will never change. In interviews, 302 is the safer default.
Not found or expired: 404.
Analytics
GET /api/urls/:shortCode/analytics
Response:
{
"shortCode": "jt-url",
"clickCount": 18420,
"createdAt": "2026-09-19T08:00:00Z",
"expiresAt": "2027-09-19T00:00:00Z"
}
Redirects should not wait for analytics. Count clicks after the user has been redirected.
6. Data Model
urls
| Field | Type | Notes |
|---|---|---|
id | bigint | Internal primary key |
short_code | varchar(16) | Unique. Used in the public URL |
long_url | text | Original destination |
user_id | bigint, nullable | Owner, if logged in |
expires_at | timestamp, nullable | Null means no expiry |
is_active | boolean | Soft delete / disable spam links |
click_count | bigint | Approximate is fine |
created_at | timestamp | |
updated_at | timestamp |
Indexes:
- Unique index on
short_code(the lookup key for redirects) - Index on
expires_at(cleanup job) - Index on
user_id(user's link list)
click_events or url_analytics
Do not put one row per click in urls. Use a separate store.
Option A — simple counter: keep click_count on urls and update it in the background.
Option B — event table (better for later reports):
| Field | Type | Notes |
|---|---|---|
id | bigint | |
short_code | varchar(16) | |
clicked_at | timestamp | |
referrer | varchar, nullable | |
country | char(2), nullable |
Index (short_code, clicked_at).
In a Java interview you can say: Postgres or MySQL for urls, Redis for hot redirects, and Kafka plus a column store for click events when volume grows.
7. High-Level Design
Redirect traffic should take the shortest path: cache first, database second.
URL Shortener architecture
Components:
- Client: browser, mobile app, or another backend.
- API Gateway / Load Balancer: TLS, routing, rate limits, auth for create APIs.
- URL Shortener Service: validates the long URL, generates or accepts a short code, writes to the database.
- Database: source of truth for mappings.
- Cache: Redis (or similar) stores
short_code -> long_urlfor hot links. - Redirect Service: can be the same app as create, but treat it as a fast read path.
- Analytics / Event Queue: redirect publishes a click event and returns immediately.
Create flow:
- Validate URL (https, not empty, not a known malware host).
- If custom alias is sent, check uniqueness.
- Else generate a short code.
- Insert into
urls. - Optionally write-through to cache.
- Return the short URL.
Redirect flow:
- Look up short code in cache.
- If miss, load from database and fill cache.
- If missing or expired, return 404.
- Return 302.
- Push a click event to the queue.
8. Deep Dives
How to generate short codes
We need a short, URL-safe string. A common alphabet is Base62:
0-9, a-z, A-Z → 62 characters
Length 7 gives:
62^7 ≈ 3.5 trillion codes
That is more than enough for 100 million new URLs per month.
Base62 encoding
If you have a numeric id 123456789, convert it to Base62 the same way you convert to hex, but with 62 symbols. Example idea in Java:
static final char[] BASE62 =
"0123456789abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ".toCharArray();
String toBase62(long n) {
StringBuilder sb = new StringBuilder();
while (n > 0) {
sb.append(BASE62[(int) (n % 62)]);
n /= 62;
}
return sb.reverse().toString();
}
Pad to 7 characters if you want a stable length.
Random code vs ID-based code
Random code: generate 7 random Base62 chars, then insert. Simple. Codes are not guessable in order. You must handle collisions.
ID-based code: database auto-increment (or a ticket server) gives a unique number. Encode that number in Base62. No collision if ids are unique. Codes can be guessed (n, then n+1). Attackers could scan sequential links.
Interview choice:
- If privacy of links matters, prefer random codes or hash ids.
- If simplicity matters, ID + Base62 is easy to explain.
A balanced answer: use a unique id, encode it in Base62, then shuffle with a secret mapping so codes are not obviously sequential.
Handling collisions
For random codes:
- Generate a code.
- Insert with unique constraint on
short_code. - If insert fails, generate again (retry 2–3 times).
- If custom alias collides, return 409. Do not overwrite.
For ID-based codes, collisions should not happen unless two services mint the same id. Use one id generator, or a Snowflake-style id.
Cache strategy
Redirect is the hot path.
- Cache key:
url:plus the short code - Cache value: long URL, expiry time, active flag
- TTL: start with 24 hours. Popular links stay hot because they get hits and you can refresh TTL.
- Read: cache-aside. Miss → DB → set cache.
- Write: when a new link is created, you can write-through. When a link is disabled or expired, delete the cache key.
Do not cache 404 forever. A short negative TTL (30–60 seconds) is enough to protect the database from bad bots, without blocking a code that is created a moment later.
Redirect flow
Keep this path tiny:
- Validate short code format.
- Redis GET.
- If miss, SELECT by
short_code. - Check
is_activeandexpires_at. - Emit analytics event.
No joins. No user lookup. No disk write on the request thread.
Analytics processing
Redirect service sends the short code and timestamp to Kafka or another queue. Workers batch-update click_count or insert into click_events.
If the queue is slow, the user still got redirected. That is the right tradeoff.
If you need a rough live counter, increment a Redis counter on redirect and flush to DB every minute.
Expiry cleanup
Options:
- Check
expires_aton every redirect (required). - A daily job marks expired rows
is_active = falseand deletes cache keys. - Optionally archive old rows to cheaper storage.
Do not rely only on the job. Always check expiry during redirect.
Abuse prevention
- Rate limit create APIs per IP and per API key. See rate limiting.
- Validate URL format and scheme (
httpspreferred). - Block known malware domains with a blocklist.
- Require login after a daily create quota for anonymous users.
- CAPTCHA or extra checks for bulk creation.
- Allow reporting a bad link and disable it quickly (
is_active = false+ cache delete).
9. Bottlenecks
| Bottleneck | What happens | What you do |
|---|---|---|
| Database hot reads | Popular links all hit the same rows | Cache first. Read replicas if needed. |
| Popular URLs | One celebrity link can be millions of QPS | That key must live in cache. Pin it. Add CDN/edge cache later. |
| Cache misses | Traffic falls through to the database | Watch miss rate. Warm cache after deploys. |
| Analytics write volume | 10 billion clicks/month will crush the urls table | Async queue. Separate analytics store. |
| Collision handling | Too many random retries slow creates | 7+ chars. Unique index. Retry with backoff. |
| Regional latency | Users far from your region wait on DNS + app + DB | Put redirect + cache closer to users. Later: multi-region reads. |
The first bottleneck to name in an interview is hot keys on the redirect path.
10. Tradeoffs
SQL vs NoSQL
- SQL (Postgres/MySQL): strong unique constraint on
short_code, simple transactions, easy expiry queries. Good default for this problem. - NoSQL (DynamoDB/Cassandra): easier to scale writes if you later create billions of links. Unique aliases need extra care.
Start with SQL. Mention NoSQL if create volume becomes huge.
Random code vs sequential ID
- Random: harder to guess, needs collision handling.
- Sequential ID: simple uniqueness, easier to enumerate.
Many teams use unique ids internally and random-looking public codes.
Synchronous vs async analytics
- Sync: accurate count, slower redirect, database under write pressure.
- Async: fast redirect, count can lag by seconds.
Choose async. Interviewers want the redirect to stay fast.
Strong consistency vs eventual consistency
The mapping short_code -> long_url should be strongly consistent. After create returns, the short link must work.
Click counts can be eventually consistent. A delay of a few seconds is fine.
Cache TTL choices
- Short TTL: safer if links get disabled, more DB load.
- Long TTL: faster, but a disabled link may redirect until TTL ends.
Compromise: TTL of hours, plus explicit cache delete when a link is updated, expired, or marked spam.
11. Failure Modes
Database down
Creates fail. Redirects can still work from cache for hot links. For a short outage that is acceptable. For a long outage, serve stale cache and show 503 for misses. Do not invent mappings.
Cache down
Redirects go to the database. Availability stays, latency and DB load get worse. Autoscale read replicas if this lasts. Keep cache as an optimization, not the only copy of data.
Queue down
Redirects still work. You lose some click events unless you have a local fallback (write to disk or increment Redis). Say you prefer losing some analytics over delaying users.
Duplicate short code
Unique index rejects the write. Retry with a new code, or return 409 for custom aliases. Never attach two long URLs to one code.
Expired URL
Return 404 or a small "this link expired" page. Remove cache entry.
Malicious/spam URL
Scan on create, allow reports after create, and disable fast. The redirect path should check is_active so a takedown is immediate after cache delete.
12. Interview Answer in 10 Minutes
Here is a version you can speak out loud.
"I would design a URL shortener as a read-heavy system. The two APIs are create short URL and redirect.
For scale, assume 100 million new links a month and 10 billion redirects. That is about 40 writes per second and about 4,000 reads per second, maybe 20,000 at peak. Storage for mappings is only tens of gigabytes per month, so the hard part is fast redirects, not disk.
I would keep a urls table with a unique short_code, the long URL, expiry, and an active flag. Short codes can be a unique id encoded in Base62, or a random 7-character Base62 string with a unique constraint.
The redirect path is cache first. Look up the short code in Redis. If it is missing, load it from the database and fill the cache. Then return HTTP 302. I would not write analytics on that request. I would push a click event to a queue and update counts in the background.
Creates go through an API gateway with rate limiting. Custom aliases check uniqueness and return 409 if taken. Expired or disabled links return 404.
If Redis is down, I fall back to the database. If the queue is down, redirects still work and analytics can lag. If one link becomes extremely popular, it must stay in cache so the database does not melt.
That gives us a simple system: SQL as source of truth, Redis for hot reads, and async analytics."
Practice this until it is under 10 minutes, then use the leftover time for deep dives the interviewer cares about.
13. Interview Talking Points
- This is a read-heavy design. Redirects matter more than creates.
- Cache-first redirects. Database is the source of truth, not the hot path.
- Unique short codes with a unique index. Collisions are handled, never overwritten.
- Horizontal scaling behind a load balancer. Stateless app servers.
- Async analytics so click counting cannot slow redirects.
- Rate limiting and URL checks to reduce spam.
- Monitoring: redirect latency, cache hit ratio, create error rate, queue lag.
- Clear split between strong consistency for mappings and eventual consistency for counts.
14. Follow-up Questions
Interviewers often push further. Short answers:
How would you handle 1 billion redirects per day?
That is about 12,000 QPS average, maybe 50,000+ at peak. Add more redirect servers, shard Redis, put a CDN or edge cache in front of the hottest codes, and keep analytics fully async.
How would you prevent spam links?
Rate limit creates, require auth after a quota, validate URLs, use a malware blocklist, and give ops a kill switch that sets is_active = false and deletes the cache key.
How would you support custom aliases?
Accept an optional alias, check format and uniqueness, and store it as short_code. If it is taken, return 409. Do not auto-rename the user's alias.
How would you make redirects faster globally?
Serve redirects from the closest region. Replicate cache and read-only mapping data to other regions. Keep creates in the primary region, or use a global unique id generator.
How would analytics work without slowing redirects?
Return 302 first. Publish an event to Kafka. Workers aggregate counts. The user should never wait on analytics storage.
Other useful probes: 301 vs 302, hash collisions, GDPR deletion, and how you would shard the urls table by short_code.
15. Internal Links
Related JavaThoughts reading:
- System Design Interview Preparation — top 20 questions
- Design a Rate Limiter
- Design a Distributed Cache
- Java
- Spring Boot
- Kafka
- Microservices
- Distributed Systems
- API Gateway
- Caching
- Event-driven architecture
- E-commerce platform system design
- 20 system design concepts
Next in this series: Design a Rate Limiter.
