Javathoughts Logo
Javathoughts
Published on
Views

Design a URL Shortener System Design Interview Guide

Authors
  • avatar
    Name
    Javed Shaikh
    Twitter

← System Design Interview Preparation

This guide walks through Design a URL Shortener the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.


1. Problem

A URL shortener takes a long web address and returns a short one.

Example:

  • Long URL: https://www.javathoughts.com/system-design/url-shortener
  • Short URL: https://short.ly/aB3x9K

When a user opens the short URL, the system should send them to the original long URL. Products like Bitly and TinyURL do this. Companies also use short links in SMS, emails, ads, and QR codes because a short link is easier to share.

In this design we are building a service that:

  1. Creates a short code for a long URL
  2. Redirects that short code to the original URL
  3. Stores a few extra features that interviewers often ask about: custom alias, expiry, and click count

We are not building a full marketing platform. We are designing the core create + redirect system.


2. Functional Requirements / FR

RequirementWhat it means
Create short URLA user (or API client) sends a long URL and gets a short URL back.
RedirectOpening GET /:shortCode sends the browser to the original long URL.
Custom aliasOptional. The user can ask for https://short.ly/javathoughts if that alias is free.
ExpiryOptional. A link can expire after a date or after N days. Expired links should not redirect.
Basic analyticsStore a click count. A later version can store referrer and country.

Out of scope for a 45-minute interview:

  • User dashboards and billing
  • Full A/B testing
  • Custom domains for every customer
  • Preview pages before redirect

Confirm scope with the interviewer. Then lock it and move on.


3. Non-Functional Requirements / NFR

RequirementWhy it matters
Low latency redirectUsers click short links and expect an instant jump. Redirects should usually complete in a few milliseconds on the server side.
High availabilityIf the redirect path is down, every shared link looks broken.
Read-heavyPeople create a link once and click it many times. Redirects will dominate traffic.
Scalable storageNew links keep arriving every month. Storage must grow without a redesign.
Reliable redirectThe same short code should always map to the same long URL until it expires.
Basic abuse protectionAttackers will try to create millions of links, hide malware URLs, or flood redirects.

A good interview sentence: optimize the redirect path first. Creation can be a bit slower. Redirects cannot.


4. Back-of-the-Envelope Calculation

Say these assumptions out loud. Interviewers care more about the method than the exact number.

Traffic assumptions

  • 100 million new URLs per month
  • 10 billion redirects per month
  • Read/write ratio ≈ 10,000,000,000 / 100,000,000 = 100 : 1

This is clearly a read-heavy system.

Seconds in a month ≈ 30 × 24 × 3600 = 2,592,000

QPS

Write QPS (create short URL):

100,000,000 / 2,592,000 ≈ 39 writes/sec

Read QPS (redirect):

10,000,000,000 / 2,592,000 ≈ 3,860 reads/sec

Peak is often 3× to 5× average. Use 5× to be safe:

  • Peak writes ≈ 200 QPS
  • Peak redirects ≈ 20,000 QPS

Storage

Average long URL size ≈ 100 bytes
Short code ≈ 7 bytes
Metadata (id, timestamps, expiry, click count, user id) ≈ 400 bytes

Per URL ≈ 500 bytes

New storage per month:

100,000,000 × 500 bytes ≈ 50 GB/month

Five years, ignoring deletes:

50 GB × 12 × 5 ≈ 3 TB

That fits in a normal database cluster. Storage is not the scary part. Redirect QPS is.

Click events are heavier if you store every click as a row:

10 billion clicks/month × 50 bytes ≈ 500 GB/month of event data

Do not write every click into the same table used for redirects. Send events to a queue and store analytics separately.

Cache estimate

If 20% of links get 80% of clicks, cache the hot keys.

Assume we cache 100 million hot mappings. Each cache entry is short code + long URL + a little metadata ≈ 200 bytes.

100,000,000 × 200 bytes ≈ 20 GB

One Redis instance with 20–32 GB RAM can hold this. In production you would shard Redis, but in the interview one cache cluster is enough to mention.

Server estimate

A small Java/Spring Boot instance that mostly hits cache can handle a few thousand redirects per second. Use a conservative 2,000 QPS per server.

Peak 20,000 QPS / 2,000 ≈ 10 redirect servers

Add a few more for create APIs, deploys, and failure. About 12–15 app servers is a reasonable interview answer.

These are rough numbers. The point is: cache-first redirects, async analytics, and horizontal scaling.


5. APIs

Keep the API small.

Create a short URL

POST /api/urls

Request:

{
  "longUrl": "https://www.javathoughts.com/system-design/url-shortener",
  "customAlias": "jt-url",
  "expiresAt": "2027-09-19T00:00:00Z"
}

customAlias and expiresAt are optional.

Success response (201 Created):

{
  "shortCode": "jt-url",
  "shortUrl": "https://short.ly/jt-url",
  "longUrl": "https://www.javathoughts.com/system-design/url-shortener",
  "expiresAt": "2027-09-19T00:00:00Z"
}

Error examples:

  • 400 invalid URL
  • 409 custom alias already taken
  • 429 too many create requests

Redirect

GET /:shortCode

Success: 302 Found with header:

Location: https://www.javathoughts.com/system-design/url-shortener

Use 302 if links can change or expire. Use 301 only if you promise the mapping will never change. In interviews, 302 is the safer default.

Not found or expired: 404.

Analytics

GET /api/urls/:shortCode/analytics

Response:

{
  "shortCode": "jt-url",
  "clickCount": 18420,
  "createdAt": "2026-09-19T08:00:00Z",
  "expiresAt": "2027-09-19T00:00:00Z"
}

Redirects should not wait for analytics. Count clicks after the user has been redirected.


6. Data Model

urls

FieldTypeNotes
idbigintInternal primary key
short_codevarchar(16)Unique. Used in the public URL
long_urltextOriginal destination
user_idbigint, nullableOwner, if logged in
expires_attimestamp, nullableNull means no expiry
is_activebooleanSoft delete / disable spam links
click_countbigintApproximate is fine
created_attimestamp
updated_attimestamp

Indexes:

  • Unique index on short_code (the lookup key for redirects)
  • Index on expires_at (cleanup job)
  • Index on user_id (user's link list)

click_events or url_analytics

Do not put one row per click in urls. Use a separate store.

Option A — simple counter: keep click_count on urls and update it in the background.

Option B — event table (better for later reports):

FieldTypeNotes
idbigint
short_codevarchar(16)
clicked_attimestamp
referrervarchar, nullable
countrychar(2), nullable

Index (short_code, clicked_at).

In a Java interview you can say: Postgres or MySQL for urls, Redis for hot redirects, and Kafka plus a column store for click events when volume grows.


7. High-Level Design

Redirect traffic should take the shortest path: cache first, database second.

URL Shortener architecture

URL Shortener architectureCreates go to the database. Redirects check cache first, then fall back to the database. Click analytics travel on a queue so the user is not delayed.createredirectmissasync click👤Client🌐API Gateway⚙️URL Shortener Service🧠Cache🗄️Database📩Queue👷Analytics Worker📊Analytics Store
Creates go to the database. Redirects check cache first, then fall back to the database. Click analytics travel on a queue so the user is not delayed.

Components:

  • Client: browser, mobile app, or another backend.
  • API Gateway / Load Balancer: TLS, routing, rate limits, auth for create APIs.
  • URL Shortener Service: validates the long URL, generates or accepts a short code, writes to the database.
  • Database: source of truth for mappings.
  • Cache: Redis (or similar) stores short_code -> long_url for hot links.
  • Redirect Service: can be the same app as create, but treat it as a fast read path.
  • Analytics / Event Queue: redirect publishes a click event and returns immediately.

Create flow:

  1. Validate URL (https, not empty, not a known malware host).
  2. If custom alias is sent, check uniqueness.
  3. Else generate a short code.
  4. Insert into urls.
  5. Optionally write-through to cache.
  6. Return the short URL.

Redirect flow:

  1. Look up short code in cache.
  2. If miss, load from database and fill cache.
  3. If missing or expired, return 404.
  4. Return 302.
  5. Push a click event to the queue.

8. Deep Dives

How to generate short codes

We need a short, URL-safe string. A common alphabet is Base62:

0-9, a-z, A-Z  →  62 characters

Length 7 gives:

62^7 ≈ 3.5 trillion codes

That is more than enough for 100 million new URLs per month.

Base62 encoding

If you have a numeric id 123456789, convert it to Base62 the same way you convert to hex, but with 62 symbols. Example idea in Java:

static final char[] BASE62 =
    "0123456789abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ".toCharArray();

String toBase62(long n) {
    StringBuilder sb = new StringBuilder();
    while (n > 0) {
        sb.append(BASE62[(int) (n % 62)]);
        n /= 62;
    }
    return sb.reverse().toString();
}

Pad to 7 characters if you want a stable length.

Random code vs ID-based code

Random code: generate 7 random Base62 chars, then insert. Simple. Codes are not guessable in order. You must handle collisions.

ID-based code: database auto-increment (or a ticket server) gives a unique number. Encode that number in Base62. No collision if ids are unique. Codes can be guessed (n, then n+1). Attackers could scan sequential links.

Interview choice:

  • If privacy of links matters, prefer random codes or hash ids.
  • If simplicity matters, ID + Base62 is easy to explain.

A balanced answer: use a unique id, encode it in Base62, then shuffle with a secret mapping so codes are not obviously sequential.

Handling collisions

For random codes:

  1. Generate a code.
  2. Insert with unique constraint on short_code.
  3. If insert fails, generate again (retry 2–3 times).
  4. If custom alias collides, return 409. Do not overwrite.

For ID-based codes, collisions should not happen unless two services mint the same id. Use one id generator, or a Snowflake-style id.

Cache strategy

Redirect is the hot path.

  • Cache key: url: plus the short code
  • Cache value: long URL, expiry time, active flag
  • TTL: start with 24 hours. Popular links stay hot because they get hits and you can refresh TTL.
  • Read: cache-aside. Miss → DB → set cache.
  • Write: when a new link is created, you can write-through. When a link is disabled or expired, delete the cache key.

Do not cache 404 forever. A short negative TTL (30–60 seconds) is enough to protect the database from bad bots, without blocking a code that is created a moment later.

Redirect flow

Keep this path tiny:

  1. Validate short code format.
  2. Redis GET.
  3. If miss, SELECT by short_code.
  4. Check is_active and expires_at.
  5. Emit analytics event.

No joins. No user lookup. No disk write on the request thread.

Analytics processing

Redirect service sends the short code and timestamp to Kafka or another queue. Workers batch-update click_count or insert into click_events.

If the queue is slow, the user still got redirected. That is the right tradeoff.

If you need a rough live counter, increment a Redis counter on redirect and flush to DB every minute.

Expiry cleanup

Options:

  • Check expires_at on every redirect (required).
  • A daily job marks expired rows is_active = false and deletes cache keys.
  • Optionally archive old rows to cheaper storage.

Do not rely only on the job. Always check expiry during redirect.

Abuse prevention

  • Rate limit create APIs per IP and per API key. See rate limiting.
  • Validate URL format and scheme (https preferred).
  • Block known malware domains with a blocklist.
  • Require login after a daily create quota for anonymous users.
  • CAPTCHA or extra checks for bulk creation.
  • Allow reporting a bad link and disable it quickly (is_active = false + cache delete).

9. Bottlenecks

BottleneckWhat happensWhat you do
Database hot readsPopular links all hit the same rowsCache first. Read replicas if needed.
Popular URLsOne celebrity link can be millions of QPSThat key must live in cache. Pin it. Add CDN/edge cache later.
Cache missesTraffic falls through to the databaseWatch miss rate. Warm cache after deploys.
Analytics write volume10 billion clicks/month will crush the urls tableAsync queue. Separate analytics store.
Collision handlingToo many random retries slow creates7+ chars. Unique index. Retry with backoff.
Regional latencyUsers far from your region wait on DNS + app + DBPut redirect + cache closer to users. Later: multi-region reads.

The first bottleneck to name in an interview is hot keys on the redirect path.


10. Tradeoffs

SQL vs NoSQL

  • SQL (Postgres/MySQL): strong unique constraint on short_code, simple transactions, easy expiry queries. Good default for this problem.
  • NoSQL (DynamoDB/Cassandra): easier to scale writes if you later create billions of links. Unique aliases need extra care.

Start with SQL. Mention NoSQL if create volume becomes huge.

Random code vs sequential ID

  • Random: harder to guess, needs collision handling.
  • Sequential ID: simple uniqueness, easier to enumerate.

Many teams use unique ids internally and random-looking public codes.

Synchronous vs async analytics

  • Sync: accurate count, slower redirect, database under write pressure.
  • Async: fast redirect, count can lag by seconds.

Choose async. Interviewers want the redirect to stay fast.

Strong consistency vs eventual consistency

The mapping short_code -> long_url should be strongly consistent. After create returns, the short link must work.

Click counts can be eventually consistent. A delay of a few seconds is fine.

Cache TTL choices

  • Short TTL: safer if links get disabled, more DB load.
  • Long TTL: faster, but a disabled link may redirect until TTL ends.

Compromise: TTL of hours, plus explicit cache delete when a link is updated, expired, or marked spam.


11. Failure Modes

Database down
Creates fail. Redirects can still work from cache for hot links. For a short outage that is acceptable. For a long outage, serve stale cache and show 503 for misses. Do not invent mappings.

Cache down
Redirects go to the database. Availability stays, latency and DB load get worse. Autoscale read replicas if this lasts. Keep cache as an optimization, not the only copy of data.

Queue down
Redirects still work. You lose some click events unless you have a local fallback (write to disk or increment Redis). Say you prefer losing some analytics over delaying users.

Duplicate short code
Unique index rejects the write. Retry with a new code, or return 409 for custom aliases. Never attach two long URLs to one code.

Expired URL
Return 404 or a small "this link expired" page. Remove cache entry.

Malicious/spam URL
Scan on create, allow reports after create, and disable fast. The redirect path should check is_active so a takedown is immediate after cache delete.


12. Interview Answer in 10 Minutes

Here is a version you can speak out loud.

"I would design a URL shortener as a read-heavy system. The two APIs are create short URL and redirect.

For scale, assume 100 million new links a month and 10 billion redirects. That is about 40 writes per second and about 4,000 reads per second, maybe 20,000 at peak. Storage for mappings is only tens of gigabytes per month, so the hard part is fast redirects, not disk.

I would keep a urls table with a unique short_code, the long URL, expiry, and an active flag. Short codes can be a unique id encoded in Base62, or a random 7-character Base62 string with a unique constraint.

The redirect path is cache first. Look up the short code in Redis. If it is missing, load it from the database and fill the cache. Then return HTTP 302. I would not write analytics on that request. I would push a click event to a queue and update counts in the background.

Creates go through an API gateway with rate limiting. Custom aliases check uniqueness and return 409 if taken. Expired or disabled links return 404.

If Redis is down, I fall back to the database. If the queue is down, redirects still work and analytics can lag. If one link becomes extremely popular, it must stay in cache so the database does not melt.

That gives us a simple system: SQL as source of truth, Redis for hot reads, and async analytics."

Practice this until it is under 10 minutes, then use the leftover time for deep dives the interviewer cares about.


13. Interview Talking Points

  • This is a read-heavy design. Redirects matter more than creates.
  • Cache-first redirects. Database is the source of truth, not the hot path.
  • Unique short codes with a unique index. Collisions are handled, never overwritten.
  • Horizontal scaling behind a load balancer. Stateless app servers.
  • Async analytics so click counting cannot slow redirects.
  • Rate limiting and URL checks to reduce spam.
  • Monitoring: redirect latency, cache hit ratio, create error rate, queue lag.
  • Clear split between strong consistency for mappings and eventual consistency for counts.

14. Follow-up Questions

Interviewers often push further. Short answers:

How would you handle 1 billion redirects per day?
That is about 12,000 QPS average, maybe 50,000+ at peak. Add more redirect servers, shard Redis, put a CDN or edge cache in front of the hottest codes, and keep analytics fully async.

How would you prevent spam links?
Rate limit creates, require auth after a quota, validate URLs, use a malware blocklist, and give ops a kill switch that sets is_active = false and deletes the cache key.

How would you support custom aliases?
Accept an optional alias, check format and uniqueness, and store it as short_code. If it is taken, return 409. Do not auto-rename the user's alias.

How would you make redirects faster globally?
Serve redirects from the closest region. Replicate cache and read-only mapping data to other regions. Keep creates in the primary region, or use a global unique id generator.

How would analytics work without slowing redirects?
Return 302 first. Publish an event to Kafka. Workers aggregate counts. The user should never wait on analytics storage.

Other useful probes: 301 vs 302, hash collisions, GDPR deletion, and how you would shard the urls table by short_code.


Related JavaThoughts reading:


Next in this series: Design a Rate Limiter.