Javathoughts Logo
Javathoughts
Published on
Views

Design a Web Crawler System Design Interview Guide

Authors
  • avatar
    Name
    Javed Shaikh
    Twitter

← System Design Interview Preparation

This guide walks through Design a Web Crawler the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.


1. Problem

A crawler starts from a few URLs, downloads pages, extracts links, and repeats — politely.

Example:

  • Seed https://www.javathoughts.com
  • Fetch HTML, store it
  • Parse <a href>
  • Skip URLs you already saw
  • Honor robots.txt and per-domain rate limits
  • Hand useful pages to a search index later (indexing is a sibling system)

You are building a fetch pipeline, not Google.


2. Functional Requirements / FR

RequirementWhat it means
Seed URLsStart set from config.
URL frontierQueue of URLs to visit.
Fetcher workersHTTP GET with timeouts.
robots.txtCache per host, skip disallowed paths.
Per-domain rate limitDo not hammer one site.
Duplicate detectionCanonical URL + content hash.
Parse linksResolve relative URLs.
Content storageHTML or extracted text.
Search index handoffPush documents to an indexer queue.
RetriesTransient 5xx/timeout with backoff.

Out of scope: full ranking, JavaScript-heavy rendering farm (mention headless browsers as a later tier).


3. Non-Functional Requirements / NFR

RequirementWhy it matters
PolitenessLegal and operational.
Scalable workersThroughput via more fetchers, not one thread.
FreshnessRecrawl important hosts more often.
RobustnessBad HTML, traps, infinite calendars.
Dedup efficiencyDo not store the web twice.

Interview line: frontier + politeness + dedup. Throughput is easy. Not being a DDoS bot is the design.


4. Back-of-the-Envelope Calculation

Traffic assumptions

  • Target 100 million pages/day
  • Average page 100 KB downloaded
  • Read/write: almost all writes to storage; reads are robots cache and Bloom filter

QPS

100,000,000 fetches/day / 86,400 seconds = around 1,160 QPS.
If peak is 5x, design for around 6,000 fetches/sec.

That 6,000 is across many domains. One domain might allow 1 req/sec.

Storage

100,000,000 × 100 KB ≈ 10 TB/day raw

You will not keep raw HTML forever. Interview: store compressed HTML for 7 days (~20 TB with compression ~3:1 → ~3 TB/week) plus extracted text longer.

URL metadata: 100M × 200 bytes ≈ 20 GB/day.

Cache / memory

Bloom filter for seen URLs: 10 billion URLs, 1% FPR ≈ a few GB (say 10 GB). Robots cache in Redis.

Servers

If one fetcher does 20 QPS (network bound), 6,000 QPS needs ~300 workers. Interview estimate, not exact production numbers. Frontier and storage scale separately.


5. APIs

Mostly internal. A small control plane:

POST /v1/seeds            { "urls": [] }
GET  /v1/jobs/{id}
POST /v1/hosts/{host}/pause
GET  /v1/stats

Workers pull from the frontier; they do not expose public GET of the web.


6. Data Model

url_frontier

url, host, priority, next_fetch_at, attempts.

Partition by host so one crawler owns politeness for that host (or a host-level token bucket).

seen_urls

Canonical URL, first_seen, last_fetched. Plus Bloom in memory.

content

url_hash, storage_key, content_hash, fetched_at, status_code.

robots_cache

host, rules, fetched_at, TTL ~1 day.


7. High-Level Design

The frontier is the heartbeat.

Web Crawler architecture

Web Crawler architectureSeed URLs enter the frontier. Workers fetch politely, store pages, and push new links after duplicate checks. Useful pages go to a search index.checkextractnew URLshandoff🌱Seed URLs📩URL Frontier Queue👷Crawler Workers🛡️Robots / Rate📦Content Store🧠Dedup Service🔍Search Index🕷️Parsed Links
Seed URLs enter the frontier. Workers fetch politely, store pages, and push new links after duplicate checks. Useful pages go to a search index.

Components:

  • Seed URLs → URL Frontier Queue
  • Crawler Workers check robots / rate limiter
  • Fetched pages → Content Store
  • Parsed links → Dedup → Frontier
  • Important content → Search Index

Worker loop:

  1. Take URL (respect next_fetch_at).
  2. Load robots.
  3. Fetch with timeout, cap size.
  4. Store if new content_hash.
  5. Extract links, canonicalize, Bloom check, enqueue.

8. Deep Dives

Canonicalization

Lowercase host, strip tracking query params you choose, default port, trailing slash policy. Otherwise duplicates explode.

Duplicate detection

URL Bloom + exact DB. Content hash skips recrawl of identical pages (mirrors).

Politeness

Token bucket per host. robots.txt Crawl-delay. Shuffle hosts so one worker is not stuck on a slow site (host-per-queue).

Traps

Max path depth, max pages per host per day, no unbounded calendar query params. Deny-list.

Retries

Retry 3 times on timeout/5xx. 404: mark seen, do not retry. 429: honor Retry-After.

Index handoff

Emit {url, title, text} to Kafka for the indexer. Crawler should not run Elasticsearch queries.


9. Bottlenecks

  • Frontier DB hot if one table
  • DNS resolver
  • Slow hosts blocking threads — use async HTTP
  • Disk for 10 TB/day
  • Seed bias (only crawl popular graph)

10. Tradeoffs

ChoiceUpsideDownside
BFSBroad webSlow to important pages
Priority (PageRank-like)Better indexExtra system
Headless browserJS sites10–100× cost
Store raw HTMLReparse laterDisk

Pick: per-host queues, Bloom dedup, robots+rate limit, async fetch, handoff to index.


11. Failure Modes

FailureHandling
Worker crashURL not acked; frontier retry
Poison HTMLSize limit, parser timeout
robots fetch failConservative: skip or retry host
Bloom false positiveMiss a URL (acceptable) or check DB
Storage fullStop enqueue, alert

12. Interview Answer in 10 Minutes

"A crawler is a polite BFS over URLs.

100 million pages a day is about 1,160 fetches per second, about 6,000 at 5× peak. That is 10 TB raw, so we compress and expire. I would keep a frontier partitioned by host, workers that check robots and a per-host rate limit, then store content and push new links through a Bloom filter.

Duplicates are canonical URLs plus content hashes. Important documents go to a search index queue. We cap depth and pages per host so we do not crawl a calendar forever.

If a worker dies, the URL is retried. If a host is slow, we do not stall other hosts."


13. Interview Talking Points

  • Politeness first.
  • Frontier by host.
  • Bloom + canonical URLs.
  • Async HTTP.
  • Separate indexer.
  • Metrics: fetch QPS, 4xx/5xx, robots blocks, frontier size, bytes stored.
  • Rate limiter per domain.

14. Follow-up Questions

How do you recrawl?
Priority = change frequency. News sites hourly; static docs weekly.

Sitemap.xml?
Treat as extra seeds, still honor robots.

Distributed Bloom?
One Redis cluster or partitioned filters by URL hash.

Legal?
robots.txt, rate limits, ToS — say it in the interview.


Related JavaThoughts reading:


Next in this series: Design a Distributed Queue.