- Published on
- Views
Design a Web Crawler System Design Interview Guide
- Authors

- Name
- Javed Shaikh
← System Design Interview Preparation
This guide walks through Design a Web Crawler the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.
1. Problem
A crawler starts from a few URLs, downloads pages, extracts links, and repeats — politely.
Example:
- Seed
https://www.javathoughts.com - Fetch HTML, store it
- Parse
<a href> - Skip URLs you already saw
- Honor
robots.txtand per-domain rate limits - Hand useful pages to a search index later (indexing is a sibling system)
You are building a fetch pipeline, not Google.
2. Functional Requirements / FR
| Requirement | What it means |
|---|---|
| Seed URLs | Start set from config. |
| URL frontier | Queue of URLs to visit. |
| Fetcher workers | HTTP GET with timeouts. |
| robots.txt | Cache per host, skip disallowed paths. |
| Per-domain rate limit | Do not hammer one site. |
| Duplicate detection | Canonical URL + content hash. |
| Parse links | Resolve relative URLs. |
| Content storage | HTML or extracted text. |
| Search index handoff | Push documents to an indexer queue. |
| Retries | Transient 5xx/timeout with backoff. |
Out of scope: full ranking, JavaScript-heavy rendering farm (mention headless browsers as a later tier).
3. Non-Functional Requirements / NFR
| Requirement | Why it matters |
|---|---|
| Politeness | Legal and operational. |
| Scalable workers | Throughput via more fetchers, not one thread. |
| Freshness | Recrawl important hosts more often. |
| Robustness | Bad HTML, traps, infinite calendars. |
| Dedup efficiency | Do not store the web twice. |
Interview line: frontier + politeness + dedup. Throughput is easy. Not being a DDoS bot is the design.
4. Back-of-the-Envelope Calculation
Traffic assumptions
- Target 100 million pages/day
- Average page 100 KB downloaded
- Read/write: almost all writes to storage; reads are robots cache and Bloom filter
QPS
100,000,000 fetches/day / 86,400 seconds = around 1,160 QPS.
If peak is 5x, design for around 6,000 fetches/sec.
That 6,000 is across many domains. One domain might allow 1 req/sec.
Storage
100,000,000 × 100 KB ≈ 10 TB/day raw
You will not keep raw HTML forever. Interview: store compressed HTML for 7 days (~20 TB with compression ~3:1 → ~3 TB/week) plus extracted text longer.
URL metadata: 100M × 200 bytes ≈ 20 GB/day.
Cache / memory
Bloom filter for seen URLs: 10 billion URLs, 1% FPR ≈ a few GB (say 10 GB). Robots cache in Redis.
Servers
If one fetcher does 20 QPS (network bound), 6,000 QPS needs ~300 workers. Interview estimate, not exact production numbers. Frontier and storage scale separately.
5. APIs
Mostly internal. A small control plane:
POST /v1/seeds { "urls": [] }
GET /v1/jobs/{id}
POST /v1/hosts/{host}/pause
GET /v1/stats
Workers pull from the frontier; they do not expose public GET of the web.
6. Data Model
url_frontier
url, host, priority, next_fetch_at, attempts.
Partition by host so one crawler owns politeness for that host (or a host-level token bucket).
seen_urls
Canonical URL, first_seen, last_fetched. Plus Bloom in memory.
content
url_hash, storage_key, content_hash, fetched_at, status_code.
robots_cache
host, rules, fetched_at, TTL ~1 day.
7. High-Level Design
The frontier is the heartbeat.
Web Crawler architecture
Components:
- Seed URLs → URL Frontier Queue
- Crawler Workers check robots / rate limiter
- Fetched pages → Content Store
- Parsed links → Dedup → Frontier
- Important content → Search Index
Worker loop:
- Take URL (respect
next_fetch_at). - Load robots.
- Fetch with timeout, cap size.
- Store if new
content_hash. - Extract links, canonicalize, Bloom check, enqueue.
8. Deep Dives
Canonicalization
Lowercase host, strip tracking query params you choose, default port, trailing slash policy. Otherwise duplicates explode.
Duplicate detection
URL Bloom + exact DB. Content hash skips recrawl of identical pages (mirrors).
Politeness
Token bucket per host. robots.txt Crawl-delay. Shuffle hosts so one worker is not stuck on a slow site (host-per-queue).
Traps
Max path depth, max pages per host per day, no unbounded calendar query params. Deny-list.
Retries
Retry 3 times on timeout/5xx. 404: mark seen, do not retry. 429: honor Retry-After.
Index handoff
Emit {url, title, text} to Kafka for the indexer. Crawler should not run Elasticsearch queries.
9. Bottlenecks
- Frontier DB hot if one table
- DNS resolver
- Slow hosts blocking threads — use async HTTP
- Disk for 10 TB/day
- Seed bias (only crawl popular graph)
10. Tradeoffs
| Choice | Upside | Downside |
|---|---|---|
| BFS | Broad web | Slow to important pages |
| Priority (PageRank-like) | Better index | Extra system |
| Headless browser | JS sites | 10–100× cost |
| Store raw HTML | Reparse later | Disk |
Pick: per-host queues, Bloom dedup, robots+rate limit, async fetch, handoff to index.
11. Failure Modes
| Failure | Handling |
|---|---|
| Worker crash | URL not acked; frontier retry |
| Poison HTML | Size limit, parser timeout |
| robots fetch fail | Conservative: skip or retry host |
| Bloom false positive | Miss a URL (acceptable) or check DB |
| Storage full | Stop enqueue, alert |
12. Interview Answer in 10 Minutes
"A crawler is a polite BFS over URLs.
100 million pages a day is about 1,160 fetches per second, about 6,000 at 5× peak. That is 10 TB raw, so we compress and expire. I would keep a frontier partitioned by host, workers that check robots and a per-host rate limit, then store content and push new links through a Bloom filter.
Duplicates are canonical URLs plus content hashes. Important documents go to a search index queue. We cap depth and pages per host so we do not crawl a calendar forever.
If a worker dies, the URL is retried. If a host is slow, we do not stall other hosts."
13. Interview Talking Points
- Politeness first.
- Frontier by host.
- Bloom + canonical URLs.
- Async HTTP.
- Separate indexer.
- Metrics: fetch QPS, 4xx/5xx, robots blocks, frontier size, bytes stored.
- Rate limiter per domain.
14. Follow-up Questions
How do you recrawl?
Priority = change frequency. News sites hourly; static docs weekly.
Sitemap.xml?
Treat as extra seeds, still honor robots.
Distributed Bloom?
One Redis cluster or partitioned filters by URL hash.
Legal?
robots.txt, rate limits, ToS — say it in the interview.
15. Internal Links
Related JavaThoughts reading:
- System Design Interview Preparation
- Design Search Autocomplete
- Design a Rate Limiter
- Design a Distributed Queue
- Java
- Spring Boot
- Kafka
- Microservices
- Distributed Systems
- Message queues
- 20 system design concepts
Next in this series: Design a Distributed Queue.
