Javathoughts Logo
Javathoughts
Published on
Views

Design an API Gateway System Design Interview Guide

Authors
  • avatar
    Name
    Javed Shaikh
    Twitter

← System Design Interview Preparation

This guide walks through Design an API Gateway the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.


1. Problem

Clients should not know 40 microservice URLs. They call one host. The gateway routes, authenticates, and protects.

Example:

  • Mobile app calls https://api.company.com/v1/orders
  • Gateway terminates TLS, checks JWT, rate limits, finds Order Service, proxies the HTTP call
  • Logs status and latency
  • If Order Service is dead, circuit breaker fails fast

See also API Gateway concept. This article is the interview design.


2. Functional Requirements / FR

RequirementWhat it means
Request routingPath/host → service.
AuthenticationValidate token/session; reject 401.
Rate limitingPer key/IP.
Load balancingAcross service instances.
Request/response transformStrip headers, add correlation id, maybe protocol map.
Service discoveryWhere is Order Service now?
TLS terminationCerts at the edge.
Logging / metricsStatus, latency, route.
Circuit breakingStop calling a sick dependency.
Versioning/v1 vs /v2 routes.

Out of scope: full service mesh sidecar story (mention as east-west; gateway is north-south).


3. Non-Functional Requirements / NFR

RequirementWhy it matters
Low added latencyA few milliseconds, not 100ms.
High availabilityGateway down = company down.
Horizontal scaleStateless (or sticky only if needed).
Safe config reloadNew route without dropping all connections.

Interview line: stateless reverse proxy plus policy plugins. Business logic stays in services.


4. Back-of-the-Envelope Calculation

Traffic assumptions

  • 100 million API requests/day at the edge
  • Mix of reads and writes; gateway treats them the same
  • Payload average 2 KB in + 4 KB out

QPS

100,000,000 / 86,400 ≈ 1,160 QPS average
If peak is 5x, design for around 6,000 QPS.

Large apps: 10× that. Same design, more boxes.

Storage

Gateway is not a system of record. Access logs: 6,000 QPS × 200 byte log line ≈ 1.2 MB/s ≈ 100 GB/day. Ship to observability, do not keep on the box.

Cache / memory

Hot JWKS keys, rate-limit counters in Redis, discovery cache. A few GB per node + Redis cluster.

Servers

If one gateway process does 2,000 QPS comfortably, peak 6,000 needs ~6–8 instances across AZs plus extras. Interview estimates, not exact production numbers.


5. APIs

The gateway is the API. Control plane:

PUT  /admin/routes
{ "pathPrefix": "/v1/orders", "service": "order", "auth": "required", "rateLimit": "100/min" }

GET  /admin/health
GET  /metrics

Data plane: any https://api.example.com/... matching routes.


6. Data Model

Route table

id, match (path, method), upstream, timeout_ms, auth_mode, rate_limit_policy, version.

Rate limit policy

key (ip / user / api_key), limit, window.

Discovery

service_name → list of {host, port, healthy}.

Config in etcd/Consul/DB; nodes watch and hot reload.


7. High-Level Design

Keep the data path short.

API Gateway architecture

API Gateway architectureClients hit one gateway. Auth, rate limits, and routing happen before traffic reaches services. Logs and discovery stay off the happy path where possible.logs / metricsdiscover👤Client🌐API Gateway🔐Auth Check🛡️Rate Limit⚙️User Service⚙️Order Service⚙️Payment Service📊Observability🗺️Service Registry
Clients hit one gateway. Auth, rate limits, and routing happen before traffic reaches services. Logs and discovery stay off the happy path where possible.

Components:

  • Client → API Gateway
  • Checks Auth, Rate Limit, routing rules
  • Routes to User / Order / Payment services
  • Observability for logs/metrics
  • Service registry for instance lists

Request path:

  1. TLS
  2. Match route
  3. Rate limit (rate limiter)
  4. Auth (auth/session JWT/session)
  5. Discover + pick instance (load balancing)
  6. Proxy with timeout
  7. Circuit stats
  8. Log asynchronously

8. Deep Dives

Routing

Longest prefix match. /v2/orders vs /v1/orders. Header X-Canary optional.

Auth

Validate JWT locally with cached JWKS. Do not call Auth Service on every request if the JWT is self-contained. Session cookies: Redis lookup.

Rate limiting

Token bucket in Redis. Fail open or closed: for login, fail closed; for public read, maybe fail open.

Transform

Add X-Request-Id. Strip Cookie to internal services if they should only see a user id header you set after auth.

Circuit breaking

Per-instance error rate. Open → 503 fast. Half-open probe. See circuit breaker.

Versioning

URL version is simplest. Header version is messier for caches.

Service discovery

Watch registry. Cache 5–10s. Do not DNS every request without caching.


9. Bottlenecks

  • Redis for limits
  • Huge request bodies (cap size)
  • Logging sync on disk
  • One giant Lua/plugin chain
  • Hot single instance if LB is broken

10. Tradeoffs

ChoiceUpsideDownside
Gateway does authOne placeGateway becomes critical + complex
Each service does authSimple gatewayInconsistent
Mesh insteadEast-westDoes not replace public edge
Shared vs dedicated gatewaysIsolationCost

Pick: edge gateway for north-south, local JWT, Redis limits, timeouts everywhere, async logs.


11. Failure Modes

FailureHandling
Auth JWKS fetch failUse cached keys; expire → 503
Redis downFail policy (declare it)
Upstream timeout504, circuit
Config bad deployCanary routes, instant rollback
Traffic spikeAutoscale + rate limit

12. Interview Answer in 10 Minutes

"An API gateway is a stateless reverse proxy with policies.

100 million requests a day is about 1,160 QPS, about 6,000 at 5× peak. I would run several gateway instances, terminate TLS, match a route table, rate-limit with Redis, validate JWTs with cached keys, then load-balance to instances from a registry.

I would add a request id, enforce timeouts, and trip a circuit breaker on a sick instance. Logs go async to the observability pipeline. Business rules stay in Order and Payment services.

If Redis is down I would say fail closed on login routes and fail open on public GETs, or I would keep a local token bucket as fallback."


13. Interview Talking Points

  • North-south vs mesh.
  • Stateless data plane.
  • Timeouts + circuit breaker.
  • JWT local validation.
  • Do not put domain logic here.
  • Metrics: gateway p99, 4xx/5xx, upstream p99, limit rejections.
  • Microservices + Spring Boot (Spring Cloud Gateway is a valid Java mention).

14. Follow-up Questions

GraphQL gateway?
BFF pattern: still a gateway, plus query cost limits.

WebSockets?
Separate upgrade path; sticky or pub/sub. See chat design.

mTLS to services?
Yes for internal; gateway presents client identity.

Why not Nginx only?
Fine for routing. Interview still wants auth, limits, discovery.


Related JavaThoughts reading:


Next in this series: Design Authentication and Session System.