- Published on
- Views
Design Logging, Metrics, and Observability Platform System Design Interview Guide
- Authors

- Name
- Javed Shaikh
← System Design Interview Preparation
This guide walks through Design a Logging, Metrics, and Observability Platform the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.
1. Problem
When production breaks, engineers need logs (what happened), metrics (how much), and a little tracing (which hop).
Example:
- Payment service emits
payment.failedlogs and apayments_failed_totalcounter - An agent on the box or sidecar ships data
- A pipeline indexes logs and stores time series
- A dashboard shows error rate
- An alert pages Slack at 2am
This is the platform behind monitoring & observability, not an APM vendor clone of every feature.
2. Functional Requirements / FR
| Requirement | What it means |
|---|---|
| App logs | Structured JSON with level, service, trace id. |
| Metrics | Counters, gauges, histograms. |
| Traces | Optional: span per RPC, sampled. |
| Agents / SDKs | Push from JVM, sidecar, or both. |
| Ingestion pipeline | Buffer so apps do not block on storage. |
| Stream processing | Parse, drop PII, sample, enrich. |
| Storage / indexing | Search logs; query metrics by time. |
| Dashboards | Graphs and log search UI. |
| Alerts | Thresholds and notifications. |
| Retention / sampling | Hot 7–14 days, cheap archive later. |
Out of scope: building a full BI warehouse, log-based billing product.
3. Non-Functional Requirements / NFR
| Requirement | Why it matters |
|---|---|
| Ingestion durability | Do not lose the error logs that explain an outage. |
| Query latency | Dashboards in seconds, not minutes. |
| App isolation | Telemetry must not take down checkout. |
| Multi-tenant isolation | One team cannot fill the cluster. |
| Cost control | Logs grow faster than traffic. |
Interview line: async ingest, two stores (logs vs metrics), sample traces, alerts off metrics.
4. Back-of-the-Envelope Calculation
Traffic assumptions
- 500 service instances
- Each emits 1,000 logs/sec at peak? That is huge. Interview scale down:
- 100 million log lines/day across the company
- Metrics: 50,000 time series, scrape every 15s
- Read/write: write-heavy ingest; query is bursty (incidents)
QPS
Logs:
100,000,000 / 86,400 ≈ 1,160 log lines/sec average
If peak is 5x, design for around 6,000 lines/sec ingest.
Metrics scrapes:
50,000 series / 15s ≈ 3,300 points/sec write
Storage
Log line ≈ 500 bytes
100,000,000 × 500 bytes ≈ 50 GB/day raw logs
14 days hot ≈ 700 GB indexed (often 2× with index overhead) → say 1–2 TB
Metrics: 50k series × 2 bytes compact × 4/min × 60 × 24 ≈ order of few GB/day. Metrics are cheap. Logs are expensive.
Cache / memory
Query cache for dashboard panels. Ingest buffers in Kafka.
Servers
A few ingest workers per 6k lines/sec. Index cluster sized for 2 TB. Interview estimates, not exact production numbers.
5. APIs
POST /v1/ingest/logs (agent, batched)
POST /v1/ingest/metrics
POST /v1/ingest/traces
GET /v1/logs/search?q=&from=&to=&service=
GET /v1/metrics/query PromQL-like { "query", "start", "end" }
POST /v1/alerts
GET /v1/dashboards/{id}
Auth: service tokens for ingest, user SSO for query.
6. Data Model
Log document
timestamp, service, host, level, message, trace_id, fields JSON.
Index: time + service + level.
Metric sample
metric_name, labels{}, timestamp, value.
Trace span
trace_id, span_id, parent, name, duration_ms.
Alert rule
query, threshold, for_duration, notify_channel.
Retention metadata per tenant: hot days, sample rate.
7. High-Level Design
Never write Elasticsearch on the request thread of the product app.
Observability platform architecture
Components:
- App Services → Collector/Agent → Ingestion Queue
- Workers → Log Index and Metrics Store
- Engineer → Dashboard/API → Query Service
- Alert worker → Notification
Why two stores? Log search wants inverted index. Metrics want compressed time series. One database is a weak interview answer.
Traces: sample 1% of requests, 100% of errors.
8. Deep Dives
Agents / SDKs
Java: Micrometer + logback JSON. Agent batches 1–5 seconds. Drop logs if the queue is full on the agent with a counter so you see loss.
Stream processing
Parse JSON, strip password, add region from host inventory, downsample debug logs.
Query
Logs: filter by time first (partition by day). Metrics: range scan. Trace: lookup by trace_id then show waterfall.
Alerts
Evaluate every minute. Do not alert on raw log volume spikes without a metric. Page from error rate and latency p99.
Retention and sampling
Debug 3 days, info 14, error 90. Sample health-check logs. See distributed systems ops reality: cost is a requirement.
9. Bottlenecks
- Hot index shards (one noisy service)
- Cardinality explosion (
user_idas metric label) - Query that scans 14 days of
* - Kafka lag during incidents (when you need logs most)
Mitigate: per-service quotas, label allowlists, forced time range, dedicated ingest vs query nodes.
10. Tradeoffs
| Choice | Upside | Downside |
|---|---|---|
| Push metrics | Simple | Stampede |
| Pull/scrape | Control | Miss dead instances |
| 100% traces | Perfect | Cost |
| Sampled traces | Cheap | Miss rare bugs unless errors kept |
Pick: queue ingest, split stores, sample traces, quota noisy tenants.
11. Failure Modes
| Failure | Handling |
|---|---|
| Index down | Buffer in queue; drop debug first |
| Agent down | Local disk spool |
| Alert worker down | Dead man's switch (watchdog) |
| PII leak | Redaction pipeline + drop fields |
| Wrong clock | Require NTP; reject far-future timestamps |
Product apps must timeout ingest so observability never deadlocks checkout. Tie to circuit breakers.
12. Interview Answer in 10 Minutes
"I would design observability as an ingest pipeline, not as apps writing to Elasticsearch.
100 million logs a day is about 1,160 lines per second, about 6,000 at 5× peak, about 50 GB/day. Metrics are far smaller. I would put SDKs and agents in front of a queue, then workers that write a log index and a time-series store.
Engineers query through a dashboard API. Alerts run on metrics, then notify. Traces are sampled except errors. Retention and sampling exist because logs will bankrupt you.
If the index is down, the queue holds. If the agent cannot send, it counts drops. The payment service never blocks on this platform."
13. Interview Talking Points
- Don't block the product.
- Logs vs metrics vs traces.
- Cardinality is the metrics killer.
- Quotas per service.
- Dead man's switch for alerting.
- Metrics: ingest lag, drop rate, query p95, alert MTTA.
- Kafka + Java workers are a natural story.
14. Follow-up Questions
OpenTelemetry?
Say SDK emits OTLP to the collector. You still own storage.
How is this different from a data lake?
This is operational: last 14 days, fast, indexed. Lake is cheap and slow.
Multi-region?
Ingest regional, query federates or ships to one query region.
Log-based metrics?
Extract counters in stream processing to avoid scanning logs for alerts.
15. Internal Links
Related JavaThoughts reading:
- System Design Interview Preparation
- Design a Distributed Queue
- Design an API Gateway
- Monitoring & observability
- Message queues
- Java
- Spring Boot
- Kafka
- Microservices
- Distributed Systems
- Scaling Java event-driven systems
- 20 system design concepts
Next in this series: Design a Web Crawler.
