Javathoughts Logo
Javathoughts
Published on
Views

Design Logging, Metrics, and Observability Platform System Design Interview Guide

Authors
  • avatar
    Name
    Javed Shaikh
    Twitter

← System Design Interview Preparation

This guide walks through Design a Logging, Metrics, and Observability Platform the way you would in a backend or Java interview. The numbers are interview estimates. They help you show your thinking. They are not a production capacity plan.


1. Problem

When production breaks, engineers need logs (what happened), metrics (how much), and a little tracing (which hop).

Example:

  • Payment service emits payment.failed logs and a payments_failed_total counter
  • An agent on the box or sidecar ships data
  • A pipeline indexes logs and stores time series
  • A dashboard shows error rate
  • An alert pages Slack at 2am

This is the platform behind monitoring & observability, not an APM vendor clone of every feature.


2. Functional Requirements / FR

RequirementWhat it means
App logsStructured JSON with level, service, trace id.
MetricsCounters, gauges, histograms.
TracesOptional: span per RPC, sampled.
Agents / SDKsPush from JVM, sidecar, or both.
Ingestion pipelineBuffer so apps do not block on storage.
Stream processingParse, drop PII, sample, enrich.
Storage / indexingSearch logs; query metrics by time.
DashboardsGraphs and log search UI.
AlertsThresholds and notifications.
Retention / samplingHot 7–14 days, cheap archive later.

Out of scope: building a full BI warehouse, log-based billing product.


3. Non-Functional Requirements / NFR

RequirementWhy it matters
Ingestion durabilityDo not lose the error logs that explain an outage.
Query latencyDashboards in seconds, not minutes.
App isolationTelemetry must not take down checkout.
Multi-tenant isolationOne team cannot fill the cluster.
Cost controlLogs grow faster than traffic.

Interview line: async ingest, two stores (logs vs metrics), sample traces, alerts off metrics.


4. Back-of-the-Envelope Calculation

Traffic assumptions

  • 500 service instances
  • Each emits 1,000 logs/sec at peak? That is huge. Interview scale down:
  • 100 million log lines/day across the company
  • Metrics: 50,000 time series, scrape every 15s
  • Read/write: write-heavy ingest; query is bursty (incidents)

QPS

Logs:

100,000,000 / 86,400 ≈ 1,160 log lines/sec average
If peak is 5x, design for around 6,000 lines/sec ingest.

Metrics scrapes:

50,000 series / 15s ≈ 3,300 points/sec write

Storage

Log line ≈ 500 bytes

100,000,000 × 500 bytes ≈ 50 GB/day raw logs
14 days hot ≈ 700 GB indexed (often 2× with index overhead) → say 1–2 TB

Metrics: 50k series × 2 bytes compact × 4/min × 60 × 24 ≈ order of few GB/day. Metrics are cheap. Logs are expensive.

Cache / memory

Query cache for dashboard panels. Ingest buffers in Kafka.

Servers

A few ingest workers per 6k lines/sec. Index cluster sized for 2 TB. Interview estimates, not exact production numbers.


5. APIs

POST /v1/ingest/logs     (agent, batched)
POST /v1/ingest/metrics
POST /v1/ingest/traces

GET  /v1/logs/search?q=&from=&to=&service=
GET  /v1/metrics/query   PromQL-like { "query", "start", "end" }

POST /v1/alerts
GET  /v1/dashboards/{id}

Auth: service tokens for ingest, user SSO for query.


6. Data Model

Log document

timestamp, service, host, level, message, trace_id, fields JSON.

Index: time + service + level.

Metric sample

metric_name, labels{}, timestamp, value.

Trace span

trace_id, span_id, parent, name, duration_ms.

Alert rule

query, threshold, for_duration, notify_channel.

Retention metadata per tenant: hot days, sample rate.


7. High-Level Design

Never write Elasticsearch on the request thread of the product app.

Observability platform architecture

Observability platform architectureApps ship logs and metrics through agents into a queue. Workers index logs and write metrics. Engineers query dashboards. Alerts fire from the same stores.⚙️App Services📥Collector / Agent📩Ingestion Queue👷Workers🔍Log Index📊Metrics Store👤Engineer⚙️Query / Dashboard👷Alert Worker🔔Alert
Apps ship logs and metrics through agents into a queue. Workers index logs and write metrics. Engineers query dashboards. Alerts fire from the same stores.

Components:

  • App Services → Collector/Agent → Ingestion Queue
  • Workers → Log Index and Metrics Store
  • Engineer → Dashboard/API → Query Service
  • Alert worker → Notification

Why two stores? Log search wants inverted index. Metrics want compressed time series. One database is a weak interview answer.

Traces: sample 1% of requests, 100% of errors.


8. Deep Dives

Agents / SDKs

Java: Micrometer + logback JSON. Agent batches 1–5 seconds. Drop logs if the queue is full on the agent with a counter so you see loss.

Stream processing

Parse JSON, strip password, add region from host inventory, downsample debug logs.

Query

Logs: filter by time first (partition by day). Metrics: range scan. Trace: lookup by trace_id then show waterfall.

Alerts

Evaluate every minute. Do not alert on raw log volume spikes without a metric. Page from error rate and latency p99.

Retention and sampling

Debug 3 days, info 14, error 90. Sample health-check logs. See distributed systems ops reality: cost is a requirement.


9. Bottlenecks

  • Hot index shards (one noisy service)
  • Cardinality explosion (user_id as metric label)
  • Query that scans 14 days of *
  • Kafka lag during incidents (when you need logs most)

Mitigate: per-service quotas, label allowlists, forced time range, dedicated ingest vs query nodes.


10. Tradeoffs

ChoiceUpsideDownside
Push metricsSimpleStampede
Pull/scrapeControlMiss dead instances
100% tracesPerfectCost
Sampled tracesCheapMiss rare bugs unless errors kept

Pick: queue ingest, split stores, sample traces, quota noisy tenants.


11. Failure Modes

FailureHandling
Index downBuffer in queue; drop debug first
Agent downLocal disk spool
Alert worker downDead man's switch (watchdog)
PII leakRedaction pipeline + drop fields
Wrong clockRequire NTP; reject far-future timestamps

Product apps must timeout ingest so observability never deadlocks checkout. Tie to circuit breakers.


12. Interview Answer in 10 Minutes

"I would design observability as an ingest pipeline, not as apps writing to Elasticsearch.

100 million logs a day is about 1,160 lines per second, about 6,000 at 5× peak, about 50 GB/day. Metrics are far smaller. I would put SDKs and agents in front of a queue, then workers that write a log index and a time-series store.

Engineers query through a dashboard API. Alerts run on metrics, then notify. Traces are sampled except errors. Retention and sampling exist because logs will bankrupt you.

If the index is down, the queue holds. If the agent cannot send, it counts drops. The payment service never blocks on this platform."


13. Interview Talking Points

  • Don't block the product.
  • Logs vs metrics vs traces.
  • Cardinality is the metrics killer.
  • Quotas per service.
  • Dead man's switch for alerting.
  • Metrics: ingest lag, drop rate, query p95, alert MTTA.
  • Kafka + Java workers are a natural story.

14. Follow-up Questions

OpenTelemetry?
Say SDK emits OTLP to the collector. You still own storage.

How is this different from a data lake?
This is operational: last 14 days, fast, indexed. Lake is cheap and slow.

Multi-region?
Ingest regional, query federates or ships to one query region.

Log-based metrics?
Extract counters in stream processing to avoid scanning logs for alerts.


Related JavaThoughts reading:


Next in this series: Design a Web Crawler.