Skip to content
Reliable Data Engineering
Practice problem hard logsmetricstime-seriestiered-storagecardinality
Practise with timer, notes and rubric

Design a Logs and Metrics Observability Platform

Problem

Your company runs 5,000 microservices across three regions. Engineers need to search logs, view metrics dashboards and receive alerts within seconds. Today’s vendor bill is exploding. Design an in-house observability data platform for logs and metrics (traces optional), with retention of 30 days searchable and 1 year archived.


Clarifying questions

QuestionAssumed answer
Log volume?5M lines/s peak, ~500 bytes average → 2.5 GB/s, ~150 TB/day raw
Metrics?50M samples/s, ~20M active time series (after cardinality controls)
Query patterns?Logs: recent (last 1-24 h) search by service, level, trace_id, free text. Metrics: dashboards over 1 h-30 d, alert rules evaluated every 30 s
Retention?Logs: 7 days hot, 30 days warm, 1 year cold (compliance). Metrics: 15 days raw, 13 months downsampled
Latency?Logs searchable < 30 s after emission; alerts evaluated on data < 60 s old
Multi-tenancy?Teams as tenants with quotas and chargeback

1. Requirements

Functional: collect logs and metrics; parse and enrich (service, region, Kubernetes metadata); log search (filters + full text); metrics queries (PromQL-like); alerting; tiered retention and archive; per-team quotas.

Non-functional: ingestion must never block applications (drop or sample rather than back up into apps); the alerting path must be more reliable than the platform it monitors; cost per GB well below the vendor’s; isolation between tenants.

2. Estimates

3. Architecture

flowchart LR
    subgraph HOSTS[Services / nodes]
        APP[Apps] --> AG[Agent: OpenTelemetry Collector / Vector<br/>batching, sampling, local buffer]
    end
    AG --> GW[Ingest gateways<br/>auth, quotas, rate limits]
    GW --> KL[(Kafka: logs<br/>key = tenant:service)]
    GW --> KM[(Kafka: metrics)]
    KL --> PROC[Stream processing<br/>parse, enrich, redact PII,<br/>route by tenant & level]
    PROC --> HOT[(Hot log store<br/>ClickHouse / OpenSearch<br/>7 days, SSD)]
    PROC --> OBJ[(Object storage: Parquet<br/>partitioned by tenant/date/hour)]
    OBJ --> WARM[Warm query engine<br/>Trino/Spark over Parquet, 30 days]
    OBJ --> COLD[(Archive tier, 1 year)]
    KM --> TSDB[(Metrics TSDB<br/>Prometheus-compatible: Mimir / VictoriaMetrics / M3)]
    TSDB --> DOWN[Downsampling: 1m → 5m → 1h]
    TSDB --> ALERT[Alert evaluator<br/>rules every 30 s]
    HOT --> UI[Query UI: log search, dashboards]
    WARM --> UI
    TSDB --> UI
    ALERT --> PAGE[Paging / incident tools]

4. Data model

Logs (columnar): timestamp, tenant, service, env, region, host/pod, level, trace_id, span_id, message, plus a attributes map with promoted columns for frequently filtered keys (e.g. http.status, user_tier). Sort/primary key (tenant, service, timestamp) for locality; skip indexes (bloom filters, token indexes) on trace_id and message tokens.

Metrics: a series = metric name + label set (http_requests_total{service="checkout", status="500", region="eu"}); samples (timestamp, value) compressed per series in blocks (delta-of-delta timestamps, XOR-encoded floats).

5. Deep dives

5.1 Ingestion that never hurts applications

5.2 Cost control for logs

5.3 Metrics cardinality

5.4 Reliable alerting

5.5 PII and security

5.6 Multi-region

6. Trade-offs

DecisionChoiceAlternative
Hot log storeColumnar (ClickHouse) with skip indexesInverted-index search (OpenSearch): faster full-text, much higher storage/indexing cost
Warm/coldParquet on object storage + query engineKeep everything in the hot store (cost explodes)
Metrics storePrometheus-compatible horizontally scalable TSDBLogs-as-metrics (expensive queries)
Overload behaviourSample/drop low-priority data at agentsBackpressure into apps (outages)
AlertingSeparate, simpler evaluation path + meta-monitoringAlerts on the same query cluster as dashboards

7. Failure modes

8. What separates a senior answer

9. Follow-up questions

How do you add distributed tracing without tripling cost?

Tail-based sampling in collectors: keep 100% of traces with errors or high latency and a small percentage of normal ones; store spans in the same columnar store keyed by trace_id with a skip index; link logs and metrics via trace_id/exemplars so engineers can pivot from a metric spike to example traces.

A team wants to keep 1 year of searchable logs for security investigations.

Keep them in the warm/cold Parquet tier with partitioning by tenant/date and bloom filters on IP, user and trace fields; provide an asynchronous query interface (minutes, not seconds) and charge the team for the extra storage and scans. Security events specifically can be routed to a dedicated SIEM table with longer hot retention.


Self-assessment rubric