Skip to content
Reliable Data Engineering
Practice problem medium llm-opstracingevaluationstreaminglakehouse
Practise with timer, notes and rubric

Design an LLM Observability and Evaluation Data Platform

Problem

A company runs 40 internal and customer-facing LLM applications (chatbots, RAG assistants, agents that call tools). Leadership asks: What do we spend per app? Which apps got worse after last week’s model or prompt change? Where are users unhappy? Design the platform that captures, stores and analyses LLM traces and evaluations.


Clarifying questions

QuestionAssumed answer
Volume?20 M LLM calls/day; agents produce 5–30 spans per task
Payload sizes?Prompts/completions average 6 KB, up to 200 KB
Privacy?Prompts may contain PII/customer data; retention 30 days for raw text, metrics forever
Real-time needs?Cost/error dashboards within minutes; quality evals can lag hours
Eval methods?Heuristics, LLM-as-judge, human labels, user feedback

1. Estimates

20M calls × ~3 spans avg = 60M spans/day ≈ 700/s avg, ~3k/s peak
Raw payload: 20M × 6 KB ≈ 120 GB/day (compresses ~4–5×) → 30-day raw retention ≈ 1 TB compressed
Metrics rows (no text): 60M × 300 B = 18 GB/day → keep forever
LLM-as-judge on 100% of traffic would cost as much as the apps themselves → sample (e.g. 5%) + target low-feedback traces

2. Architecture

flowchart LR
    APPS[LLM apps + agents<br/>OpenTelemetry GenAI SDK] --> COLL[Trace collector<br/>OTLP endpoint]
    COLL --> RED[PII redaction +<br/>payload offload]
    RED --> K[[Kafka: spans]]
    RED --> OBJ[(Object store:<br/>large payloads by hash)]
    K --> STR[Streaming job: cost calc,<br/>error/latency aggregates]
    STR --> TS[(Metrics store / OLAP<br/>dashboards, alerts)]
    K --> BR[(Bronze: spans)]
    BR --> SIL[(Silver: traces assembled,<br/>requests, tool calls)]
    SIL --> SAMP[Sampler: random + low feedback<br/>+ new prompt versions]
    SAMP --> JUDGE[Eval workers:<br/>heuristics, LLM-judge, PII checks]
    JUDGE --> EV[(silver.evaluations)]
    FB[User feedback events] --> SIL
    HUM[Human labelling queue] --> EV
    EV --> GOLD[(Gold: app × version × day<br/>quality, cost, latency)]
    SIL --> GOLD
    GOLD --> DASH[Dashboards + regression alerts]
    EV --> GS[(Golden datasets<br/>for offline CI evals)]

3. Data model

silver.spans (trace_id, span_id, parent_span_id, app, env, span_kind,  -- llm | retrieval | tool | agent_step
              model, prompt_version, start_ts, end_ts, latency_ms,
              input_tokens, output_tokens, cached_tokens, cost_usd,
              status, error_type, payload_ref, user_hash, session_id)
silver.traces  (trace_id, app, root latency, total cost, n_llm_calls, n_tool_calls, outcome, feedback)
silver.evaluations (trace_id, evaluator, evaluator_version, metric, score, rationale_ref, ts)
gold.app_daily (app, date, prompt_version, model, requests, p50/p95 latency, cost, error_rate,
                faithfulness_avg, thumbs_down_rate, eval_sample_size)

4. Deep dives

4.1 Instrumentation and trace assembly

4.2 Cost attribution

Price table (model, token type, effective date) as an SCD2 dimension → cost = input × price_in + output × price_out − cached discounts, computed in streaming for live dashboards and recomputed in batch (prices change, corrections). Tag by app/team/customer for chargeback.

4.3 Evaluation pipeline

flowchart LR
    T[Trace] --> H[Cheap heuristics on 100%:<br/>empty answer, refusal, JSON valid,<br/>latency, tool errors]
    T --> S{Sampled?}
    S -->|"5% random<br/>+ all thumbs-down<br/>+ new versions 20%"| J[LLM-as-judge:<br/>faithfulness vs retrieved context,<br/>relevance, policy]
    J --> CAL[Calibrate vs human labels<br/>monthly agreement check]
    H --> E[(evaluations)]
    J --> E

4.4 Privacy

Redaction at the collector (PII detectors) before data lands; raw text retained 30 days with restricted access; metrics and evaluations (no raw text) kept long-term; customer-facing apps can opt out of payload capture entirely (metadata only).

5. Trade-offs

DecisionChoiceAlternative
StorageLakehouse (Delta) for analytics + OLAP/metrics store for live dashboardsVendor tool (LangSmith, Langfuse, Arize): fast start; build when you need cross-app analytics and governance
Eval coverageHeuristics on all + sampled judgesJudge everything (cost explodes)
PayloadsOffloaded by hashInline in spans (bloat, harder to purge)

6. What separates a senior answer

7. Follow-up questions

An agent's average cost doubled this week. How do you find out why?

Drill down in gold/silver: by app version, model, prompt_version, span_kind. Look at tokens per trace and LLM calls per trace (agent loops increasing?), cache hit rate drop, longer retrieved context, retry storms on tool errors. Trace-level examples of the most expensive tasks usually reveal the loop or prompt bloat.

How do you know your LLM judge is trustworthy?

Measure agreement with human labels on a stratified sample (Cohen’s kappa / correlation), check for biases (length, position, self-preference), version and freeze judge prompts, re-calibrate when the judge model changes, and use pairwise comparisons for A/B decisions.


Self-assessment rubric