Skip to content
Reliable Data Engineering
Practice problem hard identity-resolutionstreamingprofilessegmentationreverse-etl
Practise with timer, notes and rubric

Design a Real-Time Customer Data Platform

Problem

Marketing and product teams want a single, up-to-date profile per customer: identities across devices, traits (country, lifetime value), recent behaviour and segment memberships. They want to trigger messages within seconds (“abandoned cart after 30 minutes”) and sync audiences to ad and email platforms. Data comes from web/app event SDKs, the CRM, the order database and the support tool. Consent and GDPR deletion must be enforced everywhere.


Clarifying questions

QuestionAssumed answer
Users?150M known + anonymous profiles; 40M monthly active
Events?2B events/day (page views, product views, cart, purchases)
Latency?Real-time segments and triggers < 10 s; batch traits daily
Destinations?Email/push service, ad platforms (audiences), CRM, in-app personalisation API
Identity sources?Anonymous device IDs, login user IDs, emails, phone numbers, loyalty IDs
Consent?Per purpose (marketing, personalisation, analytics), per region

1. Requirements

Functional: ingest events and records; resolve identities into persistent profiles (merge anonymous browsing into the customer after login); compute traits and segments (streaming and batch); a profile lookup API; audience sync to destinations; triggers; consent and deletion propagation.

Non-functional: profile reads < 20 ms p99; segment membership updates < 10 s; deterministic, explainable merges (and un-merges); privacy by design; destinations rate-limited and retried.

2. Estimates

3. Architecture

flowchart LR
    subgraph SRC[Sources]
        SDK[Web / app SDKs]
        CRM[(CRM)]
        ORD[(Orders DB)]
        SUP[Support tool]
    end
    SDK --> COL[Collector: validation, consent check,<br/>schema registry]
    COL --> K[(Kafka: events<br/>key = anonymous_id or user_id)]
    CRM -- CDC/connector --> K2[(Kafka: records)]
    ORD -- CDC --> K2
    SUP --> K2
    K & K2 --> IDR[Identity resolution service<br/>deterministic rules, identity graph]
    IDR --> IDG[(Identity graph store<br/>identifier → profile_id)]
    IDR --> PS[Profile stream processor<br/>traits, rolling counters, segment rules]
    PS --> PROF[(Profile store: KV<br/>traits, segments, consent)]
    PS --> TRIG[Trigger engine<br/>timers: abandoned cart]
    TRIG --> MSG[Email / push service]
    PROF --> API[Profile API<br/>personalisation]
    K & K2 --> LH[(Lakehouse: bronze → silver → gold)]
    LH --> BATCH[Batch traits & ML scores<br/>LTV, churn propensity]
    BATCH --> PROF
    PROF --> AUD[Audience sync<br/>diffs, rate limits, consent filter]
    AUD --> ADS[Ad platforms / CRM]
    DEL[Consent & deletion service] --> PROF & IDG & LH & AUD

4. Data model

5. Deep dives

5.1 Identity resolution

See the Customer 360 data model and the accounts merge problem.

5.2 Streaming traits and segments

5.3 Activation (audience sync)

6. Trade-offs

DecisionChoiceAlternative
IdentityDeterministic rules + provenance, probabilistic optionalProbabilistic everywhere (more reach, more wrong merges)
Profile storeKV with versioned writesWarehouse-native CDP (simpler, slower activation)
SegmentsStreaming for real-time segments, batch for heavy onesAll batch (no real-time triggers) / all streaming (expensive for complex ML traits)
ActivationDiff-based syncsFull audience uploads (rate limits, cost)

7. Failure modes

8. What separates a senior answer

9. Follow-up questions

Marketing wants a "warehouse-native" CDP instead. What changes?

Profiles, identity resolution and segments are computed as models in the lakehouse/warehouse (SQL + scheduled jobs), and activation uses reverse ETL. It’s simpler, cheaper and governed in one place, but freshness is minutes to hours. Keep a small streaming path only for the few triggers that need seconds.

How do you measure identity resolution quality?

Track merge sizes, the distribution of identifiers per profile, rates of blocklisted identifiers, un-merge requests, and precision on a labelled sample (manual review of merged pairs). Monitor match rates at destinations as a downstream signal.


Self-assessment rubric