Architecture and Streaming Interview Questions
Breadth-and-depth questions on data architecture. Learn the concepts in learn/architecture.
Tags: [core] = expected at every level · [senior] = expected at senior/staff level.
Lakehouse and storage
[core] Data warehouse vs data lake vs lakehouse?
Warehouse: structured, schema-on-write, ACID, great BI, proprietary storage. Lake: cheap open files of anything, schema-on-read, no ACID → swamps. Lakehouse: open table formats on object storage adding ACID, schema enforcement, time travel and governance, so BI and ML share one copy.
[core] Explain the medallion architecture.
Bronze (raw, append-only, replayable), silver (cleaned, deduplicated, conformed entities, SCD), gold (business-level marts/aggregates). Each layer has a contract, owners and consumers; it enables replay, debugging and decoupling of ingestion from modeling.
[core] Why Parquet over CSV/JSON for analytics?
Columnar: read only needed columns; strong compression; per-row-group statistics enable predicate pushdown; typed schema. CSV/JSON must be parsed fully and have no stats.
[senior] Delta vs Iceberg vs Hudi?
All provide ACID, time travel, schema evolution on object storage. Delta: ordered JSON log, tight Spark/Databricks integration, CDF, liquid clustering. Iceberg: snapshot/manifest tree, engine-neutral, hidden partitioning and partition evolution. Hudi: upsert-heavy pipelines, merge-on-read, indexing. Choose by engine ecosystem and workload.
[senior] What causes small files and how do you fix them?
Frequent streaming commits, over-partitioning, many writers, too many shuffle partitions. Fix: compaction (OPTIMIZE), optimized writes/auto-compaction, coarser partitioning or clustering instead, longer triggers, fewer output partitions.
[senior] Partitioning vs Z-ordering vs liquid clustering?
Partitioning creates directories by a low-cardinality column for pruning (risk: small files, hard to change). Z-order co-locates multi-column values within files via OPTIMIZE rewrite. Liquid clustering clusters incrementally with changeable keys: the default for new Delta tables.
Batch vs streaming
[core] When should you choose streaming over batch?
When acting on fresher data creates value (fraud, pricing, alerting, personalisation, user-facing analytics). Otherwise batch or incremental batch is cheaper and simpler. Use the slowest approach that meets the freshness SLA.
[core] Lambda vs Kappa architecture?
Lambda: separate batch (correct) and speed (fresh) layers merged at query time, with two codebases that drift. Kappa: one streaming pipeline; reprocess by replaying the log. Lakehouse hybrid: one codebase run as stream or batch, replaying from bronze tables.
[senior] Micro-batch vs continuous processing?
Micro-batch groups records per trigger (simple, efficient, exactly-once per batch, latency ≥ trigger). Continuous processes records individually with checkpoint barriers (ms latency, more complex). Matters only below ~1 s latency budgets.
Kafka
[core] How does Kafka guarantee ordering?
Only within a partition. Messages with the same key go to the same partition (hash(key) % partitions), so per-key order holds. No global ordering across partitions.
[core] What is a consumer group?
A set of consumers sharing a topic’s partitions; each partition is consumed by one member at a time. Different groups each receive all messages independently (fan-out). Max parallelism = number of partitions.
[senior] How do you choose the number of partitions?
Peak throughput ÷ per-partition throughput (plan ~5–10 MB/s), max consumer parallelism needed, headroom for growth (adding partitions later breaks key→partition mapping), and broker limits. Typically dozens to low hundreds.
[senior] What does acks=all with min.insync.replicas=2 guarantee?
A write is acknowledged only when at least 2 in-sync replicas have it, so losing one broker doesn’t lose acknowledged data. If fewer than 2 replicas are in sync, producers get errors instead of silently losing durability.
[senior] What is log compaction and when is it used?
Kafka retains at least the latest value per key (older values cleaned up; tombstones delete keys). Used for changelog/CDC topics and state restore, where the latest state per key matters more than full history.
Streaming semantics
[core] Event time vs processing time?
Event time is when the event happened (in the payload); processing time is when the system sees it. Aggregate on event time for correctness; processing time varies with delays and replays.
[core] What is a watermark?
The engine’s estimate that all events up to time T have arrived (e.g. max event time − 10 min). It finalises windows, evicts state, and defines how late data can be before it’s dropped or side-output.
[core] Tumbling vs sliding vs session windows?
Tumbling: fixed, non-overlapping. Sliding/hopping: fixed size, overlapping by slide interval (an event lands in several windows). Session: dynamic per key, closed after an inactivity gap.
[senior] How do you achieve exactly-once end to end?
Replayable source (offsets) + engine checkpointing offsets and state atomically + idempotent or transactional sink (MERGE by key, Kafka transactions, Delta txn ids). External side effects need idempotency keys. One at-least-once hop breaks the guarantee.
[senior] How do you join two streams without unbounded state?
Watermarks on both sides plus a time-range join condition (e.g. click within 15 minutes after impression), so the engine can evict state older than the bound. Outer joins require both.
[senior] How do you handle very late data (days)?
Stream with a modest watermark for freshness; late events still land in bronze; a batch restatement job recomputes affected periods (MERGE into aggregates); dashboards label recent periods as preliminary.
Ingestion and CDC
[core] Full load vs incremental vs CDC?
Full: simple, heavy, catches deletes via diff. Incremental watermark (updated_at): light, misses deletes and intermediate states, depends on app discipline. Log-based CDC: every change incl. deletes in commit order, minimal source load, more infrastructure.
[senior] What can go wrong with Debezium on Postgres?
Replication slot lag retains WAL and can fill the source disk; schema changes; TOAST columns missing in updates; initial snapshot load on large tables; ordering only per key; failover handling of slots.
[senior] What is the outbox pattern?
Write the business change and an event row in the same local DB transaction; CDC publishes the outbox to Kafka. Solves the dual-write problem: events are published if and only if the transaction committed.
Orchestration and reliability
[core] What makes a pipeline idempotent?
Re-running with the same input yields the same output: overwrite partitions for a logical date or MERGE on keys, deterministic logic (no now()), deterministic keys, no side effects in transformations.
[senior] How do you design a safe backfill?
Confirm raw data retention, run partition-parallel idempotent jobs into a shadow table, cap concurrency, validate vs prod, swap atomically, rebuild downstream in lineage order, communicate restated numbers.
[senior] What is write-audit-publish?
Write new data to a staging location/branch, run quality checks, and only then atomically publish to the production table, so consumers never see unvalidated data.