Skip to content
Reliable Data Engineering
Overview

Spark and Databricks Practice Problems

#ProblemDifficulty
1Fix the skewed joinMedium
2Incremental load with watermarks and MERGEMedium
3SCD Type 2 with Delta MERGEHard
4Handle extreme data skew with saltingHard
5Fix the small files problemEasy
6Read the physical plan and fix the queryMedium
7Diagnose spill from Spark UI metrics and size the executorsHard
8Rewrite slow Python UDFs with native functions and pandas UDFsMedium
9Partitioning, dynamic partition pruning and caching decisionsHard
10Design an end-to-end Databricks lakehouse for a retailerHard
11Debug a nightly job: driver OOM, broadcast timeouts and a stale cacheHard

Concepts: Spark & Databricks learning path ยท Storage & table formats