Cloud Platforms: Learning Path
Many senior data engineering roles name a platform in the job description, and the interview probes it in depth: not just “have you used it?” but how it works, what it costs and how you’d operate it safely. This track covers platform-specific knowledge. It starts with Databricks; AWS, Azure and GCP data services will follow.
Databricks
| # | Module | You will be able to |
|---|---|---|
| 1 | Platform architecture & compute | Explain control vs compute plane, pick the right compute (all-purpose, jobs, serverless, SQL warehouses), reason about Photon and DBU cost |
| 2 | Delta Lake internals | Explain the transaction log, optimistic concurrency and conflicts, MERGE internals, VACUUM and time travel, liquid clustering, deletion vectors, CDF and clones |
| 3 | Unity Catalog & governance | Design catalogs and permissions, implement row filters, masks and ABAC, use lineage and system tables, share data with Delta Sharing and federation |
| 4 | Ingestion, pipelines, jobs & streaming | Choose Auto Loader vs COPY INTO, build Declarative Pipelines with expectations and AUTO CDC, orchestrate with Lakeflow Jobs, run Structured Streaming well |
| 5 | DevOps, security, serving & AI | Ship with Asset Bundles and CI/CD, secure the deployment, serve BI with SQL warehouses, support MLflow, Model Serving and Vector Search |
Then: Databricks interview questions · End-to-end Databricks design scenario · Related: Spark internals & tuning
Databricks evolves quickly and renames products often (Delta Live Tables → Lakeflow Declarative Pipelines, Workflows → Lakeflow Jobs). Interviewers accept either name; knowing both shows you’re current.