Skip to content
Reliable Data Engineering
Lesson
Open in the interactive app

Databricks 5: DevOps, Security, Serving and AI

Senior Databricks roles are judged on how you ship and operate the platform: code promotion, secure networking, serving data to BI and applications, and supporting ML/AI workloads. This module covers what to say about each.


1. Code as the source of truth: Git folders and Asset Bundles

# databricks.yml
bundle:
  name: sales_pipelines
targets:
  dev:
    mode: development            # prefixes resources with the developer's name, pauses schedules
    workspace: { host: https://dev.cloud.databricks.com }
  prod:
    mode: production
    workspace: { host: https://prod.cloud.databricks.com }
    run_as: { service_principal_name: sp-sales-prod }
resources:
  jobs:
    daily_sales:
      name: daily_sales
      tasks:
        - task_key: ingest
          pipeline_task: { pipeline_id: ${resources.pipelines.sales_ingest.id} }
        - task_key: gold
          depends_on: [{ task_key: ingest }]
          notebook_task: { notebook_path: ./src/gold.py, base_parameters: { catalog: "${var.catalog}" } }

CI/CD flow:

  1. PR → unit tests (pytest with local Spark or Databricks Connect), linting, databricks bundle validate.
  2. Merge → bundle deploy -t staging → integration tests on staging data.
  3. Approval → bundle deploy -t prod running as a service principal.

Infrastructure (workspaces, metastores, networking, UC grants) is usually managed with Terraform; application resources with bundles. No manual changes in prod.

Environment isolation: separate workspaces per environment bound to environment catalogs (dev, prod), parameterised code (catalog as a job parameter), and service principals with least privilege per environment.


2. Security and networking

ControlPurpose
SSO + SCIM from the identity providerCentral identity; groups synced automatically
Service principals for automationNo personal tokens in production; OAuth M2M
Secure cluster connectivity (no public IPs)Compute nodes have no public IPs; outbound-only connection to the control plane
Customer-managed VNet/VPCYour network controls (firewalls, routes, egress filtering) for classic compute
Private Link (front-end and back-end)Private connectivity between users/compute and the control plane
IP access listsRestrict workspace access to corporate networks
Serverless network policies / egress controlRestrict which destinations serverless compute can reach
Customer-managed keysEncrypt managed services data and storage with your keys
Secrets (secret scopes, backed by Key Vault/Secrets Manager)Keep credentials out of code; redacted in notebook output
Unity CatalogData access, auditing and lineage (module 3)
Compliance security profile / enhanced security monitoringHardened images and monitoring for regulated workloads (HIPAA, PCI)

Interview framing: “Defence in depth: identity (SSO, SCIM, service principals), network (no public IPs, private link, egress control), data (UC permissions, ABAC, encryption with customer keys), and monitoring (audit logs in system tables, alerts).“


3. Serving data: SQL warehouses and BI


4. ML and AI on Databricks (what data engineers should know)


5. Operating the platform


Interview questions

How do you implement CI/CD for Databricks pipelines?

Keep code in Git; define jobs and pipelines as Databricks Asset Bundles with dev/staging/prod targets; run unit tests and bundle validate on PRs; deploy to staging on merge and run integration tests; promote to prod with approval, running as a service principal. Infrastructure and permissions go in Terraform. Use parameterised catalogs per environment and never edit prod resources manually.

How do you secure a Databricks deployment for a regulated company?

SSO and SCIM groups, service principals for automation, secure cluster connectivity with no public IPs in a customer-managed VNet, Private Link for front-end and back-end traffic, IP access lists, egress controls for serverless, customer-managed keys, secrets in a vault-backed scope, Unity Catalog with least-privilege group grants, row filters/column masks or ABAC for PII, the compliance security profile where required, and audit logs in system tables monitored with alerts.

Dashboards on a SQL warehouse are slow at 9am. What do you check?

Query profiles for the slow queries (scan size, spill, Photon coverage, join types); whether warehouses are queueing (scale max clusters up or use serverless with intelligent workload management); whether queries hit the result/disk cache; table layout (liquid clustering on filter columns, file sizes, stats); and whether the dashboard should read a materialized view or aggregate table instead of raw facts. Also stagger heavy scheduled refreshes.

What's the data engineer's role in a RAG application on Databricks?

Build reliable ingestion of source documents into volumes/Delta, parse and chunk them with metadata and access-control attributes, compute embeddings in a pipeline, keep a Vector Search index in sync with the source table (delta sync), handle updates and deletes (including permission changes), log and process inference tables for quality monitoring, and govern everything in Unity Catalog.