Python for Data Engineering: Interview Questions
Conceptual Python questions. Coding problems are in practice/python.
Tags: [core] = expected at every level · [senior] = expected at senior/staff level.
Language
[core] List vs tuple vs set vs dict: when each?
List: ordered, mutable sequence. Tuple: immutable (hashable if contents are), good as dict keys/records. Set: unique members, O(1) membership. Dict: key→value, O(1) lookup, insertion-ordered since 3.7.
[core] What is a generator and why do data engineers care?
A function using yield that produces values lazily, one at a time. Constant memory for arbitrarily large inputs (stream a 50 GB file line by line), composable pipelines (parse → filter → batch).
[core] Explain mutable default arguments.
Defaults are evaluated once at function definition, so def f(x, acc=[]) shares the same list across calls. Use acc=None and create inside.
[core] Shallow vs deep copy?
Shallow copy (list(x), dict.copy(), copy.copy) copies the container but shares nested objects; deep copy (copy.deepcopy) recursively copies. Matters when mutating nested records.
[senior] What are decorators and a real DE use case?
Functions that wrap other functions to add behaviour. Uses: retries with backoff, timing/metrics, logging, caching (functools.lru_cache), input validation. Use functools.wraps to preserve metadata.
[senior] Context managers?
Objects with __enter__/__exit__ (or @contextmanager) guaranteeing setup/teardown: files, DB connections, locks, temp tables, timing blocks. Ensures cleanup on exceptions.
Performance and concurrency
[core] What is the GIL and how does it affect data pipelines?
The Global Interpreter Lock lets only one thread execute Python bytecode at a time (in CPython builds with the GIL). Threads still help for I/O-bound work (API calls, S3 downloads); CPU-bound work needs multiprocessing, vectorised libraries (NumPy, Arrow, Polars) or distributed engines (Spark).
[senior] threading vs multiprocessing vs asyncio?
threading: I/O concurrency with shared memory, GIL-limited for CPU. multiprocessing: true parallel CPU, separate memory, serialisation overhead. asyncio: single-threaded cooperative concurrency for many I/O-bound tasks (thousands of HTTP calls) with low overhead.
[senior] How do you process a file larger than memory in Python?
Stream it (iterate lines/chunks, pandas.read_csv(chunksize=), pyarrow.dataset batches), aggregate incrementally, use external sort for ordering, or switch to DuckDB/Polars (out-of-core) or Spark.
[senior] Why are Python UDFs slow in Spark and what are the alternatives?
Rows are serialised between the JVM and Python worker processes and executed row by row. Use built-in functions (Catalyst-optimised), Pandas UDFs (Arrow, vectorised batches), or Scala/SQL functions.
Ecosystem
[core] pandas vs Polars vs DuckDB vs PySpark?
pandas: ubiquitous, single-threaded, in-memory. Polars: fast multi-threaded DataFrame library (Rust, Arrow), lazy optimisation, larger-than-memory streaming. DuckDB: in-process analytical SQL, extremely fast on local files. PySpark: distributed for data beyond one machine. Use the smallest tool that fits the data.
[core] What is Apache Arrow?
A columnar in-memory format standard enabling zero-copy data exchange between systems (pandas ↔ Spark ↔ DuckDB ↔ Polars), vectorised processing, and fast Parquet I/O.
[senior] How do you test data pipelines in Python?
pytest unit tests on pure transformation functions with small fixture DataFrames (local Spark session or DuckDB), property-based tests for edge cases, schema/contract tests, integration tests on sampled data, and data diffs between versions. Keep I/O at the edges so logic is testable.
[senior] How do you structure a production Python data project?
Package (src/ layout, pyproject), typed modules for transformations separate from I/O and orchestration, config via env/params, structured logging, tests + linting (ruff, mypy) in CI, dependency pinning, and deployment as wheels/containers.