170 cards
This deck offers a broad introduction to the core concepts of data engineering, covering the foundational ideas you'll encounter in the field. You'll find questions about how data moves and is stored, from pipelines and orchestration tools like Apache Airflow, to storage architectures such as data warehouses, data lakes, and the schema designs that organize them. It also touches on key technologies like Apache Spark and Apache Kafka, as well as important distinctions like batch versus stream processing and ETL versus ELT.
It's a great fit if you're starting out as a data engineer, a data analyst looking to understand the systems behind your reports, or a software developer exploring a new specialization. The questions are also useful for anyone preparing for technical interviews, since many of the topics here come up frequently in screening conversations about data infrastructure and design choices.
To get the most out of these cards, try to connect each concept to a practical scenario as you review, for example imagining how a particular tool would fit into a real workflow. Because the topics build on one another, it's worth spacing your review sessions out over several days rather than cramming, so the definitions and comparisons have time to settle. Revisiting the more abstract ideas, like schema design or data quality, after a day or two often makes them stick much better than a single long session.