Modern Data & AI Fundamentals + Ingestion at Scale
Objective: Establish the foundations of the Databricks ecosystem and Lakehouse architecture, and master batch and streaming ingestion techniques with Auto Loader and Spark Structured Streaming.
- Lakehouse vs Data Warehouse architecture
- Databricks clusters and runtime
- Notebooks, repos and version control
- Introduction to Delta Lake
- Auto Loader and Structured Streaming
- Data sources: S3, ADLS, GCS
- Schema evolution and enforcement
- Checkpointing and error handling
Hands-on Lab: Set up a Databricks workspace and build a continuous ingestion pipeline from S3 to a Delta table.
Outcome: Solid understanding of the Databricks environment and the ability to design robust ingestion pipelines for data in motion.