This intensive three-day workshop is dedicated to designing and fine-tuning high-efficiency data-processing workloads utilizing PySpark, Pandas, and Polars within Kubernetes-based environments.
Learners will gain a hands-on understanding of how Spark applications operate on Kubernetes and how application-level configuration choices impact performance, scalability, resource utilization, and overall cost. The curriculum explores critical optimization domains, including executor sizing, memory management, dynamic allocation, partitioning methodologies, shuffle mechanics, mitigating small-file issues, and optimizing Parquet processing.
The training also tackles frequent obstacles encountered when using Pandas, such as memory constraints and out-of-memory errors, while introducing Polars as a high-performance alternative for specific data-processing tasks. Through practical exercises, participants will learn to diagnose performance and memory bottlenecks, evaluate various configuration strategies, and implement optimization techniques in realistic ETL and machine learning contexts.
The primary focus of the course is on practical decision-making: equipping participants with the ability to identify performance bottlenecks, select the right tools, configure Spark efficiently, and strike a balance between performance and infrastructure resource costs.
Read more...