This intensive three-day workshop is designed to help you build and fine-tune high-performance data-processing pipelines using PySpark, Pandas, and Polars within Kubernetes ecosystems.
Attendees will gain a practical grasp of Spark execution mechanics on Kubernetes, learning how application-level configuration impacts performance, scalability, resource efficiency, and operational costs. The curriculum addresses critical optimization strategies, such as executor sizing, memory management, dynamic allocation, partitioning techniques, shuffle optimization, and efficient Parquet handling, while also tackling the small-file problem.
The program also explores common pitfalls associated with Pandas, such as memory constraints and out-of-memory errors, and positions Polars as a high-performance alternative for specific workload scenarios. Through hands-on lab exercises, participants will learn to diagnose performance and memory bottlenecks, evaluate different configuration strategies, and apply optimization techniques to real-world ETL and machine learning projects.
The core focus of the training is on practical decision-making: mastering the skills to identify performance bottlenecks, select the right tools, configure Spark effectively, and strike the ideal balance between processing performance and infrastructure resource consumption and cost.
Read more...