Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction:
- Apache Spark within the Hadoop Ecosystem
- Brief Overview of Python and Scala
Core Concepts (Theoretical Foundation):
- Spark Architecture
- RDD (Resilient Distributed Dataset)
- Transformations vs. Actions
- Stages, Tasks, and Dependencies
Databricks Environment: Fundamental Exploration (Hands-On Workshop):
- Practical exercises using the RDD API
- Implementing basic action and transformation functions
- Working with PairRDD
- Join Operations
- Caching Strategies
- Practical exercises using the DataFrame API
- SparkSQL Integration
- DataFrame Operations: select, filter, group, sort
- UDF (User Defined Functions)
- Exploring the Dataset API
- Stream Processing
AWS Environment: Deployment Strategies (Hands-On Workshop):
- Foundations of AWS Glue
- Comparing AWS EMR and AWS Glue
- Implementing sample jobs in both environments
- Evaluating advantages and trade-offs
Additional Topics:
- Introduction to Orchestration with Apache Airflow
Requirements
Programming proficiency (Python and Scala preferred)
Fundamental knowledge of SQL
21 Hours
Testimonials (3)
Having hands on session / assignments
Poornima Chenthamarakshan - Intelligent Medical Objects
Course - Apache Spark in the Cloud
1. Right balance between high level concepts and technical details. 2. Andras is very knowledgeable about his teaching. 3. Exercise
Steven Wu - Intelligent Medical Objects
Course - Apache Spark in the Cloud
Get to learn spark streaming , databricks and aws redshift