Get in Touch

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • Brief Overview of Python and Scala

Core Concepts (Theoretical Foundation):

  • Spark Architecture
  • RDD (Resilient Distributed Dataset)
  • Transformations vs. Actions
  • Stages, Tasks, and Dependencies

Databricks Environment: Fundamental Exploration (Hands-On Workshop):

  • Practical exercises using the RDD API
  • Implementing basic action and transformation functions
  • Working with PairRDD
  • Join Operations
  • Caching Strategies
  • Practical exercises using the DataFrame API
  • SparkSQL Integration
  • DataFrame Operations: select, filter, group, sort
  • UDF (User Defined Functions)
  • Exploring the Dataset API
  • Stream Processing

AWS Environment: Deployment Strategies (Hands-On Workshop):

  • Foundations of AWS Glue
  • Comparing AWS EMR and AWS Glue
  • Implementing sample jobs in both environments
  • Evaluating advantages and trade-offs

Additional Topics:

  • Introduction to Orchestration with Apache Airflow

Requirements

Programming proficiency (Python and Scala preferred)

Fundamental knowledge of SQL

 21 Hours

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories