Get in Touch

Course Outline

Introduction

  • Overview of Spark and Hadoop features and architecture
  • Concepts and insights into big data
  • Fundamentals of Python programming

Getting Started

  • Setting up Python, Spark, and Hadoop
  • Exploring data structures in Python
  • Understanding the PySpark API
  • Overview of HDFS and MapReduce

Integrating Spark and Hadoop with Python

  • Implementing Spark RDD in Python
  • Processing data using MapReduce
  • Creating distributed datasets in HDFS

Machine Learning with Spark MLlib

Processing Big Data with Spark Streaming

Working with Recommender Systems

Working with Kafka, Sqoop, Kafka, and Flume

Apache Mahout with Spark and Hadoop

Troubleshooting

Summary and Next Steps

Requirements

  • Familiarity with Spark and Hadoop
  • Proficiency in Python programming

Target Audience

  • Data scientists
  • Software developers
 21 Hours

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories