Get in Touch

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • Brief overview of Python and Scala

Core Concepts (Theoretical):

  • System Architecture
  • RDD (Resilient Distributed Dataset)
  • Transformations and Actions
  • Stages, Tasks, and Dependencies

Practical Basics in Databricks (Hands-on Workshop):

  • Practical exercises with the RDD API
  • Implementing basic action and transformation functions
  • Utilizing PairRDDs
  • Executing Joins
  • Implementing caching strategies
  • Practical exercises with the DataFrame API
  • Working with SparkSQL
  • DataFrame operations: select, filter, group, and sort
  • Creating UDFs (User Defined Functions)
  • Exploring the DataSet API
  • Implementing Streaming

Deployment Strategies in AWS (Hands-on Workshop):

  • Fundamentals of AWS Glue
  • Analyzing the differences between AWS EMR and AWS Glue
  • Running sample jobs in both environments
  • Evaluating the advantages and disadvantages

Additional Topics:

  • Overview of Apache Airflow orchestration

Requirements

Programming proficiency (ideally in Python or Scala)

Fundamental knowledge of SQL

 21 Hours

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories