Get in Touch

Course Outline

Introduction

This module offers a comprehensive overview of when to apply 'machine learning', key considerations, and its broader implications, including advantages and limitations. It covers datatypes (structured/unstructured/static/streamed), data validity and volume, the distinction between data-driven and user-driven analytics, the comparison between statistical and machine learning models, the challenges of unsupervised learning, the bias-variance trade-off, iterative evaluation, cross-validation strategies, and the paradigms of supervised, unsupervised, and reinforcement learning.

MAJOR TOPICS

1. Exploring Naive Bayes

  • Fundamentals of Bayesian methods
  • Probability
  • Joint probability
  • Conditional probability using Bayes' theorem
  • The Naive Bayes algorithm
  • Classification with Naive Bayes
  • The Laplace estimator
  • Integrating numeric features into Naive Bayes

2. Exploring Decision Trees

  • Divide and conquer strategies
  • The C5.0 decision tree algorithm
  • Selecting optimal splits
  • Pruning decision trees

3. Exploring Neural Networks

  • Transitioning from biological to artificial neurons
  • Activation functions
  • Network architecture
  • Determining the number of layers
  • Information flow direction
  • Node distribution per layer
  • Training neural networks via backpropagation
  • Deep Learning

4. Exploring Support Vector Machines

  • Classification using hyperplanes
  • Maximizing margins
  • Handling linearly separable data
  • Addressing non-linearly separable data
  • Utilizing kernels for non-linear spaces

5. Exploring Clustering

  • Clustering as a machine learning objective
  • The k-means clustering algorithm
  • Assigning and updating clusters based on distance
  • Determining the optimal number of clusters

6. Assessing classification performance

  • Processing classification prediction data
  • In-depth analysis of confusion matrices
  • Evaluating performance with confusion matrices
  • Performance metrics beyond accuracy
  • The Kappa statistic
  • Sensitivity and specificity
  • Precision and recall
  • The F-measure
  • Visualizing performance compromises
  • ROC curves
  • Predicting future model performance
  • The holdout method
  • Cross-validation
  • Bootstrap sampling

7. Optimizing standard models for enhanced results

  • Leveraging caret for automated parameter tuning
  • Developing a basic tuned model
  • Customizing the tuning workflow
  • Enhancing performance through meta-learning
  • Comprehending ensembles
  • Bagging
  • Boosting
  • Random forests
  • Training random forests
  • Assessing random forest performance

MINOR TOPICS

8. Classification via nearest neighbors

  • The kNN algorithm
  • Distance calculation
  • Selecting an appropriate k
  • Data preparation for kNN
  • The lazy nature of the kNN algorithm

9. Classification rules

  • Separate and conquer
  • The One Rule algorithm
  • The RIPPER algorithm
  • Deriving rules from decision trees

10. Regression

  • Simple linear regression
  • Ordinary least squares estimation
  • Correlations
  • Multiple linear regression

11. Regression trees and model trees

  • Integrating regression into trees

12. Association rules

  • The Apriori algorithm for association rule mining
  • Measuring rule relevance via support and confidence
  • Generating rule sets using the Apriori principle

Extras

  • Spark/PySpark/MLlib and Multi-armed bandits

Requirements

Python proficiency

 21 Hours

Number of participants


Price per participant

Testimonials (7)

Upcoming Courses

Related Categories