Get in Touch
 Duration 21 hours

Course Outline

Comprehensive training outline

  1. Introduction to NLP
    • Concepts of NLP
    • NLP frameworks
    • Commercial use cases for NLP
    • Web data scraping
    • Utilising various APIs for text data retrieval
    • Managing text corpora: saving content and relevant metadata
    • Benefits of Python and an NLTK quick start
  2. Practical Understanding of a Corpus and Dataset
    • The necessity of a corpus
    • Corpus analysis techniques
    • Categorisation of data attributes
    • File formats for corpora
    • Preparing datasets for NLP applications
  3. Understanding the Structure of a Sentence
    • Core NLP components
    • Natural language understanding
    • Morphological analysis: stems, words, tokens, and speech tags
    • Syntactic analysis
    • Semantic analysis
    • Addressing ambiguity
  4. Text Data Preprocessing
    • Raw text corpus
      • Sentence tokenization
      • Stemming raw text
      • Lemmatisation of raw text
      • Stop word removal
    • Raw sentence corpus
      • Word tokenization
      • Word lemmatisation
    • Constructing Term-Document and Document-Term matrices
    • Tokenizing text into n-grams and sentences
    • Customised and practical preprocessing methods
  5. Analyzing Text Data
    • Fundamental NLP features
      • Parsers and parsing
      • POS tagging and taggers
      • Named entity recognition
      • N-grams
      • Bag of words
    • Statistical NLP features
      • Linear algebra concepts for NLP
      • Probabilistic theory for NLP
      • TF-IDF
      • Vectorization
      • Encoders and Decoders
      • Normalization
      • Probabilistic Models
    • Advanced feature engineering in NLP
      • Word2vec fundamentals
      • Word2vec model components
      • Logic behind the word2vec model
      • Extending the word2vec concept
      • Applications of the word2vec model
    • Case study: Applying bag of words for automatic text summarization using simplified and true Luhn's algorithms
  6. Document Clustering, Classification, and Topic Modeling
    • Document clustering and pattern mining (including hierarchical clustering and k-means)
    • Comparing and classifying documents using TFIDF, Jaccard, and cosine distance metrics
    • Document classification using Naïve Bayes and Maximum Entropy
  7. Identifying Important Text Elements
    • Dimensionality reduction: PCA, SVD, and Non-negative Matrix Factorization
    • Topic modelling and information retrieval via Latent Semantic Analysis
  8. Entity Extraction, Sentiment Analysis, and Advanced Topic Modeling
    • Positive vs. negative sentiment: measuring intensity
    • Item Response Theory
    • Part-of-speech tagging applications: identifying people, places, and organizations
    • Advanced topic modelling: Latent Dirichlet Allocation
  9. Case Studies
    • Extracting insights from unstructured user reviews
    • Sentiment classification and visualisation of product review data
    • Analysing search logs for usage patterns
    • Text classification
    • Topic modelling

Requirements

Familiarity with NLP principles and an understanding of AI applications in business contexts.

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories