Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps with Open Source Tools
- Overview of AIOps principles and advantages.
- The role of Prometheus and Grafana within the observability stack.
- The place of ML in AIOps: contrasting predictive and reactive analytics.
Setting Up Prometheus and Grafana
- Installation and configuration of Prometheus for time series data collection.
- Building dashboards in Grafana utilizing real-time metrics.
- Investigating exporters, relabeling, and service discovery mechanisms.
Data Preprocessing for ML
- Extraction and transformation of Prometheus metrics.
- Preparing datasets suitable for anomaly detection and forecasting models.
- Utilizing Grafana’s transformation features or Python-based pipelines.
Applying Machine Learning for Anomaly Detection
- Foundational ML models for outlier detection, such as Isolation Forest and One-Class SVM.
- Training and assessing models using time series data.
- Visualizing detected anomalies within Grafana dashboards.
Forecasting Metrics with ML
- Developing simple forecasting models, including ARIMA, Prophet, and introductory LSTM concepts.
- Anticipating system load or resource utilization trends.
- Leveraging predictions for early alerting and scaling decisions.
Integrating ML with Alerting and Automation
- Establishing alert rules based on ML outputs or defined thresholds.
- Configuring Alertmanager and notification routing strategies.
- Initiating scripts or automation workflows upon anomaly detection.
Scaling and Operationalizing AIOps
- Integration with external observability platforms like the ELK stack, Moogsoft, or Dynatrace.
- Operationalizing ML models within observability pipelines.
- Best practices for implementing AIOps at scale.
Summary and Next Steps
Requirements
- A solid grasp of system monitoring and observability concepts.
- Practical experience with Grafana or Prometheus.
- Proficiency in Python and an understanding of fundamental machine learning principles.
Target Audience
- Observability engineers.
- Infrastructure and DevOps teams.
- Monitoring platform architects and Site Reliability Engineers (SREs).