Get in Touch
 Duration 14 hours

Course Outline

Introduction to AIOps with Open Source Solutions

  • Key concepts and advantages of AIOps
  • The role of Prometheus and Grafana within the observability architecture
  • Positioning ML in AIOps: distinguishing between predictive and reactive analytics

Configuring Prometheus and Grafana

  • Deployment and configuration of Prometheus for time series data acquisition
  • Designing Grafana dashboards utilizing live metrics
  • In-depth exploration of exporters, relabeling mechanisms, and service discovery

Data Preparation for Machine Learning

  • Extraction and transformation of Prometheus metrics
  • Curating datasets suitable for anomaly detection and forecasting tasks
  • Utilizing Grafana’s built-in transformations or Python-based data pipelines

Machine Learning for Anomaly Detection

  • Implementation of basic ML models for outlier identification (e.g., Isolation Forest, One-Class SVM)
  • Model training and evaluation processes on time series data
  • Visualization of detected anomalies within Grafana interfaces

Metrics Forecasting via Machine Learning

  • Construction of foundational forecasting models (including ARIMA, Prophet, and an introduction to LSTM)
  • Prediction of system load and resource consumption patterns
  • Leveraging forecasts for proactive alerting and scaling strategies

Integrating ML with Alerting and Automation

  • Formulation of alert rules based on ML outputs or defined thresholds
  • Configuration of Alertmanager and notification routing paths
  • Execution of scripts or automation workflows in response to detected anomalies

Scaling and Operationalizing AIOps

  • Integration with external observability platforms (e.g., ELK stack, Moogsoft, Dynatrace)
  • Operational deployment of ML models within observability pipelines
  • Best practices for managing AIOps at scale

Summary and Future Steps

Requirements

  • A solid grasp of system monitoring and observability principles
  • Practical experience with Grafana or Prometheus
  • Proficiency in Python and foundational knowledge of machine learning concepts

Target Audience

  • Observability engineers
  • Infrastructure and DevOps teams
  • Monitoring platform architects and Site Reliability Engineers (SREs)

Number of participants


Price per participant

Upcoming Courses

Related Categories