Get in Touch

Course Outline

Core Concepts of Cloud Operations on AWS

  • Defining operational roles and responsibilities in a cloud context.
  • Structuring AWS accounts, organizations, and multi-account strategies.
  • Key operational services: CloudWatch, CloudTrail, and AWS Config.

Infrastructure as Code (IaC) and Provisioning

  • Understanding IaC principles and the concept of immutable infrastructure.
  • Executing provisioning tasks with Terraform and AWS CloudFormation.
  • Managing state, modules, and promoting environments.

CI/CD and Deployment Methodologies

  • Architecting CI/CD pipelines for cloud-native applications.
  • Implementing blue/green, canary, and rolling deployment strategies.
  • Automating rollbacks, health checks, and release validation processes.

Monitoring, Observability, and Alerting

  • Handling metrics, logs, and traces: ingestion, storage, and analysis.
  • Utilizing CloudWatch, X-Ray, and third-party observability solutions.
  • Establishing SLOs/SLIs, alerting policies, and on-call workflows.

Security Operations and Identity Management

  • Applying IAM best practices, least privilege principles, and cross-account access controls.
  • Managing secrets, KMS, and secure parameter stores.
  • Operational security measures: patching strategies, vulnerability scanning, and audit trails.

Resilience, Backup, and Disaster Recovery

  • Designing systems for fault tolerance and high availability.
  • Developing backup strategies, automating snapshots, and defining restore procedures.
  • Planning disaster recovery and creating comprehensive runbooks.

Cost Optimization and Governance

  • Enhancing cost visibility through billing analysis, tagging, and allocation strategies.
  • Rightsizing resources, leveraging reserved instances/savings plans, and enforcing budget controls.
  • Implementing governance through policies, guardrails, and compliance automation.

Containers, Serverless, and Runtime Operations

  • Addressing operational needs for ECS, EKS, and Lambda.
  • Managing service discovery, autoscaling, and resource constraints.
  • Logging, tracing, and debugging containerized workloads.

Incident Response, Playbooks, and Chaos Engineering

  • Driving incident response through runbooks and conducting postmortems.
  • Automating remediation and implementing self-healing patterns.
  • Introduction to chaos engineering experiments for resilience validation.

Practical Workshop: Managing a Sample Workload

  • Deploying a sample application via IaC and a CI/CD pipeline.
  • Configuring monitoring, alerts, and automated remediation scripts.
  • Simulating incidents and practicing runbook-based response actions.

Conclusion and Path Forward

Requirements

  • Foundational knowledge of cloud computing concepts and networking principles.
  • Proficiency with the Linux command line and scripting languages.
  • Practical experience with source control (Git) and fundamental CI/CD concepts.

Target Audience

  • Cloud operations engineers.
  • SREs and platform engineers.
  • DevOps engineers and technical team leads.
 21 Hours

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories