Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of Cloud Operations on AWS
- Defining operational roles and responsibilities within the cloud landscape
- Understanding AWS account structures, organizations, and multi-account strategies
- Exploring core operational services including CloudWatch, CloudTrail, and AWS Config
Infrastructure as Code and Provisioning
- Core principles of IaC and the benefits of immutable infrastructure
- Executing provisioning tasks with Terraform and AWS CloudFormation
- Managing state files, modules, and environment promotion workflows
CI/CD and Deployment Strategies
- Implementing blue/green, canary, and rolling deployment models
- Automating rollback procedures, health checks, and release validation
Monitoring, Observability, and Alerting
- Handling metrics, logs, and traces: shipping, storage, and analysis techniques
- Utilizing CloudWatch, X-Ray, and third-party observability tools effectively
- Establishing SLOs/SLIs, alerting policies, and on-call operational practices
Security Operations and Identity Management
- Applying IAM best practices, enforcing least privilege, and managing cross-account access
- Implementing secrets management, KMS, and secure parameter stores
- Enhancing operational security through patching strategies, vulnerability scanning, and audit trail maintenance
Resilience, Backup, and Disaster Recovery
- Designing systems for fault tolerance and high availability
- Developing backup strategies, automating snapshots, and defining restore procedures
- Formulating disaster recovery plans and creating comprehensive runbooks
Cost Optimization and Governance
- Achieving cost efficiency through rightsizing, reserved instances/savings plans, and budgeting controls
- Enforcing governance via policies, guardrails, and compliance automation
Containers, Serverless, and Runtime Operations
- Addressing operational considerations for ECS, EKS, and Lambda
- Managing service discovery, autoscaling, and resource limits
- Logging, tracing, and debugging techniques for containerized workloads
Incident Response, Playbooks, and Chaos Engineering
- Executing runbook-driven incident response and conducting postmortem analyses
- Automating remediation processes and implementing self-healing patterns
- Introduction to chaos experiments for validating system resilience
Hands-on Workshop: Operate a Sample Workload
- Deploying a sample application using IaC and a CI/CD pipeline
- Setting up monitoring, alerts, and automated remediation scripts
- Simulating incidents to practice runbook-based response strategies
Summary and Next Steps
Requirements
- Foundational knowledge of cloud computing concepts and networking principles
- Proficiency with the Linux command line and basic scripting capabilities
- Practical experience with source control systems (Git) and an understanding of basic CI/CD workflows
Target Audience
- Cloud operations engineers
- Site Reliability Engineers (SREs) and platform engineers
- DevOps engineers and technical team leads
21 Hours
Testimonials (1)
I've find out new interesting things about Lambda and Serverless