Amazon Web Services Cloud Operations Training Course
| Course code | SD-IT-017 |
|---|---|
| Duration | 5 days |
| Level | Foundation to Intermediate |
| Category | Information Technology |
| Delivery | Classroom or live online |
| Language | English |
| Certificate | Certificate of completion |
Course overview
Cloud operations teams must keep AWS workloads available, secure, observable and cost-controlled while responding to incidents without creating further risk. Many administrators can launch services in the AWS Management Console but lack a repeatable operating model for access control, monitoring, patching, backups, change management and recovery. This course gives participants the practical operating discipline needed to run production AWS environments and communicate operational status, risks and actions clearly to technical and business stakeholders.
Participants work across the core services and practices used in day-to-day AWS operations: IAM identity and permission design, Amazon EC2 administration, Amazon VPC connectivity, Amazon CloudWatch monitoring and alarms, AWS CloudTrail audit trails, AWS Systems Manager fleet operations, AWS Backup, AWS Config and cost-management controls. They learn to investigate alerts, use the AWS CLI for routine administration, define operational runbooks, apply tags, assess resource configuration, automate common remediation tasks and plan incident escalation.
Instructor-led demonstrations are followed by guided labs in an AWS training environment, operational case studies and team-based troubleshooting exercises. Each participant builds an AWS Operations Runbook Pack containing a monitoring dashboard specification, alert-response workflow, instance maintenance procedure, backup and recovery checklist, tagging standard and cost-control action plan. This provides a practical set of templates that can be adapted to the participant's own AWS account structure and operating procedures.
The course is suited to professionals moving into AWS support or platform operations, as well as experienced infrastructure staff who need a structured approach to managing AWS services. Managers gain staff who can translate cloud operational requirements into observable controls, documented procedures and measurable service improvements.
Course objectives
By the end of this course, participants will be able to:
- Configure IAM users, groups, roles and least-privilege policies for operational access
- Operate Amazon EC2 instances using launch templates, security groups, key management and instance lifecycle controls
- Create Amazon CloudWatch metrics, logs, alarms and dashboards for workload monitoring
- Investigate operational events using AWS CloudTrail records, CloudWatch Logs and resource metadata
- Use AWS Systems Manager to inventory, patch and execute controlled commands across managed instances
- Design backup, retention and recovery procedures using AWS Backup and Amazon EBS snapshots
- Apply resource tags, AWS Config rules and cost-allocation methods to improve governance
- Produce an AWS Operations Runbook Pack with alert, incident, maintenance and recovery procedures
Benefits of attending
For you
- Build evidence-based confidence for responding to AWS alerts, failed instances and access issues
- Gain practical experience with CloudWatch, CloudTrail, Systems Manager and AWS Backup in realistic scenarios
- Create reusable runbook and checklist templates for an AWS operations portfolio
- Improve credibility in conversations about availability, security controls, recovery objectives and cloud cost ownership
- Prepare for operational responsibilities in cloud support, infrastructure administration and platform engineering roles
For your organisation
- Reduce avoidable service disruption through documented monitoring thresholds, escalation paths and response procedures
- Improve audit readiness by strengthening IAM controls, CloudTrail review practices and configuration visibility
- Standardise EC2 patching and maintenance using AWS Systems Manager rather than inconsistent manual administration
- Increase recovery readiness through defined backup schedules, retention decisions and restoration testing steps
- Improve cloud cost accountability through tagging standards, cost allocation and operational review actions
Target competencies
Who should attend
- Cloud Operations Engineers — who need repeatable methods for monitoring, maintaining and recovering AWS workloads
- Systems Administrators — who are moving Windows or Linux server operations into Amazon EC2
- IT Support Analysts — who must diagnose AWS alerts and escalate incidents with useful evidence
- DevOps Engineers — who need stronger operational controls around deployed AWS services
- Infrastructure Engineers — who are responsible for secure connectivity, availability and lifecycle management in AWS
- IT Operations Managers — who need to establish measurable AWS support processes and team runbooks
Requirements and prerequisites
Participants should be comfortable using a web browser, command line and basic file management, and should understand fundamental IT concepts including IP addressing, DNS, virtual machines, operating systems and user permissions. Familiarity with Linux or Windows server administration is helpful because exercises use Amazon EC2 instances. Prior AWS experience is not required; the course begins with AWS accounts, regions, IAM and the Management Console before progressing to operational tools. Participants do not need programming, infrastructure-as-code experience, cloud certification or prior use of the AWS CLI, although willingness to work through guided command-line tasks is expected.
Training methodology
The five days combine instructor-led explanation with live AWS demonstrations and guided labs in a controlled training account. Participants configure services through the AWS Management Console and AWS CLI, then interpret the operational evidence those tools produce. Case studies simulate common conditions such as an unreachable EC2 instance, an unexpected cost increase, a failed backup and an unauthorised access event. Small groups compare response decisions against service impact and escalation criteria. The final workshop turns completed lab work into an AWS Operations Runbook Pack and a 30-day application plan.
Course outline
Day 1: AWS operational foundations and access control
- AWS global infrastructure, Regions, Availability Zones and shared responsibility
- AWS account structure, root-user protection and operational account hygiene
- IAM users, groups, roles and federation concepts
- Least-privilege policy construction with IAM policy evaluation logic
- Multi-factor authentication and credential lifecycle management
- AWS Management Console navigation and Resource Groups
- AWS CLI configuration, profiles and secure credential handling
Workshop: Participants configure a least-privilege operations role, MFA controls and AWS CLI profile, then document an access-control procedure for their runbook.
Day 2: Compute, networking and routine service operations
- Amazon EC2 instance families, AMIs, launch templates and lifecycle states
- Amazon EBS volumes, snapshots, encryption and performance considerations
- Amazon VPC architecture, subnets, route tables and internet gateways
- Security groups and network ACLs for operational troubleshooting
- Elastic Load Balancing health checks and target status investigation
- Amazon Route 53 records and DNS fault isolation
- Resource tagging standards and ownership metadata
Workshop: Participants diagnose an inaccessible EC2-hosted application by tracing security group, route table, load balancer and DNS configuration, then record the remediation steps.
Day 3: Monitoring, logging and incident investigation
- Amazon CloudWatch metrics, namespaces, dimensions and statistics
- CloudWatch alarms, thresholds, missing-data treatment and notification actions
- CloudWatch dashboards for service health and operational reporting
- CloudWatch Logs groups, streams, retention and Logs Insights queries
- AWS CloudTrail event history, trails and management event analysis
- Amazon EventBridge event patterns for operational automation
- Incident triage workflows, severity classification and escalation evidence
Workshop: Participants build a CloudWatch dashboard and alarm set for an EC2 application, investigate a simulated service issue and produce an incident timeline with escalation evidence.
Day 4: Maintenance, governance, backup and cost control
- AWS Systems Manager managed nodes, inventory and session access
- Systems Manager Run Command and Automation documents
- Patch Manager baselines, maintenance windows and patch compliance
- AWS Backup plans, vaults, lifecycle rules and restore testing
- AWS Config resource recording, configuration history and managed rules
- AWS Cost Explorer, cost allocation tags and budget alerts
- Operational change control, maintenance communication and rollback planning
Workshop: Participants create a maintenance and backup plan for a server fleet, define compliance checks and prepare a cost-review action list for tagged resources.
Day 5: Operational resilience and runbook implementation
- Availability design across Availability Zones and operational failure domains
- Recovery point objective and recovery time objective definition
- Backup restoration validation and post-recovery checks
- Operational runbook structure, decision points and ownership fields
- Service-level indicators, service-level objectives and operational reporting
- Root-cause analysis using timelines, contributing factors and corrective actions
- Thirty-day AWS operations improvement planning
Workshop: Participants complete and present an AWS Operations Runbook Pack for a simulated production environment, including monitoring, incident, maintenance, backup and improvement actions.
Tools & standards covered
AWS Management Console, AWS Command Line Interface, Amazon CloudWatch, AWS Systems Manager
A typical training day
| 08:30 – 10:30 | First session |
| 10:30 – 10:45 | Refreshment break |
| 10:45 – 12:30 | Second session |
| 12:30 – 13:30 | Lunch and networking |
| 13:30 – 15:00 | Third session |
| 15:00 – 15:15 | Refreshment break |
| 15:15 – 16:30 | Workshop and daily review |
Live online deliveries follow the same structure in the East Africa Time zone, with shorter screen blocks and longer breaks.
What the fee includes
- Instruction by a practitioner facilitator
- Full course workbook and materials
- Exercise files, templates and case studies
- Certificate of completion
- Refreshments and lunch (classroom deliveries)
- Post-course application plan
- Facilitator follow-up on request
- Group rates from five participants
How you can take this course
Classroom
Scheduled sessions in Nairobi, Mombasa, Kigali, Dar es Salaam, Dubai and Cape Town.
Live online
The same facilitator and materials, delivered live for distributed teams and individuals.
In-house
Delivered privately for your team, at your offices or a venue of your choice, tailored to your context. Request a proposal.
Certification
Participants who complete the full five days receive the Skillset Development Certificate of Completion, stating the course title, course code, dates and delivery format — suitable for professional-development records and employer reimbursement.
Frequently asked questions
Upcoming sessions
New dates are being scheduled. Ask us about the next session or an in-house delivery for your team.
Ask about datesGroup of 5+?
Request in-house delivery or group rates →Related courses in Information Technology
Advanced IT Strategy and Governance Training Course
Technology leaders are expected to defend investment decisions, govern risk, and show how platforms, data, sourcing, and delivery programmes…
Technology Project Delivery for IT Project Managers Training Course
IT project managers are expected to deliver technology change while balancing uncertain requirements, constrained technical capacity, vendor…
IT Asset Lifecycle Management for Asset Managers Training Course
IT asset managers are expected to produce reliable answers to questions that often sit across disconnected systems: what technology the orga…
IT Controls Testing for Internal Auditors Training Course
Internal auditors are increasingly expected to provide assurance over access management, change control, IT operations, cloud services and a…