Datadog Cloud Monitoring and Observability Training Course
| Course code | SD-CC-014 |
|---|---|
| Duration | 5 days |
| Level | Intermediate |
| Category | Cloud Computing |
| Delivery | Classroom or live online |
| Language | English |
| Certificate | Certificate of completion |
Course overview
Cloud teams often have metrics in one dashboard, logs in another, incomplete service ownership, and alerts that fire without identifying the customer impact or responsible dependency. This makes incident response slower, drives alert fatigue, and leaves engineering leaders unable to distinguish genuine reliability risks from normal operational noise. This five-day Datadog Cloud Monitoring and Observability Training Course equips practitioners to build an observability approach that connects infrastructure health, application performance, user experience, security signals, and business-facing service objectives.
Participants configure and use Datadog across Infrastructure Monitoring, Log Management, APM, Real User Monitoring, Synthetic Monitoring, dashboards, monitors, Service Catalog, and incident workflows. They learn to instrument services with OpenTelemetry, correlate traces with logs and host or container metrics, define service-level objectives, tune monitor thresholds, manage tags, investigate incidents with Watchdog and Analytics views, and design role-appropriate dashboards. The course also addresses Kubernetes and cloud monitoring patterns, cost-aware telemetry collection, and configuration-as-code practices using Terraform.
Teaching combines instructor-led demonstrations with guided labs in a Datadog sandbox, using realistic cloud application failures including latency regressions, Kubernetes resource contention, failed deployments, and third-party dependency outages. Participants complete an end-of-course observability implementation plan and dashboard-and-monitor pack for a representative service, including a tag model, service health dashboard, monitor specifications, SLOs, escalation routes, and investigation runbook. This gives both the attendee and their manager a practical artefact that can be adapted for production adoption.
The course is designed for intermediate engineers and operations professionals who already work with cloud-hosted applications and need to standardise how their teams detect, investigate, and prevent service degradation using Datadog.
Course objectives
By the end of this course, participants will be able to:
- Configure Datadog integrations and agents to collect cloud, host, container, and application telemetry
- Apply a consistent tag taxonomy to support service ownership, environment filtering, and cost allocation
- Correlate metrics, logs, traces, and user-session data during Datadog investigations
- Build APM service maps, trace queries, and latency analyses to isolate application bottlenecks
- Create actionable monitors with threshold logic, notification routing, recovery conditions, and message templates
- Define SLOs and error budgets that measure service reliability against customer-facing targets
- Design role-based Datadog dashboards for service health, Kubernetes operations, and incident command
- Produce an observability implementation plan with monitor specifications, runbooks, and Terraform configuration patterns
Benefits of attending
For you
- Gain practical evidence of Datadog capability through a completed dashboard, monitor, SLO, and investigation pack
- Improve incident diagnosis by tracing a user-facing symptom through infrastructure, logs, and distributed services
- Develop the ability to challenge noisy alerts and replace them with service-impact-focused monitor logic
- Build credibility for SRE, DevOps, cloud operations, and platform engineering responsibilities
- Learn to translate technical telemetry into reliability measures and risk reports for engineering leadership
For your organisation
- Reduce alert fatigue through better monitor thresholds, routing rules, recovery settings, and ownership tags
- Shorten incident investigation by standardising correlation between logs, metrics, traces, and service dependencies
- Improve reliability governance with defined SLOs, error budgets, and service health dashboards
- Establish reusable Datadog implementation patterns for cloud accounts, Kubernetes clusters, and application teams
- Control observability spend through purposeful telemetry collection, tag governance, and retention-aware design
Target competencies
Who should attend
- Site Reliability Engineers — who need to reduce mean time to detect and diagnose cloud service failures
- Cloud Engineers — who operate AWS, Azure, or Google Cloud workloads and need unified operational visibility
- DevOps Engineers — who build deployment and monitoring practices across application delivery pipelines
- Platform Engineers — who provide reusable observability standards and self-service monitoring for product teams
- Application Support Engineers — who investigate production incidents across logs, traces, and infrastructure signals
- Engineering Managers — who need service health measures, ownership visibility, and defensible reliability targets
Requirements and prerequisites
Participants should be comfortable navigating a Linux command line, reading application logs, and working with basic cloud concepts such as virtual machines, containers, load balancers, managed databases, IAM roles, and environments. Experience operating applications on AWS, Microsoft Azure, or Google Cloud is useful, as is familiarity with Kubernetes concepts including pods, deployments, namespaces, and resource limits. Participants should understand basic HTTP behaviour and application response times. Prior Datadog experience is not required, and attendees do not need to be software developers, Terraform specialists, or OpenTelemetry experts before attending.
Training methodology
The course is delivered through short instructor-led technical briefings followed by guided Datadog labs in a realistic cloud application environment. Participants configure integrations, query telemetry, build dashboards, create monitors, and investigate staged failures using metrics, logs, traces, RUM, and Synthetic Monitoring. Small-group case work focuses on alert quality, service ownership, and SLO decisions. Daily review sessions connect configuration choices to operational outcomes. On day five, each participant assembles an observability plan and receives structured instructor feedback on its implementation priorities.
Course outline
Day 1: Datadog foundations and telemetry architecture
- Datadog organisation structure, roles, teams, and access controls
- Datadog Agent deployment models for hosts, containers, and cloud environments
- Metrics types, metric submission, retention, and query fundamentals
- Tagging strategy for service, environment, team, region, and cost attribution
- Cloud integration patterns for AWS, Microsoft Azure, and Google Cloud
- Infrastructure Monitoring views, host maps, process monitoring, and network context
- Service Catalog ownership metadata and dependency documentation
Workshop: Participants design a tag taxonomy and configure infrastructure visibility for a multi-service cloud application, producing an ownership and telemetry coverage map.
Day 2: Logs, APM, and distributed service diagnosis
- Log collection pipelines, parsing rules, facets, and sensitive-data handling
- Log Explorer queries, patterns, analytics measures, and saved views
- APM service instrumentation and Unified Service Tagging
- OpenTelemetry concepts, semantic conventions, and Datadog trace ingestion
- Trace search, flame graphs, span tags, and latency decomposition
- Service Map analysis for upstream, downstream, and external dependencies
- Log-trace-metric correlation workflows for incident investigation
Workshop: Participants investigate a simulated checkout latency incident, producing a root-cause evidence trail from a user request through traces, logs, and database metrics.
Day 3: Cloud, Kubernetes, and digital experience monitoring
- Kubernetes Cluster Agent, kube-state metrics, and container tagging
- Kubernetes dashboards for pods, deployments, nodes, namespaces, and resource requests
- Container resource saturation, restart analysis, and workload health signals
- Cloud service monitoring for load balancers, databases, queues, and serverless functions
- Real User Monitoring application setup, session views, and frontend error analysis
- Synthetic API and browser tests for critical user journeys
- Network Performance Monitoring and dependency traffic analysis
Workshop: Participants build a Kubernetes and customer-journey dashboard, then identify the operational impact of a staged pod capacity and API failure.
Day 4: Alert engineering, SLOs, and incident response
- Monitor types, evaluation windows, thresholds, anomaly detection, and forecast alerts
- Composite monitors, dependency-aware alert logic, and monitor grouping
- Notification templates, escalation routing, recovery messages, and runbook links
- Alert quality review using false-positive, duplicate, and actionable-signal criteria
- Service-level indicators, SLO target selection, and error budget calculations
- SLO status dashboards, burn-rate monitoring, and reliability reporting
- Incident Management workflows, timelines, roles, and post-incident evidence capture
Workshop: Participants replace a noisy alert set with a monitored service objective, routed monitors, and an incident-response runbook for a payment service.
Day 5: Operationalising Datadog at scale
- Dashboard design for executives, service owners, on-call teams, and platform operations
- Notebook investigations, shared diagnostic context, and operational reporting
- Datadog Watchdog insights and automated anomaly investigation
- Terraform provider patterns for monitors, dashboards, SLOs, and notification configuration
- Telemetry governance, naming standards, data access, and audit considerations
- Observability cost management through metric volume, log indexes, and retention choices
- Phased Datadog adoption roadmap, success measures, and operating model decisions
Workshop: Participants present an observability implementation plan and deliver a service dashboard, monitor set, SLO definition, tag standard, and prioritised 90-day adoption roadmap.
Tools & standards covered
Datadog, OpenTelemetry, Kubernetes, Terraform
A typical training day
| 08:30 – 10:30 | First session |
| 10:30 – 10:45 | Refreshment break |
| 10:45 – 12:30 | Second session |
| 12:30 – 13:30 | Lunch and networking |
| 13:30 – 15:00 | Third session |
| 15:00 – 15:15 | Refreshment break |
| 15:15 – 16:30 | Workshop and daily review |
Live online deliveries follow the same structure in the East Africa Time zone, with shorter screen blocks and longer breaks.
What the fee includes
- Instruction by a practitioner facilitator
- Full course workbook and materials
- Exercise files, templates and case studies
- Certificate of completion
- Refreshments and lunch (classroom deliveries)
- Post-course application plan
- Facilitator follow-up on request
- Group rates from five participants
How you can take this course
Classroom
Scheduled sessions in Nairobi, Mombasa, Kigali, Dar es Salaam, Dubai and Cape Town.
Live online
The same facilitator and materials, delivered live for distributed teams and individuals.
In-house
Delivered privately for your team, at your offices or a venue of your choice, tailored to your context. Request a proposal.
Certification
Participants who complete the full five days receive the Skillset Development Certificate of Completion, stating the course title, course code, dates and delivery format — suitable for professional-development records and employer reimbursement.
Frequently asked questions
Upcoming sessions
New dates are being scheduled. Ask us about the next session or an in-house delivery for your team.
Ask about datesGroup of 5+?
Request in-house delivery or group rates →Related courses in Cloud Computing
Cloud Services for Public Sector Digital Teams Training Course
Public-sector digital teams must modernise citizen-facing services while protecting sensitive data, sustaining continuity, meeting procureme…
AWS Well-Architected Framework Implementation Training Course
AWS workloads often grow faster than their governance, documentation and operational controls. Teams inherit accounts with inconsistent tagg…
Pulumi Cloud Infrastructure Automation Training Course
Cloud teams often inherit manually created environments, inconsistent naming, undocumented access settings and deployment scripts that canno…
Cloud Security Alliance CCM Controls Implementation Training Course
Cloud security programmes often contain sound policies but lack a consistent way to translate cloud risks, provider assurances and technical…