Datadog Cloud Monitoring and Observability Training Course

5 days Cloud Computing Certificate on completion
Course codeSD-CC-014
Duration5 days
LevelIntermediate
CategoryCloud Computing
DeliveryClassroom or live online
LanguageEnglish
CertificateCertificate of completion

Course overview

Cloud teams often have metrics in one dashboard, logs in another, incomplete service ownership, and alerts that fire without identifying the customer impact or responsible dependency. This makes incident response slower, drives alert fatigue, and leaves engineering leaders unable to distinguish genuine reliability risks from normal operational noise. This five-day Datadog Cloud Monitoring and Observability Training Course equips practitioners to build an observability approach that connects infrastructure health, application performance, user experience, security signals, and business-facing service objectives.

Participants configure and use Datadog across Infrastructure Monitoring, Log Management, APM, Real User Monitoring, Synthetic Monitoring, dashboards, monitors, Service Catalog, and incident workflows. They learn to instrument services with OpenTelemetry, correlate traces with logs and host or container metrics, define service-level objectives, tune monitor thresholds, manage tags, investigate incidents with Watchdog and Analytics views, and design role-appropriate dashboards. The course also addresses Kubernetes and cloud monitoring patterns, cost-aware telemetry collection, and configuration-as-code practices using Terraform.

Teaching combines instructor-led demonstrations with guided labs in a Datadog sandbox, using realistic cloud application failures including latency regressions, Kubernetes resource contention, failed deployments, and third-party dependency outages. Participants complete an end-of-course observability implementation plan and dashboard-and-monitor pack for a representative service, including a tag model, service health dashboard, monitor specifications, SLOs, escalation routes, and investigation runbook. This gives both the attendee and their manager a practical artefact that can be adapted for production adoption.

The course is designed for intermediate engineers and operations professionals who already work with cloud-hosted applications and need to standardise how their teams detect, investigate, and prevent service degradation using Datadog.

Course objectives

By the end of this course, participants will be able to:

  • Configure Datadog integrations and agents to collect cloud, host, container, and application telemetry
  • Apply a consistent tag taxonomy to support service ownership, environment filtering, and cost allocation
  • Correlate metrics, logs, traces, and user-session data during Datadog investigations
  • Build APM service maps, trace queries, and latency analyses to isolate application bottlenecks
  • Create actionable monitors with threshold logic, notification routing, recovery conditions, and message templates
  • Define SLOs and error budgets that measure service reliability against customer-facing targets
  • Design role-based Datadog dashboards for service health, Kubernetes operations, and incident command
  • Produce an observability implementation plan with monitor specifications, runbooks, and Terraform configuration patterns

Benefits of attending

For you

  • Gain practical evidence of Datadog capability through a completed dashboard, monitor, SLO, and investigation pack
  • Improve incident diagnosis by tracing a user-facing symptom through infrastructure, logs, and distributed services
  • Develop the ability to challenge noisy alerts and replace them with service-impact-focused monitor logic
  • Build credibility for SRE, DevOps, cloud operations, and platform engineering responsibilities
  • Learn to translate technical telemetry into reliability measures and risk reports for engineering leadership

For your organisation

  • Reduce alert fatigue through better monitor thresholds, routing rules, recovery settings, and ownership tags
  • Shorten incident investigation by standardising correlation between logs, metrics, traces, and service dependencies
  • Improve reliability governance with defined SLOs, error budgets, and service health dashboards
  • Establish reusable Datadog implementation patterns for cloud accounts, Kubernetes clusters, and application teams
  • Control observability spend through purposeful telemetry collection, tag governance, and retention-aware design

Target competencies

Datadog telemetry designDistributed tracing analysisAlert engineeringSLO managementKubernetes observabilityIncident investigation

Who should attend

  • Site Reliability Engineers — who need to reduce mean time to detect and diagnose cloud service failures
  • Cloud Engineers — who operate AWS, Azure, or Google Cloud workloads and need unified operational visibility
  • DevOps Engineers — who build deployment and monitoring practices across application delivery pipelines
  • Platform Engineers — who provide reusable observability standards and self-service monitoring for product teams
  • Application Support Engineers — who investigate production incidents across logs, traces, and infrastructure signals
  • Engineering Managers — who need service health measures, ownership visibility, and defensible reliability targets

Requirements and prerequisites

Participants should be comfortable navigating a Linux command line, reading application logs, and working with basic cloud concepts such as virtual machines, containers, load balancers, managed databases, IAM roles, and environments. Experience operating applications on AWS, Microsoft Azure, or Google Cloud is useful, as is familiarity with Kubernetes concepts including pods, deployments, namespaces, and resource limits. Participants should understand basic HTTP behaviour and application response times. Prior Datadog experience is not required, and attendees do not need to be software developers, Terraform specialists, or OpenTelemetry experts before attending.

Training methodology

The course is delivered through short instructor-led technical briefings followed by guided Datadog labs in a realistic cloud application environment. Participants configure integrations, query telemetry, build dashboards, create monitors, and investigate staged failures using metrics, logs, traces, RUM, and Synthetic Monitoring. Small-group case work focuses on alert quality, service ownership, and SLO decisions. Daily review sessions connect configuration choices to operational outcomes. On day five, each participant assembles an observability plan and receives structured instructor feedback on its implementation priorities.

Course outline

Day 1: Datadog foundations and telemetry architecture

  • Datadog organisation structure, roles, teams, and access controls
  • Datadog Agent deployment models for hosts, containers, and cloud environments
  • Metrics types, metric submission, retention, and query fundamentals
  • Tagging strategy for service, environment, team, region, and cost attribution
  • Cloud integration patterns for AWS, Microsoft Azure, and Google Cloud
  • Infrastructure Monitoring views, host maps, process monitoring, and network context
  • Service Catalog ownership metadata and dependency documentation

Workshop: Participants design a tag taxonomy and configure infrastructure visibility for a multi-service cloud application, producing an ownership and telemetry coverage map.

Day 2: Logs, APM, and distributed service diagnosis

  • Log collection pipelines, parsing rules, facets, and sensitive-data handling
  • Log Explorer queries, patterns, analytics measures, and saved views
  • APM service instrumentation and Unified Service Tagging
  • OpenTelemetry concepts, semantic conventions, and Datadog trace ingestion
  • Trace search, flame graphs, span tags, and latency decomposition
  • Service Map analysis for upstream, downstream, and external dependencies
  • Log-trace-metric correlation workflows for incident investigation

Workshop: Participants investigate a simulated checkout latency incident, producing a root-cause evidence trail from a user request through traces, logs, and database metrics.

Day 3: Cloud, Kubernetes, and digital experience monitoring

  • Kubernetes Cluster Agent, kube-state metrics, and container tagging
  • Kubernetes dashboards for pods, deployments, nodes, namespaces, and resource requests
  • Container resource saturation, restart analysis, and workload health signals
  • Cloud service monitoring for load balancers, databases, queues, and serverless functions
  • Real User Monitoring application setup, session views, and frontend error analysis
  • Synthetic API and browser tests for critical user journeys
  • Network Performance Monitoring and dependency traffic analysis

Workshop: Participants build a Kubernetes and customer-journey dashboard, then identify the operational impact of a staged pod capacity and API failure.

Day 4: Alert engineering, SLOs, and incident response

  • Monitor types, evaluation windows, thresholds, anomaly detection, and forecast alerts
  • Composite monitors, dependency-aware alert logic, and monitor grouping
  • Notification templates, escalation routing, recovery messages, and runbook links
  • Alert quality review using false-positive, duplicate, and actionable-signal criteria
  • Service-level indicators, SLO target selection, and error budget calculations
  • SLO status dashboards, burn-rate monitoring, and reliability reporting
  • Incident Management workflows, timelines, roles, and post-incident evidence capture

Workshop: Participants replace a noisy alert set with a monitored service objective, routed monitors, and an incident-response runbook for a payment service.

Day 5: Operationalising Datadog at scale

  • Dashboard design for executives, service owners, on-call teams, and platform operations
  • Notebook investigations, shared diagnostic context, and operational reporting
  • Datadog Watchdog insights and automated anomaly investigation
  • Terraform provider patterns for monitors, dashboards, SLOs, and notification configuration
  • Telemetry governance, naming standards, data access, and audit considerations
  • Observability cost management through metric volume, log indexes, and retention choices
  • Phased Datadog adoption roadmap, success measures, and operating model decisions

Workshop: Participants present an observability implementation plan and deliver a service dashboard, monitor set, SLO definition, tag standard, and prioritised 90-day adoption roadmap.

Tools & standards covered

Datadog, OpenTelemetry, Kubernetes, Terraform

A typical training day

08:30 – 10:30First session
10:30 – 10:45Refreshment break
10:45 – 12:30Second session
12:30 – 13:30Lunch and networking
13:30 – 15:00Third session
15:00 – 15:15Refreshment break
15:15 – 16:30Workshop and daily review

Live online deliveries follow the same structure in the East Africa Time zone, with shorter screen blocks and longer breaks.

What the fee includes

  • Instruction by a practitioner facilitator
  • Full course workbook and materials
  • Exercise files, templates and case studies
  • Certificate of completion
  • Refreshments and lunch (classroom deliveries)
  • Post-course application plan
  • Facilitator follow-up on request
  • Group rates from five participants

How you can take this course

Classroom

Scheduled sessions in Nairobi, Mombasa, Kigali, Dar es Salaam, Dubai and Cape Town.

Live online

The same facilitator and materials, delivered live for distributed teams and individuals.

In-house

Delivered privately for your team, at your offices or a venue of your choice, tailored to your context. Request a proposal.

Certification

Participants who complete the full five days receive the Skillset Development Certificate of Completion, stating the course title, course code, dates and delivery format — suitable for professional-development records and employer reimbursement.

Frequently asked questions

No previous Datadog experience is required. You should, however, understand basic cloud infrastructure, application logs, HTTP behaviour, and the operational purpose of monitoring.

A laptop capable of accessing a modern web browser and a command-line terminal is required for live labs. Training access to a Datadog sandbox and the required exercise materials are provided; a personal production account is not needed.

Yes. The course includes Kubernetes Cluster Agent concepts, container metrics, pod and deployment health, cloud integrations, distributed tracing, and service ownership practices. It is particularly useful for teams operating microservices or managed cloud platforms.

This course is specifically built around Datadog configuration, investigation workflows, monitor design, SLOs, dashboards, and governance. Rather than comparing monitoring products at a high level, participants work directly with Datadog features and implementation decisions.

Participants practise the tasks performed during real service operations: finding relevant telemetry, isolating dependencies, tuning alerts, reviewing error budgets, and communicating service health. The methods can be applied immediately to existing dashboards, monitors, and incident runbooks.

You will leave with a completed observability implementation plan for a representative service, including tagging standards, dashboard designs, monitor specifications, SLOs, investigation steps, and a 90-day adoption roadmap. You will also receive a certificate on completion.

Upcoming sessions

New dates are being scheduled. Ask us about the next session or an in-house delivery for your team.

Ask about dates

Group of 5+?

Request in-house delivery or group rates →

Related courses in Cloud Computing

5 Days Certificate

Cloud Services for Public Sector Digital Teams Training Course

Public-sector digital teams must modernise citizen-facing services while protecting sensitive data, sustaining continuity, meeting procureme…

5 Days Certificate

AWS Well-Architected Framework Implementation Training Course

AWS workloads often grow faster than their governance, documentation and operational controls. Teams inherit accounts with inconsistent tagg…

5 Days Certificate

Pulumi Cloud Infrastructure Automation Training Course

Cloud teams often inherit manually created environments, inconsistent naming, undocumented access settings and deployment scripts that canno…

5 Days Certificate

Cloud Security Alliance CCM Controls Implementation Training Course

Cloud security programmes often contain sound policies but lack a consistent way to translate cloud risks, provider assurances and technical…