MLOps for Production AI Systems Training Course

5 days Artificial Intelligence Certificate on completion
Course codeSD-AI-014
Duration5 days
LevelFoundation to Intermediate
CategoryArtificial Intelligence
DeliveryClassroom or live online
LanguageEnglish
CertificateCertificate of completion

Course overview

Production AI failures rarely begin with a poor model alone. They emerge when training data cannot be traced, experiments are recorded inconsistently, deployment environments differ from development, model behaviour is not monitored, or no team owns the response when quality degrades. This course equips data and engineering professionals to build an operational MLOps workflow that makes machine-learning systems reproducible, deployable, observable and governable. It addresses the practical hand-off between data science, software engineering, platform teams and risk stakeholders.

Participants design the lifecycle of a production AI service, from dataset and code versioning through experiment tracking, model packaging, automated testing, deployment and post-release monitoring. They work with Git, Docker, MLflow and Kubeflow to structure repositories, record model lineage, register approved model versions, build containers, orchestrate training pipelines and define promotion criteria. The course also covers data drift, model performance decay, alert thresholds, rollback procedures, access controls and documentation needed for audit-ready AI operations.

Instruction combines concise technical briefings with guided labs, architecture reviews and a running business case involving a prediction service moving from notebook prototype to managed deployment. Each participant produces an MLOps implementation blueprint: a target architecture, model lifecycle workflow, CI/CD control points, monitoring specification, model release checklist and 90-day adoption plan. The result is a practical artefact that can be adapted for a team’s current AI use case rather than a collection of disconnected tool demonstrations.

The course is suited to professionals who contribute to, operate or govern machine-learning products and need a shared operating model across development and production.

Course objectives

By the end of this course, participants will be able to:

  • Map an end-to-end MLOps lifecycle from data acquisition to model retirement
  • Structure a Git repository for reproducible training code, configuration and deployment assets
  • Record experiments, parameters, metrics and model artefacts in MLflow
  • Package a model inference service into a tested Docker container
  • Design a Kubeflow pipeline for repeatable data preparation, training and evaluation
  • Define model promotion gates using validation metrics, approval criteria and release evidence
  • Create a monitoring specification for data drift, prediction quality, latency and service errors
  • Produce an MLOps implementation blueprint with roles, controls, architecture and a 90-day rollout plan

Benefits of attending

For you

  • Build the ability to move from notebook-based experimentation to a controlled model release workflow
  • Gain hands-on evidence of MLOps capability through an architecture and implementation blueprint
  • Communicate model reliability, drift and release risk using operational metrics rather than vague technical assurances
  • Collaborate more effectively with platform, software, data governance and security teams on AI delivery
  • Strengthen readiness for machine learning engineering, AI platform and technical product leadership responsibilities

For your organisation

  • Reduce rework caused by untracked datasets, undocumented experiments and environment differences between training and deployment
  • Establish repeatable model release gates that improve accountability for performance, approval and rollback decisions
  • Shorten the path from validated model to deployable service through standardised pipelines and container packaging
  • Detect data drift, quality degradation and service failures earlier through defined monitoring signals and ownership
  • Create a shared MLOps operating model that aligns data science, engineering, platform and governance teams

Target competencies

Model lifecycle designExperiment trackingPipeline orchestrationContainerised deploymentModel observabilityRelease governance

Who should attend

  • Machine Learning Engineers — who must turn trained models into reliable, repeatable production services
  • Data Scientists — who need to hand models to engineering teams with reproducible experiments and clear release evidence
  • Data Engineers — who build dependable data pipelines and need to manage training-data lineage
  • Software Engineers — who integrate model services into applications and maintain CI/CD delivery controls
  • Cloud and Platform Engineers — who provide the container, pipeline and runtime foundations for AI workloads
  • AI Product Managers and Technical Delivery Leads — who must coordinate model release decisions, operating ownership and risk controls

Requirements and prerequisites

Participants should be comfortable navigating files and a command line, reading basic Python code, and using Git for commits, branches and pull requests. Familiarity with the machine-learning workflow—training a model, evaluating metrics such as accuracy or precision, and saving a model artefact—is assumed. Experience deploying web services, containers or cloud platforms is helpful but not essential. Participants do not need prior Kubeflow or MLflow experience, advanced statistics, deep-learning expertise, Kubernetes administration or a production ML system. Complete beginners to coding should first gain basic Python and Git practice before attending.

Training methodology

The instructor uses a single production-AI case throughout the week, beginning with an unstructured notebook model and progressively adding operational controls. Short instructor-led demonstrations introduce Git workflows, MLflow tracking, Docker packaging and Kubeflow pipeline concepts; participants then complete guided labs in small teams. Architecture critiques test trade-offs around batch versus real-time inference, approval gates and monitoring ownership. Daily outputs feed an end-of-course working session in which participants assemble a tailored MLOps blueprint and adoption plan for a realistic organisational use case.

Course outline

Day 1: MLOps foundations and operating model

  • MLOps lifecycle stages from problem framing to model retirement
  • Failure modes in notebook-to-production AI delivery
  • Roles and hand-offs across data science, engineering, platform and governance
  • Repository structure for code, configuration, tests and deployment manifests
  • Git branching, pull requests and code review controls for ML projects
  • Data lineage, feature definitions and reproducibility requirements
  • MLOps maturity assessment and target operating model design

Workshop: Assess a prototype prediction service, identify its production gaps, and produce a lifecycle map with named ownership for each stage.

Day 2: Reproducible experimentation and model lineage

  • Experiment design using parameter, metric and artefact logging
  • MLflow Tracking runs, tags and experiment comparisons
  • MLflow Model Registry stages, versions and model aliases
  • Training-data snapshots and dataset versioning strategies
  • Configuration-driven training and deterministic execution practices
  • Validation datasets, baseline metrics and acceptance thresholds
  • Model cards and release evidence documentation

Workshop: Run and compare multiple training experiments in MLflow, then register a selected model with validation evidence and a model card.

Day 3: Packaging and automated model delivery

  • Inference patterns for batch, asynchronous and real-time prediction
  • Dockerfiles for Python model inference services
  • Dependency pinning, environment variables and secret-handling principles
  • Unit, integration and contract tests for model-serving APIs
  • Continuous integration checks for code, data schemas and model artefacts
  • Continuous delivery stages and environment promotion rules
  • Deployment strategies including shadow, canary and blue-green releases

Workshop: Package a model API in Docker and define a CI/CD release pipeline with automated tests, approval gates and a rollback route.

Day 4: Pipeline orchestration and production monitoring

  • Kubeflow Pipelines components, parameters and pipeline artefacts
  • Orchestrating data preparation, training, evaluation and registration steps
  • Pipeline caching, scheduling and repeatable retraining runs
  • Feature, schema and data-quality monitoring signals
  • Concept drift, data drift and model performance decay measures
  • Operational telemetry for latency, throughput, errors and resource use
  • Alert thresholds, incident triage and rollback decision procedures

Workshop: Design a Kubeflow training pipeline and create a monitoring dashboard specification with alerts, owners and remediation actions.

Day 5: Governance, scale and implementation planning

  • Model risk classification and proportional control selection
  • Access control, audit trails and approval segregation
  • Bias, explainability and human-review checkpoints in operations
  • Model release checklists and change-management records
  • Cost, capacity and reliability considerations for AI platform operations
  • MLOps reference architectures for managed cloud and hybrid environments
  • 90-day roadmap, adoption metrics and capability prioritisation

Workshop: Present an MLOps implementation blueprint containing target architecture, release controls, monitoring plan, accountable roles and a 90-day rollout roadmap.

Tools & standards covered

Git, Docker, MLflow, Kubeflow

A typical training day

08:30 – 10:30First session
10:30 – 10:45Refreshment break
10:45 – 12:30Second session
12:30 – 13:30Lunch and networking
13:30 – 15:00Third session
15:00 – 15:15Refreshment break
15:15 – 16:30Workshop and daily review

Live online deliveries follow the same structure in the East Africa Time zone, with shorter screen blocks and longer breaks.

What the fee includes

  • Instruction by a practitioner facilitator
  • Full course workbook and materials
  • Exercise files, templates and case studies
  • Certificate of completion
  • Refreshments and lunch (classroom deliveries)
  • Post-course application plan
  • Facilitator follow-up on request
  • Group rates from five participants

How you can take this course

Classroom

Scheduled sessions in Nairobi, Mombasa, Kigali, Dar es Salaam, Dubai and Cape Town.

Live online

The same facilitator and materials, delivered live for distributed teams and individuals.

In-house

Delivered privately for your team, at your offices or a venue of your choice, tailored to your context. Request a proposal.

Certification

Participants who complete the full five days receive the Skillset Development Certificate of Completion, stating the course title, course code, dates and delivery format — suitable for professional-development records and employer reimbursement.

Frequently asked questions

No. The course introduces the relevant features of each tool through guided labs. You should, however, understand basic Git use, Python code structure and the idea of training and evaluating a machine-learning model.

Participants need a laptop able to run a modern browser, a code editor and Docker Desktop or an approved equivalent. Pre-course joining instructions provide repository access, installation steps and any hosted lab credentials required for MLflow and Kubeflow exercises.

Yes. Data scientists learn how to make experiments reproducible, prepare release evidence and define monitoring requirements. Platform and engineering participants gain equal value by learning how these requirements translate into pipelines, containers and operating controls.

This course does not focus on selecting algorithms or tuning models for maximum accuracy. It focuses on operating a model after development: lineage, packaging, release automation, monitoring, incident response and governance.

The course distinguishes practices from products, so participants can begin with repository standards, experiment records, release checklists and monitoring specifications before adopting a full platform. The final 90-day plan helps sequence controls according to current maturity and business risk.

You leave with a completed MLOps implementation blueprint covering architecture, roles, lifecycle stages, promotion gates, monitoring signals and a rollout plan. You also retain practical templates for model cards, release checklists and incident-response decisions.

Upcoming sessions

  • 21 – 25 Sep 2026
    Dar es Salaam · USD 3,500
    Book
  • 05 – 09 Oct 2026
    Live Online · USD 1,500
    Book
  • 05 – 09 Oct 2026
    Cape Town · USD 4,200
    Book
  • 12 – 16 Oct 2026
    Live Online · USD 1,500
    Book
  • 26 – 30 Oct 2026
    Dar es Salaam · USD 3,500
    Book
  • 26 – 30 Oct 2026
    Dubai · USD 4,500
    Book
  • 16 – 20 Nov 2026
    Cape Town · USD 4,200
    Book
  • 30 Nov – 04 Dec 2026
    Dubai · USD 4,500
    Book

49 more dates — ask us.


Group of 5+?

Request in-house delivery or group rates →

Related courses in Artificial Intelligence

5 Days Certificate

IBM watsonx AI Platform Administration Training Course

IBM watsonx administrators must make the platform usable for data scientists and application teams without creating uncontrolled access to s…

5 Days Certificate

Google Vertex AI Generative AI Development Training Course

Organisations are moving from generative AI experiments to applications that must be secure, observable, cost-controlled and useful to real …

5 Days Certificate

Advanced Computer Vision Model Deployment Training Course

Computer vision models that perform well in notebooks can fail under production conditions: camera feeds vary, object sizes shift, latency e…

5 Days Certificate

Artificial Intelligence for Data Analysts Training Course

Data analysts are increasingly expected to do more than produce dashboards and retrospective reports. They must identify patterns in large, …