MLOps for Production AI Systems Training Course
| Course code | SD-AI-014 |
|---|---|
| Duration | 5 days |
| Level | Foundation to Intermediate |
| Category | Artificial Intelligence |
| Delivery | Classroom or live online |
| Language | English |
| Certificate | Certificate of completion |
Course overview
Production AI failures rarely begin with a poor model alone. They emerge when training data cannot be traced, experiments are recorded inconsistently, deployment environments differ from development, model behaviour is not monitored, or no team owns the response when quality degrades. This course equips data and engineering professionals to build an operational MLOps workflow that makes machine-learning systems reproducible, deployable, observable and governable. It addresses the practical hand-off between data science, software engineering, platform teams and risk stakeholders.
Participants design the lifecycle of a production AI service, from dataset and code versioning through experiment tracking, model packaging, automated testing, deployment and post-release monitoring. They work with Git, Docker, MLflow and Kubeflow to structure repositories, record model lineage, register approved model versions, build containers, orchestrate training pipelines and define promotion criteria. The course also covers data drift, model performance decay, alert thresholds, rollback procedures, access controls and documentation needed for audit-ready AI operations.
Instruction combines concise technical briefings with guided labs, architecture reviews and a running business case involving a prediction service moving from notebook prototype to managed deployment. Each participant produces an MLOps implementation blueprint: a target architecture, model lifecycle workflow, CI/CD control points, monitoring specification, model release checklist and 90-day adoption plan. The result is a practical artefact that can be adapted for a team’s current AI use case rather than a collection of disconnected tool demonstrations.
The course is suited to professionals who contribute to, operate or govern machine-learning products and need a shared operating model across development and production.
Course objectives
By the end of this course, participants will be able to:
- Map an end-to-end MLOps lifecycle from data acquisition to model retirement
- Structure a Git repository for reproducible training code, configuration and deployment assets
- Record experiments, parameters, metrics and model artefacts in MLflow
- Package a model inference service into a tested Docker container
- Design a Kubeflow pipeline for repeatable data preparation, training and evaluation
- Define model promotion gates using validation metrics, approval criteria and release evidence
- Create a monitoring specification for data drift, prediction quality, latency and service errors
- Produce an MLOps implementation blueprint with roles, controls, architecture and a 90-day rollout plan
Benefits of attending
For you
- Build the ability to move from notebook-based experimentation to a controlled model release workflow
- Gain hands-on evidence of MLOps capability through an architecture and implementation blueprint
- Communicate model reliability, drift and release risk using operational metrics rather than vague technical assurances
- Collaborate more effectively with platform, software, data governance and security teams on AI delivery
- Strengthen readiness for machine learning engineering, AI platform and technical product leadership responsibilities
For your organisation
- Reduce rework caused by untracked datasets, undocumented experiments and environment differences between training and deployment
- Establish repeatable model release gates that improve accountability for performance, approval and rollback decisions
- Shorten the path from validated model to deployable service through standardised pipelines and container packaging
- Detect data drift, quality degradation and service failures earlier through defined monitoring signals and ownership
- Create a shared MLOps operating model that aligns data science, engineering, platform and governance teams
Target competencies
Who should attend
- Machine Learning Engineers — who must turn trained models into reliable, repeatable production services
- Data Scientists — who need to hand models to engineering teams with reproducible experiments and clear release evidence
- Data Engineers — who build dependable data pipelines and need to manage training-data lineage
- Software Engineers — who integrate model services into applications and maintain CI/CD delivery controls
- Cloud and Platform Engineers — who provide the container, pipeline and runtime foundations for AI workloads
- AI Product Managers and Technical Delivery Leads — who must coordinate model release decisions, operating ownership and risk controls
Requirements and prerequisites
Participants should be comfortable navigating files and a command line, reading basic Python code, and using Git for commits, branches and pull requests. Familiarity with the machine-learning workflow—training a model, evaluating metrics such as accuracy or precision, and saving a model artefact—is assumed. Experience deploying web services, containers or cloud platforms is helpful but not essential. Participants do not need prior Kubeflow or MLflow experience, advanced statistics, deep-learning expertise, Kubernetes administration or a production ML system. Complete beginners to coding should first gain basic Python and Git practice before attending.
Training methodology
The instructor uses a single production-AI case throughout the week, beginning with an unstructured notebook model and progressively adding operational controls. Short instructor-led demonstrations introduce Git workflows, MLflow tracking, Docker packaging and Kubeflow pipeline concepts; participants then complete guided labs in small teams. Architecture critiques test trade-offs around batch versus real-time inference, approval gates and monitoring ownership. Daily outputs feed an end-of-course working session in which participants assemble a tailored MLOps blueprint and adoption plan for a realistic organisational use case.
Course outline
Day 1: MLOps foundations and operating model
- MLOps lifecycle stages from problem framing to model retirement
- Failure modes in notebook-to-production AI delivery
- Roles and hand-offs across data science, engineering, platform and governance
- Repository structure for code, configuration, tests and deployment manifests
- Git branching, pull requests and code review controls for ML projects
- Data lineage, feature definitions and reproducibility requirements
- MLOps maturity assessment and target operating model design
Workshop: Assess a prototype prediction service, identify its production gaps, and produce a lifecycle map with named ownership for each stage.
Day 2: Reproducible experimentation and model lineage
- Experiment design using parameter, metric and artefact logging
- MLflow Tracking runs, tags and experiment comparisons
- MLflow Model Registry stages, versions and model aliases
- Training-data snapshots and dataset versioning strategies
- Configuration-driven training and deterministic execution practices
- Validation datasets, baseline metrics and acceptance thresholds
- Model cards and release evidence documentation
Workshop: Run and compare multiple training experiments in MLflow, then register a selected model with validation evidence and a model card.
Day 3: Packaging and automated model delivery
- Inference patterns for batch, asynchronous and real-time prediction
- Dockerfiles for Python model inference services
- Dependency pinning, environment variables and secret-handling principles
- Unit, integration and contract tests for model-serving APIs
- Continuous integration checks for code, data schemas and model artefacts
- Continuous delivery stages and environment promotion rules
- Deployment strategies including shadow, canary and blue-green releases
Workshop: Package a model API in Docker and define a CI/CD release pipeline with automated tests, approval gates and a rollback route.
Day 4: Pipeline orchestration and production monitoring
- Kubeflow Pipelines components, parameters and pipeline artefacts
- Orchestrating data preparation, training, evaluation and registration steps
- Pipeline caching, scheduling and repeatable retraining runs
- Feature, schema and data-quality monitoring signals
- Concept drift, data drift and model performance decay measures
- Operational telemetry for latency, throughput, errors and resource use
- Alert thresholds, incident triage and rollback decision procedures
Workshop: Design a Kubeflow training pipeline and create a monitoring dashboard specification with alerts, owners and remediation actions.
Day 5: Governance, scale and implementation planning
- Model risk classification and proportional control selection
- Access control, audit trails and approval segregation
- Bias, explainability and human-review checkpoints in operations
- Model release checklists and change-management records
- Cost, capacity and reliability considerations for AI platform operations
- MLOps reference architectures for managed cloud and hybrid environments
- 90-day roadmap, adoption metrics and capability prioritisation
Workshop: Present an MLOps implementation blueprint containing target architecture, release controls, monitoring plan, accountable roles and a 90-day rollout roadmap.
Tools & standards covered
Git, Docker, MLflow, Kubeflow
A typical training day
| 08:30 – 10:30 | First session |
| 10:30 – 10:45 | Refreshment break |
| 10:45 – 12:30 | Second session |
| 12:30 – 13:30 | Lunch and networking |
| 13:30 – 15:00 | Third session |
| 15:00 – 15:15 | Refreshment break |
| 15:15 – 16:30 | Workshop and daily review |
Live online deliveries follow the same structure in the East Africa Time zone, with shorter screen blocks and longer breaks.
What the fee includes
- Instruction by a practitioner facilitator
- Full course workbook and materials
- Exercise files, templates and case studies
- Certificate of completion
- Refreshments and lunch (classroom deliveries)
- Post-course application plan
- Facilitator follow-up on request
- Group rates from five participants
How you can take this course
Classroom
Scheduled sessions in Nairobi, Mombasa, Kigali, Dar es Salaam, Dubai and Cape Town.
Live online
The same facilitator and materials, delivered live for distributed teams and individuals.
In-house
Delivered privately for your team, at your offices or a venue of your choice, tailored to your context. Request a proposal.
Certification
Participants who complete the full five days receive the Skillset Development Certificate of Completion, stating the course title, course code, dates and delivery format — suitable for professional-development records and employer reimbursement.
Frequently asked questions
Upcoming sessions
-
21 – 25 Sep 2026Book
Dar es Salaam · USD 3,500 -
05 – 09 Oct 2026Book
Live Online · USD 1,500 -
05 – 09 Oct 2026Book
Cape Town · USD 4,200 -
12 – 16 Oct 2026Book
Live Online · USD 1,500 -
26 – 30 Oct 2026Book
Dar es Salaam · USD 3,500 -
26 – 30 Oct 2026Book
Dubai · USD 4,500 -
16 – 20 Nov 2026Book
Cape Town · USD 4,200 -
30 Nov – 04 Dec 2026Book
Dubai · USD 4,500
49 more dates — ask us.
Group of 5+?
Request in-house delivery or group rates →Related courses in Artificial Intelligence
IBM watsonx AI Platform Administration Training Course
IBM watsonx administrators must make the platform usable for data scientists and application teams without creating uncontrolled access to s…
Google Vertex AI Generative AI Development Training Course
Organisations are moving from generative AI experiments to applications that must be secure, observable, cost-controlled and useful to real …
Advanced Computer Vision Model Deployment Training Course
Computer vision models that perform well in notebooks can fail under production conditions: camera feeds vary, object sizes shift, latency e…
Artificial Intelligence for Data Analysts Training Course
Data analysts are increasingly expected to do more than produce dashboards and retrospective reports. They must identify patterns in large, …