Databricks Data Science and MLflow Training Course

5 days Data Science Certificate on completion
Course codeSD-DS-035
Duration5 days
LevelIntermediate
CategoryData Science
DeliveryClassroom or live online
LanguageEnglish
CertificateCertificate of completion

Course overview

Data science teams often lose time moving between notebooks, data preparation scripts, model experiments and deployment handovers. Inconsistent environments, untracked parameters, unclear data lineage and models that cannot be reproduced make it difficult to turn promising analysis into governed production assets. This five-day Databricks Data Science and MLflow Training Course helps practitioners build a repeatable workflow for preparing data, developing models, recording experiments and handing validated models to engineering or MLOps teams.

Participants work in Databricks using Apache Spark, Delta Lake and MLflow to create an end-to-end machine learning workflow. They learn to structure notebooks and repositories, ingest and transform data with Spark DataFrames, build reliable Delta tables, engineer features, train and compare models, and register model versions with documented metrics and parameters. The course also covers MLflow Tracking, Model Registry workflows, batch inference, model lifecycle controls and practical approaches to collaboration through Databricks workspaces and Unity Catalog.

Teaching combines instructor demonstrations with guided labs built around a realistic predictive modelling case. Participants progressively develop a documented Databricks project: a Delta-based feature dataset, trained and evaluated model, MLflow experiment record, registered model version and batch scoring notebook. They leave with reusable notebook patterns, a project structure and an implementation plan that can be adapted to their organisation's data platform, governance rules and deployment process.

Course objectives

By the end of this course, participants will be able to:

  • Configure a Databricks workspace workflow using notebooks, clusters, repositories and environment-aware parameters
  • Transform raw datasets into curated Delta Lake tables using Apache Spark DataFrames and SQL
  • Engineer reproducible features with documented transformations and train-test validation logic
  • Train and evaluate machine learning models in Databricks using suitable metrics and baseline comparisons
  • Track parameters, metrics, artefacts and model signatures with MLflow Tracking
  • Compare experiment runs and select a defensible candidate model using MLflow experiment evidence
  • Register, version and transition models through an MLflow Model Registry lifecycle workflow
  • Produce a batch inference notebook and end-to-end Databricks ML project handover pack

Benefits of attending

For you

  • Build evidence of practical Databricks capability through a completed Delta Lake, MLflow and batch-scoring project
  • Replace ad hoc notebook experimentation with a repeatable workflow for training, comparison and model handover
  • Gain confidence explaining model metrics, experiment evidence and version choices to technical stakeholders
  • Develop working knowledge of MLflow Model Registry processes used in collaborative ML delivery
  • Strengthen readiness for data scientist, ML engineer and MLOps-focused roles using Databricks

For your organisation

  • Reduce duplicated model development effort through standardised notebooks, experiments and reusable feature datasets
  • Improve model reproducibility by recording data inputs, parameters, metrics and artefacts in MLflow
  • Create clearer audit trails for model selection and version approval through registry-based lifecycle controls
  • Shorten handovers between data science, data engineering and deployment teams with shared Databricks patterns
  • Increase the value of existing Databricks investment by enabling staff to deliver governed batch ML workflows

Target competencies

Spark data preparationDelta Lake designFeature engineeringMLflow experiment trackingModel registry managementBatch inference delivery

Who should attend

  • Data Scientists — who need to operationalise experiments and make model results reproducible in Databricks
  • Machine Learning Engineers — who build training and inference workflows on the Databricks platform
  • Data Engineers — who prepare Delta Lake datasets and support feature pipelines for modelling teams
  • Analytics Engineers — who want to create governed, reusable data assets for predictive analytics
  • MLOps Engineers — who need practical MLflow tracking, registry and model lifecycle patterns
  • Technical Data Leads — who oversee collaborative data science delivery and platform standards

Requirements and prerequisites

Participants should be comfortable writing basic Python and SQL, including variables, functions, joins, filtering and aggregation. They should understand tabular data concepts and have practical exposure to exploratory analysis or machine learning workflows, such as splitting data, fitting a model and reviewing evaluation metrics. Prior use of Databricks is helpful but not essential. Participants need access to a laptop with a modern browser and should be able to install or use organisation-approved Python libraries if required. Deep learning expertise, advanced Spark administration, cloud infrastructure knowledge and prior MLflow experience are not required.

Training methodology

Each day combines focused instructor-led explanations with live Databricks demonstrations and guided individual labs. Participants work through a connected machine learning case rather than isolated commands: they prepare Delta tables, build features, train models, log MLflow runs and register a selected model. Short peer reviews are used to assess metric choices, experiment evidence and handover decisions. On the final day, participants complete a batch scoring workflow and map the course project patterns to a current organisational use case, including data, governance and operational dependencies.

Course outline

Day 1: Databricks foundations for collaborative data science

  • Databricks workspace architecture and compute options
  • Notebook authoring, execution context and notebook parameters
  • Databricks Repos and Git-based project organisation
  • Apache Spark DataFrame concepts and lazy execution
  • Reading CSV, Parquet and Delta data sources
  • Spark SQL for profiling and exploratory analysis
  • Unity Catalog concepts for data access and governance

Workshop: Participants create a structured Databricks project, inspect a supplied dataset with Spark, and produce an initial data-quality profile notebook.

Day 2: Data preparation and feature engineering with Delta Lake

  • Delta Lake tables, transaction logs and schema enforcement
  • Bronze, silver and gold data layer design
  • Data cleansing with Spark DataFrame transformations
  • Joins, window functions and aggregations for analytical features
  • Handling missing values, duplicates and invalid records
  • Feature creation and point-in-time data considerations
  • Persisting curated features as managed Delta tables

Workshop: Participants transform raw operational data into a validated Delta feature table and document the transformations used.

Day 3: Model development and MLflow experiment tracking

  • Train-validation-test splits and leakage prevention
  • Baseline model selection and evaluation criteria
  • Scikit-learn model training from Databricks notebooks
  • Classification and regression metric interpretation
  • MLflow experiments, runs and nested run structure
  • Logging parameters, metrics, artefacts and datasets
  • MLflow model signatures and input examples

Workshop: Participants train baseline and candidate models, log both runs in MLflow, and produce a comparison of their recorded metrics.

Day 4: Model selection, registration and lifecycle control

  • Interpreting MLflow run comparisons and artefact outputs
  • Model selection criteria beyond a single performance metric
  • MLflow Model Registry versions and aliases
  • Registering models from tracked MLflow runs
  • Model metadata, tags and business documentation
  • Validation gates and approval evidence for model promotion
  • Loading registered models for controlled reuse

Workshop: Participants select a candidate model, register it in MLflow, assign lifecycle metadata, and prepare a concise model decision record.

Day 5: Batch inference and production-ready handover

  • Batch inference patterns in Databricks
  • Loading MLflow models by registered alias
  • Writing prediction outputs to Delta Lake
  • Monitoring data quality and prediction distributions
  • Notebook modularisation and configuration management
  • Job scheduling concepts and operational dependencies
  • Data science handover checklist and implementation planning

Workshop: Participants build a batch scoring notebook that writes governed prediction outputs and complete an implementation plan for their own use case.

Tools & standards covered

Databricks, Apache Spark, MLflow, Delta Lake

A typical training day

08:30 – 10:30First session
10:30 – 10:45Refreshment break
10:45 – 12:30Second session
12:30 – 13:30Lunch and networking
13:30 – 15:00Third session
15:00 – 15:15Refreshment break
15:15 – 16:30Workshop and daily review

Live online deliveries follow the same structure in the East Africa Time zone, with shorter screen blocks and longer breaks.

What the fee includes

  • Instruction by a practitioner facilitator
  • Full course workbook and materials
  • Exercise files, templates and case studies
  • Certificate of completion
  • Refreshments and lunch (classroom deliveries)
  • Post-course application plan
  • Facilitator follow-up on request
  • Group rates from five participants

How you can take this course

Classroom

Scheduled sessions in Nairobi, Mombasa, Kigali, Dar es Salaam, Dubai and Cape Town.

Live online

The same facilitator and materials, delivered live for distributed teams and individuals.

In-house

Delivered privately for your team, at your offices or a venue of your choice, tailored to your context. Request a proposal.

Certification

Participants who complete the full five days receive the Skillset Development Certificate of Completion, stating the course title, course code, dates and delivery format — suitable for professional-development records and employer reimbursement.

Frequently asked questions

You should be able to write basic Python and SQL and understand the purpose of common machine learning evaluation metrics. You do not need prior MLflow experience, advanced Spark tuning knowledge or cloud administration skills.

A modern laptop with a browser is required; course access to a Databricks training environment is normally provided for the practical labs. If your organisation requires use of its own workspace, confirm access, cluster permissions and approved library policies before the course.

It is designed for data scientists, ML engineers, data engineers and technical leads who already work with data and need a practical Databricks-based ML workflow. It is not a first introduction to Python programming or statistical machine learning.

General platform courses focus primarily on workspace navigation, SQL or Spark processing. This course uses those capabilities to deliver a connected data science workflow, with particular emphasis on feature datasets, MLflow experiments, model registration and batch scoring.

The notebooks and project patterns can be adapted for common use cases such as churn prediction, demand forecasting, risk scoring or anomaly detection. Participants leave with a concrete handover and implementation plan that identifies the datasets, model controls and operational dependencies needed for their own workflow.

You will complete a Databricks project containing a curated Delta feature table, model training notebook, MLflow experiment history, registered model version and batch inference notebook. You will also have a model decision record and a practical checklist for moving the pattern into your organisation.

Upcoming sessions

New dates are being scheduled. Ask us about the next session or an in-house delivery for your team.

Ask about dates

Group of 5+?

Request in-house delivery or group rates →

Related courses in Data Science

5 Days Certificate

Jupyter Notebook Data Science Workflow Training Course

Data science teams often lose time moving between exploratory analysis, data cleaning, visualisation, model experiments and stakeholder repo…

5 Days Certificate

Data Science for Business Analysts Training Course

Business analysts are increasingly expected to move beyond static dashboards and descriptive reporting: they must test whether a pattern is …

5 Days Certificate

Data Science for Marketing Professionals Training Course

Marketing teams generate campaign, web, CRM and customer-service data every day, yet many decisions still rely on channel-level reports, las…

5 Days Certificate

Banking Data Science and Fraud Analytics Training Course

Banks hold rich transaction, customer, channel and behavioural data, yet fraud teams often face delayed alerts, high false-positive rates an…