Advanced Data Analytics for Causal Inference and Experiment Design Training Course

5 days Data Analytics Certificate on completion
Course codeSD-DA-039
Duration5 days
LevelIntermediate to Advanced
CategoryData Analytics
DeliveryClassroom or live online
LanguageEnglish
CertificateCertificate of completion

Course overview

Business teams routinely make high-stakes decisions from observational data: changing prices, targeting customers, redesigning workflows, introducing product features, or reallocating marketing spend. Standard dashboards and predictive models can show associations, but they cannot reliably answer whether an intervention caused an outcome. This course equips experienced analysts to distinguish correlation from causation, challenge weak claims, and design experiments that produce evidence decision-makers can defend.

Across five days, participants build an end-to-end causal analysis workflow. They formulate causal questions, construct directed acyclic graphs (DAGs), identify confounders and invalid controls, define estimands, and select appropriate experimental or quasi-experimental designs. The programme covers randomised controlled trials, power and sample-size calculations, stratification, cluster and sequential experiments, difference-in-differences, regression discontinuity, instrumental variables, propensity-score methods, and sensitivity analysis. Participants use Python, R, SQL and DAGitty to prepare data, analyse treatment effects and communicate uncertainty.

Teaching combines instructor-led method demonstrations with applied labs based on realistic product, marketing, operations and policy scenarios. Participants critique flawed analyses, diagnose common sources of bias, design an experiment, and complete a capstone causal-inference case. They leave with a documented causal analysis and experiment-design pack containing a DAG, estimand statement, data requirements, design choice, power calculation, analysis plan, and stakeholder-ready recommendation.

The course is designed for analysts and data professionals who already work with business data and need to move from descriptive or predictive reporting to credible intervention measurement. It is also valuable for analytics leaders responsible for approving test plans, assessing vendor claims, and setting standards for evidence-based decisions.

Course objectives

By the end of this course, participants will be able to:

  • Construct directed acyclic graphs to represent causal assumptions, confounders, mediators and colliders
  • Define treatment effects and estimands, including average treatment effects, intention-to-treat and treatment-on-the-treated
  • Design randomised controlled trials using randomisation, blocking, stratification and cluster assignment
  • Calculate statistical power, minimum detectable effect and sample size for A/B and multivariate experiments
  • Analyse experimental results with regression adjustment, confidence intervals, multiple-testing controls and heterogeneous treatment effects
  • Apply difference-in-differences, regression discontinuity and instrumental-variable methods to observational business data
  • Evaluate propensity-score matching and weighting models using covariate-balance diagnostics and sensitivity analysis
  • Produce a defensible causal analysis and experiment-design pack for stakeholder review

Benefits of attending

For you

  • Gain a repeatable framework for challenging causal claims before they influence product, marketing or operations decisions
  • Build the credibility to explain why a dashboard trend or regression coefficient is not automatically an intervention effect
  • Design statistically powered experiments rather than relying on arbitrary test durations or sample targets
  • Add quasi-experimental methods to your analytical portfolio when controlled trials are impractical
  • Leave with a reusable causal analysis and experiment-design pack that demonstrates advanced analytical capability

For your organisation

  • Reduce spend on initiatives supported only by correlations, vanity metrics or poorly controlled before-and-after comparisons
  • Improve experiment quality through explicit power calculations, assignment rules, guardrail metrics and pre-specified analysis plans
  • Create more reliable estimates of campaign, product and process incrementality for investment decisions
  • Lower analytical risk by identifying confounding, selection bias, post-treatment controls and multiple-testing errors before publication
  • Establish a common evidence standard across analytics, product, marketing and operational teams

Target competencies

Causal graph modellingExperiment power planningTreatment effect estimationQuasi-experimental designBias diagnostic methodsEvidence-based recommendations

Who should attend

  • Senior Data Analysts — who need to substantiate whether business interventions changed outcomes
  • Data Scientists — who must complement predictive models with causal estimates for product and commercial decisions
  • Product Analysts — who design feature experiments and interpret A/B test results for product roadmaps
  • Marketing Analytics Managers — who need credible incrementality evidence for campaign, channel and offer investment
  • Business Intelligence Leads — who set measurement standards and challenge correlation-based performance claims
  • Operations and Strategy Analysts — who evaluate policy, process and service changes when randomisation is constrained

Requirements and prerequisites

Participants should be comfortable working with tabular data and writing or reviewing basic SQL queries. They need practical familiarity with either Python or R, including data frames, visualisation and fitting a basic regression model. The course assumes understanding of descriptive statistics, probability, confidence intervals and hypothesis testing, plus experience interpreting business metrics. Prior exposure to A/B testing is useful but not essential. Participants do not need prior knowledge of causal inference, DAGs, econometrics, advanced machine learning or Bayesian methods. No specialist mathematics beyond applied algebra and statistical reasoning is required.

Training methodology

The instructor uses short technical briefings to introduce each method, then guides participants through notebook-based analysis in Python or R and SQL data preparation tasks. Case studies include feature roll-outs, campaign targeting and operational policy changes, allowing participants to compare randomised and observational approaches. Small groups review causal diagrams, identify invalid controls and defend design decisions to a mock steering group. Each day closes with a practical output, culminating in an individual application plan and a documented causal analysis and experiment-design pack.

Course outline

Day 1: Causal Questions, Assumptions and Data Structure

  • Correlation, prediction and causal-effect questions
  • Potential outcomes and counterfactual reasoning
  • Treatment, outcome, unit and estimand definitions
  • Directed acyclic graph construction in DAGitty
  • Confounders, mediators, colliders and selection bias
  • Backdoor criterion and valid adjustment sets
  • SQL extraction patterns for treatment and outcome datasets

Workshop: Participants map a business intervention with a DAG and produce an estimand statement, adjustment-set rationale and initial data specification.

Day 2: Randomised Experiments and Test Planning

  • Randomised controlled trial architecture
  • Unit of randomisation and interference risks
  • Simple randomisation, blocking and stratified assignment
  • Cluster randomisation and intracluster correlation
  • Primary metrics, guardrails and success criteria
  • Power, minimum detectable effect and sample-size calculation
  • Pre-registration and statistical analysis plans

Workshop: Participants create a powered A/B test plan for a product or campaign decision, including assignment logic, metrics, sample target and stopping rules.

Day 3: Experimental Analysis and Decision Rules

  • Intention-to-treat and treatment-on-the-treated estimation
  • Difference in means and regression-adjusted treatment effects
  • Confidence intervals, p-values and practical significance
  • Covariate adjustment and precision improvement
  • Multiple comparisons and false discovery control
  • Sequential testing and peeking bias
  • Heterogeneous treatment effects and subgroup analysis

Workshop: Participants analyse an A/B test in Python or R and produce a decision memo that reports effect size, uncertainty, guardrail results and limitations.

Day 4: Quasi-Experimental Methods for Observational Data

  • Identification strategy selection for non-randomised interventions
  • Propensity-score matching and inverse-probability weighting
  • Covariate-balance diagnostics and common-support checks
  • Difference-in-differences and parallel-trends assessment
  • Regression discontinuity design and bandwidth choice
  • Instrumental variables and exclusion-restriction assumptions
  • Sensitivity analysis for unobserved confounding

Workshop: Participants evaluate a non-randomised programme using two candidate quasi-experimental approaches and recommend the more credible identification strategy.

Day 5: Causal Evidence Governance and Applied Capstone

  • Data-quality checks for causal measurement
  • Missing data, attrition and non-compliance handling
  • Spillovers, novelty effects and external-validity threats
  • Causal model diagnostics and assumption documentation
  • Visual communication of treatment effects and uncertainty
  • Stakeholder challenge sessions and decision framing
  • Causal analysis and experiment-design pack structure

Workshop: Participants complete and present a capstone causal analysis and experiment-design pack containing a DAG, estimand, method choice, diagnostic evidence, findings and implementation recommendation.

Tools & standards covered

Python, R, SQL, DAGitty

A typical training day

08:30 – 10:30First session
10:30 – 10:45Refreshment break
10:45 – 12:30Second session
12:30 – 13:30Lunch and networking
13:30 – 15:00Third session
15:00 – 15:15Refreshment break
15:15 – 16:30Workshop and daily review

Live online deliveries follow the same structure in the East Africa Time zone, with shorter screen blocks and longer breaks.

What the fee includes

  • Instruction by a practitioner facilitator
  • Full course workbook and materials
  • Exercise files, templates and case studies
  • Certificate of completion
  • Refreshments and lunch (classroom deliveries)
  • Post-course application plan
  • Facilitator follow-up on request
  • Group rates from five participants

How you can take this course

Classroom

Scheduled sessions in Nairobi, Mombasa, Kigali, Dar es Salaam, Dubai and Cape Town.

Live online

The same facilitator and materials, delivered live for distributed teams and individuals.

In-house

Delivered privately for your team, at your offices or a venue of your choice, tailored to your context. Request a proposal.

Certification

Participants who complete the full five days receive the Skillset Development Certificate of Completion, stating the course title, course code, dates and delivery format — suitable for professional-development records and employer reimbursement.

Frequently asked questions

You should already be comfortable with regression, confidence intervals and working with data frames in Python or R. The course teaches causal-inference methods from first principles, but moves quickly into applied estimation, diagnostics and design decisions rather than introductory programming.

Yes. Bring a laptop capable of running Python or R notebooks and accessing a SQL environment. Exercises use Python, R, SQL and DAGitty; starter code and datasets are provided so participants can focus on causal reasoning rather than environment setup.

Yes. Randomised experiments are covered because they are the strongest design where feasible, but substantial time is devoted to quasi-experimental methods. You will practise selecting and validating difference-in-differences, regression discontinuity, instrumental variables and propensity-score approaches for constrained settings.

Standard analytics courses focus on describing data, forecasting outcomes or building predictive models. This course focuses on identification: determining whether an intervention caused an outcome, documenting the assumptions required, and testing whether those assumptions are credible.

You can use the causal-question template, DAG review process, power-calculation workflow and pre-specified analysis plan on upcoming initiatives. The course also provides a structure for reviewing existing reports that make unsupported causal claims.

Each participant leaves with a completed causal analysis and experiment-design pack based on a realistic case or an approved work problem. It includes a DAG, estimand, data requirements, proposed design, power or identification assessment, diagnostic plan and stakeholder recommendation.

Upcoming sessions

New dates are being scheduled. Ask us about the next session or an in-house delivery for your team.

Ask about dates

Group of 5+?

Request in-house delivery or group rates →

Related courses in Data Analytics

5 Days Certificate

dbt Analytics Engineering and Data Quality Testing Training Course

Analytics teams often inherit SQL transformations that run without ownership, documentation or reliable checks. A dashboard can look credibl…

5 Days Certificate

Oil and Gas Data Analytics for Production Performance Training Course

Production teams often hold years of historian, well test, allocation, maintenance, drilling and laboratory data, yet struggle to turn it in…

10 Days Certificate

ArcGIS Pro Spatial Data Analysis and Mapping Training Course

Organisations hold location-rich data in asset registers, customer systems, operational databases, spreadsheets and field surveys, yet many …

5 Days Certificate

Alteryx Data Preparation and Workflow Analytics Training Course

Operational data is often spread across spreadsheets, CRM exports, finance systems, databases and shared folders, leaving analysts to repeat…