
Data Science Statistics Course
Master the statistical foundations that power modern data science, from probability theory and hypothesis testing to regression, ANOVA, and Bayesian inference. This course gives you the rigorous, practical toolkit employers expect from serious data professionals. Every concept is grounded in real analytical workflows so you can apply what you learn immediately.
What you will learn:
You will build a complete understanding of statistics as it is actually practised in data science. The course covers descriptive statistics, probability distributions, confidence intervals, and hypothesis testing before advancing to linear and logistic regression, categorical data analysis, and experimental design. You will also explore Bayesian methods, nonparametric tests, time series forecasting, causal inference, and dimensionality reduction. Each topic connects directly to the decisions and models you will encounter in professional data work. By the end, you will be able to select the right statistical method, validate your assumptions, and communicate results with precision and confidence.
How you study in practice Data Science Statistics Course
How you practise Data Science Statistics Course
For companies looking to train their teams
With Dedika for businesses, the course includes exercises and examples tailored to your company and its specific needs.
Course content
8 Chapters • 39 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsFoundations of Statistical Thinking
Foundations of Statistical Thinking
Lesson 1 • Descriptive Statistics Essentials
Summarise datasets using measures of centre, spread, and shape. These summaries form the baseline for all inferential and predictive work in later chapters.
Lesson 2 • Visualising Distributions
Translate raw data into histograms, box plots, and density curves to reveal shape and outliers. Visualisation skills anchor every exploratory analysis in the course.
Lesson 3 • Types and Structures of Data
Distinguish nominal, ordinal, interval, and ratio scales and their analytical implications. Correct data classification prevents misapplied statistical methods downstream.
Lesson 4 • Populations, Samples, and Bias
Define populations, sampling frames, and common bias sources that distort inference. Understanding bias is prerequisite to valid hypothesis testing introduced later.
Chapter 2HideHide detailsSee detailsProbability Theory for Data Scientists
Probability Theory for Data Scientists
Lesson 1 • Conditional Probability and Independence
Compute conditional probabilities and test statistical independence between events. Independence assumptions drive many model choices in regression and classification.
Lesson 2 • Random Variables and Expectation
Formalise random variables, compute expectations, and derive variance algebraically. Mastery here enables rigorous treatment of estimators in the next chapter.
Lesson 3 • Discrete Probability Distributions
Model count-based outcomes using binomial, Poisson, and geometric distributions. Recognising the correct discrete model prevents systematic prediction errors.
Lesson 4 • Core Rules of Probability
Define sample spaces, events, and the axioms governing probability assignments. These rules are the logical foundation for all distributional models covered next.
Lesson 5 • Continuous Probability Distributions
Apply normal, exponential, and uniform distributions to continuous measurements. These distributions appear throughout regression, hypothesis testing, and simulation.
Chapter 3HideHide detailsSee detailsStatistical Estimation and Sampling
Statistical Estimation and Sampling
Lesson 1 • Confidence Intervals for Means
Construct Z-based and t-based confidence intervals for population means. Correct interval interpretation prevents the common misconception of probability statements about fixed parameters.
Lesson 2 • Bootstrap and Resampling Methods
Generate empirical sampling distributions via bootstrap resampling without distributional assumptions. Bootstrap methods extend interval estimation to complex statistics lacking closed-form solutions.
Lesson 3 • The Central Limit Theorem in Practice
Demonstrate how sample means converge to normality regardless of population shape. This theorem justifies the normal-based inference methods used throughout the course.
Lesson 4 • Point Estimation Methods
Derive estimators using method of moments and maximum likelihood estimation. Understanding estimator derivation prepares learners to evaluate data model parameter outputs.
Lesson 5 • Confidence Intervals for Proportions
Build intervals for population proportions using normal approximation and exact methods. Proportion intervals are essential for A/B testing and survey analysis covered later.
Chapter 4HideHide detailsSee detailsHypothesis Testing Fundamentals
Hypothesis Testing Fundamentals
Lesson 1 • Logic of Hypothesis Testing
Define null and alternative hypotheses, significance levels, and the decision framework. This logic governs every formal test introduced in this and subsequent chapters.
Lesson 2 • Tests for Proportions and Variances
Conduct z-tests for proportions and chi-square and F-tests for variance comparisons. These tests extend hypothesis testing to categorical outcomes and variance equality checks.
Lesson 3 • Multiple Testing and Corrections
Control family-wise error rate and false discovery rate when conducting many simultaneous tests. Multiple testing corrections are mandatory in genomics, A/B testing, and feature selection.
Lesson 4 • Errors, Power, and Sample Size
Quantify Type I error, Type II error, and statistical power to design adequately powered studies. Power analysis determines sample size requirements before data collection.
Lesson 5 • One-Sample and Two-Sample t-Tests
Apply t-tests to compare a sample mean to a target or two group means to each other. These tests are the most common inferential tools in applied data science projects.
Chapter 5HideHide detailsSee detailsCorrelation, Regression, and Prediction
Correlation, Regression, and Prediction
Lesson 1 • Regression Diagnostics and Assumptions
Verify linearity, homoscedasticity, normality of residuals, and independence through diagnostic plots. Undetected violations invalidate inference and degrade prediction accuracy.
Lesson 2 • Simple Linear Regression
Estimate slope and intercept via ordinary least squares and interpret coefficients. Simple regression establishes the geometric and algebraic framework extended to multiple predictors.
Lesson 3 • Measuring Association Between Variables
Compute Pearson and Spearman correlations and test their significance. Correlation analysis motivates regression modeling and reveals multicollinearity risks.
Lesson 4 • Regularisation and Variable Selection
Apply Ridge, Lasso, and stepwise selection to manage overfitting and choose predictors. Regularisation bridges classical regression and machine learning model building.
Lesson 5 • Multiple Linear Regression
Extend OLS to multiple predictors, interpret partial slopes, and assess model fit. Multiple regression is the backbone of explanatory modeling in business and science.
Chapter 6HideHide detailsSee detailsCategorical Data Analysis
Categorical Data Analysis
Lesson 1 • Contingency Tables and Chi-Square Tests
Construct two-way tables and test independence using chi-square statistics. Contingency analysis is foundational for survey data, A/B tests, and feature association screening.
Lesson 2 • Multinomial and Ordinal Regression
Extend logistic regression to outcomes with more than two unordered or ordered categories. These models handle survey ratings, product tiers, and multi-class classification tasks.
Lesson 3 • Model Evaluation for Classification
Assess logistic regression performance using confusion matrices, ROC curves, and AUC. Proper evaluation metrics prevent misleading accuracy claims on imbalanced datasets.
Lesson 4 • Binary Logistic Regression
Model binary outcomes using the logit link, maximum likelihood, and log-odds interpretation. Logistic regression is the primary classification baseline in data science workflows.
Lesson 5 • Measures of Association for Categorical Data
Quantify strength of association using odds ratios, relative risk, and phi coefficients. These measures translate statistical significance into practically meaningful effect estimates.
Chapter 7HideHide detailsSee detailsAnalysis of Variance and Experimental Design
Analysis of Variance and Experimental Design
Lesson 1 • Post-Hoc Multiple Comparisons
Identify which group pairs differ after a significant ANOVA using controlled comparison procedures. Post-hoc tests prevent inflated error rates from unplanned pairwise comparisons.
Lesson 2 • Principles of Experimental Design
Apply randomisation, replication, and blocking to eliminate confounding and maximise precision. Sound design is more powerful than any post-hoc statistical adjustment.
Lesson 3 • Repeated Measures and Mixed Designs
Account for within-subject correlation in longitudinal and crossover experiments. Repeated measures designs increase power by removing individual difference variance.
Lesson 4 • Two-Way and Factorial ANOVA
Analyse main effects and interactions between two or more categorical factors simultaneously. Factorial designs reveal synergistic effects invisible to one-factor-at-a-time experiments.
Lesson 5 • One-Way ANOVA
Partition total variance into between-group and within-group components to test group mean equality. ANOVA generalises the two-sample t-test to any number of independent groups.
Chapter 8HideHide detailsSee detailsBayesian Statistics and Advanced Inference
Bayesian Statistics and Advanced Inference
Lesson 1 • Markov Chain Monte Carlo Methods
Sample from intractable posteriors using Metropolis-Hastings and Gibbs sampling algorithms. MCMC enables Bayesian inference for complex hierarchical and nonconjugate models.
Lesson 2 • Bayesian Model Comparison
Compare models using Bayes factors, WAIC, and leave-one-out cross-validation. Bayesian model selection penalises complexity without requiring a fixed significance threshold.
Lesson 3 • Bayesian Inference Framework
Formalise prior, likelihood, and posterior and contrast Bayesian with frequentist logic. This framework enables probabilistic statements about parameters rather than about data.
Lesson 4 • Bayesian Estimation and Credible Intervals
Summarise posteriors using MAP estimates, posterior means, and credible intervals. Credible intervals have the direct probability interpretation that confidence intervals lack.
Lesson 5 • Hierarchical and Multilevel Models
Model grouped data with partial pooling to borrow strength across units. Hierarchical models outperform both complete pooling and no-pooling strategies in nested data.
Your valid completion certificate
This course is for you:
Aspiring data scientists: need statistical depth beyond coding bootcamp basics.
Business analysts: want to move from Excel summaries to rigorous inferential methods.
Software engineers: transitioning into machine learning roles requiring probabilistic reasoning.
Graduate students: need a practical statistics reference alongside their research coursework.
Marketing analysts: ready to run valid A/B tests and interpret results independently.
Academic researchers: seeking to strengthen quantitative methods before publishing findings.
What our students say
Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.

I like the content and the way videos are presented and transcribed, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.

Top qualifications
FAQ
Who is Dedika?
Is the certificate valid in South Africa?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















