
Analyzing Data Using R for Statistical and Predictive Modeling Training
Master the full data science pipeline in R — from wrangling raw data to deploying predictive models that drive real decisions. This comprehensive training covers statistical inference, machine learning, and professional reporting using industry-standard tools like tidyverse, XGBoost, and Shiny. Whether you're analysing business data or building forecasting systems, you'll gain the hands-on R skills that employers demand.
What you will learn:
Configure a professional R and RStudio environment with reproducible project workflows.
Build linear, logistic, and ensemble models to solve regression and classification problems.
Apply statistical hypothesis tests and confidence intervals to draw valid conclusions from data.
Engineer features and construct preprocessing pipelines that generalise across training and scoring datasets.
Create publication-quality visualisations and interactive Shiny dashboards for diverse stakeholder audiences.
Interpret and explain black-box models using SHAP values, partial dependence plots, and LIME.
How you study in practice Analyzing Data Using R for Statistical and Predictive Modeling Training
How you practise Analyzing Data Using R for Statistical and Predictive Modeling Training
For companies looking to train their teams
With Dedika for businesses, the course includes exercises and examples tailored to your company and its specific needs.
Course content
8 Chapters • 32 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsR Environment and Data Foundations
R Environment and Data Foundations
Lesson 1 • Importing and Exporting Data
Reads CSV, Excel, and database sources into R and writes results back. Connects raw data files to the analytical pipeline.
Lesson 2 • Installing and Configuring R and RStudio
Set up R, RStudio, and essential packages from scratch. Establishes the reproducible workspace used throughout the course.
Lesson 3 • Basic R Programming Constructs
Introduces control flow, functions, and vectorised operations. Enables automation of repetitive analytical tasks from the start.
Lesson 4 • Core R Data Structures
Covers vectors, matrices, lists, and data frames with creation and indexing. Provides the structural vocabulary for all subsequent data work.
Chapter 2HideHide detailsSee detailsData Wrangling with the Tidyverse
Data Wrangling with the Tidyverse
Lesson 1 • Reshaping Data with tidyr
Converts between wide and long formats and handles nested data. Prepares data structures required by statistical and visualisation functions.
Lesson 2 • Joining and Merging Datasets
Combines multiple tables using inner, left, right, and full joins. Enables integration of data from disparate sources into one analytical frame.
Lesson 3 • String and Date Manipulation
Cleans text with stringr and parses dates with lubridate. Addresses the most common data-quality issues in real-world datasets.
Lesson 4 • Manipulating Data with dplyr
Applies filter, select, mutate, group_by, and summarise to tabular data. Forms the primary toolkit for row- and column-level transformations.
Chapter 3HideHide detailsSee detailsExploratory Data Analysis and Visualisation
Exploratory Data Analysis and Visualisation
Lesson 1 • Identifying Outliers and Data Quality Issues
Detects anomalies using visual and statistical methods before modelling. Prevents biased estimates caused by unexamined data problems.
Lesson 2 • Descriptive Statistics in R
Computes measures of centre, spread, and shape for numeric and categorical variables. Grounds visual exploration in quantitative summaries.
Lesson 3 • Advanced Chart Types and Themes
Creates box plots, heat maps, and custom themes for professional presentation. Extends basic ggplot2 skills to complex multi-variable displays.
Lesson 4 • Building Plots with ggplot2
Constructs layered graphics using the grammar of graphics framework. Provides the primary visualisation engine for all analytical outputs.
Chapter 4HideHide detailsSee detailsStatistical Inference and Hypothesis Testing
Statistical Inference and Hypothesis Testing
Lesson 1 • Probability Distributions in R
Generates, samples, and visualises common distributions using R's built-in functions. Builds the probabilistic foundation required for all inferential work.
Lesson 2 • Confidence Intervals and Sampling
Constructs confidence intervals for means, proportions, and differences. Connects sample statistics to population parameters with quantified uncertainty.
Lesson 3 • Non-Parametric and Resampling Tests
Applies Wilcoxon, Kruskal-Wallis, and permutation tests when assumptions fail. Extends inferential capability to non-normal and small-sample data.
Lesson 4 • Parametric Hypothesis Tests
Executes t-tests, ANOVA, and chi-square tests with correct assumptions. Enables evidence-based decisions on group differences and associations.
Chapter 5HideHide detailsSee detailsLinear Regression Modelling
Linear Regression Modelling
Lesson 1 • Simple Linear Regression
Fits a single-predictor OLS model and interprets slope and intercept. Establishes the regression framework extended in all subsequent modelling chapters.
Lesson 2 • Multiple Linear Regression
Extends regression to multiple predictors with interaction and polynomial terms. Captures complex real-world relationships within the linear framework.
Lesson 3 • Variable Selection and Regularisation
Applies stepwise selection, ridge, and lasso to reduce overfitting. Produces parsimonious models that generalise beyond the training sample.
Lesson 4 • Regression Diagnostics
Checks linearity, homoscedasticity, normality, and independence of residuals. Validates model assumptions before drawing inferential conclusions.
Chapter 6HideHide detailsSee detailsClassification and Logistic Regression
Classification and Logistic Regression
Lesson 1 • Logistic Regression Fundamentals
Fits binary logistic models with glm() and interprets log-odds and probabilities. Provides the baseline classifier for all subsequent classification work.
Lesson 2 • Resampling and Cross-Validation
Implements k-fold and stratified cross-validation to estimate generalization error. Prevents overfitting by separating model selection from final evaluation.
Lesson 3 • Model Evaluation for Classifiers
Measures classifier performance with confusion matrices, ROC curves, and AUC. Connects model outputs to business decision criteria.
Lesson 4 • Multinomial and Ordinal Logistic Models
Extends logistic regression to outcomes with three or more categories. Handles ordered and unordered multi-class prediction tasks.
Chapter 7HideHide detailsSee detailsTree-Based and Ensemble Methods
Tree-Based and Ensemble Methods
Lesson 1 • Decision Trees with rpart
Grows, prunes, and visualizes classification and regression trees. Introduces the recursive partitioning logic underlying all tree-based ensembles.
Lesson 2 • Random Forests
Builds bagged ensembles of decorrelated trees using the randomForest package. Achieves robust predictions through variance reduction via averaging.
Lesson 3 • Model Tuning with caret and tidymodels
Automates hyperparameter search and model comparison within a unified framework. Standardizes the tuning workflow across all model types.
Lesson 4 • Gradient Boosting with XGBoost
Trains sequential boosted trees with xgboost and tunes key hyperparameters. Delivers state-of-the-art predictive accuracy on structured data.
Chapter 8HideHide detailsSee detailsAdvanced Predictive Modelling and Deployment
Advanced Predictive Modelling and Deployment
Lesson 1 • Reporting and Deploying Models
Packages models as R Markdown reports and Plumber API endpoints for production use. Closes the gap between analytical development and operational delivery.
Lesson 2 • Model Interpretability and Explainability
Applies SHAP values, partial dependence plots, and LIME to explain black-box models. Builds stakeholder trust by linking predictions to input drivers.
Lesson 3 • Time Series Forecasting Basics
Decomposes time series and fits ARIMA and exponential smoothing models. Extends predictive modelling to ordered temporal data structures.
Lesson 4 • Feature Engineering and Preprocessing Pipelines
Encodes categoricals, scales numerics, and handles missingness inside reproducible recipes. Ensures consistent transformations across training and scoring data.
Your valid completion certificate
This course is for you:
Business analysts who want to move beyond spreadsheet-based reporting.
Graduate students in social sciences needing rigorous quantitative modelling skills.
Data professionals transitioning from Python who want R fluency.
Healthcare researchers looking to apply predictive models to clinical datasets.
Marketing analysts ready to build forecasting tools from their own data.
Career changers entering data science without a formal computer science background.
What our students say
Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.

I like the content and the way videos are presented and transcribed, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.

Top qualifications
FAQ
Who is Dedika?
Is the certificate valid in South Africa?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















