Choose your language
Analyzing Data Using R for Statistical and Predictive Modeling Training
More than 2 million students worldwide

Analyzing Data Using R for Statistical and Predictive Modeling Training

Master the full data science pipeline in R — from wrangling raw data to deploying predictive models that drive real decisions. This comprehensive training covers statistical inference, machine learning, and professional reporting using industry-standard tools like tidyverse, XGBoost, and Shiny. Whether you're analysing business data or building forecasting systems, you'll gain the hands-on R skills that employers demand.

Dedika for businesses

What you will learn:

  • Configure a professional R and RStudio environment with reproducible project workflows.

  • Build linear, logistic, and ensemble models to solve regression and classification problems.

  • Apply statistical hypothesis tests and confidence intervals to draw valid conclusions from data.

  • Engineer features and construct preprocessing pipelines that generalise across training and scoring datasets.

  • Create publication-quality visualisations and interactive Shiny dashboards for diverse stakeholder audiences.

  • Interpret and explain black-box models using SHAP values, partial dependence plots, and LIME.

How you study in practice Analyzing Data Using R for Statistical and Predictive Modeling Training

How you practise Analyzing Data Using R for Statistical and Predictive Modeling Training

For companies looking to train their teams

With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.

Click here

Course content

8 Chapters • 32 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

R Environment and Data Foundations

  • Lesson 1 • Importing and Exporting Data

    Reads CSV, Excel, and database sources into R and writes results back. Connects raw data files to the analytical pipeline.

  • Lesson 2 • Installing and Configuring R and RStudio

    Set up R, RStudio, and essential packages from scratch. Establishes the reproducible workspace used throughout the course.

  • Lesson 3 • Basic R Programming Constructs

    Introduces control flow, functions, and vectorised operations. Enables automation of repetitive analytical tasks from the start.

  • Lesson 4 • Core R Data Structures

    Covers vectors, matrices, lists, and data frames with creation and indexing. Provides the structural vocabulary for all subsequent data work.

Chapter 2See details

Data Wrangling with the Tidyverse

  • Lesson 1 • Reshaping Data with tidyr

    Converts between wide and long formats and handles nested data. Prepares data structures required by statistical and visualisation functions.

  • Lesson 2 • Joining and Merging Datasets

    Combines multiple tables using inner, left, right, and full joins. Enables integration of data from disparate sources into one analytical frame.

  • Lesson 3 • String and Date Manipulation

    Cleans text with stringr and parses dates with lubridate. Addresses the most common data-quality issues in real-world datasets.

  • Lesson 4 • Manipulating Data with dplyr

    Applies filter, select, mutate, group_by, and summarise to tabular data. Forms the primary toolkit for row- and column-level transformations.

Chapter 3See details

Exploratory Data Analysis and Visualisation

  • Lesson 1 • Identifying Outliers and Data Quality Issues

    Detects anomalies using visual and statistical methods before modelling. Prevents biased estimates caused by unexamined data problems.

  • Lesson 2 • Descriptive Statistics in R

    Computes measures of centre, spread, and shape for numeric and categorical variables. Grounds visual exploration in quantitative summaries.

  • Lesson 3 • Advanced Chart Types and Themes

    Creates box plots, heat maps, and custom themes for professional presentation. Extends basic ggplot2 skills to complex multi-variable displays.

  • Lesson 4 • Building Plots with ggplot2

    Constructs layered graphics using the grammar of graphics framework. Provides the primary visualisation engine for all analytical outputs.

Chapter 4See details

Statistical Inference and Hypothesis Testing

  • Lesson 1 • Probability Distributions in R

    Generates, samples, and visualises common distributions using R's built-in functions. Builds the probabilistic foundation required for all inferential work.

  • Lesson 2 • Confidence Intervals and Sampling

    Constructs confidence intervals for means, proportions, and differences. Connects sample statistics to population parameters with quantified uncertainty.

  • Lesson 3 • Non-Parametric and Resampling Tests

    Applies Wilcoxon, Kruskal-Wallis, and permutation tests when assumptions fail. Extends inferential capability to non-normal and small-sample data.

  • Lesson 4 • Parametric Hypothesis Tests

    Executes t-tests, ANOVA, and chi-square tests with correct assumptions. Enables evidence-based decisions on group differences and associations.

Chapter 5See details

Linear Regression Modelling

  • Lesson 1 • Simple Linear Regression

    Fits a single-predictor OLS model and interprets slope and intercept. Establishes the regression framework extended in all subsequent modelling chapters.

  • Lesson 2 • Multiple Linear Regression

    Extends regression to multiple predictors with interaction and polynomial terms. Captures complex real-world relationships within the linear framework.

  • Lesson 3 • Variable Selection and Regularisation

    Applies stepwise selection, ridge, and lasso to reduce overfitting. Produces parsimonious models that generalise beyond the training sample.

  • Lesson 4 • Regression Diagnostics

    Checks linearity, homoscedasticity, normality, and independence of residuals. Validates model assumptions before drawing inferential conclusions.

Chapter 6See details

Classification and Logistic Regression

  • Lesson 1 • Logistic Regression Fundamentals

    Fits binary logistic models with glm() and interprets log-odds and probabilities. Provides the baseline classifier for all subsequent classification work.

  • Lesson 2 • Resampling and Cross-Validation

    Implements k-fold and stratified cross-validation to estimate generalization error. Prevents overfitting by separating model selection from final evaluation.

  • Lesson 3 • Model Evaluation for Classifiers

    Measures classifier performance with confusion matrices, ROC curves, and AUC. Connects model outputs to business decision criteria.

  • Lesson 4 • Multinomial and Ordinal Logistic Models

    Extends logistic regression to outcomes with three or more categories. Handles ordered and unordered multi-class prediction tasks.

Chapter 7See details

Tree-Based and Ensemble Methods

  • Lesson 1 • Decision Trees with rpart

    Grows, prunes, and visualizes classification and regression trees. Introduces the recursive partitioning logic underlying all tree-based ensembles.

  • Lesson 2 • Random Forests

    Builds bagged ensembles of decorrelated trees using the randomForest package. Achieves robust predictions through variance reduction via averaging.

  • Lesson 3 • Model Tuning with caret and tidymodels

    Automates hyperparameter search and model comparison within a unified framework. Standardizes the tuning workflow across all model types.

  • Lesson 4 • Gradient Boosting with XGBoost

    Trains sequential boosted trees with xgboost and tunes key hyperparameters. Delivers state-of-the-art predictive accuracy on structured data.

Chapter 8See details

Advanced Predictive Modelling and Deployment

  • Lesson 1 • Reporting and Deploying Models

    Packages models as R Markdown reports and Plumber API endpoints for production use. Closes the gap between analytical development and operational delivery.

  • Lesson 2 • Model Interpretability and Explainability

    Applies SHAP values, partial dependence plots, and LIME to explain black-box models. Builds stakeholder trust by linking predictions to input drivers.

  • Lesson 3 • Time Series Forecasting Basics

    Decomposes time series and fits ARIMA and exponential smoothing models. Extends predictive modelling to ordered temporal data structures.

  • Lesson 4 • Feature Engineering and Preprocessing Pipelines

    Encodes categoricals, scales numerics, and handles missingness inside reproducible recipes. Ensures consistent transformations across training and scoring data.

Certification

Your valid completion certificate

This course is for you:

  • Business analysts who want to move beyond spreadsheet-based reporting.

  • Graduate students in social sciences needing rigorous quantitative modelling skills.

  • Data professionals transitioning from Python who want R fluency.

  • Healthcare researchers looking to apply predictive models to clinical datasets.

  • Marketing analysts ready to build forecasting tools from their own data.

  • Career changers entering data science without a formal computer science background.

What our students say

Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of my interest without needing to change platforms... I thank you for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the way videos are presented and transcribed, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast, simple to use. The diversity of content and complementary videos help a lot with learning.
André Felipe
André FelipePrompt Engineering Student

Top qualifications

FAQs

Who is Dedika?

Is the certificate valid in Zimbabwe?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course