
Analyzing Data Using R for Statistical and Predictive Modeling Training
Master the full data science pipeline in R — from wrangling raw data to deploying predictive models that drive real decisions. This comprehensive training covers statistical inference, machine learning, and professional reporting using industry-standard tools like tidyverse, XGBoost, and Shiny. Whether you're analyzing business data or building forecasting systems, you'll gain the hands-on R skills that employers demand.
What you will learn:
Configure a professional R and RStudio environment with reproducible project workflows.
Build linear, logistic, and ensemble models to solve regression and classification problems.
Apply statistical hypothesis tests and confidence intervals to draw valid conclusions from data.
Engineer features and construct preprocessing pipelines that generalize across training and scoring datasets.
Create publication-quality visualizations and interactive Shiny dashboards for diverse stakeholder audiences.
Interpret and explain black-box models using SHAP values, partial dependence plots, and LIME.
How you study in practice Analyzing Data Using R for Statistical and Predictive Modeling Training
How you practice Analyzing Data Using R for Statistical and Predictive Modeling Training
For companies that want to train their team
With Dedika for Business, the course includes exercises and examples tailored to your own business and the way your company needs.
Course content
8 Chapters • 32 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsR Environment and Data Foundations
R Environment and Data Foundations
Lesson 1 • Importing and Exporting Data
Reads CSV, Excel, and database sources into R and writes results back. Connects raw data files to the analytical pipeline.
Lesson 2 • Installing and Configuring R and RStudio
Set up R, RStudio, and essential packages from scratch. Establishes the reproducible workspace used throughout the course.
Lesson 3 • Basic R Programming Constructs
Introduces control flow, functions, and vectorized operations. Enables automation of repetitive analytical tasks from the start.
Lesson 4 • Core R Data Structures
Covers vectors, matrices, lists, and data frames with creation and indexing. Provides the structural vocabulary for all subsequent data work.
Chapter 2HideHide detailsSee detailsData Wrangling with the Tidyverse
Data Wrangling with the Tidyverse
Lesson 1 • Reshaping Data with tidyr
Converts between wide and long formats and handles nested data. Prepares data structures required by statistical and visualization functions.
Lesson 2 • Joining and Merging Datasets
Combines multiple tables using inner, left, right, and full joins. Enables integration of data from disparate sources into one analytical frame.
Lesson 3 • String and Date Manipulation
Cleans text with stringr and parses dates with lubridate. Addresses the most common data-quality issues in real-world datasets.
Lesson 4 • Manipulating Data with dplyr
Applies filter, select, mutate, group_by, and summarize to tabular data. Forms the primary toolkit for row- and column-level transformations.
Chapter 3HideHide detailsSee detailsExploratory Data Analysis and Visualization
Exploratory Data Analysis and Visualization
Lesson 1 • Identifying Outliers and Data Quality Issues
Detects anomalies using visual and statistical methods before modeling. Prevents biased estimates caused by unexamined data problems.
Lesson 2 • Descriptive Statistics in R
Computes measures of center, spread, and shape for numeric and categorical variables. Grounds visual exploration in quantitative summaries.
Lesson 3 • Advanced Chart Types and Themes
Creates box plots, heat maps, and custom themes for professional presentation. Extends basic ggplot2 skills to complex multi-variable displays.
Lesson 4 • Building Plots with ggplot2
Constructs layered graphics using the grammar of graphics framework. Provides the primary visualization engine for all analytical outputs.
Chapter 4HideHide detailsSee detailsStatistical Inference and Hypothesis Testing
Statistical Inference and Hypothesis Testing
Lesson 1 • Probability Distributions in R
Generates, samples, and visualizes common distributions using R's built-in functions. Builds the probabilistic foundation required for all inferential work.
Lesson 2 • Confidence Intervals and Sampling
Constructs confidence intervals for means, proportions, and differences. Connects sample statistics to population parameters with quantified uncertainty.
Lesson 3 • Non-Parametric and Resampling Tests
Applies Wilcoxon, Kruskal-Wallis, and permutation tests when assumptions fail. Extends inferential capability to non-normal and small-sample data.
Lesson 4 • Parametric Hypothesis Tests
Executes t-tests, ANOVA, and chi-square tests with correct assumptions. Enables evidence-based decisions on group differences and associations.
Chapter 5HideHide detailsSee detailsLinear Regression Modeling
Linear Regression Modeling
Lesson 1 • Simple Linear Regression
Fits a single-predictor OLS model and interprets slope and intercept. Establishes the regression framework extended in all subsequent modeling chapters.
Lesson 2 • Multiple Linear Regression
Extends regression to multiple predictors with interaction and polynomial terms. Captures complex real-world relationships within the linear framework.
Lesson 3 • Variable Selection and Regularization
Applies stepwise selection, ridge, and lasso to reduce overfitting. Produces parsimonious models that generalize beyond the training sample.
Lesson 4 • Regression Diagnostics
Checks linearity, homoscedasticity, normality, and independence of residuals. Validates model assumptions before drawing inferential conclusions.
Chapter 6HideHide detailsSee detailsClassification and Logistic Regression
Classification and Logistic Regression
Lesson 1 • Logistic Regression Fundamentals
Fits binary logistic models with glm() and interprets log-odds and probabilities. Provides the baseline classifier for all subsequent classification work.
Lesson 2 • Resampling and Cross-Validation
Implements k-fold and stratified cross-validation to estimate generalization error. Prevents overfitting by separating model selection from final evaluation.
Lesson 3 • Model Evaluation for Classifiers
Measures classifier performance with confusion matrices, ROC curves, and AUC. Connects model outputs to business decision criteria.
Lesson 4 • Multinomial and Ordinal Logistic Models
Extends logistic regression to outcomes with three or more categories. Handles ordered and unordered multi-class prediction tasks.
Chapter 7HideHide detailsSee detailsTree-Based and Ensemble Methods
Tree-Based and Ensemble Methods
Lesson 1 • Decision Trees with rpart
Grows, prunes, and visualizes classification and regression trees. Introduces the recursive partitioning logic underlying all tree-based ensembles.
Lesson 2 • Random Forests
Builds bagged ensembles of decorrelated trees using the randomForest package. Achieves robust predictions through variance reduction via averaging.
Lesson 3 • Model Tuning with caret and tidymodels
Automates hyperparameter search and model comparison within a unified framework. Standardizes the tuning workflow across all model types.
Lesson 4 • Gradient Boosting with XGBoost
Trains sequential boosted trees with xgboost and tunes key hyperparameters. Delivers state-of-the-art predictive accuracy on structured data.
Chapter 8HideHide detailsSee detailsAdvanced Predictive Modeling and Deployment
Advanced Predictive Modeling and Deployment
Lesson 1 • Reporting and Deploying Models
Packages models as R Markdown reports and Plumber API endpoints for production use. Closes the gap between analytical development and operational delivery.
Lesson 2 • Model Interpretability and Explainability
Applies SHAP values, partial dependence plots, and LIME to explain black-box models. Builds stakeholder trust by linking predictions to input drivers.
Lesson 3 • Time Series Forecasting Basics
Decomposes time series and fits ARIMA and exponential smoothing models. Extends predictive modeling to ordered temporal data structures.
Lesson 4 • Feature Engineering and Preprocessing Pipelines
Encodes categoricals, scales numerics, and handles missingness inside reproducible recipes. Ensures consistent transformations across training and scoring data.
Your valid completion certificate
This course is for you:
Business analysts who want to move beyond spreadsheet-based reporting.
Graduate students in social sciences needing rigorous quantitative modeling skills.
Data professionals transitioning from Python who want R fluency.
Healthcare researchers looking to apply predictive models to clinical datasets.
Marketing analysts ready to build forecasting tools from their own data.
Career changers entering data science without a formal computer science background.
What our students say
Your classes are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of my interest without needing to switch platforms... I thank you for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can switch chapters and skip content I don't need.

I like the content and the presentation style and video transcription, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.

Top trainings
FAQ
Who is Dedika?
Is the certificate valid in the United States?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















