
Applied Biostatistics for Data Science Course
Master the statistical methods that drive real biomedical research, from probability and regression to survival analysis and causal inference. This course bridges rigorous theory with hands-on R and Python workflows, preparing you to analyze complex health datasets with confidence. Whether you work in clinical research, epidemiology, or health data science, you will leave with skills that translate directly into practice.
What you will learn:
You will build a complete biostatistics skill set starting from data quality and descriptive statistics, then advancing through hypothesis testing, group comparisons, and regression modeling for continuous and binary outcomes. You will learn to analyze time-to-event data using Kaplan-Meier estimators and Cox proportional hazards models. Epidemiological study designs, measures of association, and bias assessment are covered in depth. Advanced topics include mixed-effects models, Bayesian inference, and machine learning integration within a rigorous statistical framework. Throughout the course, you will apply these methods in R and Python, producing reproducible, publication-ready analyses.
How you study in a practical way Applied Biostatistics for Data Science Course
How you practice Applied Biostatistics for Data Science Course
For companies who want to train their team
With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.
Course content
8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsFoundations of Biostatistics and Data
Foundations of Biostatistics and Data
Lesson 1 • Variable Types and Measurement Scales
Distinguishes nominal, ordinal, interval, and ratio scales with biomedical examples. Correct classification drives every downstream analytical decision in the chapter.
Lesson 2 • Role of Statistics in Biomedicine
Defines biostatistics and its function in clinical and public health research. Establishes why rigorous quantitative reasoning is essential before any analysis begins.
Lesson 3 • Data Visualization Principles
Introduces histograms, box plots, and scatter plots as diagnostic and communication tools. Proper visualization reveals distributional assumptions before formal testing.
Lesson 4 • Data Quality and Preprocessing
Addresses missing data patterns, coding errors, and unit inconsistencies common in health datasets. Clean data is a prerequisite for valid inference in all subsequent chapters.
Lesson 5 • Descriptive Statistics for Health Data
Covers central tendency, dispersion, and shape metrics tailored to skewed biomedical distributions. Connects summary statistics to meaningful clinical interpretation.
Chapter 2HideHide detailsSee detailsProbability and Distributions in Health
Probability and Distributions in Health
Lesson 1 • Bayes' Theorem in Diagnostics
Applies Bayes' theorem to update disease probability given test results. Connects prior prevalence to posterior probability, foundational for diagnostic test evaluation.
Lesson 2 • Core Probability Rules
Covers addition, multiplication, and complement rules with disease prevalence examples. These rules underpin sensitivity, specificity, and predictive value calculations ahead.
Lesson 3 • Continuous Probability Distributions
Covers normal, log-normal, exponential, and Weibull distributions relevant to biomedical measurements. Selecting the right distribution ensures valid parametric inference later.
Lesson 4 • Discrete Probability Distributions
Introduces binomial, Poisson, and negative binomial distributions for count-based health outcomes. Each distribution is matched to realistic epidemiological scenarios.
Lesson 5 • Sampling Distributions and the CLT
Explains how sample statistics vary across repeated samples and why the central limit theorem enables normal-based inference. Bridges probability theory to hypothesis testing.
Chapter 3HideHide detailsSee detailsEstimation and Hypothesis Testing
Estimation and Hypothesis Testing
Lesson 1 • Multiple Testing and Error Control
Addresses family-wise error rate inflation and correction strategies for multi-endpoint studies. Proper correction prevents false discovery in high-dimensional biomedical data.
Lesson 2 • Logic of Hypothesis Testing
Defines null and alternative hypotheses, test statistics, and decision rules. Establishes the conceptual framework applied in every subsequent testing procedure.
Lesson 3 • Point Estimation and Confidence Intervals
Covers maximum likelihood and method-of-moments estimators alongside CI construction. Accurate interval estimation is the foundation of all inferential reporting.
Lesson 4 • Tests for Means and Proportions
Applies z-tests, t-tests, and proportion tests to common biomedical comparisons. Correct test selection depends on variable type and sample size covered earlier.
Lesson 5 • Power Analysis and Sample Size
Quantifies the relationship among effect size, alpha, power, and sample size for study planning. Underpowered studies waste resources and produce unreliable conclusions.
Chapter 4HideHide detailsSee detailsComparing Groups: Parametric and Nonparametric
Comparing Groups: Parametric and Nonparametric
Lesson 1 • One-Way ANOVA and Post-Hoc Tests
Extends two-group comparison to multiple independent groups using ANOVA. Post-hoc procedures identify which specific group pairs differ after a significant F-test.
Lesson 2 • Nonparametric Tests for Continuous Data
Covers Mann-Whitney U, Kruskal-Wallis, and Wilcoxon signed-rank tests for non-normal outcomes. These rank-based methods are widely used in small clinical samples.
Lesson 3 • Assumptions Checking Before Testing
Covers normality tests, homogeneity of variance, and independence checks as prerequisites. Assumption violations guide the choice between parametric and nonparametric alternatives.
Lesson 4 • Two-Way and Repeated-Measures ANOVA
Introduces factorial designs and within-subject repeated measures for longitudinal biomedical data. Interaction terms reveal how treatment effects vary across subgroups.
Lesson 5 • Chi-Square and Fisher's Exact Tests
Applies chi-square goodness-of-fit and independence tests to categorical health outcomes. Fisher's exact test handles small expected cell counts common in clinical tables.
Chapter 5HideHide detailsSee detailsEpidemiological Study Design and Measures
Epidemiological Study Design and Measures
Lesson 1 • Bias, Confounding, and Validity
Identifies selection bias, information bias, and confounding as threats to internal validity. Recognizing bias sources is essential for critically appraising published biomedical studies.
Lesson 2 • Randomized Controlled Trial Design
Covers randomization, allocation concealment, blinding, and intention-to-treat analysis in RCTs. RCTs provide the strongest evidence for causal inference when properly designed.
Lesson 3 • Observational Study Designs
Contrasts cross-sectional, cohort, and case-control designs by temporality and data structure. Design choice determines which effect measures are estimable and which biases arise.
Lesson 4 • Measures of Association and Impact
Computes risk ratio, odds ratio, rate ratio, and attributable risk for each study design. Impact measures quantify the public health burden attributable to a specific exposure.
Lesson 5 • Measures of Disease Frequency
Defines prevalence, incidence proportion, and incidence rate with denominator considerations. Correct frequency measures are prerequisites for valid association estimation.
Chapter 6HideHide detailsSee detailsRegression Modeling for Health Outcomes
Regression Modeling for Health Outcomes
Lesson 1 • Simple and Multiple Linear Regression
Derives OLS estimates and extends to multiple predictors for continuous outcomes like blood pressure. Regression coefficients quantify adjusted associations controlling for covariates.
Lesson 2 • Model Building and Variable Selection
Compares purposeful selection, stepwise, and penalized methods for covariate inclusion. Overfitting and underfitting are balanced through cross-validation and information criteria.
Lesson 3 • Confounding, Interaction, and Effect Modification
Distinguishes confounding from effect modification and shows how to model both in regression. Interaction terms reveal heterogeneous treatment effects across patient subgroups.
Lesson 4 • Logistic Regression for Binary Outcomes
Models the log-odds of a binary event such as disease presence using maximum likelihood. Odds ratios from logistic regression are the standard effect measure in epidemiology.
Lesson 5 • Regression Diagnostics and Remedies
Uses residual plots, influence statistics, and leverage measures to detect model violations. Addressing violations ensures valid inference from the fitted regression model.
Chapter 7HideHide detailsSee detailsSurvival Analysis and Time-to-Event Data
Survival Analysis and Time-to-Event Data
Lesson 1 • Log-Rank and Weighted Tests
Compares survival curves across groups using log-rank and Wilcoxon-Gehan weighted tests. Test choice depends on whether hazard differences are early, late, or proportional.
Lesson 2 • Parametric Survival Models
Fits exponential, Weibull, and log-normal accelerated failure time models to survival data. Parametric models enable extrapolation beyond observed follow-up for health economic models.
Lesson 3 • Cox Proportional Hazards Regression
Models covariate-adjusted hazard ratios using the semi-parametric Cox model. The proportional hazards assumption must be verified before interpreting any HR estimate.
Lesson 4 • Concepts of Survival and Censoring
Defines survival time, the hazard function, and right, left, and interval censoring mechanisms. Understanding censoring is essential before applying any survival estimator.
Lesson 5 • Kaplan-Meier Estimator
Constructs nonparametric survival curves and computes median survival with confidence bands. KM curves are the standard visual summary for time-to-event data in clinical reports.
Chapter 8HideHide detailsSee detailsAdvanced Biostatistical Methods in Data Science
Advanced Biostatistical Methods in Data Science
Lesson 1 • Generalized Linear and Mixed Models
Unifies regression for non-normal outcomes via link functions and exponential family distributions. GLMMs extend this to clustered count and binary outcomes in health data.
Lesson 2 • High-Dimensional and Omics Data Analysis
Addresses dimensionality reduction, penalized regression, and multiple testing in genomic and proteomic datasets. These methods handle the p >> n problem common in omics research.
Lesson 3 • Bayesian Inference for Biomedical Data
Introduces prior specification, likelihood, and posterior computation via MCMC for health models. Bayesian methods naturally incorporate prior clinical knowledge and quantify uncertainty fully.
Lesson 4 • Causal Inference Methods
Covers propensity score methods, instrumental variables, and difference-in-differences for causal estimation from observational health data. Causal framing goes beyond association.
Lesson 5 • Linear Mixed-Effects Models
Extends regression to clustered and longitudinal data using random intercepts and slopes. Mixed models account for within-subject correlation that violates standard regression assumptions.
Your valid completion certificate
This course is for you:
Epidemiologists: wanting to move beyond descriptive summaries into rigorous modeling.
Clinical research coordinators: ready to understand the statistics behind their trial data.
Health data analysts: seeking a structured foundation in biomedical statistical reasoning.
Biomedical graduate students: needing practical skills alongside their academic coursework.
Public health professionals: aiming to evaluate study evidence with greater analytical confidence.
Data scientists: transitioning into healthcare and needing domain-specific statistical grounding.
What our students say
Your classes are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of my interest without needing to change platforms... I thank you for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can switch chapters and skip content I don't need.

I like the content and the way videos are presented and transcribed, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.

Top trainings
FAQs
Who is Dedika?
Is the certificate valid in the Philippines?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















