
Basic Data Science Course
Launch your data science career with a comprehensive, hands-on programme that takes you from Python basics to deploying machine learning models. You will master data cleaning, statistical analysis, and core algorithms while working with real datasets throughout. This course gives you the practical skills employers are actively hiring for right now.
What you will learn:
You will build a solid foundation in Python programming, statistics, and the full data science workflow. You will learn to clean and preprocess messy datasets, perform exploratory data analysis, and apply supervised and unsupervised machine learning algorithms. The course also covers SQL for database querying, natural language processing, time series forecasting, and responsible AI practices. By the end, you will know how to evaluate models rigorously, deploy them as REST APIs, and communicate your results clearly to business stakeholders. You will finish with a portfolio-ready project that demonstrates end-to-end data science expertise.
How you study in practice Basic Data Science Course
How you practise Basic Data Science Course
For companies looking to train their teams
With Dedika for businesses, the course includes exercises and examples tailored to your company and its specific needs.
Course content
8 Chapters • 36 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsFoundations of Data Science
Foundations of Data Science
Lesson 1 • Types of Data and Variables
Classifies data by structure, format, and measurement scale. Enables correct method selection in later analytical chapters.
Lesson 2 • What Data Science Is
Defines data science, distinguishes it from related fields, and maps its components. Grounds all subsequent chapters in shared terminology.
Lesson 3 • The Data Science Workflow
Introduces the standard project lifecycle from problem framing to deployment. Provides a mental model students apply throughout the course.
Lesson 4 • Setting Up the Work Environment
Guides installation and configuration of Python, Jupyter, and essential libraries. Ensures every student has a functional environment before coding begins.
Chapter 2HideHide detailsSee detailsPython Programming for Data Science
Python Programming for Data Science
Lesson 1 • Working with NumPy Arrays
Introduces vectorized computation on numerical arrays. Underpins the performance-critical operations used in pandas and scikit-learn.
Lesson 2 • Control Flow and Functions
Teaches conditionals, loops, and reusable function design. Enables students to automate repetitive data tasks efficiently.
Lesson 3 • File I/O and Data Loading
Demonstrates reading and writing CSV, Excel, and JSON files. Prepares students to ingest real-world datasets in all subsequent projects.
Lesson 4 • Python Syntax and Data Types
Covers variables, operators, and built-in types used in data workflows. Forms the syntactic base for all subsequent Python-based work.
Lesson 5 • Data Manipulation with Pandas
Covers DataFrame creation, selection, filtering, and aggregation. Directly enables the data cleaning and exploration work in the next chapter.
Chapter 3HideHide detailsSee detailsExploratory Data Analysis
Exploratory Data Analysis
Lesson 1 • Descriptive Statistics
Quantifies central tendency, spread, and shape of distributions. Provides the numerical foundation for interpreting any dataset.
Lesson 2 • Data Visualization Fundamentals
Introduces Matplotlib and Seaborn for creating standard chart types. Visualization skills are applied in every remaining chapter.
Lesson 3 • Data Profiling and Quality Assessment
Systematically audits datasets for missing values, duplicates, and anomalies. Directly motivates the data cleaning techniques in the next chapter.
Lesson 4 • Univariate and Bivariate Analysis
Examines single-variable distributions and pairwise relationships between variables. Builds intuition for feature selection in modeling chapters.
Chapter 4HideHide detailsSee detailsStatistical Foundations for Modeling
Statistical Foundations for Modeling
Lesson 1 • Correlation and Covariance
Quantifies linear and rank-based relationships between variables. Directly informs feature selection and multicollinearity diagnosis in regression models.
Lesson 2 • Probability Distributions
Introduces key discrete and continuous distributions used in data science. Distribution knowledge is essential for choosing correct models and loss functions.
Lesson 3 • Probability Essentials
Covers probability rules, conditional probability, and Bayes' theorem. These concepts underpin classification algorithms and uncertainty quantification.
Lesson 4 • Hypothesis Testing
Teaches null hypothesis framing, p-values, and common statistical tests. Enables rigorous comparison of model variants and business experiments.
Chapter 5HideHide detailsSee detailsData Cleaning and Preprocessing
Data Cleaning and Preprocessing
Lesson 1 • Handling Missing Data
Covers deletion, imputation, and indicator strategies for missing values. Correct handling prevents bias in downstream models.
Lesson 2 • Feature Scaling and Transformation
Applies normalisation, standardisation, and mathematical transforms to numeric features. Ensures distance-based and gradient-based algorithms converge correctly.
Lesson 3 • Building a Preprocessing Pipeline
Combines all cleaning steps into a reproducible scikit-learn Pipeline object. Prevents data leakage and streamlines model training in later chapters.
Lesson 4 • Outlier Detection and Treatment
Identifies outliers using statistical and visual methods and applies appropriate remedies. Protects model performance from extreme value distortion.
Lesson 5 • Feature Encoding
Converts categorical variables into numeric representations suitable for algorithms. Encoding choices directly affect model accuracy and interpretability.
Chapter 6HideHide detailsSee detailsSupervised Machine Learning
Supervised Machine Learning
Lesson 1 • Machine Learning Fundamentals
Defines supervised learning, the bias-variance tradeoff, and the train-test split paradigm. Establishes the conceptual framework for all modeling sections.
Lesson 2 • Classification Algorithms
Implements logistic regression, decision trees, and k-nearest neighbours for categorical targets. Builds the classification toolkit expanded in the ensemble section.
Lesson 3 • Hyperparameter Tuning
Applies grid search, random search, and Bayesian optimisation to maximise model performance. Tuning is the final step before model selection and deployment.
Lesson 4 • Regression Algorithms
Covers linear, polynomial, and regularised regression for continuous target prediction. Regression skills transfer directly to feature importance analysis.
Lesson 5 • Ensemble Methods
Introduces bagging, boosting, and stacking to improve predictive performance. Ensemble techniques consistently outperform single models in practice.
Chapter 7HideHide detailsSee detailsUnsupervised Learning and Dimensionality Reduction
Unsupervised Learning and Dimensionality Reduction
Lesson 1 • Principal Component Analysis
Reduces feature dimensionality by projecting data onto principal components. PCA improves visualisation and reduces noise before supervised modeling.
Lesson 2 • Manifold Learning Techniques
Applies t-SNE and UMAP for nonlinear dimensionality reduction and visualisation. These methods reveal cluster structure invisible to linear methods like PCA.
Lesson 3 • Clustering Algorithms
Covers k-means, hierarchical, and density-based clustering for grouping unlabelled observations. Clustering outputs feed directly into business segmentation use cases.
Lesson 4 • Clustering Evaluation
Introduces internal and external metrics for assessing cluster quality without labels. Rigorous evaluation prevents misinterpretation of unsupervised results.
Chapter 8HideHide detailsSee detailsModel Evaluation, Deployment, and Communication
Model Evaluation, Deployment, and Communication
Lesson 1 • Communicating Data Science Results
Structures findings into executive summaries and visual dashboards for non-technical audiences. Clear communication determines whether insights drive business decisions.
Lesson 2 • Advanced Model Evaluation
Covers ROC-AUC, precision-recall curves, and calibration for thorough model assessment. Rigorous evaluation prevents costly deployment of underperforming models.
Lesson 3 • Monitoring Models in Production
Detects data drift, concept drift, and performance degradation in deployed models. Ongoing monitoring ensures sustained model reliability after launch.
Lesson 4 • Model Interpretability
Applies SHAP values and permutation importance to explain model predictions. Interpretability builds stakeholder trust and satisfies accountability requirements.
Lesson 5 • Model Serialisation and Serving
Serialises trained models and wraps them in a REST API using Flask or FastAPI. Deployment skills bridge the gap between experimentation and production use.
Your valid completion certificate
This course is for you:
Career changers: seeking a structured path into the data science field.
Marketing analysts: wanting to move beyond dashboards into predictive modelling work.
Recent graduates: looking to add in-demand technical skills to their CVs.
Business professionals: needing to understand and contribute to data-driven decisions.
Hobbyists and self-learners: curious about how machine learning actually works in practice.
Aspiring data scientists: ready to commit to a rigorous, project-based learning experience.
What our students say
Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.

I like the content and the way videos are presented and transcribed, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.

Top qualifications
FAQ
Who is Dedika?
Is the certificate valid in South Africa?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















