Choose your language
Data Science Foundation Course
More than 2 million students worldwide

Data Science Foundation Course

Launch your data science career with a comprehensive program that takes you from Python basics to deployed machine learning models. You'll master data wrangling, statistical analysis, and model building using industry-standard tools. This course gives you the practical skills employers are actively hiring for right now.

Dedika for businesses

What you will learn:

You will build a complete data science skill set, starting with Python programming and progressing through data cleaning, statistical foundations, and exploratory analysis. You will train and evaluate supervised machine learning models using scikit-learn, then extend your skills into unsupervised learning and feature engineering. The course also covers SQL, deep learning fundamentals, and natural language processing. You will learn to deploy models as APIs, monitor them in production, and communicate results clearly to business stakeholders. By the end, you will have a portfolio-ready project and the technical interview skills to compete for real data science roles.

How you study in practice Data Science Foundation Course

How you practice Data Science Foundation Course

For companies looking to train their teams

With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.

Click here

Course Content

8 Chapters • 39 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Introduction to Data Science

  • Lesson 1 • Setting Up a Data Science Environment

    Installs and configures essential tools for reproducible analysis. Ensures every student has a working environment before coding begins.

  • Lesson 2 • The Data Science Lifecycle

    Traces a project from business problem to deployed solution. Provides a mental model for organizing all subsequent technical skills.

  • Lesson 3 • What Data Science Is

    Defines data science, distinguishes it from related fields, and maps its core components. Establishes shared vocabulary used throughout the course.

  • Lesson 4 • Types of Data and Problems

    Categorizes structured, unstructured, and semi-structured data and maps them to problem types. Guides appropriate technique selection later in the course.

Chapter 2See details

Python Programming for Data Science

  • Lesson 1 • Data Manipulation with pandas

    Teaches DataFrame creation, indexing, filtering, and transformation. Directly supports the data cleaning and feature engineering chapters ahead.

  • Lesson 2 • File I/O and Data Loading

    Reads and writes CSV, JSON, and Excel files and connects to SQL databases. Prepares students to ingest real-world datasets in any format.

  • Lesson 3 • Working with NumPy Arrays

    Introduces vectorized numerical computation using NumPy. Enables fast array operations that underpin pandas and machine learning libraries.

  • Lesson 4 • Writing Reusable Python Code

    Applies functions, modules, and error handling to build maintainable scripts. Establishes coding practices that scale to full project pipelines.

  • Lesson 5 • Python Syntax and Data Structures

    Covers variables, control flow, and built-in collections. Forms the programming backbone for all subsequent data work.

Chapter 3See details

Data Wrangling and Cleaning

  • Lesson 1 • Handling Missing Data

    Compares deletion, imputation, and indicator strategies for missing values. Teaches trade-offs that affect downstream model performance.

  • Lesson 2 • Assessing Data Quality

    Identifies missing values, duplicates, and inconsistencies through profiling. Establishes a systematic audit process before any transformation begins.

  • Lesson 3 • Transforming and Standardizing Data

    Applies type conversion, normalization, and encoding to raw columns. Produces consistent formats required by statistical and machine learning models.

  • Lesson 4 • Building Reproducible Data Pipelines

    Encapsulates cleaning steps into reusable functions and pipeline objects. Ensures consistent preprocessing across training and production datasets.

  • Lesson 5 • Outlier Detection and Treatment

    Uses statistical and visual methods to find and handle extreme values. Prevents outliers from distorting model training and summary statistics.

Chapter 4See details

Statistical Foundations for Modeling

  • Lesson 1 • Common Statistical Tests

    Applies t-tests, chi-square tests, and ANOVA to real datasets. Equips students to choose and execute the correct test for each data scenario.

  • Lesson 2 • Correlation and Linear Relationships

    Quantifies linear associations using Pearson and Spearman correlation. Lays the groundwork for regression modeling in the next chapter.

  • Lesson 3 • Probability Essentials

    Covers probability rules, conditional probability, and common distributions. Provides the mathematical language used in every probabilistic model.

  • Lesson 4 • Experimental Design Principles

    Covers controlled experiments, randomization, and A/B testing frameworks. Prepares students to design valid tests for data-driven decision-making.

  • Lesson 5 • Statistical Inference Basics

    Introduces sampling distributions, confidence intervals, and hypothesis testing. Enables rigorous evaluation of whether observed patterns are statistically meaningful.

Chapter 5See details

Exploratory Data Analysis and Visualization

  • Lesson 1 • Descriptive Statistics Fundamentals

    Computes measures of central tendency, spread, and shape for numeric variables. Provides the quantitative foundation for interpreting all visualizations.

  • Lesson 2 • Univariate and Bivariate Analysis

    Examines single-variable distributions and pairwise relationships between variables. Reveals feature behavior and potential predictive signals.

  • Lesson 3 • Visualization Best Practices

    Applies design principles to create clear, accurate, and audience-appropriate charts. Prevents misleading visuals and improves stakeholder communication.

  • Lesson 4 • EDA-Driven Feature Insights

    Translates EDA findings into actionable decisions about feature engineering and modeling. Bridges exploratory work with the machine learning chapters ahead.

  • Lesson 5 • Multivariate Exploration Techniques

    Extends analysis to interactions among three or more variables simultaneously. Surfaces complex patterns that bivariate analysis misses.

Chapter 6See details

Supervised Machine Learning

  • Lesson 1 • Classification Algorithms

    Trains logistic regression, decision trees, k-nearest neighbors, and naive Bayes classifiers. Connects algorithm choice to data characteristics and business requirements.

  • Lesson 2 • Model Evaluation and Metrics

    Applies accuracy, precision, recall, F1, ROC-AUC, and RMSE to assess model quality. Teaches metric selection based on class imbalance and business cost trade-offs.

  • Lesson 3 • Regression Algorithms

    Covers linear, ridge, lasso, and decision tree regression for continuous targets. Teaches when and why to apply regularization to reduce overfitting.

  • Lesson 4 • Machine Learning Workflow

    Establishes the end-to-end process from data splitting to model evaluation. Provides a repeatable framework applied to every algorithm in this chapter.

  • Lesson 5 • Hyperparameter Tuning

    Uses grid search and random search to optimize model hyperparameters systematically. Prevents overfitting while maximizing generalization on unseen data.

Chapter 7See details

Unsupervised Learning and Feature Engineering

  • Lesson 1 • Dimensionality Reduction

    Reduces feature space using PCA and t-SNE while preserving meaningful variance. Improves visualization, speeds up training, and mitigates the curse of dimensionality.

  • Lesson 2 • Feature Selection Methods

    Removes irrelevant and redundant features using filter, wrapper, and embedded methods. Reduces overfitting and speeds up model training.

  • Lesson 3 • Feature Engineering Techniques

    Creates new informative features from existing columns through transformation and interaction. Directly improves predictive power of models built in the previous chapter.

  • Lesson 4 • Clustering Algorithms

    Applies k-means, hierarchical, and DBSCAN clustering to discover natural groupings. Enables customer segmentation, anomaly detection, and data summarization.

  • Lesson 5 • Anomaly Detection Fundamentals

    Detects rare and unusual observations using statistical and model-based approaches. Supports fraud detection, quality control, and data cleaning use cases.

Chapter 8See details

Model Deployment and Communication

  • Lesson 1 • Building a Model API

    Wraps a trained model in a REST API using Flask or FastAPI. Enables other applications and services to consume predictions programmatically.

  • Lesson 2 • Saving and Versioning Models

    Serializes trained models and tracks experiments with versioning tools. Ensures reproducibility and enables rollback when production models degrade.

  • Lesson 3 • Model Monitoring and Maintenance

    Tracks prediction drift, data drift, and performance degradation in production. Establishes processes for retraining and updating deployed models.

  • Lesson 4 • Communicating Results to Stakeholders

    Translates technical findings into clear narratives and visual dashboards for business audiences. Closes the gap between model output and organizational decision-making.

  • Lesson 5 • Containerization and Deployment Basics

    Packages the model API into a Docker container for consistent deployment. Introduces cloud deployment concepts without requiring deep infrastructure expertise.

Certification

Your valid completion certificate

This course is for you:

  • Career changers: seeking a structured path into a data-focused profession.

  • Business analysts: wanting to move beyond spreadsheets into predictive modeling work.

  • Recent graduates: looking to add applied technical skills to their academic background.

  • Marketing or operations professionals: aiming to make data-driven decisions independently.

  • Hobbyist coders: ready to channel their curiosity into a marketable data science skill set.

  • Scientists or researchers: hoping to automate analysis and apply machine learning to their domain.

What our students say

Your classes are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to switch platforms... I thank you for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can switch chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the presentation style and video transcription, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.
André Felipe
André FelipePrompt Engineering Student

Top trainings

FAQ

Who is Dedika?

Is the certificate valid in United States?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course