
Data Science Basics Course
Master the full data science workflow — from Python programming and statistical analysis to machine learning and model deployment. This course gives you the hands-on skills employers actually look for, built through real datasets and practical projects. Whether you're switching careers or leveling up, this is where your data science journey starts.
What you will learn:
You will learn how to set up a professional Python environment and manipulate data using Pandas and NumPy. You will apply exploratory data analysis techniques and use statistical foundations to make sound modeling decisions. The course covers supervised and unsupervised machine learning algorithms, including ensemble methods and dimensionality reduction. You will also clean and preprocess messy datasets using reproducible pipelines. By the end, you will build and document a complete, portfolio-ready data science project from problem definition to deployment.
How you study in practice Data Science Basics Course
How you practice Data Science Basics Course
For companies looking to train their teams
With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.
Course Content
8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsFoundations of Data Science
Foundations of Data Science
Lesson 1 • Types of Data and Their Sources
Categorizes structured, unstructured, and semi-structured data and identifies common source types. Prepares students to select appropriate tools for each data type.
Lesson 2 • What Data Science Is
Defines data science, distinguishes it from related fields, and maps its core components. Establishes shared vocabulary used throughout the course.
Lesson 3 • The Data Science Workflow
Introduces the end-to-end project lifecycle from problem framing to deployment. Provides a repeatable mental model for every subsequent chapter.
Lesson 4 • Setting Up a Data Science Environment
Guides installation and configuration of essential tools for data science work. Ensures every student has a functional workspace before coding begins.
Chapter 2HideHide detailsSee detailsPython Programming for Data Science
Python Programming for Data Science
Lesson 1 • Working with Files and Libraries
Demonstrates reading and writing files and importing third-party libraries. Connects Python basics to the data ingestion tasks common in data science.
Lesson 2 • NumPy for Numerical Computing
Introduces array-based computation with NumPy for fast numerical operations. Bridges raw Python to the vectorized workflows used in data analysis.
Lesson 3 • Python Syntax and Core Data Types
Covers variables, operators, and built-in data types essential for scripting. Forms the syntactic foundation for all subsequent Python-based work.
Lesson 4 • Control Flow and Functions
Teaches conditional statements, loops, and reusable function definitions. Enables students to write modular, readable data processing code.
Lesson 5 • Pandas for Data Manipulation
Covers DataFrame creation, selection, filtering, and aggregation with Pandas. Equips students to clean and reshape tabular data efficiently.
Chapter 3HideHide detailsSee detailsExploratory Data Analysis
Exploratory Data Analysis
Lesson 1 • Descriptive Statistics Essentials
Covers measures of central tendency, spread, and shape for numerical variables. Provides quantitative summaries that reveal data characteristics at a glance.
Lesson 2 • Identifying and Handling Outliers
Defines outlier detection methods and evaluates their impact on analysis. Teaches principled decisions about retaining, transforming, or removing anomalies.
Lesson 3 • EDA Workflow and Reporting
Structures a repeatable EDA process and documents findings in a shareable report. Prepares students to communicate data insights before any model is built.
Lesson 4 • Data Visualization with Matplotlib and Seaborn
Builds publication-quality charts using Matplotlib and Seaborn APIs. Translates statistical findings into clear visual narratives for stakeholders.
Lesson 5 • Univariate and Bivariate Analysis
Examines single-variable distributions and pairwise relationships between variables. Identifies patterns and anomalies that inform modeling strategy.
Chapter 4HideHide detailsSee detailsStatistical Foundations for Modeling
Statistical Foundations for Modeling
Lesson 1 • Hypothesis Testing
Explains null and alternative hypotheses, p-values, and significance levels. Enables students to validate data assumptions and compare group differences.
Lesson 2 • Probability Theory Essentials
Covers probability rules, conditional probability, and Bayes' theorem. Provides the mathematical language underlying all probabilistic models.
Lesson 3 • Correlation and Causation
Distinguishes correlation from causation and quantifies linear relationships. Prevents common misinterpretations that lead to flawed modeling decisions.
Lesson 4 • Confidence Intervals and Effect Sizes
Constructs confidence intervals and measures practical significance beyond p-values. Equips students to report findings with appropriate uncertainty quantification.
Lesson 5 • Probability Distributions
Introduces key discrete and continuous distributions and their parameters. Connects distribution choice to data-generating processes in real problems.
Chapter 5HideHide detailsSee detailsData Cleaning and Preprocessing
Data Cleaning and Preprocessing
Lesson 1 • Assessing Data Quality
Defines dimensions of data quality and methods for auditing completeness and consistency. Establishes a diagnostic baseline before any cleaning is applied.
Lesson 2 • Building Preprocessing Pipelines
Assembles cleaning and transformation steps into reproducible scikit-learn pipelines. Prevents data leakage and streamlines model training workflows.
Lesson 3 • Handling Missing Data
Explains missing data mechanisms and applies imputation and deletion strategies. Ensures students choose methods that preserve analytical validity.
Lesson 4 • Feature Scaling and Transformation
Applies normalization, standardization, and mathematical transforms to numeric features. Prepares data for distance-based and gradient-based algorithms.
Lesson 5 • Encoding Categorical Variables
Converts categorical features into numeric representations suitable for algorithms. Covers trade-offs between encoding strategies for nominal and ordinal data.
Chapter 6HideHide detailsSee detailsSupervised Machine Learning
Supervised Machine Learning
Lesson 1 • Model Evaluation and Cross-Validation
Applies k-fold cross-validation and performance metrics to assess generalization. Prevents overfitting to a single train-test split through robust evaluation.
Lesson 2 • Classification Algorithms
Introduces logistic regression, decision trees, and k-nearest neighbors for categorical targets. Compares algorithm assumptions and appropriate use cases.
Lesson 3 • Ensemble Methods
Explains bagging, boosting, and stacking to improve predictive performance. Demonstrates how combining weak learners reduces variance and bias.
Lesson 4 • Hyperparameter Tuning
Optimizes model hyperparameters using grid search, random search, and Bayesian methods. Balances computational cost against performance improvement systematically.
Lesson 5 • Regression Algorithms
Covers linear regression, regularization, and polynomial extensions for continuous targets. Builds intuition for coefficient interpretation and prediction intervals.
Lesson 6 • Machine Learning Fundamentals
Defines supervised learning, the bias-variance trade-off, and the train-test split. Establishes the conceptual framework applied across all algorithm sections.
Chapter 7HideHide detailsSee detailsUnsupervised Learning and Feature Engineering
Unsupervised Learning and Feature Engineering
Lesson 1 • Anomaly Detection
Detects rare observations using statistical and model-based approaches. Applies unsupervised techniques to fraud detection and quality control scenarios.
Lesson 2 • Feature Selection Methods
Identifies the most informative features using filter, wrapper, and embedded methods. Reduces noise and computational cost without sacrificing predictive power.
Lesson 3 • Dimensionality Reduction
Applies PCA and t-SNE to reduce feature space while preserving information. Improves model efficiency and enables high-dimensional data visualization.
Lesson 4 • Clustering Algorithms
Covers k-means, hierarchical, and density-based clustering for grouping unlabeled data. Teaches cluster evaluation and practical interpretation of results.
Lesson 5 • Feature Engineering Techniques
Creates new informative features from raw variables through domain-driven transformations. Demonstrates how engineered features often outperform algorithm tuning.
Chapter 8HideHide detailsSee detailsEnd-to-End Data Science Projects
End-to-End Data Science Projects
Lesson 1 • Project Documentation and Presentation
Structures technical reports, README files, and stakeholder presentations for a project. Ensures findings are reproducible and accessible to both technical and non-technical audiences.
Lesson 2 • Defining the Business Problem
Translates a vague business question into a precise, measurable data science objective. Aligns technical work with stakeholder success criteria from the outset.
Lesson 3 • Model Deployment Fundamentals
Packages a trained model as a REST API and deploys it to a cloud environment. Bridges the gap between notebook experimentation and production-ready systems.
Lesson 4 • Modeling and Iteration Strategy
Establishes a baseline model and iterates systematically toward performance targets. Applies experiment tracking to maintain reproducibility across model versions.
Lesson 5 • Data Acquisition and Storage Strategy
Plans data collection, storage formats, and access patterns for a full project. Ensures data pipelines are reliable and reproducible across environments.
Your valid completion certificate
This course is for you:
Marketing analyst: wants to move beyond dashboards into predictive modeling work.
Recent graduate: seeking practical technical skills to enter the data job market.
Business professional: needs to understand data outputs and collaborate with technical teams.
Aspiring data scientist: has motivation but lacks a structured, comprehensive starting point.
Software developer: looking to expand into machine learning and analytical problem-solving.
Researcher or academic: wants to apply computational methods to data-driven investigations.
What our students say
Your classes are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to switch platforms... I thank you for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can switch chapters and skip content I don't need.

I like the content and the presentation style and video transcription, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.

Top trainings
FAQ
Who is Dedika?
Is the certificate valid in United States?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















