Choose your language
Data Science Basics Course
Over 400,000 professionals on the platform
Exclusive for companies

Data Science Basics Course

Master the full data science workflow — from Python programming and statistical analysis to machine learning and model deployment. This course gives you the hands-on skills employers actually look for, built through real datasets and practical projects. Whether you're switching careers or leveling up, this is where your data science journey starts.

Dedika for students

What your team will master:

You will learn how to set up a professional Python environment and manipulate data using Pandas and NumPy. You will apply exploratory data analysis techniques and use statistical foundations to make sound modeling decisions. The course covers supervised and unsupervised machine learning algorithms, including ensemble methods and dimensionality reduction. You will also clean and preprocess messy datasets using reproducible pipelines. By the end, you will build and document a complete, portfolio-ready data science project from problem definition to deployment.

How your team learns in practice Data Science Basics Course

How your team practices Data Science Basics Course

Professionals from these companies study at Dedika

ActemiumFR
Nunner LogisticsNL
GT Constructora GeotécnicaCR
Sydel StarBR
Metrô de São PauloBR
Aguas AndinasCL
DSMIN
MeridianbetRS
CDHCN

Course Content

8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Foundations of Data Science

  • Lesson 1 • Types of Data and Their Sources

    Categorizes structured, unstructured, and semi-structured data and identifies common source types. Prepares students to select appropriate tools for each data type.

  • Lesson 2 • What Data Science Is

    Defines data science, distinguishes it from related fields, and maps its core components. Establishes shared vocabulary used throughout the course.

  • Lesson 3 • The Data Science Workflow

    Introduces the end-to-end project lifecycle from problem framing to deployment. Provides a repeatable mental model for every subsequent chapter.

  • Lesson 4 • Setting Up a Data Science Environment

    Guides installation and configuration of essential tools for data science work. Ensures every student has a functional workspace before coding begins.

Chapter 2See details

Python Programming for Data Science

  • Lesson 1 • Working with Files and Libraries

    Demonstrates reading and writing files and importing third-party libraries. Connects Python basics to the data ingestion tasks common in data science.

  • Lesson 2 • NumPy for Numerical Computing

    Introduces array-based computation with NumPy for fast numerical operations. Bridges raw Python to the vectorized workflows used in data analysis.

  • Lesson 3 • Python Syntax and Core Data Types

    Covers variables, operators, and built-in data types essential for scripting. Forms the syntactic foundation for all subsequent Python-based work.

  • Lesson 4 • Control Flow and Functions

    Teaches conditional statements, loops, and reusable function definitions. Enables students to write modular, readable data processing code.

  • Lesson 5 • Pandas for Data Manipulation

    Covers DataFrame creation, selection, filtering, and aggregation with Pandas. Equips students to clean and reshape tabular data efficiently.

Chapter 3See details

Exploratory Data Analysis

  • Lesson 1 • Descriptive Statistics Essentials

    Covers measures of central tendency, spread, and shape for numerical variables. Provides quantitative summaries that reveal data characteristics at a glance.

  • Lesson 2 • Identifying and Handling Outliers

    Defines outlier detection methods and evaluates their impact on analysis. Teaches principled decisions about retaining, transforming, or removing anomalies.

  • Lesson 3 • EDA Workflow and Reporting

    Structures a repeatable EDA process and documents findings in a shareable report. Prepares students to communicate data insights before any model is built.

  • Lesson 4 • Data Visualization with Matplotlib and Seaborn

    Builds publication-quality charts using Matplotlib and Seaborn APIs. Translates statistical findings into clear visual narratives for stakeholders.

  • Lesson 5 • Univariate and Bivariate Analysis

    Examines single-variable distributions and pairwise relationships between variables. Identifies patterns and anomalies that inform modeling strategy.

Chapter 4See details

Statistical Foundations for Modeling

  • Lesson 1 • Hypothesis Testing

    Explains null and alternative hypotheses, p-values, and significance levels. Enables students to validate data assumptions and compare group differences.

  • Lesson 2 • Probability Theory Essentials

    Covers probability rules, conditional probability, and Bayes' theorem. Provides the mathematical language underlying all probabilistic models.

  • Lesson 3 • Correlation and Causation

    Distinguishes correlation from causation and quantifies linear relationships. Prevents common misinterpretations that lead to flawed modeling decisions.

  • Lesson 4 • Confidence Intervals and Effect Sizes

    Constructs confidence intervals and measures practical significance beyond p-values. Equips students to report findings with appropriate uncertainty quantification.

  • Lesson 5 • Probability Distributions

    Introduces key discrete and continuous distributions and their parameters. Connects distribution choice to data-generating processes in real problems.

Chapter 5See details

Data Cleaning and Preprocessing

  • Lesson 1 • Assessing Data Quality

    Defines dimensions of data quality and methods for auditing completeness and consistency. Establishes a diagnostic baseline before any cleaning is applied.

  • Lesson 2 • Building Preprocessing Pipelines

    Assembles cleaning and transformation steps into reproducible scikit-learn pipelines. Prevents data leakage and streamlines model training workflows.

  • Lesson 3 • Handling Missing Data

    Explains missing data mechanisms and applies imputation and deletion strategies. Ensures students choose methods that preserve analytical validity.

  • Lesson 4 • Feature Scaling and Transformation

    Applies normalization, standardization, and mathematical transforms to numeric features. Prepares data for distance-based and gradient-based algorithms.

  • Lesson 5 • Encoding Categorical Variables

    Converts categorical features into numeric representations suitable for algorithms. Covers trade-offs between encoding strategies for nominal and ordinal data.

Chapter 6See details

Supervised Machine Learning

  • Lesson 1 • Model Evaluation and Cross-Validation

    Applies k-fold cross-validation and performance metrics to assess generalization. Prevents overfitting to a single train-test split through robust evaluation.

  • Lesson 2 • Classification Algorithms

    Introduces logistic regression, decision trees, and k-nearest neighbors for categorical targets. Compares algorithm assumptions and appropriate use cases.

  • Lesson 3 • Ensemble Methods

    Explains bagging, boosting, and stacking to improve predictive performance. Demonstrates how combining weak learners reduces variance and bias.

  • Lesson 4 • Hyperparameter Tuning

    Optimizes model hyperparameters using grid search, random search, and Bayesian methods. Balances computational cost against performance improvement systematically.

  • Lesson 5 • Regression Algorithms

    Covers linear regression, regularization, and polynomial extensions for continuous targets. Builds intuition for coefficient interpretation and prediction intervals.

  • Lesson 6 • Machine Learning Fundamentals

    Defines supervised learning, the bias-variance trade-off, and the train-test split. Establishes the conceptual framework applied across all algorithm sections.

Chapter 7See details

Unsupervised Learning and Feature Engineering

  • Lesson 1 • Anomaly Detection

    Detects rare observations using statistical and model-based approaches. Applies unsupervised techniques to fraud detection and quality control scenarios.

  • Lesson 2 • Feature Selection Methods

    Identifies the most informative features using filter, wrapper, and embedded methods. Reduces noise and computational cost without sacrificing predictive power.

  • Lesson 3 • Dimensionality Reduction

    Applies PCA and t-SNE to reduce feature space while preserving information. Improves model efficiency and enables high-dimensional data visualization.

  • Lesson 4 • Clustering Algorithms

    Covers k-means, hierarchical, and density-based clustering for grouping unlabeled data. Teaches cluster evaluation and practical interpretation of results.

  • Lesson 5 • Feature Engineering Techniques

    Creates new informative features from raw variables through domain-driven transformations. Demonstrates how engineered features often outperform algorithm tuning.

Chapter 8See details

End-to-End Data Science Projects

  • Lesson 1 • Project Documentation and Presentation

    Structures technical reports, README files, and stakeholder presentations for a project. Ensures findings are reproducible and accessible to both technical and non-technical audiences.

  • Lesson 2 • Defining the Business Problem

    Translates a vague business question into a precise, measurable data science objective. Aligns technical work with stakeholder success criteria from the outset.

  • Lesson 3 • Model Deployment Fundamentals

    Packages a trained model as a REST API and deploys it to a cloud environment. Bridges the gap between notebook experimentation and production-ready systems.

  • Lesson 4 • Modeling and Iteration Strategy

    Establishes a baseline model and iterates systematically toward performance targets. Applies experiment tracking to maintain reproducibility across model versions.

  • Lesson 5 • Data Acquisition and Storage Strategy

    Plans data collection, storage formats, and access patterns for a full project. Ensures data pipelines are reliable and reproducible across environments.

Certification

Your valid completion certificate

This course is for you:

  • Marketing analyst: wants to move beyond dashboards into predictive modeling work.

  • Recent graduate: seeking practical technical skills to enter the data job market.

  • Business professional: needs to understand data outputs and collaborate with technical teams.

  • Aspiring data scientist: has motivation but lacks a structured, comprehensive starting point.

  • Software developer: looking to expand into machine learning and analytical problem-solving.

  • Researcher or academic: wants to apply computational methods to data-driven investigations.

Related courses

FAQ

Who is Dedika?

Is the certificate valid in United States?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course