Choose your language
Data Science: Working with Data Course
More than 2 million learners worldwide

Data Science: Working with Data Course

Master the complete data science workflow — from raw data collection to deployed machine learning models. This course equips you with hands-on skills in Python, SQL, feature engineering, and model monitoring. Whether you are breaking into the field or levelling up professionally, you will build the technical foundation employers actually demand.

Dedika for businesses

What you will learn:

  • Build and automate end-to-end data pipelines that ingest, clean, and store data reliably.

  • Apply supervised and unsupervised machine learning algorithms to solve real business problems.

  • Engineer, select, and scale features to maximise predictive model performance.

  • Conduct thorough exploratory data analysis using statistical summaries and compelling visualisations.

  • Deploy trained models as REST APIs and monitor them for drift and degradation in production.

  • Communicate data findings and model results clearly to both technical and non-technical stakeholders.

How you study in practice Data Science: Working with Data Course

How you practise Data Science: Working with Data Course

For companies looking to train their teams

With Dedika for Businesses, the course includes exercises and examples tailored to your own business and the specific needs of your company.

Click here

Course content

8 Chapters • 39 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Foundations of Data Science

  • Lesson 1 • Tools and Environment Setup

    Introduces the standard data science toolchain and configures a reproducible working environment. Ensures all students operate from a consistent technical baseline.

  • Lesson 2 • The Data Science Lifecycle

    Traces a project from business problem to deployed insight. Provides a repeatable mental model for structuring any data science engagement.

  • Lesson 3 • Types of Data and Problems

    Categorizes structured, unstructured, and semi-structured data and maps each to problem types. Guides appropriate method selection from the start.

  • Lesson 4 • What Data Science Is

    Defines data science, distinguishes it from related fields, and maps its core components. Establishes shared vocabulary used throughout the course.

Chapter 2See details

Data Collection and Storage

  • Lesson 1 • Data Sources and Formats

    Surveys common data origins and file formats encountered in practice. Prepares students to handle any incoming data format confidently.

  • Lesson 2 • Consuming APIs and Web Data

    Covers REST API calls, authentication, and HTML scraping for data collection. Expands the range of data sources students can access programmatically.

  • Lesson 3 • Data Storage Strategies

    Compares flat files, relational databases, and columnar stores for different workloads. Students select appropriate storage based on volume and query patterns.

  • Lesson 4 • Querying Relational Databases

    Teaches SQL fundamentals for extracting and filtering data from relational stores. Directly enables data retrieval tasks in every subsequent chapter.

  • Lesson 5 • Building a Simple Data Pipeline

    Integrates collection and storage concepts into an end-to-end ingestion pipeline. Reinforces lifecycle thinking introduced in Chapter 1.

Chapter 3See details

Data Cleaning and Preparation

  • Lesson 1 • Data Type Conversion and Formatting

    Addresses parsing dates, encoding categories, and standardizing text fields. Ensures downstream tools receive correctly typed inputs.

  • Lesson 2 • Reshaping and Merging Datasets

    Teaches pivoting, melting, and joining multiple tables into a unified analytical dataset. Directly prepares data for exploratory analysis in the next chapter.

  • Lesson 3 • Detecting and Treating Outliers

    Covers statistical and visual methods for identifying anomalous values and deciding on treatment. Protects model performance from distorted inputs.

  • Lesson 4 • Handling Missing Data

    Explains mechanisms of missingness and appropriate imputation or removal strategies. Prevents bias introduced by naive missing-value treatment.

  • Lesson 5 • Assessing Data Quality

    Introduces profiling methods to quantify completeness, consistency, and accuracy. Establishes a diagnostic baseline before any transformation begins.

Chapter 4See details

Exploratory Data Analysis

  • Lesson 1 • Univariate Visual Analysis

    Teaches histograms, box plots, and density plots for single-variable distributions. Builds visual literacy applied throughout the course.

  • Lesson 2 • Bivariate and Multivariate Analysis

    Explores relationships between two or more variables using scatter plots, heatmaps, and pair plots. Surfaces correlations and interactions relevant to modeling.

  • Lesson 3 • Categorical Data Exploration

    Analyzes frequency distributions and cross-tabulations for categorical variables. Completes the EDA toolkit for mixed data types.

  • Lesson 4 • Descriptive Statistics

    Covers central tendency, dispersion, and distributional shape for numeric variables. Provides the quantitative foundation for all subsequent visual analysis.

  • Lesson 5 • Hypothesis Generation and EDA Reporting

    Translates EDA findings into testable hypotheses and a structured summary report. Bridges exploration and the modeling decisions made in later chapters.

Chapter 5See details

Feature Engineering and Selection

  • Lesson 1 • Scaling and Normalization

    Applies min-max scaling, standardization, and robust scaling to numeric features. Ensures distance-based and gradient-based models receive comparable inputs.

  • Lesson 2 • Encoding Categorical Variables

    Covers one-hot, ordinal, and target encoding for converting categories into numeric inputs. Prevents encoding-induced bias in downstream models.

  • Lesson 3 • Dimensionality Reduction

    Introduces PCA and other reduction techniques to compress high-dimensional feature spaces. Reduces overfitting risk and computational cost before modeling.

  • Lesson 4 • Creating New Features

    Derives interaction terms, polynomial features, and domain-specific aggregates. Expands the feature space with signals not present in raw columns.

  • Lesson 5 • Feature Selection Methods

    Applies filter, wrapper, and embedded methods to identify the most predictive features. Produces leaner, more interpretable models with less noise.

Chapter 6See details

Supervised Machine Learning

  • Lesson 1 • Cross-Validation and Overfitting

    Applies k-fold and stratified cross-validation to detect overfitting and estimate generalization. Produces reliable performance estimates before final model selection.

  • Lesson 2 • Classification Algorithms

    Introduces logistic regression, decision trees, random forests, and support vector machines for categorical targets. Extends the supervised workflow to discrete outcomes.

  • Lesson 3 • Hyperparameter Tuning

    Uses grid search, random search, and Bayesian optimization to find optimal model configurations. Maximizes model performance within computational constraints.

  • Lesson 4 • Model Evaluation Metrics

    Defines accuracy, precision, recall, F1, ROC-AUC, RMSE, and R-squared for assessing model quality. Enables objective comparison across algorithms and configurations.

  • Lesson 5 • Regression Algorithms

    Covers linear, ridge, lasso, and tree-based regression for continuous target prediction. Establishes the supervised learning workflow applied to all subsequent algorithms.

Chapter 7See details

Unsupervised Learning and Clustering

  • Lesson 1 • Evaluating Cluster Solutions

    Applies internal and external validation metrics to assess and compare cluster solutions. Enables defensible cluster selection without ground-truth labels.

  • Lesson 2 • Density and Hierarchical Clustering

    Applies DBSCAN for noise-tolerant clustering and hierarchical methods for dendrogram-based segmentation. Handles non-spherical and nested cluster structures.

  • Lesson 3 • Clustering Fundamentals

    Explains the unsupervised learning paradigm and the role of distance metrics in grouping. Provides the conceptual grounding for all clustering algorithms covered.

  • Lesson 4 • Interpreting and Using Cluster Outputs

    Profiles cluster characteristics and integrates cluster labels as features for supervised models. Connects unsupervised findings to business decisions and downstream pipelines.

  • Lesson 5 • Centroid-Based Clustering

    Implements k-means and k-medoids algorithms and evaluates results with inertia and silhouette scores. Covers the most widely used clustering family in practice.

Chapter 8See details

Model Deployment and Monitoring

  • Lesson 1 • Model Governance and Documentation

    Establishes model cards, audit trails, and access controls for responsible production use. Satisfies organisational and regulatory accountability requirements.

  • Lesson 2 • Monitoring Model Performance

    Tracks prediction quality, data drift, and system health metrics in production. Detects degradation before it impacts business outcomes.

  • Lesson 3 • Containerisation and Deployment Pipelines

    Packages model services in containers and automates deployment through CI/CD pipelines. Reduces manual deployment errors and accelerates iteration cycles.

  • Lesson 4 • Preparing Models for Production

    Covers serialisation, dependency management, and reproducibility requirements for production handoff. Ensures models behave identically outside the development environment.

  • Lesson 5 • Serving Models via APIs

    Builds REST endpoints to expose model predictions to applications and downstream systems. Enables real-time inference consumption by non-data-science teams.

Certification

Your valid completion certificate

This course is for you:

  • Business analyst: ready to move beyond dashboards into predictive modelling work.

  • Software developer: wants to add machine learning capabilities to their existing skill set.

  • Recent STEM graduate: seeking practical, employer-relevant data science project experience.

  • Marketing or operations professional: uses data daily but lacks formal modelling knowledge.

  • Career changer: pivoting into data science from a non-technical professional background.

  • Academic researcher: needs reproducible, code-based methods for analysing study datasets.

What our students say

Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of my interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the way videos are presented and transcribed, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast, simple to use. The diversity of content and complementary videos help a lot with learning.
André Felipe
André FelipePrompt Engineering Student

Top trainings

FAQs

Who is Dedika?

Is the certificate valid in Pakistan?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course