Choose your language
Cheminformatics Course
More than 2 million students worldwide

Cheminformatics Course

Master the computational tools and methods driving modern drug discovery and molecular design. This course takes you from molecular representations and chemical databases all the way to deep learning and molecular dynamics simulations. Whether you are a chemist stepping into computation or a data scientist entering the life sciences, you will gain the technical depth the field demands.

Dedika for businesses

What you'll learn:

You will build a complete cheminformatics skill set, starting with molecular representations, file formats, and chemical database management. You will learn to compute molecular descriptors and fingerprints, then apply them in QSAR modelling and virtual screening workflows. The course covers molecular docking, pharmacophore modelling, and MD simulation analysis using industry-standard tools. You will also implement graph neural networks and generative models for de novo molecular design. Additional modules address ADMET prediction, reaction informatics, workflow automation, and emerging AI trends in chemistry.

How you study in practice Cheminformatics Course

How you practise Cheminformatics Course

For businesses looking to train their team

With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.

Click here

Course content

8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Foundations of Cheminformatics

  • Lesson 1 • Core Software Ecosystem

    Surveys dominant open-source and commercial cheminformatics toolkits. Prepares students to select appropriate tools for specific tasks.

  • Lesson 2 • Chemical Data Types and Sources

    Surveys molecular data formats and major public databases. Connects data literacy to downstream computational workflows.

  • Lesson 3 • What Is Cheminformatics

    Defines cheminformatics and its scope within computational chemistry and biology. Positions the discipline relative to bioinformatics and data science.

  • Lesson 4 • Molecular Representation Basics

    Introduces 1D, 2D, and 3D molecular representations. Provides the conceptual basis for all subsequent encoding and modelling tasks.

  • Lesson 5 • Setting Up a Cheminformatics Workflow

    Guides students through configuring a reproducible computational environment. Establishes best practices for project organisation from the start.

Chapter 2See details

Molecular Representations and File Formats

  • Lesson 1 • Format Interconversion and Validation

    Teaches programmatic format conversion and error detection. Ensures data integrity when moving molecules between tools and pipelines.

  • Lesson 2 • SMILES and SMARTS in Depth

    Covers advanced SMILES syntax and SMARTS pattern matching. Enables precise substructure queries and molecular filtering.

  • Lesson 3 • Large-Scale Molecular Collections

    Covers efficient storage and retrieval of millions of compounds. Prepares students for virtual screening and database management tasks.

  • Lesson 4 • Structure File Formats

    Examines MOL, SDF, PDB, CIF, and XYZ formats in detail. Connects format choice to application requirements in modelling and databases.

  • Lesson 5 • Stereochemistry Representation

    Addresses chiral centres, E/Z isomerism, and atropisomerism in file formats. Critical for accurate property prediction and biological activity modelling.

Chapter 3See details

Molecular Descriptors and Fingerprints

  • Lesson 1 • Physicochemical Property Descriptors

    Covers logP, pKa, solubility, and related ADMET-relevant properties. Links computed properties to drug-likeness and bioavailability assessment.

  • Lesson 2 • Structural Fingerprints

    Explains bit-vector and count fingerprints including ECFP, MACCS, and path-based types. Provides the foundation for similarity searching and machine learning.

  • Lesson 3 • Similarity Metrics and Searching

    Applies Tanimoto, Dice, and cosine coefficients to fingerprint-based similarity. Enables nearest-neighbour searching in chemical space.

  • Lesson 4 • Constitutional and Topological Descriptors

    Introduces atom counts, molecular weight, and graph-based topological indices. Establishes the simplest numerical encodings of molecular structure.

  • Lesson 5 • Descriptor Selection and Reduction

    Addresses variance filtering, correlation removal, and dimensionality reduction for descriptor sets. Prevents overfitting and improves model interpretability.

Chapter 4See details

Chemical Database Management

  • Lesson 1 • Chemical Cartridges and Extensions

    Introduces chemistry-aware database extensions enabling substructure and similarity queries. Extends standard SQL with molecular search capabilities.

  • Lesson 2 • Compound Registration Systems

    Explains deduplication, salt stripping, and standardisation in registration workflows. Ensures data integrity across large corporate or public compound collections.

  • Lesson 3 • Data Curation and Quality Control

    Covers systematic detection and correction of errors in chemical datasets. Produces clean, analysis-ready compound libraries.

  • Lesson 4 • Public Database Integration

    Demonstrates programmatic access to PubChem, ChEMBL, and ZINC via APIs. Enables automated data retrieval for virtual screening and SAR analysis.

  • Lesson 5 • Relational Databases for Chemistry

    Covers SQL schema design tailored to chemical data storage. Connects relational concepts to compound registration and property management.

Chapter 5See details

Quantitative Structure-Activity Relationships

  • Lesson 1 • QSAR Fundamentals and History

    Traces QSAR from Hansch analysis to modern machine learning approaches. Establishes the theoretical basis for structure-activity modelling.

  • Lesson 2 • Classical QSAR Modelling Methods

    Implements multiple linear regression, PLS, and ridge regression for QSAR. Provides interpretable baseline models for SAR exploration.

  • Lesson 3 • Dataset Preparation for QSAR

    Addresses activity data curation, endpoint selection, and chemical space coverage. Directly impacts model reliability and applicability domain.

  • Lesson 4 • Machine Learning QSAR Models

    Applies random forests, SVMs, and gradient boosting to QSAR datasets. Achieves higher predictive accuracy for complex structure-activity landscapes.

  • Lesson 5 • Model Validation and Applicability Domain

    Covers cross-validation, external test sets, and applicability domain estimation. Ensures models are reliable and predictions are trustworthy.

Chapter 6See details

Virtual Screening and Docking

  • Lesson 1 • Running Docking Campaigns

    Guides high-throughput docking of large compound libraries against prepared targets. Translates docking theory into practical screening execution.

  • Lesson 2 • Virtual Screening Concepts

    Defines ligand-based and structure-based screening strategies and their trade-offs. Sets the strategic context for computational hit identification.

  • Lesson 3 • Consensus Scoring and Hit Selection

    Combines multiple scoring methods to improve hit selection reliability. Reduces false positives before expensive experimental follow-up.

  • Lesson 4 • Molecular Docking Principles

    Explains scoring functions, search algorithms, and binding site preparation. Provides the mechanistic understanding needed to run docking correctly.

  • Lesson 5 • Pharmacophore Modelling

    Builds and applies 3D pharmacophore models for ligand-based screening. Captures essential interaction features independent of scaffold.

Chapter 7See details

Molecular Dynamics and Simulation

  • Lesson 1 • Trajectory Analysis

    Analyses RMSD, RMSF, hydrogen bonds, and contact maps from trajectories. Extracts structural and dynamic insights relevant to binding.

  • Lesson 2 • Binding Free Energy Calculations

    Applies MM-PBSA, MM-GBSA, and FEP methods to estimate binding affinities. Connects simulation data to quantitative affinity predictions.

  • Lesson 3 • MD Simulation Fundamentals

    Introduces force fields, equations of motion, and ensemble types. Establishes the physical basis for interpreting simulation trajectories.

  • Lesson 4 • System Preparation for MD

    Covers protein and ligand parameterisation, solvation, and ionisation. Correct preparation is prerequisite to meaningful simulation results.

  • Lesson 5 • Running and Monitoring Simulations

    Executes equilibration and production runs using GROMACS or AMBER. Teaches real-time monitoring to detect simulation artefacts early.

Chapter 8See details

Deep Learning for Molecular Design

  • Lesson 1 • Graph Neural Networks for Properties

    Implements GCN, GAT, and MPNN architectures for molecular property prediction. Achieves state-of-the-art accuracy on benchmark datasets.

  • Lesson 2 • Generative Models for De Novo Design

    Trains VAEs, GANs, and flow-based models to generate novel drug-like molecules. Enables goal-directed molecular generation with property constraints.

  • Lesson 3 • Model Interpretability and Deployment

    Applies GradCAM, attention weights, and SHAP to explain deep learning predictions. Prepares models for integration into screening and design pipelines.

  • Lesson 4 • Molecular Graph Representations

    Encodes molecules as graphs with atom nodes and bond edges for neural networks. Provides the data structure foundation for all graph-based models.

  • Lesson 5 • Transformer Models for Chemistry

    Applies SMILES-based transformers and chemical language models to property tasks. Leverages pre-trained representations for low-data regimes.

Certification

Your valid completion certificate

This course is for you:

  • Medicinal chemist: ready to add computational methods to their toolkit.

  • Bioinformatician: looking to expand expertise into small-molecule chemical space.

  • Pharmaceutical researcher: wanting to run in-house virtual screening without outsourcing.

  • Data scientist: pivoting into life sciences and needing chemistry-specific ML skills.

  • Graduate student: building a dissertation around structure-activity or molecular modelling.

  • Computational biology professional: bridging the gap towards drug design workflows.

What our students say

Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the way videos are presented and transcribed, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast and simple to use. The diversity of content and complementary videos really help with learning.
André Felipe
André FelipePrompt Engineering Student

Top upskilling courses

FAQ

Who is Dedika?

Is the certificate valid in Australia?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course