Choose your language
Understanding Information Theory Course
More than 2 million learners worldwide

Understanding Information Theory Course

Master the mathematical framework that underlies modern communication, data compression, and machine learning. This course takes you from Shannon's foundational axioms through channel capacity, error-correcting codes, and rate-distortion theory. Whether you work in engineering, statistics, or AI, you'll gain the rigorous tools to quantify, transmit, and compress information optimally.

Dedika for businesses

What you will learn:

  • Derive Shannon entropy and apply it to measure uncertainty in discrete and continuous sources.

  • Compute mutual information, KL divergence, and cross-entropy across practical statistical problems.

  • Design optimal prefix-free codes and analyze their efficiency against theoretical entropy bounds.

  • Model noisy channels, compute capacity, and apply the channel coding theorem to real systems.

  • Construct and decode linear block codes, convolutional codes, and modern LDPC and turbo codes.

  • Apply information-theoretic principles to machine learning, feature selection, and model comparison.

How you study in a practical way Understanding Information Theory Course

How you practice Understanding Information Theory Course

For companies who want to train their team

With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.

Click here

Course content

8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Foundations of Information Theory

  • Lesson 1 • Probability Review for Information Theory

    Refreshes essential probability concepts needed to derive information measures. Connects random variables and distributions to uncertainty quantification.

  • Lesson 2 • Surprise and Self-Information

    Defines self-information as the surprise of a single outcome. Motivates the logarithmic measure and its intuitive properties.

  • Lesson 3 • What Is Information Theory

    Traces the origins of information theory from communication engineering to modern data science. Establishes why quantifying information matters across disciplines.

  • Lesson 4 • Axiomatic Foundations of Entropy

    Presents Shannon's axioms that uniquely characterize entropy. Reinforces why entropy is the correct measure of uncertainty.

  • Lesson 5 • Shannon Entropy Defined

    Derives Shannon entropy as the expected surprise of a distribution. Students compute entropy for discrete sources and interpret its meaning.

Chapter 2See details

Core Information Measures

  • Lesson 1 • KL Divergence and Relative Entropy

    Measures the cost of assuming the wrong distribution. Students apply KL divergence to model comparison and hypothesis testing.

  • Lesson 2 • Information Diagrams and Relationships

    Visualizes relationships among entropy, mutual information, and divergence using Venn-style diagrams. Consolidates all measures into a unified framework.

  • Lesson 3 • Cross-Entropy and Its Applications

    Defines cross-entropy as the average code length under a mismatched model. Links cross-entropy to KL divergence and loss functions in machine learning.

  • Lesson 4 • Joint and Conditional Entropy

    Defines entropy for pairs of variables and the residual uncertainty given side information. Builds the chain rule for entropy.

  • Lesson 5 • Mutual Information

    Quantifies shared information between two variables as reduction in uncertainty. Connects mutual information to joint and marginal entropies.

Chapter 3See details

Source Coding and Data Compression

  • Lesson 1 • Codes and Code Properties

    Introduces symbol codes, uniquely decodable codes, and prefix-free codes. Establishes the Kraft inequality as a necessary condition for efficiency.

  • Lesson 2 • Dictionary-Based and Universal Coding

    Covers LZ-family algorithms that adapt to unknown source statistics. Connects universal coding to the concept of entropy rate for stationary sources.

  • Lesson 3 • Arithmetic Coding

    Encodes entire messages as intervals to approach entropy more closely than Huffman coding. Covers encoding, decoding, and precision issues.

  • Lesson 4 • Shannon's Source Coding Theorem

    Proves that entropy is the fundamental limit of lossless compression. Students interpret the theorem's achievability and converse.

  • Lesson 5 • Huffman Coding

    Constructs optimal prefix-free codes using the Huffman algorithm. Students build trees by hand and analyze code efficiency.

Chapter 4See details

Channel Models and Capacity

  • Lesson 1 • Channel Capacity Computation Methods

    Applies the Blahut-Arimoto algorithm to compute capacity numerically. Students iterate the algorithm and verify convergence on example channels.

  • Lesson 2 • Gaussian Channels and Bandwidth

    Extends capacity analysis to additive white Gaussian noise channels. Derives the Shannon-Hartley formula relating bandwidth, power, and capacity.

  • Lesson 3 • Discrete Memoryless Channels

    Defines the discrete memoryless channel via transition probability matrices. Students compute output distributions and identify channel symmetry.

  • Lesson 4 • Shannon's Channel Coding Theorem

    States that reliable communication is possible at any rate below capacity. Covers achievability via random coding and the converse argument.

  • Lesson 5 • Channel Capacity Definition

    Defines capacity as the maximum mutual information over all input distributions. Students optimize input distributions for simple channels.

Chapter 5See details

Error-Correcting Codes

  • Lesson 1 • Convolutional Codes and Viterbi Decoding

    Encodes streams using shift-register circuits and decodes with the Viterbi algorithm. Students trace trellis diagrams to find maximum-likelihood paths.

  • Lesson 2 • Cyclic Codes and Reed-Solomon Codes

    Exploits algebraic structure for efficient encoding and burst-error correction. Covers polynomial representation and Reed-Solomon applications.

  • Lesson 3 • Modern Codes: Turbo and LDPC

    Introduces capacity-approaching codes and iterative belief-propagation decoding. Connects modern code performance to Shannon limits.

  • Lesson 4 • Error Detection and Correction Basics

    Introduces Hamming distance, error detection, and correction capabilities. Connects code parameters to the channel noise model.

  • Lesson 5 • Linear Block Codes

    Defines linear codes via generator and parity-check matrices. Students encode messages and perform syndrome decoding.

Chapter 6See details

Rate-Distortion Theory

  • Lesson 1 • Rate-Distortion Function

    Defines the rate-distortion function as the minimum rate for a given distortion level. Students derive it for Gaussian and binary sources.

  • Lesson 2 • Quantization Theory

    Applies rate-distortion principles to scalar and vector quantization design. Students analyze quantization noise and high-resolution approximations.

  • Lesson 3 • Practical Lossy Compression Standards

    Connects rate-distortion theory to transform coding used in audio and image compression. Analyzes how standards approach theoretical limits.

  • Lesson 4 • Shannon's Rate-Distortion Theorem

    Proves that any rate above the rate-distortion function is achievable. Covers the converse showing rates below the function are impossible.

  • Lesson 5 • Lossy Compression Fundamentals

    Distinguishes lossless from lossy compression and introduces distortion measures. Motivates the rate-distortion tradeoff for perceptual and practical applications.

Chapter 7See details

Information Theory in Statistics and Learning

  • Lesson 1 • Mutual Information in Feature Selection

    Uses mutual information to rank and select informative features for classification. Students apply information gain and compare it to correlation-based methods.

  • Lesson 2 • Variational Inference and Information Theory

    Frames variational inference as KL divergence minimization. Students connect the evidence lower bound to rate-distortion and mutual information.

  • Lesson 3 • Minimum Description Length Principle

    Frames model selection as a compression problem using the MDL principle. Connects MDL to Bayesian model comparison and Occam's razor.

  • Lesson 4 • Hypothesis Testing and Divergence

    Links KL divergence to error exponents in binary hypothesis testing. Students derive the Chernoff-Stein lemma and interpret its operational meaning.

  • Lesson 5 • PAC Learning and Information Bounds

    Derives generalization bounds using mutual information between training data and learned models. Connects information complexity to sample efficiency.

Chapter 8See details

Advanced Topics and Modern Applications

  • Lesson 1 • Information-Theoretic Security

    Defines perfect secrecy and the wiretap channel model. Students compute secrecy capacity and compare information-theoretic to computational security.

  • Lesson 2 • Information Theory in Neuroscience

    Applies entropy and mutual information to neural coding and sensory systems. Students analyze spike train data using information-theoretic metrics.

  • Lesson 3 • Emerging Frontiers in Information Theory

    Surveys active research areas including coded distributed computing and semantic communication. Students identify open problems and research directions.

  • Lesson 4 • Network Information Theory

    Extends single-channel results to multi-user networks including broadcast and multiple-access channels. Students compute capacity regions for two-user cases.

  • Lesson 5 • Quantum Information Theory Basics

    Introduces qubits, von Neumann entropy, and quantum channel capacity. Connects classical information measures to their quantum analogs.

Certification

Your valid completion certificate

This course is for you:

  • Electrical engineers seeking a rigorous theoretical foundation for communication systems.

  • Data scientists wanting to understand the math behind loss functions and model evaluation.

  • Computer science students ready to move beyond algorithms into information fundamentals.

  • Statisticians curious about how divergence measures connect inference to coding theory.

  • AI researchers aiming to ground deep learning intuitions in provable theoretical limits.

  • Self-taught programmers who want to close gaps in their mathematical understanding of data.

What our students say

Your classes are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of my interest without needing to change platforms... I thank you for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can switch chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the way videos are presented and transcribed, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.
André Felipe
André FelipePrompt Engineering Student

Top trainings

FAQs

Who is Dedika?

Is the certificate valid in the Philippines?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course