Choose your language
Evaluating LLM Performance and Efficiency Course
More than 2 million students worldwide

Evaluating LLM Performance and Efficiency Course

Master every dimension of LLM evaluation — from benchmark design and automated metrics to safety testing and efficiency profiling. This course gives AI practitioners and ML teams the rigorous frameworks needed to measure what models actually do, not just what they claim to do. Stop guessing and start making data-driven decisions about every model you deploy.

Dedika for businesses

What you will learn:

  • Design rigorous benchmarks and evaluation datasets that resist contamination and overfitting.

  • Apply automated metrics, semantic scoring, and LLM-as-judge methods to assess output quality.

  • Measure inference latency, throughput, memory usage, and total cost of ownership for LLM deployments.

  • Build safety and fairness test suites covering toxicity, bias, adversarial robustness, and hallucination.

  • Construct scalable, reproducible evaluation pipelines integrated into CI/CD development workflows.

  • Communicate evaluation findings clearly to engineering, product, and executive stakeholders.

How you study in practice Evaluating LLM Performance and Efficiency Course

How you practise Evaluating LLM Performance and Efficiency Course

For companies looking to train their teams

With Dedika for businesses, the course includes exercises and examples tailored to your company and its specific needs.

Click here

Course content

8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Foundations of LLM Evaluation

  • Lesson 1 • The Evaluation Mindset

    Introduces the distinction between capability, reliability, and safety as evaluation dimensions. Frames evaluation as a scientific process requiring hypotheses and controls.

  • Lesson 2 • Evaluation Scope and Constraints

    Defines the boundaries of an evaluation project, including budget, time, and access constraints. Teaches students to scope evaluations realistically before selecting methods.

  • Lesson 3 • What LLMs Are and How They Work

    Covers transformer architecture, token generation, and probability distributions at an accessible level. Grounds all later evaluation concepts in how models actually produce outputs.

  • Lesson 4 • Types of LLM Tasks and Outputs

    Catalogs generation, classification, extraction, and reasoning tasks and their distinct output types. Connects task type to the appropriate evaluation strategy used in later chapters.

  • Lesson 5 • Sources of LLM Failure

    Maps common failure modes—hallucination, refusal, drift, and bias—to their root causes. Provides a diagnostic vocabulary used throughout the course.

Chapter 2See details

Benchmark Design and Dataset Construction

  • Lesson 1 • Benchmark Maintenance Over Time

    Addresses benchmark decay, model saturation, and the need for continuous dataset refresh. Prepares students to sustain evaluation quality as models improve.

  • Lesson 2 • Preventing Data Contamination

    Explains how training data leakage invalidates benchmarks and how to detect and prevent it. Directly addresses a critical threat to evaluation validity.

  • Lesson 3 • Principles of Good Benchmark Design

    Covers validity, reliability, and coverage as the three pillars of benchmark quality. Connects design principles to the failure modes identified in Chapter 1.

  • Lesson 4 • Annotation and Labeling Workflows

    Covers human annotation pipelines, inter-annotator agreement, and quality control for labeled datasets. Produces reliable ground-truth labels needed for automated metrics.

  • Lesson 5 • Sourcing and Curating Evaluation Data

    Teaches methods for collecting, filtering, and balancing evaluation examples from real and synthetic sources. Emphasises data quality over quantity.

Chapter 3See details

Automated Metrics for Text Quality

  • Lesson 1 • Metric Selection and Combination

    Provides a decision framework for choosing and combining metrics into composite evaluation scores. Prevents over-reliance on any single metric.

  • Lesson 2 • LLM-as-Judge Evaluation

    Teaches using a separate LLM to score outputs via structured prompts, including bias risks and calibration. Represents a scalable alternative to human evaluation.

  • Lesson 3 • Task-Specific Automated Metrics

    Covers accuracy, F1, exact match, and perplexity as task-appropriate metrics for classification, QA, and generation. Connects metric choice to task type from Chapter 1.

  • Lesson 4 • Reference-Based Lexical Metrics

    Covers BLEU, ROUGE, METEOR, and ChrF as overlap-based metrics and their mathematical foundations. Establishes baseline metric literacy before introducing more complex approaches.

  • Lesson 5 • Semantic Similarity Metrics

    Introduces embedding-based metrics such as BERTScore and semantic textual similarity measures. Bridges the gap between surface-level overlap and meaning-level evaluation.

Chapter 4See details

Human Evaluation Methods

  • Lesson 1 • Evaluation Study Design

    Covers experimental designs including pairwise comparison, Likert rating, and ranking protocols. Teaches how design choices affect statistical power and bias.

  • Lesson 2 • Analysing and Reporting Human Judgments

    Covers statistical analysis of rating data, inter-rater reliability, and communicating findings. Connects human evaluation outputs to actionable model improvement decisions.

  • Lesson 3 • When Human Evaluation Is Necessary

    Identifies tasks where automated metrics fail and human judgment is irreplaceable. Sets the stage for designing cost-effective human studies.

  • Lesson 4 • Bias and Confounds in Human Studies

    Identifies anchoring, order effects, and presentation biases that distort human ratings. Teaches mitigation strategies to improve study validity.

  • Lesson 5 • Evaluator Recruitment and Management

    Addresses sourcing expert and crowd-sourced evaluators, training them, and managing quality. Directly impacts the reliability of collected judgments.

Chapter 5See details

Measuring LLM Efficiency

  • Lesson 1 • Throughput and Concurrency Testing

    Covers requests-per-second, tokens-per-second, and concurrent user load testing for LLM APIs. Connects throughput metrics to real-world serving capacity planning.

  • Lesson 2 • Memory and Hardware Utilisation

    Explains GPU memory allocation, KV cache sizing, and hardware utilisation metrics during inference. Enables students to diagnose memory bottlenecks and optimise resource use.

  • Lesson 3 • Cost Modeling for LLM Deployments

    Builds cost models covering API pricing, self-hosted infrastructure, and total cost of ownership. Equips students to make data-driven build-vs.-buy decisions.

  • Lesson 4 • Profiling Inference Latency

    Teaches time-to-first-token, inter-token latency, and end-to-end latency measurement under realistic load. Provides hands-on profiling techniques applicable to any serving stack.

  • Lesson 5 • Efficiency Dimensions and Trade-offs

    Defines latency, throughput, memory footprint, and cost as the four axes of LLM efficiency. Frames efficiency as a multi-objective optimisation problem.

Chapter 6See details

Safety, Fairness, and Robustness Evaluation

  • Lesson 1 • Robustness and Adversarial Evaluation

    Measures model sensitivity to input perturbations, paraphrases, and out-of-distribution prompts. Reveals brittleness that standard benchmarks miss.

  • Lesson 2 • Defining Safety in LLM Outputs

    Establishes a taxonomy of harmful outputs including toxicity, misinformation, and privacy leakage. Provides the conceptual foundation for all safety evaluation methods in this chapter.

  • Lesson 3 • Automated Safety Testing Methods

    Covers red-teaming prompts, adversarial probing, and automated classifiers for detecting unsafe outputs. Scales safety evaluation beyond what manual review alone can achieve.

  • Lesson 4 • Reporting Safety and Fairness Findings

    Covers model cards, safety datasheets, and structured disclosure formats for communicating risk. Prepares students to present findings to technical and non-technical audiences.

  • Lesson 5 • Bias and Fairness Measurement

    Teaches counterfactual fairness testing, demographic parity checks, and stereotype probing across groups. Connects fairness metrics to organisational equity commitments.

Chapter 7See details

Evaluation Infrastructure and Tooling

  • Lesson 1 • Experiment Tracking and Reproducibility

    Covers logging hyperparameters, prompts, seeds, and results to ensure reproducible evaluation runs. Prevents the common failure of unreproducible benchmark results.

  • Lesson 2 • Scaling Evaluation Across Models

    Addresses parallelisation, caching, and cost optimisation for evaluating many models or configurations. Prepares students to run large-scale evaluation campaigns efficiently.

  • Lesson 3 • Continuous Evaluation in CI/CD

    Integrates evaluation gates into continuous integration pipelines to catch regressions automatically. Shifts evaluation left in the development lifecycle.

  • Lesson 4 • Open-Source Evaluation Frameworks

    Surveys leading open-source evaluation libraries, their APIs, and integration patterns. Enables students to select and adopt tools without vendor lock-in.

  • Lesson 5 • Evaluation Pipeline Architecture

    Designs modular evaluation pipelines covering data ingestion, model inference, scoring, and reporting. Establishes the engineering foundation for all tooling covered in this chapter.

Chapter 8See details

Strategic Evaluation and Decision-Making

  • Lesson 1 • Evaluation-Driven Prompt Engineering

    Uses evaluation metrics as feedback signals to iteratively improve prompts and system instructions. Closes the loop between measurement and model behaviour improvement.

  • Lesson 2 • Model Selection and Comparison Frameworks

    Builds multi-criteria decision frameworks for comparing models across quality, efficiency, and risk dimensions. Synthesises all prior evaluation skills into a unified selection process.

  • Lesson 3 • Building an Evaluation Culture

    Addresses organisational practices, team roles, and incentive structures that sustain rigorous evaluation. Moves evaluation from a one-time activity to a continuous discipline.

  • Lesson 4 • Communicating Evaluation Results

    Teaches structuring evaluation reports for engineering, product, and executive audiences. Ensures evaluation insights translate into organisational action.

  • Lesson 5 • Evaluation Roadmaps and Prioritisation

    Teaches how to prioritise evaluation investments based on risk, impact, and resource availability. Produces actionable evaluation roadmaps aligned with product development cycles.

Certification

Your valid completion certificate

This course is for you:

  • ML engineers who deploy models but lack structured evaluation practices.

  • AI product managers who need to interpret model performance with confidence.

  • Data scientists moving from model training into production quality assurance.

  • Research engineers building internal LLM tools for enterprise applications.

  • AI safety professionals expanding their technical evaluation skill set.

  • Software engineers transitioning into machine learning operations roles.

What our students say

Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the way videos are presented and transcribed, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.
André Felipe
André FelipePrompt Engineering Student

Top qualifications

FAQ

Who is Dedika?

Is the certificate valid in South Africa?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course