
Evaluating LLM Performance and Efficiency Course
Master every dimension of LLM evaluation — from benchmark design and automated metrics to safety testing and efficiency profiling. This course gives AI practitioners and ML teams the rigorous frameworks needed to measure what models actually do, not just what they claim to do. Stop guessing and start making data-driven decisions about every model you deploy.
What you will learn:
Design rigorous benchmarks and evaluation datasets that resist contamination and overfitting.
Apply automated metrics, semantic scoring, and LLM-as-judge methods to assess output quality.
Measure inference latency, throughput, memory usage, and total cost of ownership for LLM deployments.
Build safety and fairness test suites covering toxicity, bias, adversarial robustness, and hallucination.
Construct scalable, reproducible evaluation pipelines integrated into CI/CD development workflows.
Communicate evaluation findings clearly to engineering, product, and executive stakeholders.
How you study in practice Evaluating LLM Performance and Efficiency Course
How you practise Evaluating LLM Performance and Efficiency Course
For companies looking to train their teams
With Dedika for businesses, the course includes exercises and examples tailored to your company and its specific needs.
Course content
8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsFoundations of LLM Evaluation
Foundations of LLM Evaluation
Lesson 1 • The Evaluation Mindset
Introduces the distinction between capability, reliability, and safety as evaluation dimensions. Frames evaluation as a scientific process requiring hypotheses and controls.
Lesson 2 • Evaluation Scope and Constraints
Defines the boundaries of an evaluation project, including budget, time, and access constraints. Teaches students to scope evaluations realistically before selecting methods.
Lesson 3 • What LLMs Are and How They Work
Covers transformer architecture, token generation, and probability distributions at an accessible level. Grounds all later evaluation concepts in how models actually produce outputs.
Lesson 4 • Types of LLM Tasks and Outputs
Catalogs generation, classification, extraction, and reasoning tasks and their distinct output types. Connects task type to the appropriate evaluation strategy used in later chapters.
Lesson 5 • Sources of LLM Failure
Maps common failure modes—hallucination, refusal, drift, and bias—to their root causes. Provides a diagnostic vocabulary used throughout the course.
Chapter 2HideHide detailsSee detailsBenchmark Design and Dataset Construction
Benchmark Design and Dataset Construction
Lesson 1 • Benchmark Maintenance Over Time
Addresses benchmark decay, model saturation, and the need for continuous dataset refresh. Prepares students to sustain evaluation quality as models improve.
Lesson 2 • Preventing Data Contamination
Explains how training data leakage invalidates benchmarks and how to detect and prevent it. Directly addresses a critical threat to evaluation validity.
Lesson 3 • Principles of Good Benchmark Design
Covers validity, reliability, and coverage as the three pillars of benchmark quality. Connects design principles to the failure modes identified in Chapter 1.
Lesson 4 • Annotation and Labeling Workflows
Covers human annotation pipelines, inter-annotator agreement, and quality control for labeled datasets. Produces reliable ground-truth labels needed for automated metrics.
Lesson 5 • Sourcing and Curating Evaluation Data
Teaches methods for collecting, filtering, and balancing evaluation examples from real and synthetic sources. Emphasises data quality over quantity.
Chapter 3HideHide detailsSee detailsAutomated Metrics for Text Quality
Automated Metrics for Text Quality
Lesson 1 • Metric Selection and Combination
Provides a decision framework for choosing and combining metrics into composite evaluation scores. Prevents over-reliance on any single metric.
Lesson 2 • LLM-as-Judge Evaluation
Teaches using a separate LLM to score outputs via structured prompts, including bias risks and calibration. Represents a scalable alternative to human evaluation.
Lesson 3 • Task-Specific Automated Metrics
Covers accuracy, F1, exact match, and perplexity as task-appropriate metrics for classification, QA, and generation. Connects metric choice to task type from Chapter 1.
Lesson 4 • Reference-Based Lexical Metrics
Covers BLEU, ROUGE, METEOR, and ChrF as overlap-based metrics and their mathematical foundations. Establishes baseline metric literacy before introducing more complex approaches.
Lesson 5 • Semantic Similarity Metrics
Introduces embedding-based metrics such as BERTScore and semantic textual similarity measures. Bridges the gap between surface-level overlap and meaning-level evaluation.
Chapter 4HideHide detailsSee detailsHuman Evaluation Methods
Human Evaluation Methods
Lesson 1 • Evaluation Study Design
Covers experimental designs including pairwise comparison, Likert rating, and ranking protocols. Teaches how design choices affect statistical power and bias.
Lesson 2 • Analysing and Reporting Human Judgments
Covers statistical analysis of rating data, inter-rater reliability, and communicating findings. Connects human evaluation outputs to actionable model improvement decisions.
Lesson 3 • When Human Evaluation Is Necessary
Identifies tasks where automated metrics fail and human judgment is irreplaceable. Sets the stage for designing cost-effective human studies.
Lesson 4 • Bias and Confounds in Human Studies
Identifies anchoring, order effects, and presentation biases that distort human ratings. Teaches mitigation strategies to improve study validity.
Lesson 5 • Evaluator Recruitment and Management
Addresses sourcing expert and crowd-sourced evaluators, training them, and managing quality. Directly impacts the reliability of collected judgments.
Chapter 5HideHide detailsSee detailsMeasuring LLM Efficiency
Measuring LLM Efficiency
Lesson 1 • Throughput and Concurrency Testing
Covers requests-per-second, tokens-per-second, and concurrent user load testing for LLM APIs. Connects throughput metrics to real-world serving capacity planning.
Lesson 2 • Memory and Hardware Utilisation
Explains GPU memory allocation, KV cache sizing, and hardware utilisation metrics during inference. Enables students to diagnose memory bottlenecks and optimise resource use.
Lesson 3 • Cost Modeling for LLM Deployments
Builds cost models covering API pricing, self-hosted infrastructure, and total cost of ownership. Equips students to make data-driven build-vs.-buy decisions.
Lesson 4 • Profiling Inference Latency
Teaches time-to-first-token, inter-token latency, and end-to-end latency measurement under realistic load. Provides hands-on profiling techniques applicable to any serving stack.
Lesson 5 • Efficiency Dimensions and Trade-offs
Defines latency, throughput, memory footprint, and cost as the four axes of LLM efficiency. Frames efficiency as a multi-objective optimisation problem.
Chapter 6HideHide detailsSee detailsSafety, Fairness, and Robustness Evaluation
Safety, Fairness, and Robustness Evaluation
Lesson 1 • Robustness and Adversarial Evaluation
Measures model sensitivity to input perturbations, paraphrases, and out-of-distribution prompts. Reveals brittleness that standard benchmarks miss.
Lesson 2 • Defining Safety in LLM Outputs
Establishes a taxonomy of harmful outputs including toxicity, misinformation, and privacy leakage. Provides the conceptual foundation for all safety evaluation methods in this chapter.
Lesson 3 • Automated Safety Testing Methods
Covers red-teaming prompts, adversarial probing, and automated classifiers for detecting unsafe outputs. Scales safety evaluation beyond what manual review alone can achieve.
Lesson 4 • Reporting Safety and Fairness Findings
Covers model cards, safety datasheets, and structured disclosure formats for communicating risk. Prepares students to present findings to technical and non-technical audiences.
Lesson 5 • Bias and Fairness Measurement
Teaches counterfactual fairness testing, demographic parity checks, and stereotype probing across groups. Connects fairness metrics to organisational equity commitments.
Chapter 7HideHide detailsSee detailsEvaluation Infrastructure and Tooling
Evaluation Infrastructure and Tooling
Lesson 1 • Experiment Tracking and Reproducibility
Covers logging hyperparameters, prompts, seeds, and results to ensure reproducible evaluation runs. Prevents the common failure of unreproducible benchmark results.
Lesson 2 • Scaling Evaluation Across Models
Addresses parallelisation, caching, and cost optimisation for evaluating many models or configurations. Prepares students to run large-scale evaluation campaigns efficiently.
Lesson 3 • Continuous Evaluation in CI/CD
Integrates evaluation gates into continuous integration pipelines to catch regressions automatically. Shifts evaluation left in the development lifecycle.
Lesson 4 • Open-Source Evaluation Frameworks
Surveys leading open-source evaluation libraries, their APIs, and integration patterns. Enables students to select and adopt tools without vendor lock-in.
Lesson 5 • Evaluation Pipeline Architecture
Designs modular evaluation pipelines covering data ingestion, model inference, scoring, and reporting. Establishes the engineering foundation for all tooling covered in this chapter.
Chapter 8HideHide detailsSee detailsStrategic Evaluation and Decision-Making
Strategic Evaluation and Decision-Making
Lesson 1 • Evaluation-Driven Prompt Engineering
Uses evaluation metrics as feedback signals to iteratively improve prompts and system instructions. Closes the loop between measurement and model behaviour improvement.
Lesson 2 • Model Selection and Comparison Frameworks
Builds multi-criteria decision frameworks for comparing models across quality, efficiency, and risk dimensions. Synthesises all prior evaluation skills into a unified selection process.
Lesson 3 • Building an Evaluation Culture
Addresses organisational practices, team roles, and incentive structures that sustain rigorous evaluation. Moves evaluation from a one-time activity to a continuous discipline.
Lesson 4 • Communicating Evaluation Results
Teaches structuring evaluation reports for engineering, product, and executive audiences. Ensures evaluation insights translate into organisational action.
Lesson 5 • Evaluation Roadmaps and Prioritisation
Teaches how to prioritise evaluation investments based on risk, impact, and resource availability. Produces actionable evaluation roadmaps aligned with product development cycles.
Your valid completion certificate
This course is for you:
ML engineers who deploy models but lack structured evaluation practices.
AI product managers who need to interpret model performance with confidence.
Data scientists moving from model training into production quality assurance.
Research engineers building internal LLM tools for enterprise applications.
AI safety professionals expanding their technical evaluation skill set.
Software engineers transitioning into machine learning operations roles.
What our students say
Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.

I like the content and the way videos are presented and transcribed, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.

Top qualifications
FAQ
Who is Dedika?
Is the certificate valid in South Africa?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















