AI Evaluation & Trust Engineering
Make AI systems prove they deserve trust.
Design evaluations, guardrails and monitoring for language-model and agent systems, and turn them into evidence that product, engineering and compliance can act on.
AI TEACHING PERSPECTIVES · Meet the team
Organisations are putting language models and agents into products faster than they can say whether those systems are good enough. A new role is forming around that gap: people who define what trustworthy behaviour means, measure it, defend it and keep measuring it in production.
This course prepares developers, QA engineers and analysts for that role. You will build evaluation sets and graders, run defensive red-team exercises, design guardrails and monitoring, and assemble the release evidence that lets a team ship with confidence.
One step closer.
With every module.
7 modules · 0 published lessons
Open a module to explore its lessons.
01LEARNING MODULE · ≈3 hoursWhat trust means for an AI system0 lessons · Week 1
Define trust as evidence of behaviour against agreed criteria, and map the harms that matter for a system.
Objectives
- Distinguish capability, reliability and safety.
- Map stakeholders and potential harms for an AI system.
- Describe the artefacts and responsibilities of an evaluation and trust role.
Topics
- Capability, reliability and safety
- Harm mapping
- Who decides what is acceptable
- Orientation to common frameworks: NIST AI RMF, ISO/IEC 42001 and the EU AI Act
- The evaluation and trust engineer’s artefacts
Concepts
Trust criterion
A specific, testable statement of behaviour the system must show before people should rely on it.
Harm map
A structured list of who could be harmed by the system, how, and how badly.
Risk appetite
How much risk an organisation is willing to accept for a given benefit, decided by accountable people.
Exercise
Create a harm map for an AI support assistant.
Deliverable
A harm map and a first set of trust criteria.
Lesson material is being prepared for this module.
02LEARNING MODULE · ≈4 hoursDesigning evaluation sets0 lessons · Week 1
Build evaluation sets that represent real use, including the hard and adversarial cases.
Objectives
- Build representative evaluation sets from real or realistic traffic.
- Write rubrics and expected behaviours.
- Establish and version baselines.
Topics
- Sampling real usage
- Golden sets and expected behaviour
- Edge and adversarial cases
- Rubrics for open-ended output
- Baselines and versioning
Concepts
Golden set
A curated set of inputs with agreed expected behaviour, used as the reference for every version.
Rubric
Explicit criteria and levels for grading output that has no single correct answer.
Baseline
The measured performance of the current version, which every change is compared against.
Exercise
Build a forty-case evaluation set for the support assistant.
Deliverable
A versioned evaluation set with rubric and a baseline run.
Lesson material is being prepared for this module.
03LEARNING MODULE · ≈4 hoursAutomated graders and their limits0 lessons · Week 2
Automate grading with code and models, and prove that the graders agree with people.
Objectives
- Implement programmatic and model-based graders.
- Calibrate a model grader against human labels.
- Detect and reduce grader bias.
Topics
- Exact, programmatic and model-graded checks
- Human labelling protocols
- Agreement between raters
- Grader drift and bias
Concepts
Model-graded evaluation
Using a language model to score outputs against a rubric. Useful at scale, only after calibration.
Calibration
Comparing an automated grader with human judgments and adjusting until they agree well enough.
Inter-rater agreement
How often independent graders reach the same judgment.
Exercise
Calibrate a model grader against thirty human labels.
Deliverable
A grader with an agreement report and known failure cases.
Lesson material is being prepared for this module.
04LEARNING MODULE · ≈4 hoursRed teaming and adversarial testing0 lessons · Week 2
Plan and run defensive adversarial tests against a system you are responsible for, in a sandbox.
Objectives
- Plan a scoped red-team exercise.
- Test for prompt injection through user input, documents and tools.
- Record, rate and triage findings.
Topics
- Threat modelling for LLM applications, for example with the OWASP Top 10 for LLM Applications
- Direct and indirect prompt injection
- Data leakage
- Excessive agency in tool-using agents
- Severity ratings and triage
Concepts
Indirect prompt injection
Instructions hidden in content the system retrieves, such as a web page or document, that try to change its behaviour.
Excessive agency
An agent able to take more actions, or more damaging actions, than its task requires.
Severity rating
A consistent judgment of a finding’s impact and likelihood, used to prioritise fixes.
Exercise
Red-team a retrieval-based assistant in a provided sandbox.
Deliverable
A findings log with severity, reproduction steps and proposed mitigations.
Lesson material is being prepared for this module.
05LEARNING MODULE · ≈4 hoursGuardrails and safe design0 lessons · Week 3
Design layered defences, and measure what they cost in usefulness.
Objectives
- Choose guardrails for each layer of an AI system.
- Scope tool permissions to the minimum a task needs.
- Design fail-closed behaviour and human approval gates.
Topics
- Defence in depth for AI systems
- Output validation and structured outputs
- Least-privilege tools
- Approval gates for irreversible actions
- Measuring false refusals and other guardrail costs
Concepts
Defence in depth
Several independent safeguards, so that one failing does not expose the system.
Fail closed
When a safeguard cannot decide, the system refuses or escalates rather than proceeding.
False refusal
A harmless request the system wrongly blocks. Guardrails that refuse too often get switched off.
Exercise
Harden the support assistant against your red-team findings.
Deliverable
A guardrail design and before-and-after evaluation results.
Lesson material is being prepared for this module.
06LEARNING MODULE · ≈4 hoursMonitoring in production0 lessons · Week 3
Keep measuring after launch: traces, online evaluation, drift and incidents.
Objectives
- Define production signals for quality and safety.
- Set up online evaluation sampling and alerts.
- Run an AI incident review.
Topics
- Tracing and logging, with privacy in mind
- Online evaluation sampling
- Drift in data, models and prompts
- User feedback loops
- Incident response for AI systems
Concepts
Online evaluation
Grading a sample of live traffic continuously, using the same criteria as offline evaluation.
Drift
A change in inputs, model behaviour or context that makes past evaluation results less valid.
Incident review
A blameless analysis of what happened, why, and what changes will prevent a repeat.
Exercise
Design monitoring for the support assistant.
Deliverable
A monitoring plan with alert thresholds and an incident runbook.
Lesson material is being prepared for this module.
07LEARNING MODULE · ≈5 hoursEvidence, governance and release0 lessons · Week 4
Assemble the evidence that lets an organisation decide, and document, that a system is ready.
Objectives
- Write a system card for an AI feature.
- Define release gates based on evaluation results.
- Communicate residual risk to non-technical decision makers.
Topics
- System and model documentation
- Release gates and sign-off
- Managing model and prompt changes
- Communicating to product, legal and leadership
- Audit trails
Concepts
System card
A document describing what an AI system does, how it was evaluated, its known limits and how it is monitored.
Release gate
A threshold on agreed evaluation results that must be met before release.
Residual risk
The risk that remains after mitigation, stated so that accountable people can accept or reject it.
Exercise
Prepare a release dossier for the support assistant.
Deliverable
A system card, a release gate checklist and a residual-risk statement.
Lesson material is being prepared for this module.
The course,
in context.
Explore the approach, expectations and practical details. These are public course notes, separate from your learning modules.
PUBLIC COURSE GUIDEWho this is for
- Developers moving into AI engineering who want to own quality and safety.
- QA engineers testing language-model and agent features.
- Security, risk and governance specialists who need hands-on evaluation skills.
- Product and data people who must decide whether an AI system is ready.
Before you begin
- Comfort reading simple Python or JavaScript; templates are provided.
- Basic understanding of how language-model applications work.
- Access to a language-model API or a local model for the exercises.
- Recommended: Agentic Systems or AI-Assisted Quality Engineering first.
What you will learn
- Define trust criteria and map harms for an AI system.
- Build evaluation sets, rubrics and baselines.
- Calibrate automated graders against human judgment.
- Run defensive red-team exercises and triage findings.
- Design layered guardrails and measure their cost.
- Monitor AI systems in production and assemble release evidence.
Applied assignments
Harm map and evaluation set
Trust criteria, a forty-case evaluation set, rubric and baseline.
Grader and red-team report
A calibrated grader with agreement report, and a findings log with severities.
Guardrails, monitoring and release dossier
Guardrail design, monitoring plan, system card and release gates.
Your capstone project
Take an AI support assistant from prototype to release
Capstone brief
You receive a working prototype of a retrieval-based support assistant that has never been evaluated. Define what trustworthy behaviour means, measure it, attack it in a sandbox, harden it, plan its monitoring, and produce the release dossier an accountable owner could sign.
Assessment
A 75-minute scenario exam: critique an evaluation set, interpret grader agreement results, triage red-team findings and decide whether a system meets its release gates.
What you will produce
- Harm map and trust-criteria template
- Evaluation set and rubric template
- Red-team plan and findings log
- Monitoring plan and incident runbook
- System card and release gate checklist
The method, explained
Trust is a claim that needs evidence
“The assistant works well” is an opinion. “The assistant answers 94% of the golden set correctly, refuses fewer than 2% of harmless requests and has no open high-severity findings” is evidence. This course is about producing the second kind of statement, and keeping it true.
The loop
- Define what trustworthy behaviour means for this system and these users.
- Measure it with evaluation sets and calibrated graders.
- Attack it, defensively and in a sandbox, to find what the evaluation missed.
- Defend it with layered guardrails, and measure what they cost.
- Watch it in production, because models, data and users change.
- Document it, so accountable people can decide with open eyes.
A role, not a gate
Evaluation and trust engineers do not replace product, engineering or legal judgment. They give those people reliable evidence early enough to change the design, not just to block a release.
Responsible scope
Adversarial exercises in this course are defensive and run only against systems you own, in a provided sandbox. Framework references are an orientation, not legal or compliance advice.
Course glossary
Trust criterion
A testable statement of behaviour required before people should rely on a system.
Golden set
Curated inputs with agreed expected behaviour.
Model-graded evaluation
Using a model to score outputs against a rubric, after calibration.
Indirect prompt injection
Instructions hidden in retrieved content that try to change system behaviour.
Excessive agency
An agent able to do more than its task requires.
Online evaluation
Continuous grading of a sample of live traffic.
System card
Documentation of an AI system’s purpose, evaluation, limits and monitoring.
Compare the approaches
Ways to answer “is this AI system good enough?”
| Approach | What it tells you | What it misses |
|---|---|---|
| Demo and gut feeling | Whether the happy path is impressive | Everything else |
| Public benchmarks | General model capability | Your users, your data, your risks |
| Offline evaluation sets | Quality on representative cases, version to version | Cases nobody collected |
| Red teaming | How the system fails under deliberate pressure | Ordinary, everyday failures |
| Online monitoring | Behaviour with real users and data | Problems before launch |
| Governance review | Whether accountable people accept the risk | Technical detail, unless evidence is provided |
Trust engineering combines them: offline evaluation as the backbone, red teaming to stress it, monitoring to keep it honest and governance to make the decision explicit.
Teaching perspective
This course is developed with our AI teaching personas and reviewed by the experienced software developers behind the lab. The personas are creative identities for AI-assisted perspectives, not human instructors.
- Aegis — Security & Responsible AI Guide: what are we trusting, and should we?
- Prism — Data & Model Analyst: what does the evidence actually support?
- Verity — Testing & Reliability Specialist: how would we notice if this were wrong?
Human responsibility stays human: we decide what belongs in the course, test the examples and correct what is wrong.
Experiences.
In their own words.
Reviews are submitted by enrolled learners and checked before publication. Ratings include approved reviews only. Critical feedback is welcome; spam, abuse and unrelated content are not.
No published reviews yet. Your experience could help the next learner.
