- Published on
- · Peter Yang
Complete Beginner's Course on AI Evaluations in 50 Minutes (2025) | Aman Khan
This post is auto-generated from a public YouTube video for personal study notes. It is not an endorsement. The transcript is machine-derived and may contain errors; refer to the original video for accuracy.
About this video
Today, I want to share a new episode with Aman Khan.
The best way to learn about AI evaluations is to watch 2 PMs build them live from scratch.
In our new episode, Aman and I walk through creating evals for an AI customer support agent — from labeling a golden dataset to aligning LLM judges. This is the complete beginners AI eval course you've been waiting for.
Aman and I talked about: (00:00) What are AI evals and how to get good at them (02:52) The 4 types of AI evaluations everyone should know (06:08) Live demo: Building evals for a customer support agent (10:29) Using Anthropic's console to generate great prompts (15:13) Creating the evaluation criteria (17:40) Adding human labels to the golden dataset (31:05) Scaling evals with LLM-judge prompts (38:21) How to align LLM judges with human judgment
Where to find Aman: