Published on
· Peter Yang

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

This post is auto-generated from a public YouTube video for personal study notes. It is not an endorsement. The transcript is machine-derived and may contain errors; refer to the original video for accuracy.

About this video

Shreya and Hamel teach the industry standard course for AI evaluations taken by 4,500+ students from OpenAI, Google, and more.

In our interview, I asked them to give live feedback on the AI evals that I built for my skills and demo how anyone can run evals with ChatGPT/Claude and their free evals skill.

Don’t miss this up to date primer on evals from the real experts. 🙂

We talked about: (00:00) Why AI evals are completely different now with the latest models (02:30) Demo: Live audit of the evals that I built for my AI skills (05:14) Top down vs. bottom up evals and why AI sucks at the latter (08:16) Spinning up AI agents to grade each eval criteria (14:32) Demo: Using Shreyas AI skill to run evals in Claude (23:46) How to review output labels to identify failures (34:37) How to turn failures into reusable evals (43:19) Where human judgement is still needed

Resources mentioned:

Thanks to our sponsors:

Where to find Shreya and Hamel:

Loading transcript…