How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel
Watch my favorite eval experts do a live audit of the evals that I built for my skills and I promise you’ll learn something new. Plus, their new free skill to build evals in Claude Code.
TL;DR
- AI evaluations have significantly changed with the latest models.
- A live audit of the author's AI skills' evaluations was performed by Shreya and Hamel.
- Top-down evals are predefined rules, while bottom-up evals are derived from comparing AI output with final edits.
- A free 'Error Discovery' skill in Claude Code/Codex can help build useful evals in 5 steps.
- Comparing 10-20 past examples is crucial before turning a failure into an eval to avoid overfitting.
- Wispr Flow is recommended for voice dictation of AI prompts to save time and increase accuracy.