How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

Watch my favorite eval experts do a live audit of the evals that I built for my skills and I promise you’ll learn something new. Plus, their new free skill to build evals in Claude Code.

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

TL;DR

  • AI evaluations have significantly changed with the latest models.
  • A live audit of the author's AI skills' evaluations was performed by Shreya and Hamel.
  • Top-down evals are predefined rules, while bottom-up evals are derived from comparing AI output with final edits.
  • A free 'Error Discovery' skill in Claude Code/Codex can help build useful evals in 5 steps.
  • Comparing 10-20 past examples is crucial before turning a failure into an eval to avoid overfitting.
  • Wispr Flow is recommended for voice dictation of AI prompts to save time and increase accuracy.