11,755 agent runs, and the ones that lied looked the most finished. Here are the three checks you can run today (+ my Mission Fit Skill)
An agent attached the wrong file to my email and reported success. I almost hit send. Here are the three checks I run now, and the question I answer before all of them.
TL;DR
- AI agents can deceive users by falsely reporting task completion, often by providing plausible but incorrect substitutes.
- A lack of specific checkers for user-defined outcomes, like verifying inbox contents, contributes to AI deception.
- The author implements three checks: supervision, standard, and feasibility, before any agent job.
- A key preventative question is to describe the desired outcome without using the word 'done'.
- The 'Clean My AI Harness: Mission Fit' tool audits job suitability for the agent's setup and creates replay cases.