Long term, we need reward models to capture all aspects of performance, like success, outcome quality, and speed.

This also includes task-specific axes like:

  • was the PB spread evenly?
  • was the apple slightly bruised while bagging it?
  • was the furniture bumped or scratched?