Sparse rewards, progress metrics, and preferences are popular, but they

  • often neglect many aspects of a task
  • collapse many axes into one measure
  • frequently yield ambiguity and disagreement across annotators

We instead propose freeform preference learning