Sparse rewards, progress metrics, and preferences are popular, but they
- often neglect many aspects of a task
- collapse many axes into one measure
- frequently yield ambiguity and disagreement across annotators
We instead propose freeform preference learning