Freeform preference learning has multiple nice properties:
(1) It works better
When controlling for the number of preference queries, learning with multi-axis preferences yields far more performant policies than single-axis rewards.

Freeform preference learning has multiple nice properties:
(1) It works better
When controlling for the number of preference queries, learning with multi-axis preferences yields far more performant policies than single-axis rewards.
