Freeform preference learning has multiple nice properties:

(1) It works better

When controlling for the number of preference queries, learning with multi-axis preferences yields far more performant policies than single-axis rewards.