With freeform preference data, we train a lang-conditioned reward that captures all axes of a task.

We then train a policy conditioned on each reward axis and the corresponding reward.

FPL allows the robot to maximally leverage and learn from each axis of supervision. https://t.co/qb83hPVjdH