With freeform preference data, we train a lang-conditioned reward that captures all axes of a task.
We then train a policy conditioned on each reward axis and the corresponding reward.
FPL allows the robot to maximally leverage and learn from each axis of supervision. https://t.co/qb83hPVjdH
