(3) Long-horizon credit assignment Most robot RL focuses on short horizon tasks b/c dense temporal rewards are hard to get.

Freeform preferences yield dense rewards for subtasks without subtask segmentation.