We’ve been approaching reward supervision for robots the wrong way.
I think freeform preferences are part of the answer.
A short 🧵 https://t.co/9DZbaf1mNX

We’ve been approaching reward supervision for robots the wrong way.
I think freeform preferences are part of the answer.
A short 🧵 https://t.co/9DZbaf1mNX
