We’ve been approaching reward supervision for robots the wrong way.

I think freeform preferences are part of the answer.

A short 🧵 https://t.co/9DZbaf1mNX