Pretraining a Q-function often doesn’t actually help RL finetuning, compared to initializing Q from scratch.

We find that pretraining Q-functions on data from diverse policies is critical to see improvements from pretraining.

Paper: https://t.co/dlj5RXVFED

@perryadong: Pretraining has worked remarkably well across domains

We show this doesn’t hold for Q-functions in online RL from a pretrained policy — and propose IPE, a more effective way to learn Q-functions for online RL fine-tuning

(1/6) https://t.co/G5SYHmgz83