Pretraining a Q-function often doesn’t actually help RL finetuning, compared to initializing Q from scratch.
We find that pretraining Q-functions on data from diverse policies is critical to see improvements from pretraining.
Paper: https://t.co/dlj5RXVFED
@perryadong: Pretraining has worked remarkably well across domains
We show this doesn’t hold for Q-functions in online RL from a pretrained policy — and propose IPE, a more effective way to learn Q-functions for online RL fine-tuning
(1/6) https://t.co/G5SYHmgz83
