Pretraining Recurrent Networks without Recurrence by @akarshkumar0101 & @phillip_isola is a great paper!
It makes an end-run around the problems of RNNs via a transformer teacher to learn good predictive state representations & supervised learning of a memory transition function https://t.co/lKGK3wfbe2
