Paper: https://t.co/uNpwJMqneq

This is one of a whole bunch of recent papers reviving study of recurrent neural networks.

One weird omission is not testing LSTM RNNs. Surely they remain the canonical successful RNN architecture?

Another completely uninvestigated thing is the failure of SMT→DMT training of GRUs in Section 3.3. No details are provided beyond the ominous sentence: “SMT→DMT is unable to train GRU RNNs, because the GRU architecture induces memory space collapse during SMT training, degrading RNN rollout.”

Finally, the paper amply references and shows the value of William Merrill’s (@lambdaviking) work from the last few years on the computational power of transformers.