The Sequence Knowledge #894: When the Student Started Talking Back: Distillation in the LLM Era
A journey through the evolution of distillation for frontier models.

TL;DR
- Initial distillation papers assumed a fixed input distribution and a teacher model producing probability vectors.
- Language models broke these assumptions, changing the focus from compression to capability transfer.
- The field moved from making smaller copies of fixed functions to teaching small models hard tasks with larger models' help.
- This conceptual shift took about five years and involved three recognizable stages.