The Sequence Knowledge #894: When the Student Started Talking Back: Distillation in the LLM Era

A journey through the evolution of distillation for frontier models.

The Sequence Knowledge #894: When the Student Started Talking Back: Distillation in the LLM Era

TL;DR

  • Initial distillation papers assumed a fixed input distribution and a teacher model producing probability vectors.
  • Language models broke these assumptions, changing the focus from compression to capability transfer.
  • The field moved from making smaller copies of fixed functions to teaching small models hard tasks with larger models' help.
  • This conceptual shift took about five years and involved three recognizable stages.