The Sequence Knowledge - 928: The Missing 5%: Why Distillation Is Harder Than It Looks

Distillation can make models smaller, faster, and cheaper. The difficult part is deciding what the student cannot afford to forget.

The Sequence Knowledge - 928: The Missing 5%: Why Distillation Is Harder Than It Looks

TL;DR

  • A new 7-billion-parameter AI model claims 95% of the performance of a 70-billion-parameter teacher model.
  • This size reduction offers benefits like lower cost and increased deployment flexibility.
  • The article questions what specific capabilities are lost in the remaining 5% of performance.
  • Potential losses include the ability to recognize confusion, recover from errors, or maintain reasoning on unfamiliar problems.