A large model like Claude can be compressed into a compact and fast version — using knowledge distillation. This is a machine learning technique. A small Student model learns to replicate the behavior of a large one: its answers, logit probabilities, and step-by-step reasoning chains. We break down how this kind of training works and exactly what the Student inherits from the large model — without fluff. The idea is simple: the model is smaller and faster, while its behavior remains similar to that of the large model.
Source: habr.com · post in Telegram