Awesome Testing

Models · reviewed · reviewed Oct 5, 2026 · 4 min

How does knowledge distillation work?

Distillation trains a student to match signals from a teacher, such as output probabilities or internal representations. The student's parameters change; the teacher's signals guide the objective. A smaller student may be cheaper to run, but copying a teacher can also copy its errors.

A teacher provides a target, not a truth certificate

Imagine a classification task with three possible answers: cat, dog, and truck. A teacher may favor cat while treating dog as a more plausible alternative than truck. Its full output distribution preserves that distinction; the single answer “cat” discards it.

Knowledge distillation uses teacher signals to guide a student’s training. The student predicts, a loss measures its disagreement with the target, and an optimizer updates the student’s parameters. The teacher supplies supervision; it is not copied into the prompt each time the deployed student answers.

A common motivation is to train a smaller student that retains useful behaviour at a lower deployment cost. Imitation is an objective, however, not a guarantee of equal capability. The architecture, training examples, signals, and evaluation determine what survives.

See what the student actually learns

The experiment uses a synthetic example with the independent reference label cat. The teacher’s three logits—the scores before softmax—are authored. The student begins with three equal zero logits and therefore equal probabilities.

Choose teacher probabilities to preserve alternatives, teacher top class to keep only its answer, or trusted label to supervise independently. Apply one update and inspect how the student changes. Then choose a mistaken teacher: it favors dog even though the reference label remains cat. Predict which target can correct that error.

Explore the mechanism

Teach a tiny student, one update at a time

The reference label is cat. Decide what supervises the student: the teacher’s distribution, its single answer, or that independent label.

Training target

Student at training temperature

Student updates
0 / 20
Training loss
1.099

Student top class at deployment temperature 1: tie. The target favors the reference class.

Inspect deployment probabilities and current logits

Student logits: [0.000, 0.000, 0.000]. Deployment uses temperature 1; training used 1.

A synthetic three-class example uses fixed teacher logits and an independent cat label. The student really updates three local logits, not a full neural network. Probabilities and loss are computed; no image classifier, training service, or model-quality benchmark runs. Every setup change resets the student.

Both teacher-derived targets inherit the mistaken teacher’s preference. Training against the independent cat label moves the student toward cat instead. A loss can fall while the student becomes more wrong relative to that reference: the loss measures agreement with the chosen target.

The widget really updates three numerical logits. It has no input encoder, learned features, full dataset, or held-out accuracy measurement. Its purpose is to expose the direction of learning, not to simulate a successful compressed LLM.

Temperature reveals relationships during training

For probability-based distillation, a higher temperature softens the teacher distribution. Less favored classes become more visible. The student uses the same temperature while matching that distribution, so the comparison remains consistent.

Try temperature 4. The target is flatter, yet the correct teacher still ranks cat ahead of dog and truck. Temperature does not turn the mistaken teacher into a correct one. It changes the supervision signal’s shape.

This training temperature is distinct from a deployment decision about sampling generated tokens. The experiment shows the student’s deployment probabilities at temperature 1 in the optional inspector. A large training-temperature probability for dog is not a permanent rule that the deployed student must sample dog that often.

The original distillation formulation can combine teacher supervision with an independent hard-label objective. This demo isolates the objectives so their effects are visible. In a real training recipe, their weights and the availability of reliable labels matter.

Distillation can match more than final answers

Teacher probabilities are one possible signal. Internal representations can also guide the student, provided the training procedure defines how differently sized layers correspond.

DistilBERT combines output distillation with language-model training and a representation-alignment objective. TinyBERT explicitly transfers information from embedding, intermediate transformer, and prediction layers. These examples show why “distilled” is not a complete description of a training recipe.

For an autoregressive language model, the signals may concern distributions over next tokens or teacher-generated sequences. Access to probabilities and internal states differs from access to text responses alone. A recipe using only generated answers cannot recover teacher signals it never observes.

A smaller student needs its own evidence

The teacher can remain useful while making mistakes. The student can reproduce those mistakes, fail to learn rare behaviours, or behave differently on inputs unlike its training examples. Agreement with the teacher and correctness for the product are separate questions.

An illustrative deployment decision makes the boundary concrete: a student that reproduces common support answers may still fail the exceptional refund rule that determines a costly action. Evaluate that rule independently; do not accept the teacher’s answer as the grading oracle merely because it supplied the training data.

Distillation changes how a model is trained. Quantization changes how numerical values are represented. MoE changes which parameter blocks run for an input. A deployment can combine these techniques, but each needs its own description and validation.

The student update used in this experiment

The teacher logits are [3.6, 2.4, 0.8] for the correct scenario and [2.4, 3.6, 0.8] for the mistaken one, in cat/dog/truck order. The trusted label is always [1, 0, 0].

With temperature T, softmax maps a score zᵢ to exp(zᵢ/T) divided by the sum across classes. Let q denote target probabilities and p student training probabilities. The loss is −T² Σ qᵢ log(pᵢ); its gradient with respect to the student logit is T(pᵢ − qᵢ).

One step updates each student logit by subtracting 0.6 × T × (pᵢ − qᵢ). Hard teacher answers and trusted labels use T=1. The teacher stays fixed; any setup change resets the student. Twenty updates end the experiment, not a claim that twenty steps train a useful real model.

Sources and further reading

  1. 01
    Distilling the Knowledge in a Neural NetworkHinton, Vinyals, and Dean · research · published Mar 9, 2015 · source checked Oct 5, 2026

    Teacher probability targets, training temperature, temperature-scaled gradients, and independent hard-label supervision.

  2. 02
    DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighterSanh et al. · research · published Oct 2, 2019 · source checked Oct 5, 2026

    A practical student recipe combines probability distillation, language modeling, and representation alignment instead of copying only final answers.

  3. 03
    TinyBERT: Distilling BERT for Natural Language UnderstandingJiao et al. · research · published Sep 23, 2019 · source checked Oct 5, 2026

    Distillation of embeddings, intermediate transformer representations, and predictions illustrates signals beyond a teacher top class.