Knowledge Distillation
Updated 2026-08-05
INTRODUCTION
English translation pending.
CORE DEFINITION
Knowledge distillation transfers the behavior of a large, complex teacher model to a small, compact student model. Rather than learning from the original data labels, the student learns from the teacher's outputs, known as soft labels. The core claim is that soft labels carry information about the relative relationships between classes that hard labels omit, so the student inherits the teacher's implicit knowledge with far fewer parameters. The qualification is that the student can only be as good as the teacher it imitates.
SCAFFOLDING EFFECT
Reduce cognitive load
- Wisdom Compression: A master's value is not in teaching every detail but in transmitting the intuition distilled from vast data. - Dimensionality Reduction: Passing on the judgment rather than the raw material lets another person grasp the core with far less capacity. - Compact Deployment: Compressing a large system into a small one is what makes the capability usable in constrained settings.
Anchor fast decisions
The teacher produces soft labels that encode relative similarity between classes, such as assigning most probability to cat and some to dog. Because these outputs carry more information than a single hard label, a student that imitates them inherits the teacher's tacit knowledge and approaches its performance with far fewer parameters.
MINIMUM ACTION
In progress 0/1Practice this model in one real situation:
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more
Source support: Explicit
- zh.wikipedia.orghttps://zh.wikipedia.org/wiki/%E7%9F%A5%E8%AD%98%E8%92%B8%E9%A4%BEverified
PRIVATE NOTES · Only visible to you
SAVED Q&A
ENTRY Q&A · Private saving available
Ask with a clear boundary
thinkingmodels answers from published entry context only.
Your question is sent to thinkingmodels. The answer uses public entry context only.
RELATED MODELS