Knowledge distillation - Wikipedia
6
公开标注数
6
参与人数
2026-07-18 10:02:04
首次 Whisper
本页的公开 Whisper
划选高亮2026-07-18 16:26:04
原文高亮摘录
“model compression is a special case of knowledge distillation”
Whisper 随想笔记
So compression is just distillation with extra steps, got it.
划选高亮2026-07-18 13:23:04
原文高亮摘录
“high temperature increases the entropy of the output”
Whisper 随想笔记
But doesn't that just blur the signal? I'd rather keep crisp targets.
划选高亮2026-07-18 13:14:04
原文高亮摘录
“high temperature increases the entropy of the output”
Whisper 随想笔记
So higher temp basically makes the model less sure but more teachable, got it.
划选高亮2026-07-18 10:20:04
原文高亮摘录
“process of transferring knowledge from a large model to a smaller one”
Whisper 随想笔记
We use this at work to shrink our models for edge devices, works great.
划选高亮2026-07-18 10:11:04
原文高亮摘录
“process of transferring knowledge from a large model to a smaller one”
Whisper 随想笔记
Doesn't the smaller one lose some nuance though? Seems inevitable.
划选高亮2026-07-18 10:02:04
原文高亮摘录
“process of transferring knowledge from a large model to a smaller one”
Whisper 随想笔记
So it's like a big brain teaching a small brain its tricks, neat.
分享本页 Whisper
短链接
https://domwhisper.com/s/618b8b2dee31嵌入代码
<iframe src="https://domwhisper.com/embed/618b8b2dee31" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>