[2001.08361] Scaling Laws for Neural Language Models
5
公开标注数
5
参与人数
2026-07-16 09:42:07
首次 Whisper
本页的公开 Whisper
划选高亮2026-07-16 13:03:07
原文高亮摘录
“empirical scaling laws for language model performance on the cross-entropy loss”
Whisper 随想笔记
But does this hold for non-English or multilingual models though?
划选高亮2026-07-16 12:54:07
原文高亮摘录
“empirical scaling laws for language model performance on the cross-entropy loss”
Whisper 随想笔记
So bigger models really do learn faster per sample, wild.
划选高亮2026-07-16 10:00:07
原文高亮摘录
“loss scales as a power-law with model size, dataset size, and the amount of compute”
Whisper 随想笔记
My tiny model hit a wall, but this suggests I should've just scaled up instead of tweaking layers.
划选高亮2026-07-16 09:51:07
原文高亮摘录
“loss scales as a power-law with model size, dataset size, and the amount of compute”
Whisper 随想笔记
But doesn't this only hold for cross-entropy? Real tasks might behave differently.
划选高亮2026-07-16 09:42:07
原文高亮摘录
“loss scales as a power-law with model size, dataset size, and the amount of compute”
Whisper 随想笔记
So if I just throw more GPUs at it, loss goes down predictably? That's kinda reassuring.
分享本页 Whisper
短链接
https://domwhisper.com/s/85f66a6419d1嵌入代码
<iframe src="https://domwhisper.com/embed/85f66a6419d1" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>