arxiv.org favicon

[2001.08361] Scaling Laws for Neural Language Models

5
公开标注数
5
参与人数
2026-07-16 09:42:07
首次 Whisper

本页的公开 Whisper

划选高亮2026-07-16 13:03:07
原文高亮摘录
empirical scaling laws for language model performance on the cross-entropy loss
Whisper 随想笔记
But does this hold for non-English or multilingual models though?
划选高亮2026-07-16 12:54:07
原文高亮摘录
empirical scaling laws for language model performance on the cross-entropy loss
Whisper 随想笔记
So bigger models really do learn faster per sample, wild.
划选高亮2026-07-16 10:00:07
原文高亮摘录
loss scales as a power-law with model size, dataset size, and the amount of compute
Whisper 随想笔记
My tiny model hit a wall, but this suggests I should've just scaled up instead of tweaking layers.
划选高亮2026-07-16 09:51:07
原文高亮摘录
loss scales as a power-law with model size, dataset size, and the amount of compute
Whisper 随想笔记
But doesn't this only hold for cross-entropy? Real tasks might behave differently.
划选高亮2026-07-16 09:42:07
原文高亮摘录
loss scales as a power-law with model size, dataset size, and the amount of compute
Whisper 随想笔记
So if I just throw more GPUs at it, loss goes down predictably? That's kinda reassuring.

分享本页 Whisper

分享到 X
短链接
https://domwhisper.com/s/85f66a6419d1
嵌入代码
<iframe src="https://domwhisper.com/embed/85f66a6419d1" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>

看看大家在 arxiv.org 上讨论了什么

安装 DomWhisper,浏览网页时实时查看 whisper,也可以加入讨论。

获取插件