[2001.08361] Scaling Laws for Neural Language Models
5
Public whispers
5
Contributors
2026-07-16 09:42:07
First whispered
Public whispers on this page
Text Highlight2026-07-16 13:03:07
Original Highlight Excerpt
"empirical scaling laws for language model performance on the cross-entropy loss"
Whisper Note
But does this hold for non-English or multilingual models though?
Text Highlight2026-07-16 12:54:07
Original Highlight Excerpt
"empirical scaling laws for language model performance on the cross-entropy loss"
Whisper Note
So bigger models really do learn faster per sample, wild.
Text Highlight2026-07-16 10:00:07
Original Highlight Excerpt
"loss scales as a power-law with model size, dataset size, and the amount of compute"
Whisper Note
My tiny model hit a wall, but this suggests I should've just scaled up instead of tweaking layers.
Text Highlight2026-07-16 09:51:07
Original Highlight Excerpt
"loss scales as a power-law with model size, dataset size, and the amount of compute"
Whisper Note
But doesn't this only hold for cross-entropy? Real tasks might behave differently.
Text Highlight2026-07-16 09:42:07
Original Highlight Excerpt
"loss scales as a power-law with model size, dataset size, and the amount of compute"
Whisper Note
So if I just throw more GPUs at it, loss goes down predictably? That's kinda reassuring.
Share this page's whispers
Short link
https://domwhisper.com/s/85f66a6419d1Embed snippet
<iframe src="https://domwhisper.com/embed/85f66a6419d1" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>See what people are discussing on arxiv.org
Install DomWhisper to view live whispers as you browse, and join the discussion.
Get the extension