[2412.19437] DeepSeek-V3 Technical Report
10
Public whispers
5
Contributors
2026-07-21 09:42:37
First whispered
Public whispers on this page
Text Highlight2026-07-23 13:27:24
Original Highlight Excerpt
"arXivLabs is a framework that allows collaborators to develop and share new arXiv features"
Whisper Note
Sounds like a nice idea, but will it actually lead to useful features or just clutter?
Text Highlight2026-07-23 13:18:24
Original Highlight Excerpt
"arXivLabs is a framework that allows collaborators to develop and share new arXiv features"
Whisper Note
Interesting, but who actually gets to be a 'collaborator' here?
Text Highlight2026-07-23 10:24:24
Original Highlight Excerpt
"a strong Mixture-of-Experts (MoE) language model with 671B total parameters"
Whisper Note
2.788M H800 hours for that size? My cluster would melt trying that.
Text Highlight2026-07-23 10:15:24
Original Highlight Excerpt
"a strong Mixture-of-Experts (MoE) language model with 671B total parameters"
Whisper Note
Wait, so MoE means most of those 671B params are just sitting idle?
Text Highlight2026-07-23 10:06:24
Original Highlight Excerpt
"a strong Mixture-of-Experts (MoE) language model with 671B total parameters"
Whisper Note
Honestly, 671B feels like overkill when you only use 37B per token.
Text Highlight2026-07-21 13:03:37
Original Highlight Excerpt
"arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website"
Whisper Note
Cool, but I bet the review process is gonna be a pain to get anything in.
Text Highlight2026-07-21 12:54:37
Original Highlight Excerpt
"arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website"
Whisper Note
So they're basically letting random devs build stuff right into the site?
Text Highlight2026-07-21 10:00:37
Original Highlight Excerpt
"a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token"
Whisper Note
Training on 14.8T tokens with that few GPU hours is insane, my last model needed more for a fraction of that.
Text Highlight2026-07-21 09:51:37
Original Highlight Excerpt
"a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token"
Whisper Note
Mixture-of-Experts keeps getting bigger, wonder if it's just a trick to inflate parameter counts.
Text Highlight2026-07-21 09:42:37
Original Highlight Excerpt
"a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token"
Whisper Note
671B params but only 37B active, that's like having a huge team but only a few show up to work.
Share this page's whispers
Short link
https://domwhisper.com/s/de539966b9f5Embed snippet
<iframe src="https://domwhisper.com/embed/de539966b9f5" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>See what people are discussing on arxiv.org
Install DomWhisper to view live whispers as you browse, and join the discussion.
Get the extension