[2401.04088] Mixtral of Experts
6
Public whispers
6
Contributors
2026-07-20 09:42:31
First whispered
Public whispers on this page
Text Highlight2026-07-20 16:06:31
Original Highlight Excerpt
"arXivLabs is a framework that allows collaborators to develop and share new arXiv features"
Whisper Note
Oh cool, so people can actually build stuff right into arXiv now?
Text Highlight2026-07-20 13:03:31
Original Highlight Excerpt
"each layer is composed of 8 feedforward blocks"
Whisper Note
8 experts per layer sounds heavy, but if it beats Llama 2 70B with less compute, I'm sold.
Text Highlight2026-07-20 12:54:31
Original Highlight Excerpt
"each layer is composed of 8 feedforward blocks"
Whisper Note
So only 2 of the 8 experts are active per token? That's clever but seems like a lot of idle capacity.
Text Highlight2026-07-20 10:00:31
Original Highlight Excerpt
"Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model"
Whisper Note
Beats GPT-3.5 on benchmarks but I'm still skeptical about real-world multilingual tasks.
Text Highlight2026-07-20 09:51:31
Original Highlight Excerpt
"Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model"
Whisper Note
47B params but only 13B active? That explains why it runs on my laptop, kinda wild.
Text Highlight2026-07-20 09:42:31
Original Highlight Excerpt
"Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model"
Whisper Note
So it's like having 8 mini-brains but only 2 wake up per word—efficiency nerds must be thrilled.
Share this page's whispers
Short link
https://domwhisper.com/s/31ff4e3be2e0Embed snippet
<iframe src="https://domwhisper.com/embed/31ff4e3be2e0" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>See what people are discussing on arxiv.org
Install DomWhisper to view live whispers as you browse, and join the discussion.
Get the extension