[2401.04088] Mixtral of Experts
6
公开标注数
6
参与人数
2026-07-20 09:42:31
首次 Whisper
本页的公开 Whisper
划选高亮2026-07-20 16:06:31
原文高亮摘录
“arXivLabs is a framework that allows collaborators to develop and share new arXiv features”
Whisper 随想笔记
Oh cool, so people can actually build stuff right into arXiv now?
划选高亮2026-07-20 13:03:31
原文高亮摘录
“each layer is composed of 8 feedforward blocks”
Whisper 随想笔记
8 experts per layer sounds heavy, but if it beats Llama 2 70B with less compute, I'm sold.
划选高亮2026-07-20 12:54:31
原文高亮摘录
“each layer is composed of 8 feedforward blocks”
Whisper 随想笔记
So only 2 of the 8 experts are active per token? That's clever but seems like a lot of idle capacity.
划选高亮2026-07-20 10:00:31
原文高亮摘录
“Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model”
Whisper 随想笔记
Beats GPT-3.5 on benchmarks but I'm still skeptical about real-world multilingual tasks.
划选高亮2026-07-20 09:51:31
原文高亮摘录
“Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model”
Whisper 随想笔记
47B params but only 13B active? That explains why it runs on my laptop, kinda wild.
划选高亮2026-07-20 09:42:31
原文高亮摘录
“Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model”
Whisper 随想笔记
So it's like having 8 mini-brains but only 2 wake up per word—efficiency nerds must be thrilled.
分享本页 Whisper
短链接
https://domwhisper.com/s/31ff4e3be2e0嵌入代码
<iframe src="https://domwhisper.com/embed/31ff4e3be2e0" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>