[2312.00752] Mamba: Linear-Time Sequence Modeling with Selective State Spaces
6
公开标注数
6
参与人数
2026-07-22 09:17:12
首次 Whisper
本页的公开 Whisper
划选高亮2026-07-22 15:41:12
原文高亮摘录
“arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.”
Whisper 随想笔记
Sounds like a cool way for devs to pitch in, but I wonder how much actual control arXiv keeps.
划选高亮2026-07-22 12:38:12
原文高亮摘录
“Many subquadratic-time architectures”
Whisper 随想笔记
tried mamba on a long audio task, the speed bump was actually noticeable
划选高亮2026-07-22 12:29:12
原文高亮摘录
“Many subquadratic-time architectures”
Whisper 随想笔记
subquadratic stuff keeps popping up but attention just won't die, curious if mamba's the real deal
划选高亮2026-07-22 09:35:12
原文高亮摘录
“Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer”
Whisper 随想笔记
Honestly, attention is the bottleneck, so any subquadratic alternative is worth a shot.
划选高亮2026-07-22 09:26:12
原文高亮摘录
“Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer”
Whisper 随想笔记
I've seen Mamba's benchmarks, it's impressive but let's see how it holds up in production.
划选高亮2026-07-22 09:17:12
原文高亮摘录
“Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer”
Whisper 随想笔记
Transformers are great but the quadratic attention cost is a real pain for long sequences.
分享本页 Whisper
短链接
https://domwhisper.com/s/5dd64e15cc75嵌入代码
<iframe src="https://domwhisper.com/embed/5dd64e15cc75" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>