arxiv.org favicon

[2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

6
Public whispers
6
Contributors
2026-07-20 10:06:07
First whispered

Public whispers on this page

Text Highlight2026-07-20 16:30:07
Original Highlight Excerpt
"arXivLabs: experimental projects with community collaborators"
Whisper Note
Oh cool, so arXiv is basically a playground for devs now, that's kinda neat.
Text Highlight2026-07-20 13:27:07
Original Highlight Excerpt
"time and memory complexity of self-attention are quadratic in sequence length"
Whisper Note
Honestly, the real bottleneck is memory bandwidth, not just compute.
Text Highlight2026-07-20 13:18:07
Original Highlight Excerpt
"time and memory complexity of self-attention are quadratic in sequence length"
Whisper Note
This quadratic blowup is why my Transformer training always dies on long docs.
Text Highlight2026-07-20 10:24:07
Original Highlight Excerpt
"Transformers are slow and memory-hungry on long sequences"
Whisper Note
I wonder if this also helps with memory on consumer GPUs, or just for the big clusters.
Text Highlight2026-07-20 10:15:07
Original Highlight Excerpt
"Transformers are slow and memory-hungry on long sequences"
Whisper Note
Quadratic complexity is brutal, but that IO-aware trick sounds like the real fix.
Text Highlight2026-07-20 10:06:07
Original Highlight Excerpt
"Transformers are slow and memory-hungry on long sequences"
Whisper Note
Yeah no kidding, my GPU starts crying the moment I go past 1k tokens.

Share this page's whispers

Share to X
Short link
https://domwhisper.com/s/5727fed78bdd
Embed snippet
<iframe src="https://domwhisper.com/embed/5727fed78bdd" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>

See what people are discussing on arxiv.org

Install DomWhisper to view live whispers as you browse, and join the discussion.

Get the extension