arxiv.org favicon

[1901.02860] Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

5
Public whispers
5
Contributors
2026-07-15 10:05:41
First whispered

Public whispers on this page

Text Highlight2026-07-15 13:26:41
Original Highlight Excerpt
"We propose a novel neural architecture Transformer-XL"
Whisper Note
80% longer than RNNs? I need to see the code before I believe that speedup claim.
Text Highlight2026-07-15 13:17:41
Original Highlight Excerpt
"We propose a novel neural architecture Transformer-XL"
Whisper Note
Wait, so this is basically making transformers remember longer than their window? That's huge for language models.
Text Highlight2026-07-15 10:23:41
Original Highlight Excerpt
"Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context"
Whisper Note
80% longer than RNNs? I'd like to see that benchmarked on real code, not just wikitext
Text Highlight2026-07-15 10:14:41
Original Highlight Excerpt
"Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context"
Whisper Note
Honestly this was the paper that made me switch from LSTMs, the speed boost is unreal
Text Highlight2026-07-15 10:05:41
Original Highlight Excerpt
"Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context"
Whisper Note
Yeah but isn't that just because they can't see beyond the window? seems kinda obvious

Share this page's whispers

Share to X
Short link
https://domwhisper.com/s/5b39694ffd7e
Embed snippet
<iframe src="https://domwhisper.com/embed/5b39694ffd7e" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>

See what people are discussing on arxiv.org

Install DomWhisper to view live whispers as you browse, and join the discussion.

Get the extension