arxiv.org favicon

[2211.17192] Fast Inference from Transformers via Speculative Decoding

5
Public whispers
5
Contributors
2026-07-23 09:17:19
First whispered

Public whispers on this page

Text Highlight2026-07-23 12:38:19
Original Highlight Excerpt
"an algorithm to sample from autoregressive models faster"
Whisper Note
Sounds cool but does it really keep the exact same outputs?
Text Highlight2026-07-23 12:29:19
Original Highlight Excerpt
"an algorithm to sample from autoregressive models faster"
Whisper Note
Saving this for later, my inference costs are killing me.
Text Highlight2026-07-23 09:35:19
Original Highlight Excerpt
"decoding K tokens takes K serial runs of the model"
Whisper Note
My training runs spend more time decoding than actually training sometimes.
Text Highlight2026-07-23 09:26:19
Original Highlight Excerpt
"decoding K tokens takes K serial runs of the model"
Whisper Note
Makes sense, but how do they keep the outputs identical if you're guessing ahead?
Text Highlight2026-07-23 09:17:19
Original Highlight Excerpt
"decoding K tokens takes K serial runs of the model"
Whisper Note
That's the whole bottleneck right there, every token is a serial dependency.

Share this page's whispers

Share to X
Short link
https://domwhisper.com/s/951958eafe57
Embed snippet
<iframe src="https://domwhisper.com/embed/951958eafe57" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>

See what people are discussing on arxiv.org

Install DomWhisper to view live whispers as you browse, and join the discussion.

Get the extension