[2306.15595] Extending Context Window of Large Language Models via Positional Interpolation
6
公开标注数
6
参与人数
2026-07-26 09:43:09
首次 Whisper
本页的公开 Whisper
划选高亮2026-07-26 16:07:09
原文高亮摘录
“arXivLabs is a framework that allows collaborators to develop and share new arXiv features”
Whisper 随想笔记
Wait, so anyone can build stuff on arXiv now? That's kinda cool actually.
划选高亮2026-07-26 13:04:09
原文高亮摘录
“minimal fine-tuning (within 1000 steps)”
Whisper 随想笔记
I tried this with LLaMA 7B and it actually worked for my summarization task, wild.
划选高亮2026-07-26 12:55:09
原文高亮摘录
“minimal fine-tuning (within 1000 steps)”
Whisper 随想笔记
Only 1000 steps to get 32k context? That sounds too good to be true honestly.
划选高亮2026-07-26 10:01:09
原文高亮摘录
“extends the context window sizes of RoPE-based pretrained LLMs such as LLaMA models to up to 32768”
Whisper 随想笔记
tried something like this once, the attention scores blew up, glad they found a fix
划选高亮2026-07-26 09:52:09
原文高亮摘录
“extends the context window sizes of RoPE-based pretrained LLMs such as LLaMA models to up to 32768”
Whisper 随想笔记
but does it actually keep quality on shorter tasks too? paper says yes but i'm skeptical
划选高亮2026-07-26 09:43:09
原文高亮摘录
“extends the context window sizes of RoPE-based pretrained LLMs such as LLaMA models to up to 32768”
Whisper 随想笔记
so we can finally read longer docs without the model losing its mind, nice
分享本页 Whisper
短链接
https://domwhisper.com/s/e8632a3cba9f嵌入代码
<iframe src="https://domwhisper.com/embed/e8632a3cba9f" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>