[2403.05530] Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
5
Public whispers
5
Contributors
2026-08-12 10:14:25
First whispered
Public whispers on this page
Text Highlight2026-08-12 13:35:25
Original Highlight Excerpt
"capable of recalling and reasoning over fine-grained information"
Whisper Note
Recalling fine-grained stuff is cool, but I'd rather see it not hallucinate the details first.
Text Highlight2026-08-12 13:26:25
Original Highlight Excerpt
"capable of recalling and reasoning over fine-grained information"
Whisper Note
10M tokens is wild, but can it actually find that one specific line in a 500-page PDF?
Text Highlight2026-08-12 10:32:25
Original Highlight Excerpt
"next generation of highly compute-efficient multimodal models"
Whisper Note
Still waiting for the open source model that does this without breaking the bank.
Text Highlight2026-08-12 10:23:25
Original Highlight Excerpt
"next generation of highly compute-efficient multimodal models"
Whisper Note
Million token context is wild, finally can throw whole codebases at it.
Text Highlight2026-08-12 10:14:25
Original Highlight Excerpt
"next generation of highly compute-efficient multimodal models"
Whisper Note
Compute-efficient but what about the energy cost for training these things?
Share this page's whispers
Short link
https://domwhisper.com/s/79a8d5beefadEmbed snippet
<iframe src="https://domwhisper.com/embed/79a8d5beefad" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>See what people are discussing on arxiv.org
Install DomWhisper to view live whispers as you browse, and join the discussion.
Get the extension