GitHub - vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs · GitHub
5
公开标注数
5
参与人数
2026-07-23 09:18:46
首次 Whisper
本页的公开 Whisper
划选高亮2026-07-23 12:39:46
原文高亮摘录
“Seamless integration with popular Hugging Face models”
Whisper 随想笔记
That's the main reason I switched from other engines honestly.
划选高亮2026-07-23 12:30:46
原文高亮摘录
“Seamless integration with popular Hugging Face models”
Whisper 随想笔记
Yeah, but sometimes the custom models still break, you know?
划选高亮2026-07-23 09:36:46
原文高亮摘录
“Efficient management of attention key and value memory with PagedAttention”
Whisper 随想笔记
I remember struggling with OOM errors before this—glad it's open source.
划选高亮2026-07-23 09:27:46
原文高亮摘录
“Efficient management of attention key and value memory with PagedAttention”
Whisper 随想笔记
But does it actually work well with existing models out of the box?
划选高亮2026-07-23 09:18:46
原文高亮摘录
“Efficient management of attention key and value memory with PagedAttention”
Whisper 随想笔记
PagedAttention sounds like a game changer for long context windows.
分享本页 Whisper
短链接
https://domwhisper.com/s/e8ba7172314e嵌入代码
<iframe src="https://domwhisper.com/embed/e8ba7172314e" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>