arxiv.org favicon

[2103.00020] Learning Transferable Visual Models From Natural Language Supervision

6
公开标注数
6
参与人数
2026-07-18 10:04:42
首次 Whisper

本页的公开 Whisper

划选高亮2026-07-18 16:28:42
原文高亮摘录
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.
Whisper 随想笔记
Nice to see they actually vet partners for privacy, not just say it.
划选高亮2026-07-18 13:25:42
原文高亮摘录
This restricted form of supervision limits their generality and usability
Whisper 随想笔记
Eh, but labeled data still works fine for most practical tasks though.
划选高亮2026-07-18 13:16:42
原文高亮摘录
This restricted form of supervision limits their generality and usability
Whisper 随想笔记
So true, fixed categories just don't scale to the real world's mess.
划选高亮2026-07-18 10:22:42
原文高亮摘录
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories.
Whisper 随想笔记
So they basically taught it to read captions and it figured out images? That's wild.
划选高亮2026-07-18 10:13:42
原文高亮摘录
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories.
Whisper 随想笔记
This is exactly why I can never use those models for niche stuff, always retraining from scratch.
划选高亮2026-07-18 10:04:42
原文高亮摘录
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories.
Whisper 随想笔记
Yeah but most real-world tasks don't need endless categories, fixed sets work fine for a lot of stuff.

分享本页 Whisper

分享到 X
短链接
https://domwhisper.com/s/3d7d444bcdb3
嵌入代码
<iframe src="https://domwhisper.com/embed/3d7d444bcdb3" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>

看看大家在 arxiv.org 上讨论了什么

安装 DomWhisper,浏览网页时实时查看 whisper,也可以加入讨论。

获取插件