[2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Model
6
公开标注数
6
参与人数
2026-07-22 09:42:45
首次 Whisper
本页的公开 Whisper
划选高亮2026-07-22 16:06:45
原文高亮摘录
“achieving precise control of their behavior is difficult due to the completely unsupervised nature”
Whisper 随想笔记
Unsupervised training is both a blessing and a curse—models know so much but can't follow simple instructions.
划选高亮2026-07-22 13:03:45
原文高亮摘录
“large-scale unsupervised language models (LMs) learn broad world knowledge and some reasoning skills”
Whisper 随想笔记
But 'some reasoning skills' is doing a lot of heavy lifting, haha.
划选高亮2026-07-22 12:54:45
原文高亮摘录
“large-scale unsupervised language models (LMs) learn broad world knowledge and some reasoning skills”
Whisper 随想笔记
So true — the unsupervised part is what makes them smart but also so hard to steer.
划选高亮2026-07-22 10:00:45
原文高亮摘录
“Your Language Model is Secretly a Reward Model”
Whisper 随想笔记
This is the kind of paper title that makes you stop scrolling.
划选高亮2026-07-22 09:51:45
原文高亮摘录
“Your Language Model is Secretly a Reward Model”
Whisper 随想笔记
Wait, so my model has been doing double duty without me knowing? Neat.
划选高亮2026-07-22 09:42:45
原文高亮摘录
“Your Language Model is Secretly a Reward Model”
Whisper 随想笔记
Finally, a title that calls out the impostor syndrome in every LLM.
分享本页 Whisper
短链接
https://domwhisper.com/s/d5a5216fcde8嵌入代码
<iframe src="https://domwhisper.com/embed/d5a5216fcde8" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>