arxiv.org favicon

[2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Model

6
Public whispers
6
Contributors
2026-07-22 09:42:45
First whispered

Public whispers on this page

Text Highlight2026-07-22 16:06:45
Original Highlight Excerpt
"achieving precise control of their behavior is difficult due to the completely unsupervised nature"
Whisper Note
Unsupervised training is both a blessing and a curse—models know so much but can't follow simple instructions.
Text Highlight2026-07-22 13:03:45
Original Highlight Excerpt
"large-scale unsupervised language models (LMs) learn broad world knowledge and some reasoning skills"
Whisper Note
But 'some reasoning skills' is doing a lot of heavy lifting, haha.
Text Highlight2026-07-22 12:54:45
Original Highlight Excerpt
"large-scale unsupervised language models (LMs) learn broad world knowledge and some reasoning skills"
Whisper Note
So true — the unsupervised part is what makes them smart but also so hard to steer.
Text Highlight2026-07-22 10:00:45
Original Highlight Excerpt
"Your Language Model is Secretly a Reward Model"
Whisper Note
This is the kind of paper title that makes you stop scrolling.
Text Highlight2026-07-22 09:51:45
Original Highlight Excerpt
"Your Language Model is Secretly a Reward Model"
Whisper Note
Wait, so my model has been doing double duty without me knowing? Neat.
Text Highlight2026-07-22 09:42:45
Original Highlight Excerpt
"Your Language Model is Secretly a Reward Model"
Whisper Note
Finally, a title that calls out the impostor syndrome in every LLM.

Share this page's whispers

Share to X
Short link
https://domwhisper.com/s/d5a5216fcde8
Embed snippet
<iframe src="https://domwhisper.com/embed/d5a5216fcde8" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>

See what people are discussing on arxiv.org

Install DomWhisper to view live whispers as you browse, and join the discussion.

Get the extension