BERT (language model) - Wikipedia
5
Public whispers
5
Contributors
2026-07-19 10:03:17
First whispered
Public whispers on this page
Text Highlight2026-07-19 13:24:17
Original Highlight Excerpt
"BERT is an "encoder-only" transformer architecture."
Whisper Note
Encoder-only, huh? Explains why it's so good at understanding but not generating text.
Text Highlight2026-07-19 13:15:17
Original Highlight Excerpt
"BERT is an "encoder-only" transformer architecture."
Whisper Note
So it's basically just the encoder part, no decoder? That's simpler than I thought.
Text Highlight2026-07-19 10:21:17
Original Highlight Excerpt
"BERT is trained by masked token prediction and next sentence prediction."
Whisper Note
I tried fine-tuning it once and yeah, those two tasks really do teach it a lot about language.
Text Highlight2026-07-19 10:12:17
Original Highlight Excerpt
"BERT is trained by masked token prediction and next sentence prediction."
Whisper Note
Next sentence prediction always felt like the weaker half, but masked tokens carried the whole thing.
Text Highlight2026-07-19 10:03:17
Original Highlight Excerpt
"BERT is trained by masked token prediction and next sentence prediction."
Whisper Note
So it's basically a fill-in-the-blank game plus sentence ordering—makes sense why it's so good at context.
Share this page's whispers
Short link
https://domwhisper.com/s/c90ebf9172b1Embed snippet
<iframe src="https://domwhisper.com/embed/c90ebf9172b1" width="100%" height="480" style="border:0;border-radius:16px" loading="lazy"></iframe>See what people are discussing on en.wikipedia.org
Install DomWhisper to view live whispers as you browse, and join the discussion.
Get the extension