16
Public whispers
7
Contributors
2026-07-21 09:44:41
First whispered
Discussion activity
Public whispers on distill.pub over the last 17 weeks
LessMore
Recent public whispers
Text Highlight2026-08-08 15:09:18
Original Highlight Excerpt
"this is sounding awfully familiar."
Whisper Note
baseline choice is like picking what 'missing' means for your model, tricky stuff.
Text Highlight2026-08-07 09:09:18
Original Highlight Excerpt
"it separates the things causing behavior from things that merely correlate with the causes"
Whisper Note
Good point—optimization shows real causes, not just correlations. That's a big deal.
Text Highlight2026-07-23 13:05:50
Original Highlight Excerpt
"With barely a modicum of extra effort, we have essentially square rooted the condition number!"
Whisper Note
Wait, so it's like a free speedup? Why didn't they teach this in class?
Text Highlight2026-07-23 12:56:50
Original Highlight Excerpt
"With barely a modicum of extra effort, we have essentially square rooted the condition number!"
Whisper Note
So basically just crank beta to 1 and hope for the best, haha.
Text Highlight2026-07-23 10:02:50
Original Highlight Excerpt
"Momentum is a heavy ball rolling down the same hill."
Whisper Note
I've seen that image before, still helps me explain it to my students.
Text Highlight2026-07-23 09:53:50
Original Highlight Excerpt
"Momentum is a heavy ball rolling down the same hill."
Whisper Note
Heavy ball also overshoots, which is why we tune the friction.
Text Highlight2026-07-23 09:44:50
Original Highlight Excerpt
"Momentum is a heavy ball rolling down the same hill."
Whisper Note
Kinda poetic but that ball analogy breaks when you hit a flat region.
Text Highlight2026-07-22 16:08:44
Original Highlight Excerpt
"use a kernel size that is divided by your stride"
Whisper Note
Wait, so kernel divisible by stride just means sub-pixel conv? That's neat but seems like a band-aid.
Text Highlight2026-07-22 13:05:44
Original Highlight Excerpt
"deconvolution can easily have “uneven overlap,”"
Whisper Note
I've seen this in my own models—the checkerboard pattern is so annoying, now I know why.
Text Highlight2026-07-22 12:56:44
Original Highlight Excerpt
"deconvolution can easily have “uneven overlap,”"
Whisper Note
Wait, so the kernel size just needs to divide the stride? That's a simple fix, why doesn't everyone do that?
Text Highlight2026-07-22 10:02:44
Original Highlight Excerpt
"a large fraction of recent models exhibit this"
Whisper Note
Makes sense, the upsampling layers are the usual suspects.
Text Highlight2026-07-22 09:53:44
Original Highlight Excerpt
"a large fraction of recent models exhibit this"
Whisper Note
Maybe it's just a few models, but I've seen it too.
Text Highlight2026-07-22 09:44:44
Original Highlight Excerpt
"a large fraction of recent models exhibit this"
Whisper Note
Yep, I've noticed that in a bunch of GAN outputs, it's everywhere.
Text Highlight2026-07-21 10:02:41
Original Highlight Excerpt
"Interpretability techniques are normally studied in isolation."
Whisper Note
I tried combining saliency and activation maps once, got a mess.
Text Highlight2026-07-21 09:53:41
Original Highlight Excerpt
"Interpretability techniques are normally studied in isolation."
Whisper Note
But isolation helps understand each method first, no?
Text Highlight2026-07-21 09:44:41
Original Highlight Excerpt
"Interpretability techniques are normally studied in isolation."
Whisper Note
Kinda true, most papers only test one method at a time.
Ranked nearby
See what people are saying on distill.pub
Install DomWhisper to view live whispers as you browse, and join the discussion.
Get the extension