What do your logits know?
Quick summary
arXiv:2604.09885v2 Announce Type: replace Abstract: Recent work has shown that probing model internals can reveal a wealth of information not apparent from the model generations. This poses a risk of unintentional or malicious information leakage, where model users are able to learn information that the model owner assumed was inaccessible. Using vision-language models as a testbed, we present the first systematic comparison of information retained at different representational levels as it is compressed from the rich information encoded in the residual stream through two natural bottlenecks:
Key takeaways
- arXiv:2604.09885v2 Announce Type: replace Abstract: Recent work has shown that probing model internals can reveal a wealth of information not apparent from the model generations.
- This poses a risk of unintentional or malicious information leakage, where model users are able to learn information that the model owner assumed was inaccessible.
- Using vision-language models as a testbed, we present the first systematic comparison of information retained at different representational levels as it is compressed from the rich information encoded in the residual stream through two natural bottlenecks:
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments