ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement
Quick summary
arXiv:2609.16284v1 Announce Type: cross Abstract: Query-conditioned vision--language models enable fine-grained interpretation by revealing how visual evidence changes with textual queries. However, evidence conditioned on complete descriptions does not necessarily resolve into object-specific evidence, nor does an exposed evidence map necessarily identify the evidence that constitutes the model's prediction. Across multiple VLM architectures and independent benchmarks, we find that object-level queries often retain evidence from co-occurring objects and shared context. In this paper, we intro
Key takeaways
- arXiv:2609.16284v1 Announce Type: cross Abstract: Query-conditioned vision--language models enable fine-grained interpretation by revealing how visual evidence changes with textual queries.
- However, evidence conditioned on complete descriptions does not necessarily resolve into object-specific evidence, nor does an exposed evidence map necessarily identify the evidence that constitutes the model's prediction.
- Across multiple VLM architectures and independent benchmarks, we find that object-level queries often retain evidence from co-occurring objects and shared context.
Why it matters
“ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments