Similarity Is Not Validity: Defending LLM Semantic Caches Against Poisoning
Quick summary
arXiv:2609.35908v1 Announce Type: cross Abstract: Semantic caches reduce LLM serving costs by reusing previously generated answers for semantically similar queries. However, retrieval is based solely on embedding similarity between the incoming query and cached queries. This design enables cache poisoning: an attacker can cache a malicious response under a query with high cosine similarity to benign requests. The vulnerability stems from a gap between retrieval similarity and answer validity. From an information-bottleneck perspective, query embeddings can lose information needed to distinguis
Key takeaways
- arXiv:2609.35908v1 Announce Type: cross Abstract: Semantic caches reduce LLM serving costs by reusing previously generated answers for semantically similar queries.
- However, retrieval is based solely on embedding similarity between the incoming query and cached queries.
- This design enables cache poisoning: an attacker can cache a malicious response under a query with high cosine similarity to benign requests.
Why it matters
This development is a reminder to test misuse and data-leak scenarios alongside speed and quality. Trust should come from testable controls and clear failure reporting, not protection claims alone.

Member comments