ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs
Quick summary
arXiv:2609.23314v1 Announce Type: cross Abstract: Modern LLMs with QK-normalization, gated attention, learned attention sinks, or logit softcapping exhibit weaker persistent attention sinks, on which existing KV cache eviction methods primarily rely. We observe that across these models, weaker sinks co-occur with greater value-vector dispersion relative to key-vector dispersion. Motivated by this value-side dispersion, we present ValueDiff, a value-geometric eviction that ranks tokens by the L2 deviation of their value vectors from the cache mean. The same score arises as the minimal-disturban
Key takeaways
- arXiv:2609.23314v1 Announce Type: cross Abstract: Modern LLMs with QK-normalization, gated attention, learned attention sinks, or logit softcapping exhibit weaker persistent attention sinks, on which existing KV cache eviction methods primarily rely.
- We observe that across these models, weaker sinks co-occur with greater value-vector dispersion relative to key-vector dispersion.
- Motivated by this value-side dispersion, we present ValueDiff, a value-geometric eviction that ranks tokens by the L2 deviation of their value vectors from the cache mean.
Why it matters
The importance of “ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Member comments