arXiv Artificial Intelligence

AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning

AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning

Quick summary

arXiv:2610.03007v1 Announce Type: cross Abstract: Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace. Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced. This routing-only view overlooks two effects: low-attention entries can carry large value payloads whose removal changes future predictions, and newly generated states can appear stale before later queries have had a chance to read them. We introduce AvoKV-E, a training-free ev

Key takeaways

  • arXiv:2610.03007v1 Announce Type: cross Abstract: Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace.
  • Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced.
  • This routing-only view overlooks two effects: low-attention entries can carry large value payloads whose removal changes future predictions, and newly generated states can appear stale before later queries have had a chance to read them.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗