arXiv Artificial Intelligence

Monte Carlo Estimation for KV Cache Eviction

Monte Carlo Estimation for KV Cache Eviction

Quick summary

arXiv:2610.07643v1 Announce Type: cross Abstract: Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt? We instead ask, which memory will matter while answering? Since decoding queries are unavailable at eviction time, prior future-aware methods rely on pseudo-responses or synthetic future-query estimates. We cast fixed-budget future-aware eviction as distributional estimation over plausible model-conditional query trajectories and introduce LORE-KV (Lookahead Output-perturbation with Reliability-weighted Ensembles for Key-Value caches), a tra

Key takeaways

  • arXiv:2610.07643v1 Announce Type: cross Abstract: Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt?
  • We instead ask, which memory will matter while answering?
  • Since decoding queries are unavailable at eviction time, prior future-aware methods rely on pseudo-responses or synthetic future-query estimates.

Why it matters

The significance goes beyond a temporary access problem: “Monte Carlo Estimation for KV Cache Eviction” exposes the operational cost of depending on one AI provider. Critical tasks need predefined fallback, queueing and human-continuation paths.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗