arXiv Artificial Intelligence

TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning

TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning

Quick summary

arXiv:2608.03276v1 Announce Type: new Abstract: Long-context inference with large language models is constrained by the linear growth of the key-value cache to sequence length. While pruning offers mitigation, prevailing methods determine query-specific token importance that cannot be reused across unseen queries. In contrast, we introduce TaskPress, a framework for task-guided, query-agnostic KV cache eviction. Instead of optimizing the cache for a single query, TaskPress constructs a reusable memory representation conditioned on a high-level task guide. The guide functions as a meta-query du

Key takeaways

  • arXiv:2608.03276v1 Announce Type: new Abstract: Long-context inference with large language models is constrained by the linear growth of the key-value cache to sequence length.
  • While pruning offers mitigation, prevailing methods determine query-specific token importance that cannot be reused across unseen queries.
  • In contrast, we introduce TaskPress, a framework for task-guided, query-agnostic KV cache eviction.

Why it matters

“TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗