arXiv Artificial Intelligence

DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving

DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving

Quick summary

arXiv:2608.30386v1 Announce Type: cross Abstract: Hybrid linear-attention architectures have recently scaled to large open-weight models, offering quality competitive with full attention while substantially reducing key/value (KV) cache growth. However, their in-place recurrent-state updates complicate cache management: prefix reuse requires state checkpoints alongside full-attention KV, while storing state checkpoints in full increases memory pressure, leading to more evictions and repeated prefill. By analyzing the decay structure of Gated DeltaNet (GDN) and Kimi Delta Attention (KDA), we fi

Key takeaways

  • arXiv:2608.30386v1 Announce Type: cross Abstract: Hybrid linear-attention architectures have recently scaled to large open-weight models, offering quality competitive with full attention while substantially reducing key/value (KV) cache growth.
  • However, their in-place recurrent-state updates complicate cache management: prefix reuse requires state checkpoints alongside full-attention KV, while storing state checkpoints in full increases memory pressure, leading to more evictions and repeated prefill.
  • By analyzing the decay structure of Gated DeltaNet (GDN) and Kimi Delta Attention (KDA), we fi

Why it matters

The importance of “DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗