DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization
Quick summary
arXiv:2608.27513v2 Announce Type: replace-cross Abstract: Complex reasoning and agentic applications increasingly rely on long-context inference, where growing KV caches increase both memory usage and decoding overhead. Hybrid models reduce these costs by combining Softmax Attention with Gated DeltaNet (GDN) or Kimi Delta Attention (KDA), which maintain fixed-size recurrent states. These states are commonly stored in FP32 and consume substantial GPU memory, while their updates are limited by memory bandwidth. Quantization can reduce both storage footprint and memory traffic, but we find that u
Key takeaways
- arXiv:2608.27513v2 Announce Type: replace-cross Abstract: Complex reasoning and agentic applications increasingly rely on long-context inference, where growing KV caches increase both memory usage and decoding overhead.
- Hybrid models reduce these costs by combining Softmax Attention with Gated DeltaNet (GDN) or Kimi Delta Attention (KDA), which maintain fixed-size recurrent states.
- These states are commonly stored in FP32 and consume substantial GPU memory, while their updates are limited by memory bandwidth.
Why it matters
“DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization” exposes the compute, energy and supply-chain layer behind model competition. Capacity shifts can influence model costs, service availability and the ability of smaller companies to compete.

Member comments