arXiv Artificial Intelligence

Block-Sparse Attention with Semantic-Geometric Decoupled Routing

Block-Sparse Attention with Semantic-Geometric Decoupled Routing

Quick summary

arXiv:2609.22884v1 Announce Type: cross Abstract: Long-context inference has become a defining capability of large language models, but exact dense attention remains costly due to its quadratic scaling with sequence length. Block-sparse attention offers a hardware-friendly alternative by routing each query block to a small set of relevant key blocks, yet accurate training-free block routing remains difficult. Existing routers often pool post-RoPE token representations, which entangles semantic aggregation with RoPE-induced geometry and attenuates local positional cues through high-frequency ph

Key takeaways

  • arXiv:2609.22884v1 Announce Type: cross Abstract: Long-context inference has become a defining capability of large language models, but exact dense attention remains costly due to its quadratic scaling with sequence length.
  • Block-sparse attention offers a hardware-friendly alternative by routing each query block to a small set of relevant key blocks, yet accurate training-free block routing remains difficult.
  • Existing routers often pool post-RoPE token representations, which entangles semantic aggregation with RoPE-induced geometry and attenuates local positional cues through high-frequency ph

Why it matters

The importance of “Block-Sparse Attention with Semantic-Geometric Decoupled Routing” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗