arXiv Artificial Intelligence

Layer-wise Positional Bias in Short-Context Language Modeling

Layer-wise Positional Bias in Short-Context Language Modeling

Quick summary

arXiv:2601.04098v2 Announce Type: replace-cross Abstract: Transformer language models systematically prefer tokens at specific input positions regardless of semantic relevance---a phenomenon known as positional bias. Prior work characterizes this bias in model behavior through performance drops in long-context tasks or in model architecture through attention-based analyses. However, it remains unmeasured how input positions actually drive predictions layer by layer. We introduce a layer conductance framework within a sliding-window design, applied to short-context next-word prediction to isola

Key takeaways

  • arXiv:2601.04098v2 Announce Type: replace-cross Abstract: Transformer language models systematically prefer tokens at specific input positions regardless of semantic relevance---a phenomenon known as positional bias.
  • Prior work characterizes this bias in model behavior through performance drops in long-context tasks or in model architecture through attention-based analyses.
  • However, it remains unmeasured how input positions actually drive predictions layer by layer.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗