How Context Attribution Handles What the Model Already Knows
Quick summary
arXiv:2607.23804v2 Announce Type: replace-cross Abstract: Context attribution methods for large language models (LLMs) identify which input context contributes to the model response. Recent works show the initial success in attributing the con- tributive score of the contexts. However, we observe that when the context overlaps with the training data, these methods can- not disentangle in-context from in-weight (IW) contributions, producing unreliable scores. Based on this observation, in this work, we introduce: 1) an evaluation protocol that relies on four new metrics (base-model context attr
Key takeaways
- arXiv:2607.23804v2 Announce Type: replace-cross Abstract: Context attribution methods for large language models (LLMs) identify which input context contributes to the model response.
- Recent works show the initial success in attributing the con- tributive score of the contexts.
- However, we observe that when the context overlaps with the training data, these methods can- not disentangle in-context from in-weight (IW) contributions, producing unreliable scores.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments