arXiv Artificial Intelligence

Can Interpretation Predict Behavior on Unseen Data?

Can Interpretation Predict Behavior on Unseen Data?

Quick summary

arXiv:2507.06445v4 Announce Type: replace-cross Abstract: Interpretability research often predicts model responses to targeted mechanistic interventions. But can we predict responses to unseen input data? We propose and demonstrate this alternate objective by using model internals to predict their out-of-distribution (OOD) behavior. We train hundreds of Transformers on simple synthetic tasks, where perfect in-distribution accuracy is compatible with multiple OOD generalization rules. We successfully use attention patterns -- observed only on in-distribution data -- to predict which rule each m

Key takeaways

  • arXiv:2507.06445v4 Announce Type: replace-cross Abstract: Interpretability research often predicts model responses to targeted mechanistic interventions.
  • But can we predict responses to unseen input data?
  • We propose and demonstrate this alternate objective by using model internals to predict their out-of-distribution (OOD) behavior.

Why it matters

“Can Interpretation Predict Behavior on Unseen Data?” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗