arXiv Artificial Intelligence

Explaining Attention with Program Synthesis

Explaining Attention with Program Synthesis

Quick summary

arXiv:2606.19317v3 Announce Type: replace-cross Abstract: A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic descriptions. In this paper, we propose an approach for approximating the behavior of components of deep networks with executable programs. We focus on attention heads in transformer language models. For a given head, we first compute its associated attention matrices on a collection of randomly selected training examples. Next, we prompt a pre-trained language model with a summary of these matrices, and

Key takeaways

  • arXiv:2606.19317v3 Announce Type: replace-cross Abstract: A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic descriptions.
  • In this paper, we propose an approach for approximating the behavior of components of deep networks with executable programs.
  • We focus on attention heads in transformer language models.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗