arXiv Artificial Intelligence

Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

Quick summary

arXiv:2608.19338v1 Announce Type: cross Abstract: Mechanistic interpretability seeks quantities that models do not expose directly: represented states, component effects, interactions, and responses to interventions. Patching, gradients, Hessian-vector products, and subset interventions provide different measurements under different access assumptions and may target different quantities. We formulate their shared measurement structure as mechanistic tomography: designed measurement for recovering internal mechanisms and intervention effects. For a chosen basis and intervention family, measurem

Key takeaways

  • arXiv:2608.19338v1 Announce Type: cross Abstract: Mechanistic interpretability seeks quantities that models do not expose directly: represented states, component effects, interactions, and responses to interventions.
  • Patching, gradients, Hessian-vector products, and subset interventions provide different measurements under different access assumptions and may target different quantities.
  • We formulate their shared measurement structure as mechanistic tomography: designed measurement for recovering internal mechanisms and intervention effects.

Why it matters

The importance of “Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗