Xeno-Interpretability: Investigating the Alien Minds of LLMs
Quick summary
arXiv:2609.20408v1 Announce Type: cross Abstract: Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also represent and use distinctions for which no adequate human concept exists. We call such internal structures xeno-representations, and their study xeno-interpretability. We distinguish the human-interpretable semantic space from the xeno-semantic space: the region of model-native representations for which no adequate human conceptual
Key takeaways
- arXiv:2609.20408v1 Announce Type: cross Abstract: Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories.
- This paper asks whether models may also represent and use distinctions for which no adequate human concept exists.
- We call such internal structures xeno-representations, and their study xeno-interpretability.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments