Can LLMs Introspect? A Reality Check
Quick summary
arXiv:2605.26242v2 Announce Type: replace Abstract: Can large language models detect and report their own internal states? A number of recent studies have argued that they can. Drawing on lessons from human metacognition research, we argue that this conclusion may be premature. We identify two conditions that a paradigm needs to meet in order to establish introspection. First, the test needs to require privileged access: it should not be solvable using cues available in the input. Second, it needs to require second-order computation: second-order, meta-representations of first-order, task-rela
Key takeaways
- arXiv:2605.26242v2 Announce Type: replace Abstract: Can large language models detect and report their own internal states?
- A number of recent studies have argued that they can.
- Drawing on lessons from human metacognition research, we argue that this conclusion may be premature.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments