arXiv Artificial Intelligence

From Plausible to Actionable: A Position on LLM Self-Explanations

From Plausible to Actionable: A Position on LLM Self-Explanations

Quick summary

arXiv:2607.15957v3 Announce Type: cross Abstract: Large Language Models (LLMs) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self-explanations. Such explanations have emerged as a promising direction for explainable artificial intelligence (XAI), particularly for interpreting LLM behavior. However, while self-explanations often appear plausible, whether they faithfully reflect a model's underlying reasoning process remains an open question. In this opinion paper, we argue that self-explanations can be highly plausible, que

Key takeaways

  • arXiv:2607.15957v3 Announce Type: cross Abstract: Large Language Models (LLMs) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self-explanations.
  • Such explanations have emerged as a promising direction for explainable artificial intelligence (XAI), particularly for interpreting LLM behavior.
  • However, while self-explanations often appear plausible, whether they faithfully reflect a model's underlying reasoning process remains an open question.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗