ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs
Quick summary
arXiv:2603.18579v2 Announce Type: replace-cross Abstract: Evaluating whether explanations faithfully reflect a model's reasoning remains an open problem. Existing benchmarks use single interventions without statistical testing, making it impossible to distinguish genuine faithfulness from chance-level performance. We show that faithfulness is not a fixed property but an operator-dependent quantity that changes with the intervention method used to measure it. We introduce ICE (Intervention-Consistent Explanation), a framework that evaluates explanations against random baselines of equal size un
Key takeaways
- arXiv:2603.18579v2 Announce Type: replace-cross Abstract: Evaluating whether explanations faithfully reflect a model's reasoning remains an open problem.
- Existing benchmarks use single interventions without statistical testing, making it impossible to distinguish genuine faithfulness from chance-level performance.
- We show that faithfulness is not a fixed property but an operator-dependent quantity that changes with the intervention method used to measure it.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments