Beyond Trained Models: Compiling GNNs for a Sound Explainer Benchmark
Quick summary
arXiv:2610.03526v1 Announce Type: cross Abstract: Explainers for Graph Neural Networks (GNNs) are commonly evaluated by their plausibility, i.e., how well their explanations recover a predefined ground truth, such as a motif planted in the data. This protocol implicitly assumes that a GNN trained on such data relies on the intended motif. Although prior work has questioned this assumption, plausibility remains widespread. First, we show that the assumption is violated on several widely used benchmarks, where, e.g., degree statistics alone suffice to solve the task. Then, we remove this confoun
Key takeaways
- arXiv:2610.03526v1 Announce Type: cross Abstract: Explainers for Graph Neural Networks (GNNs) are commonly evaluated by their plausibility, i.e., how well their explanations recover a predefined ground truth, such as a motif planted in the data.
- This protocol implicitly assumes that a GNN trained on such data relies on the intended motif.
- Although prior work has questioned this assumption, plausibility remains widespread.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments