Reasoning Models Are Accurate but Unsound on Identification
Quick summary
arXiv:2610.03519v1 Announce Type: new Abstract: A reasoning model asked whether a causal effect is recoverable from observational data can fail in two ways: it refuses an identifiable query or answers a nonidentifiable one. The latter is more consequential, as no observational data can validate the claimed formula. Measuring this failure requires queries that are provably non-identifiable, which prior evaluations lack, and grading that accepts correct formulas in any equivalent form, which string matching cannot provide. We build CERTID, a formal identification pipeline that addresses both lim
Key takeaways
- arXiv:2610.03519v1 Announce Type: new Abstract: A reasoning model asked whether a causal effect is recoverable from observational data can fail in two ways: it refuses an identifiable query or answers a nonidentifiable one.
- The latter is more consequential, as no observational data can validate the claimed formula.
- Measuring this failure requires queries that are provably non-identifiable, which prior evaluations lack, and grading that accepts correct formulas in any equivalent form, which string matching cannot provide.
Why it matters
“Reasoning Models Are Accurate but Unsound on Identification” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments