arXiv Artificial Intelligence

Are Stated Reasoning Steps Causally Load-Bearing?

Are Stated Reasoning Steps Causally Load-Bearing?

Quick summary

arXiv:2609.27038v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring assumes that the reasoning a model writes reflects the computation that directly produces its answer. Previous faithfulness metrics have been predominantly behavioral, as they simply edit the reasoning text and observe the resulting answer. However, our methodology aims to measure faithfulness causally at the activation level, specifically on self-generated reasoning. Unlike previous causal audits, which measure degradation, our interventions carry a known predicted target. In this way, each patch should switch t

Key takeaways

  • arXiv:2609.27038v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring assumes that the reasoning a model writes reflects the computation that directly produces its answer.
  • Previous faithfulness metrics have been predominantly behavioral, as they simply edit the reasoning text and observe the resulting answer.
  • However, our methodology aims to measure faithfulness causally at the activation level, specifically on self-generated reasoning.

Why it matters

“Are Stated Reasoning Steps Causally Load-Bearing?” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗