arXiv Artificial Intelligence

</think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination

</think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination

Quick summary

arXiv:2609.03633v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often produces long, redundant traces. Recent training-free early-exit methods shorten these traces by choosing an intermediate point to stop reasoning. We study one such strategy that injects an end-of-think token (EoT, ) at this point to trigger the reasoning-to-answering transition, and find that the injected EoT does not always induce a clean answering phase. Answering-phase generation can continue before the model regenerates another EoT, with the

Key takeaways

  • arXiv:2609.03633v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often produces long, redundant traces.
  • Recent training-free early-exit methods shorten these traces by choosing an intermediate point to stop reasoning.
  • We study one such strategy that injects an end-of-think token (EoT, ) at this point to trigger the reasoning-to-answering transition, and find that the injected EoT does not always induce a clean answering phase.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗