arXiv Artificial Intelligence

Length Penalties Make Chain-of-Thought Less Monitorable

Length Penalties Make Chain-of-Thought Less Monitorable

Quick summary

arXiv:2607.09786v4 Announce Type: replace Abstract: Recent work trains reasoning models with length penalties to curb overthinking and cut inference cost. We show that these penalties make the chain of thought less monitorable. A length-compressed model still lets misleading hints steer its answers, but it less often verbalizes their influence. We train Qwen3-4B and Qwen3-14B with reinforcement learning under length penalties targeting 60% down to 30% of baseline chain-of-thought length, then evaluate them with nine types of biasing hints on held-out MMLU-Pro-R and four transfer benchmarks. A

Key takeaways

  • arXiv:2607.09786v4 Announce Type: replace Abstract: Recent work trains reasoning models with length penalties to curb overthinking and cut inference cost.
  • We show that these penalties make the chain of thought less monitorable.
  • A length-compressed model still lets misleading hints steer its answers, but it less often verbalizes their influence.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗