Length Penalties Make Chain-of-Thought Less Monitorable
Quick summary
arXiv:2607.09786v4 Announce Type: replace Abstract: Recent work trains reasoning models with length penalties to curb overthinking and cut inference cost. We show that these penalties make the chain of thought less monitorable. A length-compressed model still lets misleading hints steer its answers, but it less often verbalizes their influence. We train Qwen3-4B and Qwen3-14B with reinforcement learning under length penalties targeting 60% down to 30% of baseline chain-of-thought length, then evaluate them with nine types of biasing hints on held-out MMLU-Pro-R and four transfer benchmarks. A
Key takeaways
- arXiv:2607.09786v4 Announce Type: replace Abstract: Recent work trains reasoning models with length penalties to curb overthinking and cut inference cost.
- We show that these penalties make the chain of thought less monitorable.
- A length-compressed model still lets misleading hints steer its answers, but it less often verbalizes their influence.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments