arXiv Artificial Intelligence

Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning

Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning

Quick summary

arXiv:2609.31121v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is a promising safety technique for reasoning models, enabling detection of problematic reasoning before models act. A key concern is encoded reasoning, where models hide their true reasoning in ways that monitors and humans cannot interpret. Optimization pressure from CoT monitors during reinforcement learning is considered a likely driver of such behavior. We investigate this by training reasoning models to perform a main task and a side task, while penalizing them when a monitor detects reasoning about the sid

Key takeaways

  • arXiv:2609.31121v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is a promising safety technique for reasoning models, enabling detection of problematic reasoning before models act.
  • A key concern is encoded reasoning, where models hide their true reasoning in ways that monitors and humans cannot interpret.
  • Optimization pressure from CoT monitors during reinforcement learning is considered a likely driver of such behavior.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗