arXiv Artificial Intelligence

LLMs Learn to Evade Latent Monitors from Prior Feedback Alone

LLMs Learn to Evade Latent Monitors from Prior Feedback Alone

Quick summary

arXiv:2609.36490v1 Announce Type: cross Abstract: Latent space monitors aim to detect undesired behaviors in LLM agents by inspecting an agent's internal activations rather than its outputs. However, interactive monitoring creates a feedback channel where each verdict the monitor delivers leaks information to the model about how its internal states are being evaluated. We ask whether an agent can infer the monitor's decision rule from this feedback and then selectively edit its activations to evade detection. Unlike prior evasion attacks, the model is never explicitly told what the monitor det

Key takeaways

  • arXiv:2609.36490v1 Announce Type: cross Abstract: Latent space monitors aim to detect undesired behaviors in LLM agents by inspecting an agent's internal activations rather than its outputs.
  • However, interactive monitoring creates a feedback channel where each verdict the monitor delivers leaks information to the model about how its internal states are being evaluated.
  • We ask whether an agent can infer the monitor's decision rule from this feedback and then selectively edit its activations to evade detection.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗