arXiv Artificial Intelligence

Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning

Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning

Quick summary

arXiv:2609.00351v1 Announce Type: cross Abstract: Large language models can hide hidden behaviors that activate only under narrow conditions, such as backdoor triggers, sleeper-agent deployment cues, sandbagging, or topic-conditioned censorship. Such behaviors are difficult to detect without prior knowledge what to look for. We present activation-matched finetuning, an unsupervised detection method that assumes no knowledge of the trigger or the target behavior. Given a suspect model and a publicly available anchor, we finetune the anchor to reproduce the suspect's activations on a small benig

Key takeaways

  • arXiv:2609.00351v1 Announce Type: cross Abstract: Large language models can hide hidden behaviors that activate only under narrow conditions, such as backdoor triggers, sleeper-agent deployment cues, sandbagging, or topic-conditioned censorship.
  • Such behaviors are difficult to detect without prior knowledge what to look for.
  • We present activation-matched finetuning, an unsupervised detection method that assumes no knowledge of the trigger or the target behavior.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗