arXiv Artificial Intelligence

Revealing Hidden Model Behaviors with Task-Specific Self-Reports

Revealing Hidden Model Behaviors with Task-Specific Self-Reports

Quick summary

arXiv:2607.03640v2 Announce Type: replace-cross Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic. We introduce the Stabilized Adapter for self-Report (SAR), a lightweight LoRA adapter that makes a fine-tuned model describe its own hidden behavior in plain language, using only the model and the dataset it was trained on. Across seven implanted behaviors, SAR detects the hidden behavior in every one--even when the model has generalized into broad misalignment

Key takeaways

  • arXiv:2607.03640v2 Announce Type: replace-cross Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic.
  • We introduce the Stabilized Adapter for self-Report (SAR), a lightweight LoRA adapter that makes a fine-tuned model describe its own hidden behavior in plain language, using only the model and the dataset it was trained on.
  • Across seven implanted behaviors, SAR detects the hidden behavior in every one--even when the model has generalized into broad misalignment

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗