Asymmetries in Spontaneous and Instructed Deception
Quick summary
arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instruct. We compared these two deception settings through direction geometry, cross-setting classifiers, and cross-setting steering. We found the two deception settings share a component of direction (cosine of approximately 0.5) and an asymmetry in the transfer between settings regar
Key takeaways
- arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to.
- However, much of the study on deception in models involves instructed deception.
- We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instruct.
Why it matters
“Asymmetries in Spontaneous and Instructed Deception” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments