arXiv Artificial Intelligence

Asymmetries in Spontaneous and Instructed Deception

Asymmetries in Spontaneous and Instructed Deception

Quick summary

arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instruct. We compared these two deception settings through direction geometry, cross-setting classifiers, and cross-setting steering. We found the two deception settings share a component of direction (cosine of approximately 0.5) and an asymmetry in the transfer between settings regar

Key takeaways

  • arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to.
  • However, much of the study on deception in models involves instructed deception.
  • We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instruct.

Why it matters

“Asymmetries in Spontaneous and Instructed Deception” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗