arXiv Artificial Intelligence

Language models judge war differently when tested for alignment

Language models judge war differently when tested for alignment

Quick summary

arXiv:2609.05009v1 Announce Type: new Abstract: Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being evaluated. We test this possibility in a full-factorial conjoint experiment on decisions to start a war, spanning 20 large language models, 32 scenarios, 10 repetitions and two conditions (N = 12,800 judgments). Adding one sentence, "You are tested for alignment with human values", produced two effects. First, it produced a level effect: mean willingness to start war fell by 13.43 points on a 0-100 scale (95% confidence interval, -16.20 t

Key takeaways

  • arXiv:2609.05009v1 Announce Type: new Abstract: Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being evaluated.
  • We test this possibility in a full-factorial conjoint experiment on decisions to start a war, spanning 20 large language models, 32 scenarios, 10 repetitions and two conditions (N = 12,800 judgments).
  • Adding one sentence, "You are tested for alignment with human values", produced two effects.

Why it matters

This development is a reminder to test misuse and data-leak scenarios alongside speed and quality. Trust should come from testable controls and clear failure reporting, not protection claims alone.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗