arXiv Artificial Intelligence

Decomposing and Measuring Evaluation Awareness

Decomposing and Measuring Evaluation Awareness

Quick summary

arXiv:2605.23055v3 Announce Type: replace-cross Abstract: Frontier language models sometimes recognize that they are under evaluation and adjust their behavior which can undermine validity of benchmark results. Yet the field studies it without a shared foundation, conflating flaws of the evaluation with capabilities of the model, and detection with behavioral response. We ground evaluation awareness in social psychology, decomposing it into an environment component and a model component that separates recognition from propensity. We operationalize the environment component through eight catego

Key takeaways

  • arXiv:2605.23055v3 Announce Type: replace-cross Abstract: Frontier language models sometimes recognize that they are under evaluation and adjust their behavior which can undermine validity of benchmark results.
  • Yet the field studies it without a shared foundation, conflating flaws of the evaluation with capabilities of the model, and detection with behavioral response.
  • We ground evaluation awareness in social psychology, decomposing it into an environment component and a model component that separates recognition from propensity.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗