AI Revealed Preferences
Quick summary
arXiv:2608.26178v2 Announce Type: replace Abstract: There is growing interest in whether language models have stable preferences, for technical, safety, and philosophical reasons. We test 20 language models and find a range of preferences---stable dispositions to choose certain kinds of tasks. We run three forced-choice experiments on revealed rather than stated preferences, requiring models not only to rank tasks, but to actually perform them. Headline findings include evidence that models are tedium-averse, "leisure"-seeking, and covertly sycophantic. Tedium aversion means that, when tasks a
Key takeaways
- arXiv:2608.26178v2 Announce Type: replace Abstract: There is growing interest in whether language models have stable preferences, for technical, safety, and philosophical reasons.
- We test 20 language models and find a range of preferences---stable dispositions to choose certain kinds of tasks.
- We run three forced-choice experiments on revealed rather than stated preferences, requiring models not only to rank tasks, but to actually perform them.
Why it matters
“AI Revealed Preferences” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Member comments