Probing Persona-Dependent Preferences in Language Models
Quick summary
arXiv:2605.13339v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can be said to have preferences: they reliably pick certain tasks and outputs over others, and preferences shaped by post-training and prompting appear to influence much of their behaviour. But models can also adopt different personas which have radically different preferences. How is this implemented internally? Does each persona use its own preference representations, or are some representations shared? We train linear probes on residual-stream activations of Gemma-3-27B and Qwen-3.5-122B to predict reveal
Key takeaways
- arXiv:2605.13339v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can be said to have preferences: they reliably pick certain tasks and outputs over others, and preferences shaped by post-training and prompting appear to influence much of their behaviour.
- But models can also adopt different personas which have radically different preferences.
- How is this implemented internally?
Why it matters
“Probing Persona-Dependent Preferences in Language Models” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Member comments