Evaluating Steering Techniques using Human Similarity Judgments
Quick summary
arXiv:2505.19333v2 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) steering techniques focus on task-specific performance, overlooking how well steered representations align with human cognition. Using a well-established triadic similarity judgment task, we assessed steered LLMs on their ability to flexibly judge similarity between concepts based on size or kind, two central dimensions organizing human mental representations. We found that prompt-based steering methods outperformed other methods both in terms of steering accuracy and model-to-human alignment.
Key takeaways
- arXiv:2505.19333v2 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) steering techniques focus on task-specific performance, overlooking how well steered representations align with human cognition.
- Using a well-established triadic similarity judgment task, we assessed steered LLMs on their ability to flexibly judge similarity between concepts based on size or kind, two central dimensions organizing human mental representations.
- We found that prompt-based steering methods outperformed other methods both in terms of steering accuracy and model-to-human alignment.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments