Interpreting and Steering LLM Agents for Social Simulations
Quick summary
arXiv:2609.16436v1 Announce Type: cross Abstract: Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. However, LLMs are ultimately black boxes based on deep neural networks which limits their value for social science. This is because of a lack of (i) interpretability: i.e. the ability to assign clear mechanisms driving observed behavior; and a lack of (ii) steerability: i.e. the ability to mute or amplify specific theoretically meaningful mechanisms of action to drive spe
Key takeaways
- arXiv:2609.16436v1 Announce Type: cross Abstract: Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit.
- However, LLMs are ultimately black boxes based on deep neural networks which limits their value for social science.
- This is because of a lack of (i) interpretability: i.e.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments