10:15 - 11:45
Location: LAU 5-203
Chair/s:
Zeyu Lyu
Discussant/s:
Vincent Xiaohui Wang
Quang Phuc Phung - Do Personas That Sound Right Also Choose Right?
David Broska - Fine-Tuned LLMs Forecast Survey Experiment Effect Sizes
Zeyu Lyu - Controlling LLM Agent Personality and Preferences through Representation Engineering
Yi Ching Victoria LAI (presenter), Shanshan ZHENAI - Hyperrealism: A Behavioural and Neural Investigation of Real and Synthetic Human Faces
Submission 101
Controlling LLM Agent Personality and Preferences Through Representation Engineering
Panel 2-LAU 5-203-03
Presented by: Zeyu Lyu
Zeyu Lyu, Zhichao Wang
Graduate School of Arts and Letters, Tohoku University
Large language model (LLM) agents offer a flexible approach to constructing "silicon samples,'' with considerable potential for social science research. However, their behavior is often sensitive to prompt wording, difficult to interpret, and biased toward model-default preferences. We investigate activation steering as a means of controlling agent personalities and preferences by intervening in the model’s internal representations. Using Llama-3.1-8B-Instruct, we construct steering vectors and apply them during inference at varying scalar coefficients. We illustrate this approach through two proof-of-concept studies. First, we manipulate a latent personality trait and show that steering can systematically shift agents’ behavior across multiple behavioral games. Second, we manipulate a model-default preference and show that steering can alter both individual choices and collective norm trajectories. Our findings highlight the potential of representation engineering to enhance the interpretability and reproducibility of LLM-based social science research.