15:30 - 17:00
Location: Multi-Function Room 1 (19/F LAU)
Chair/s:
Songfa Zhong
Songfa Zhong - Understanding the Mechanism of Altruism in Large Language Models
Yuwen Zhou - Fairness Versus Efficiency in AI Advice: Evidence from Human and LLM Responses
Yiting Chen - Social Identity and Human-AI Task Allocation
Xiaoli Guo - Salience of Disadvantaged Position Reduces AI Aversion in Moral Delegation
Submission 82
Understanding the Mechanism of Altruism in Large Language Models
Panel 1-Multi-Function Room 1 (19/F LAU)-01
Presented by: Songfa Zhong
Songfa Zhong
Hong Kong University of Science and Technology
Altruism is fundamental to human societies, fostering cooperation and social cohesion. Recent studies suggest that large language models (LLMs) can display human-like prosocial behavior, but the internal computations that produce such behavior remain poorly understood. We investigate the mechanisms underlying LLM altruism using sparse autoencoders (SAEs). In a standard Dictator Game, minimal-pair prompts that differ only in social stance (generous versus selfish) induce large, economically meaningful shifts in allocations. Leveraging this contrast, we identify a set of SAE features (0.024% of all features across the model’s layers) whose activations are strongly associated with the behavioral shift. To interpret these features, we examine their activation profiles on benchmark tasks that exemplify either heuristic (System 1) or deliberative (System 2) processing. Causal interventions validate their functional role: activation patching in this feature direction reliably shifts allocation distributions, with System 2 features generally exerting a more proximal influence than System 1 features. The same steering direction generalizes across multiple social-preference games. Together, these results enhance our understanding of artificial cognition by translating altruistic behaviors into identifiable network states and provide a framework for aligning LLM behavior with human values, thereby informing more transparent and value-aligned deployment.