Submission 82
Understanding the Mechanism of Altruism in Large Language Models
Panel 1-Multi-Function Room 1 (19/F LAU)-01
Presented by: Songfa Zhong
Altruism is fundamental to human societies, fostering cooperation and social cohesion. Recent studies suggest that large language models (LLMs) can display human-like prosocial behavior, but the internal computations that produce such behavior remain poorly understood. We investigate the mechanisms underlying LLM altruism using sparse autoencoders (SAEs). In a standard Dictator Game, minimal-pair prompts that differ only in social stance (generous versus selfish) induce large, economically meaningful shifts in allocations. Leveraging this contrast, we identify a set of SAE features (0.024% of all features across the model’s layers) whose activations are strongly associated with the behavioral shift. To interpret these features, we examine their activation profiles on benchmark tasks that exemplify either heuristic (System 1) or deliberative (System 2) processing. Causal interventions validate their functional role: activation patching in this feature direction reliably shifts allocation distributions, with System 2 features generally exerting a more proximal influence than System 1 features. The same steering direction generalizes across multiple social-preference games. Together, these results enhance our understanding of artificial cognition by translating altruistic behaviors into identifiable network states and provide a framework for aligning LLM behavior with human values, thereby informing more transparent and value-aligned deployment.