Submission 45
Can Government Chatbots Be Equitable Without Becoming Rigid? Evidence from Germany
Panel 4-LI-2308-02
Presented by: ming ma
While large language models (LLMs) have been rapidly integrated into public service delivery, existing scholarship presents competing accounts of whether it will replicate societal biases embedded in its training data or whether such biases can be mitigated through institutional and technical interventions. Yet these mitigation efforts may lead to overly rigid systems that fail to adapt to legitimate citizen needs. To examine this tension between inequality and rigidity, we develop an algorithmic administrative language framework and conduct a large-scale audit experiment comprising 19,200 queries across 12 municipal and job center chatbots in Germany. We systematically manipulated user personas (gender and ethnicity) and language proficiency signals to test for equality and responsiveness. We find that disparities in chatbot responses are manifested primarily in the informational quality of responses and are concentrated among specific identity groups. Conversely, contrary to expectations of algorithmic rigidity, chatbots demonstrate responsiveness to explicit requests for simplified language. These findings suggest that while LLMs possess the adaptive capability to reduce administrative burdens, reproduction of identity-based biases is a persistent challenge that deserves more attention.