10:15 - 11:45
Location: Multi-Function Room 2 (19/F LAU)
Chair/s:
Sascha Riaz
Songpo Yang - An Expert–Coder Multi-Agent System for Codebook-Based Text Annotation in Social Science
Saera Lee - Can Generative AI Reliably Code Treaty Texts? Evidence from Nuclear Non-Proliferation Treaties
Sascha Riaz - The Disappearing Conflict: Tracking Political Attention to the Israeli-Palestinian Conflict Across Five Million Speeches
Yusuf Evirgen - From Reports to Data: Harnessing LLMs for Fine-Grained Human Rights Data Collection
Submission 51
An Expert–Coder Multi-Agent System for Codebook-Based Text Annotation in Social Science
Panel 2-Multi-Function Room 2 (19/F LAU)-01
Presented by: Songpo Yang
Songpo Yang 1, Yang Wu 2, Zhicheng Zhang 2, Xun Pang 1
1 Peking University
2 University of Chinese Academy of Sciences
Codebook-based text coding is a foundational task in empirical social science, yet remains bottlenecked by its dependence on trained human coders who must process lengthy, multilingual documents under heavy cognitive load. Large language models (LLMs) offer a promising alternative, but direct application produces unreliable results when codebook categories involve ambiguous conceptual boundaries or require contextual judgment. This paper proposes a multi-agent LLM framework that replicates the division of labor between human experts and coders. The framework transforms conventional codebooks into sequenced Query Workflows, implements a dual-agent architecture with a Coder Agent for extraction and an Expert Agent for interpretive oversight, and introduces a Query-Feedback loop in which conceptual ambiguities are escalated for adjudication. A task scheduler leveraging asynchronous calls and KV cache management ensures scalability. We validate the framework on two tasks: coding investment screening mechanisms using the PRISM Dataset (Princeton University), covering 38 OECD countries and 130 legal texts, and coding interstate interactions from UNFCCC climate negotiation reports building on Castro et al. (2025). These cases test factual precision on structured policy documents and semantic inference on deliberately ambiguous diplomatic language, respectively. Preliminary results indicate that the framework achieves inter-coder reliability comparable to trained human coders while reducing costs by an order of magnitude. We further find that both the Query Workflow transformation and Expert Agent consultation contribute independently to coding quality, and that naive single-pass prompting produces substantially lower reliability.