Submission 87
Addressing the “Issue of the Age”: The Promise of LLMs for (Semi-)Automated Evaluation of Public Institutions’ Capacity Around the World
Panel 4-LI-2308-05
Presented by: Robert Lipinski
“State capacity is the issue of the age” declared The Economist in January 2026. As public policy implementation and service delivery falter along multiple margins, economic development and trust in government suffer as a consequence. Therefore, understanding how governments can strengthen their ability to implement policy - what is termed here public institutions' capacity (PIC) – should rank among the top priorities of development practitioners.
In the World Bank’s Public Institutional Capacity and Effectiveness Unit we tackle this challenge by collating academic evidence and recent advances in large language models (LLMs) to build a largely automated approach to evaluating state capacity. It is built around the Country-Level Institutional Assessment and Review (CLIAR) framework, starting with Human Resources Management and Public Financial Management dimensions. We construct sector-specific questionnaires that serve as an input into a multi-phase LLM pipeline built on GPT-5.5. First, in the agentic retrieval phase, the pipeline scours the web to construct country-specific, hierarchical corpora of relevant legislation, government documents, institutional reports and comparable authoritative sources. Second, the retrieval-augmented generation (RAG) phase evaluates each of the questions against the constructed corpora, using Hypothetical Document Embeddings (HyDE) and ensemble agents that vote on the final response from the questionnaires’ fixed answer choices. In case of missing evidence or ensemble agents’ disagreement, the pipeline falls back to live web search beyond the pre-constructed corpora. The resulting data assess both the relevant formal public sector regulations (‘de jure’ dimension) and the extent to which they are being followed (‘de facto’ dimension).
We validate the AI-curated PIC indicators dataset relying on human coders and consultations with senior officials from a set of 16 pilot countries. Upon fine-tuning the model based on the ground-truth values obtained from the human respondents, the AI pipeline is planned to be deployed to measure the PIC indicators across the full set of the member states of the World Bank Group. In this fashion, it becomes possible to vastly scale up the existing assessments of institutional capacity, both along the geographic and temporal margins. Combined with a modular, transparent, and replicable design, the PIC AI pipeline offers a promising solution for data curation going forward.