Large language models are increasingly becoming research infrastructure for the social sciences, yet their methodological role remains unsettled: can LLM agents serve not only as assistants, but as controlled analyst populations for measuring the robustness of empirical findings? This paper uses LLM agents to scale the many-analyst paradigm in applied microeconomics. Human many-analyst studies show that independent researchers often reach substantially different estimates when analysing the same data and hypothesis, but such studies are expensive, slow, and difficult to decompose because human teams differ on many analytical choices at once. We replace human teams with independently prompted LLM agents and apply the design to twenty-four published papers across difference-in-differences, instrumental variables, randomised controlled trials, and regression discontinuity designs.
For each paper, we extract testable hypotheses from the published text and assign multiple agents to analyse them under six information conditions. Three tiers progressively restrict degrees of freedom by providing raw data, processed data, or a prescribed method. Additional directed conditions ask agents to test the original claim or to search for the most strongly supportive and most strongly rejecting defensible specifications. This design makes analyst discretion observable, auditable, and decomposable into data-preparation, specification, and estimator components.
We synthesise agent estimates using standardised effect sizes and a multilevel random-effects model that separates sampling variation, analyst-within-paper variation, and between-paper heterogeneity. The resulting estimands quantify how much uncertainty in published empirical results is attributable to analytical choice, how strongly LLM agents agree under fixed information conditions, and how far a motivated but defensible analyst could move a result. A pilot on two papers and roughly 320 agent analyses shows that the pipeline produces interpretable variance decompositions, robustness ladders, specification curves, and agent evidence ratings. The project contributes a reproducible benchmark and infrastructure for evaluating LLM-based social-science workflows, shifting AI-assisted replication from occasional case studies toward routine, scalable robustness reproduction.