Submission 93
Fine-Tuned LLMs Forecast Survey Experiment Effect Sizes
Panel 2-LAU 5-203-02
Presented by: David Broska
Pilot studies help researchers identify promising treatments and estimate effect sizes before conducting full experiments. Yet piloting is often too slow or costly for many labs. Large language models could lower that cost. Prompted frontier models already forecast treatment effects that correlate strongly with observed ones (r = .85; Ashokkumar et al. 2026), enough to rank candidate treatments, but they overestimate effect magnitudes and compress outcome variance, which rules out uses that depend on calibrated forecasts. We test whether fine-tuning closes this calibration gap.
We test whether fine-tuning an LLM on past survey experiments can provide useful, low-cost forecasts of treatment effects. We built an LLM-based simulator that reconstructs the Qualtrics survey each participant saw, including treatments, questions, and prior answers, then fine-tunes an open-weight LLM to predict responses sequentially. Training used 80 text-based between-subjects survey experiments with 56,177 participants. We evaluated the model on 13 held-out studies by walking it through each survey and comparing simulated with observed effects for every post-treatment outcome. Across 515 effects, simulated and observed Cohen’s d estimates were highly correlated (r = .94) and closely aligned in magnitude; among statistically significant human-sample effects, the simulator recovered the observed direction in 95% of cases. These results suggest that fine-tuned LLM simulations can support affordable piloting, power analysis, and prioritization of interventions before full human data collection.
The next project phase tests whether the approach generalizes beyond a single lab. The next step is to expand the training and test archive with datasets from other researchers. For training, this will let the model learn from a broader range of experimental designs, populations, and outcomes it has not yet seen. For testing, private data play an even more distinctive role. Because modern LLMs train on the public internet and retrieve published results through web search, experiments that never appeared online are the cleanest test of whether LLMs forecast results rather than remember them. We are therefore inviting a limited group of 30 social scientists to contribute study materials and response data from survey experiments. The talk presents the fine-tuning approach, the proof-of-concept results, and the research design of this team science project.