15:30 - 17:00
Location: Senate Room (19/F LAU)
Chair/s:
Le Bao
Discussant/s:
David Broska
Na Liu - Information conditions govern the validity of AI-generated experimental data across inferential targets
Plamen Akaliyski - When AI Thinks about Culture: Comparing Language-Based and Empirical Measures of Individualism-Collectivism
Leo Yang - Scaling Reproducibility: An AI-Assisted Workflow for Large-Scale Replication and Reanalysis
Simon Maier - Replication and Specification Range Analysis with AI Agents
Submission 78
Information Conditions Govern the Validity of AI-generated Experimental Data Across Inferential Targets
Panel 1-Senate Room (19/F LAU)-01
Presented by: Na Liu
Na Liu, Yue Wang, Lingling Hou
Peking University, China Center for Agricultural Policy, School of Advanced Agricultural Sciences, Beijing, 100871, China.
Whether AI-generated experimental data supports valid statistical inference depends on what the researcher intends to estimate and what information the model receives. Existing evaluations report contradictory findings because they assess different inferential targets against different benchmarks. Aggregate replication succeeds while individual prediction fails, but this reflects the target's statistical demand, not a disagreement about AI capability. Here we develop a multi-target evaluation framework tested on an incentive-compatible economic experiment with individual-level ground truth (N = 911, randomized treatment assignment). Seven large language models are assessed under four information conditions against four targets of increasing demand: distributional fidelity, average treatment effects, individual prediction, and heterogeneous treatment effect identification. Three regularities emerge. First, information dominates architecture. Varying what models know changes performance by 45 to 60 percentage points, while varying which model is used changes it by at most 12 points. Second, validity degrades monotonically across the target hierarchy, with one exception: reasoning models detect average treatment effects without individual data where standard models cannot. Third, optimizing for one target impairs another. Fine-tuning improves distributional matching but degrades treatment effect detection. These results resolve the apparent contradiction in existing literature and provide a replicable evaluation protocol for any domain where individual-level ground truth is available.