As traditional surveys grow more expensive and response rates collapse, large language models offer a provocative alternative: inferring public opinion from the digital traces people already leave behind. This hands-on workshop walks participants through PoSSUM — our Protocol for Surveying Social Media Users with Multimodal LLMs — the AI polling method whose state-by-state forecasts of the 2024 US presidential election tracked, and at times outperformed, the leading poll aggregators. We move through the full pipeline: designing the digital interview, building a “silicon” sample of social-media users, and producing bias-corrected estimates with multilevel regression and post-stratification (MrP) in R, validated against ground-truth election results. A featured sub-theme is our Swiss “Silicon Politicians” study, which predicts how individual politicians and citizens vote in referendums from their social-media traces alone — and showcases the app we built around it. Throughout, we keep the harder question in view, drawing on the recent APSA report Public Opinion in the Age of AI: when does simulating respondents enrich measurement, and when does it risk manufacturing the opinion it claims to observe?
Aimed at pollsters, political scientists, and survey methodologists comfortable with R; no machine-learning background required.
| Time | Session |
|---|---|
| 09:00–09:30 | Why AI polling? |
| 09:30–10:10 | PoSSUM and the 2024 US election case |
| 10:25–11:30 | Hands-on: Build a silicon sample and run MrP |
| 11:30–12:05 | Swiss 'Silicon Politicians' demo |
| 12:05–12:30 | Discussion & Q&A |
Researchers in academia and industry increasingly propose using LLM predictions of human behavior to pilot, augment, or even replace human data collection. When do these predictions support valid inferences, and when do they mislead? This workshop introduces a practical toolkit for answering that question and for using them in scientifically defensible ways. Drawing on recent frameworks for using and validating LLM predictions as behavioral evidence (Broska et al. 2025; Hullman et al. 2026), participants will learn to assess when predicted responses can be trusted, recognise common failure modes, and combine human and synthetic samples for greater precision without introducing bias. Used carefully, predictions do not replace human samples but help researchers design more informative studies.
Attendees will leave with the concepts and hands-on experience needed to decide whether and how to incorporate LLM predictions into their own research.
The afternoon extends the morning’s method beyond elections. The same architecture — unobtrusively harvesting someone’s public posts, repeatedly interviewing an LLM “digital twin” built from them, and applying panel econometrics to the result — turns scattered, self-selected commentary into a balanced, forward-looking panel of on-demand forecasters. We work through three applications. In financial markets, we build digital twins of “finfluencers” and interview them daily, recovering their stock-level beliefs even on days they post nothing, and show that these signals predict the cross-section of S&P 500 returns without look-ahead bias (drawing on our Talking to Digital Twins study, Bowles et al., 2026). We then turn to forecasting and prediction-market applications, and to a measurement problem the method is unusually suited to: eliciting views on sensitive topics people are reluctant to volunteer — treating silence itself as a belief state. Participants build a twin, run a repeated-interview protocol, assemble the panel, and evaluate a simple forecast.
Aimed at financial economists, quantitative and survey researchers, and data scientists comfortable with R; the morning session is helpful but not required.
| Time | Session |
|---|---|
| 13:30–13:55 | From polls to markets |
| 13:55–14:30 | Finfluencer twins and S&P 500 signals |
| 14:30–15:15 | Hands-on: Build a digital twin |
| 15:15–15:30 | Backtesting and prediction markets |
This capstone session turns the digital twin from a measurement instrument into a design tool. The same architecture that reads opinion can be run as a closed loop: propose a message, have a panel of LLM twins evaluate it, and search for the version that resonates best with a target audience. We work the idea on two grounds. First political campaigns — generating and optimising party platforms against a twin panel. Then the courtroom, where the same machinery is re-pointed at legal argument: twins of jurors and judges, a synthetic mock trial, and verdict probabilities read as the case is re-argued with different opening framing, evidence order and narrative emphasis.
Participants work hands-on with a pre-built twin panel, putting a set of candidate messages to it, aggregating the evaluations and identifying which performs best — one pass of the optimisation loop, done by hand. We then demonstrate the full closed loop running over a much larger candidate space, and explain how the search works. Throughout we keep validity and ethics in view: how to validate synthetic jurors against the experimental mock-trial literature, how the contamination of well-known cases threatens prediction, and why these are tools for stress-testing arguments before they are made, never for replacing adjudication.
Aimed at researchers interested in message testing, persuasion and strategy — across political communication, legal research and computational social science. Participants should be comfortable running R notebooks; no legal or machine-learning background required. The morning modules are helpful but not required.
| Time | Session |
|---|---|
| 15:45–16:10 | Closed-loop optimisation |
| 16:10–16:50 | Political platform optimisation |
| 16:50–17:30 | Juror and judge twins |
| 17:30–17:45 | Ethics, validation & Q&A |