科学发现旗舰工作 颠覆级 暂无讲解视频
发表时间
2026-07-08
DOI
10.1038/s41586-026-10742-x

收录解读

社会科学实验昂贵且周期较长,研究者通常需要通过小规模预实验或人类预测来筛选干预方案、估计效应并确定后续实验优先级。本文研究大语言模型能否在无法直接接触实验结果的条件下,预测随机调查实验的处理效应,从而辅助实验设计与资源分配。

作者没有要求模型直接猜测论文结论,而是给模型提供虚拟参与者的人口统计资料、实验刺激和结果问题,让其模拟不同条件下的个体回答,再按照真实实验的分析方式计算条件间处理效应。每个实验条件与结果问题使用120个提示词,每个提示采样5次,并将所得虚拟样本预测与真实实验结果、人类预测及多个模型进行比较。

在70项实验、469个效应和119,330名参与者组成的主档案中,GPT-4预测与真实效应达到r=0.85,校正后为0.92;在训练截止后才公开或仍未公开的研究子集上,相关性仍达到0.90以上。其表现接近汇总人类预测,与人类预测组合后进一步提升,但模型系统性地将效应估大约两倍,调整后RMSE仍约为10.59个百分点;在包含606个效应的第二档案中,表现也有所下降。

该研究提供了一个可复用的“虚拟实验参与者”方法和大规模评估档案,显示LLM可用于干预排序、预实验和复制优先级判断,可能改变部分社会科学实验的前期工作流。但证据主要来自美国、文本刺激、自我报告和短期调查实验,且绝对效应校准明显不足,不能替代真实实验,也不能直接外推到实际行为、长期干预、现场实验或其他文化群体。

原始摘要

There is growing interest in how large language models (LLMs) can advance social and behavioral science [1–5]. Prior work has assessed LLMs’ ability to predict survey responses [6–9], but less is known about whether they can predict the outcomes of social science experiments [10], particularly those absent from training data. Here, we built an archive of 70 pre-registered, nationally representative, U.S. survey experiments, involving 469 experimental effects and 119,330 participants. We prompted an LLM to simulate how representative samples of Americans would respond to experimental stimuli, then inferred treatment effects by comparing simulated responses across conditions. Predictions derived from GPT-4, whose training-data cutoff predated the publication of many studies in our archive, were strongly correlated with actual treatment effects, achieving accuracy similar to pooled human forecasts. Correlations remained high for studies not published or publicly posted by the model’s training-data cutoff date, and for predictions from prominent open-weight models. Despite high correlations, predictions systematically overestimated effect sizes. In a secondary archive of 15 megastudies featuring 606 effects, correlations were lower but comparable to pooled expert forecasters. To assess implications for scientific practice, we surveyed 460 social scientists about likely uses and perceived risks, and used our archives to assess several applications (pilot testing, intervention selection, identifying effects needing replication) and risks (bias, misuse). Together, these results suggest LLMs can augment experimental methods in science and practice while raising important considerations for responsible use.

相关论文

链接