Lezhi Tan, Tijana Zrnic
Introduces the condition of 'task exchangeability' to guarantee the validity of statistical inference using synthetic data, and develops valid inference methods when this condition holds.
The use of synthetic data in scientific research is increasing, such as LLM-generated 'silicon samples', 'LLM-as-a-judge' in AI evaluation, and generative models for protein structures. However, synthetic data can lead to unreliable inference due to bias, noise, and model misspecification. Previously, there was no general principle ensuring the validity of inference using synthetic data.
The authors introduce a new condition called 'task exchangeability'. This requires that the researcher can identify historical tasks for which real data is available, such that the current task of interest is exchangeable (i.e., identically distributed in an appropriate mathematical sense) with the historical tasks. Under this condition, they develop valid inference methods using synthetic data, and also provide extensions that work even when exchangeability is violated. Empirically, they apply the framework to public opinion surveys with silicon samples and AI evaluation with autoraters.
The proposed framework provides provable validity guarantees for statistical inference using synthetic data. It establishes a theoretical foundation for the scientific use of synthetic data and demonstrates effectiveness in real applications such as opinion surveys and AI evaluation. It is expected to contribute to increasing the reliability of research methodologies.